Compare commits

...
Author SHA1 Message Date
Chris Lu 1cb659019e shell: reset noLock in the admin script dispatcher too
Three dispatchers reuse one CommandEnv: the interactive shell, the master's
maintenance script runner, and the plugin worker's admin script handler. The
third is exempt anyway because it calls ForceNoLock, so this changes nothing
today -- it keeps the rule the same everywhere rather than resting on that
exemption staying in place.
2026-08-31 11:39:09 -07:00
Chris Lu 9ca9dd2379 shell: keep noLock to the command that set it
noLock says "this invocation changes nothing" -- volume.balance, volume.move,
volume.copy, volume.merge and volume.fix.replication all set it for a dry run,
and none clears it. The CommandEnv is created once and reused by both
dispatchers, the interactive shell and the master's maintenance script runner,
so a simulation left every later command unlocked:

    volume.balance -noLock          # changes nothing
    volume.move ...                 # mutates, and skips its lock

Reset before dispatch in both, where the invocation begins. forceNoLock is
untouched: that is set once, deliberately, for a trusted path.
2026-08-31 11:17:04 -07:00
Chris LuandGitHub 87474c2f21 s3: let attached policies authorize CreateBucket (#11049)
* s3: resolve admin bucket subresources to their specific S3 actions

Encryption, requestPayment, publicAccessBlock and ownershipControls
requests reached the policy engines as s3:*, so only a policy granting
all of s3 could authorize them. Map each subresource to its AWS action,
with DELETE sharing the PUT permission as AWS does.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ

* s3: authorize CreateBucket as s3:CreateBucket in the policy engine

A plain bucket-level PUT is registered with ACTION_ADMIN, which resolved
to s3:*, so no attached policy short of s3:* could match it. Federated
sessions whose policy explicitly allowed s3:CreateBucket were always
denied while the same policy worked for object operations. Resolve it to
s3:CreateBucket, like DeleteBucket already resolves.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ
2026-08-31 10:24:22 -07:00
Chris LuandGitHub b8049bc633 shell: keep fs.mergeVolumes from spinning past the finished moves (#11050)
* filer_pb: walk a re-delivered directory only once in TraverseBfs

A directory handed back twice by a listing (a page-boundary race with
concurrent renames, or a store whose ordering misbehaves) was enqueued
twice; the second walk re-lists the same subtree and can keep the
traversal from ever terminating.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ

* filer_pb: fail a directory listing whose pagination stops advancing

A full page ending on the very name the cursor started from re-fetches
the same page forever; a store whose listing order does not advance past
the cursor turns any full-directory read into a silent infinite loop.
Return an error naming the stuck cursor instead.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ

* shell: skip foreign-collection manifests in fs.mergeVolumes

Every manifest chunk in the namespace was resolved, downloading its
manifest needle, even when the merge plan only touches one collection.
Sub-chunks live in the manifest's own collection, so a manifest on a
volume outside the plan's collections cannot reference a source volume;
skip it and spare a cluster-wide download pass that looks like a hang
after the real moves finish.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ
2026-08-31 10:23:35 -07:00
Chris LuandGitHub 23d424248d sts: session duration no longer clamped to the web identity token exp (#11048)
* sts: session duration no longer clamped to the web identity token exp

The assumed-role session lifetime is governed by DurationSeconds and the
configured tokenDuration/maxSessionLength, matching AWS. Clamping to the
already-verified token's exp made short-lived id_tokens (GitLab issues
~2-minute ones) yield unusable sessions regardless of configuration.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ

* sts: cover session duration against short-lived web identity tokens

The mock OIDC provider now carries the token exp through to the identity
like the real provider, so the integration test would catch the clamp.

Claude-Session: https://claude.ai/code/session_01XH7iM88ZqWMEvsLB8tkWPQ
2026-08-31 10:22:06 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
49ee13635b build(deps): bump google.golang.org/grpc from 1.84.0-dev.0.20260723093437-b6eac429d7b6 to 1.85.0-dev (#11043)
build(deps): bump google.golang.org/grpc

Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.84.0-dev.0.20260723093437-b6eac429d7b6 to 1.85.0-dev.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/commits/v1.85.0-dev)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.85.0-dev
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-08-31 10:06:36 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
8bd5ec37d1 build(deps): bump google.golang.org/api from 0.293.0 to 0.294.0 (#11044)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.293.0 to 0.294.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.293.0...v0.294.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.294.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:50:08 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1971fb8cb0 build(deps): bump github.com/aws/aws-sdk-go-v2 from 1.43.5 to 1.45.1 (#11046)
Bumps [github.com/aws/aws-sdk-go-v2](https://github.com/aws/aws-sdk-go-v2) from 1.43.5 to 1.45.1.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.43.5...v1.45.1)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2
  dependency-version: 1.45.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:49:58 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
8c3695fd4a build(deps): bump actions/setup-java from 5 to 6 (#11047)
Bumps [actions/setup-java](https://github.com/actions/setup-java) from 5 to 6.
- [Release notes](https://github.com/actions/setup-java/releases)
- [Commits](https://github.com/actions/setup-java/compare/v5...v6)

---
updated-dependencies:
- dependency-name: actions/setup-java
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:49:46 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3a3d513cb8 build(deps): bump github/codeql-action from 4.37.8 to 4.37.9 (#11045)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.8 to 4.37.9.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.8...v4.37.9)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.9
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:49:16 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
86ee10a080 build(deps): bump modernc.org/sqlite from 1.56.0 to 1.57.0 (#11042)
Bumps [modernc.org/sqlite](https://gitlab.com/cznic/sqlite) from 1.56.0 to 1.57.0.
- [Changelog](https://gitlab.com/cznic/sqlite/blob/master/CHANGELOG.md)
- [Commits](https://gitlab.com/cznic/sqlite/compare/v1.56.0...v1.57.0)

---
updated-dependencies:
- dependency-name: modernc.org/sqlite
  dependency-version: 1.57.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:48:56 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
909f5cabc7 build(deps): bump github.com/ydb-platform/ydb-go-sdk/v3 from 3.147.1 to 3.151.1 (#11041)
build(deps): bump github.com/ydb-platform/ydb-go-sdk/v3

Bumps [github.com/ydb-platform/ydb-go-sdk/v3](https://github.com/ydb-platform/ydb-go-sdk) from 3.147.1 to 3.151.1.
- [Release notes](https://github.com/ydb-platform/ydb-go-sdk/releases)
- [Changelog](https://github.com/ydb-platform/ydb-go-sdk/blob/master/CHANGELOG.md)
- [Commits](https://github.com/ydb-platform/ydb-go-sdk/compare/v3.147.1...v3.151.1)

---
updated-dependencies:
- dependency-name: github.com/ydb-platform/ydb-go-sdk/v3
  dependency-version: 3.151.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-31 09:48:47 -07:00
f740210235 get volume topology info without volume details (#11036)
* get volume topology info without volume details

Signed-off-by: lou <alex1988@outlook.com>

* master: rename VolumeListRequest.without_volumes to topology_only

The field shapes the reply rather than selecting volumes, and it leaves
out the ec shards too, which the old name denied. Match the message's
*_only style and say what a master that predates the field does with it.

Claude-Session: https://claude.ai/code/session_01QHnaNRgxnjzZsiz7WTFML5

* master: refuse topology_only combined with a volume selector

A topology_only request that also names a collection or volume ids
contradicts itself, and answering either half in silence surprises the
caller. Answer InvalidArgument from both VolumeList and its stream,
before the stream sends its header.

Claude-Session: https://claude.ai/code/session_01QHnaNRgxnjzZsiz7WTFML5

---------

Signed-off-by: lou <alex1988@outlook.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-31 09:45:22 -07:00
Chris LuandGitHub 32df246a81 mq: fix idle-cleanup shard deadlock that permanently wedges the broker's topic map (#11051)
mq: remove emptied topics after the cleanup iteration, not inside it

cleanupIdlePartitions called manager.topics.Remove from inside
manager.topics.IterCb. IterCb holds the shard's read lock while running
the callback, and Remove takes the same shard's write lock, so removing
an emptied topic self-deadlocked the cleanup goroutine. The pending
writer then blocked every later reader of that shard, permanently
hanging ListTopicsInMemory and, for shard-mates, TopicExistsInMemory.

On the Kafka gateway this surfaced as flaky e2e consumer-group tests:
one minute after any earlier topic went idle, the broker's first
'Removing empty topic' wedged the map, every gateway
ListTopics/TopicExists RPC burned its full 5s timeout, Metadata could no
longer finish inside kafka-go's 5s coordinator deadline, and consumer
groups looped in PreparingRebalance until the test timed out.

Collect the emptied topic keys during the iteration and remove them
afterwards via RemoveCb, re-checking emptiness under the shard lock so a
topic that just gained a partition is kept.

Claude-Session: https://claude.ai/code/session_014yA6c8JQcY6MqPXCT13yYA
2026-08-31 09:44:34 -07:00
Chris LuandGitHub d3b8030a69 master: shed assigns retryably until volume servers register capacity (#11032)
An assign arriving before any volume server has heartbeated saw zero
available space and failed outright with a plain error no client retries,
so the first write to a fresh bucket answered 500 while the cluster was
still starting. Distinguish a topology with no registered capacity from a
genuinely full one: fail fast only when registered capacity is exhausted,
and shed ResourceExhausted otherwise so the client's retry budget rides
out the startup window.

Claude-Session: https://claude.ai/code/session_018G9kWFgy8BaBAEkYV3YL9n
2026-08-30 11:08:53 -07:00
Chris LuandGitHub 9bafeb6139 ec: refuse to mount a 0-byte shard file when the index has entries (#11030)
* ec: refuse to mount a 0-byte shard file when the index has entries

The startup scan already skips (and eventually deletes) zero-sized shard
files as residue of a failed copy, but the mount RPC path opens the file
directly with no size check, so an explicit VolumeEcShardsMount over a
truncated file registers a size-0 claim. A registered empty shard serves
nothing while advertising ownership: with placement pinned to the owning
disk, it would keep attracting re-copies to a file that was never valid.

The one legitimate 0-byte shard is the empty volume's: encoding a volume
with no live needles produces a 0-byte .ecx and 0-byte shards, and that
mount must keep working (TestMountEcShards_EmptyEcxMountsSuccessfully).
So the gate compares against the index: AddEcVolumeShard (Go) and
EcVolume::add_shard (Rust) refuse a 0-byte shard file only when the
volume's .ecx has entries. Go's AddEcVolumeShard grows an error return
for this; the loader cleans up the refused shard and, when it just
created the EcVolume, unregisters that too. The mount loop already
collects non-ENOENT failures per disk and keeps scanning, so a sibling
disk holding a real copy still wins.

Regression tests in both trees: an empty shard beside an index with
entries is refused and leaves nothing registered; an empty shard of an
empty volume still mounts.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9

* ec: release the duplicate shard when a mount retry re-loads it

Review follow-up: AddEcVolumeShard keeps the existing shard and reports
added=false for a shard this disk already registered, but the loader
discarded that result, so every retried LoadEcShard leaked the duplicate
it had just opened — an fd and a mount-gauge increment per retry. Release
both and return the existing volume. Regression test pins the gauge.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9

* ec: close the test DiskLocation instead of only its EC volumes

Review follow-up: DiskLocation.Close() also stops the background
goroutine NewDiskLocation starts; closeEcVolumes left it running for the
rest of the test process. Both uses are this PR's own tests.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9

* rust: unregister the just-created EcVolume when its first mount is refused

Review follow-up: when the first mount of a volume rejects its shard
(e.g. the new 0-byte-beside-nonempty-index refusal), the Rust mount path
had already inserted the EcVolume and propagated the error without
removing it — a zero-shard registration advertising a mount that serves
no data while pinning the .ecx/.ecj descriptors (and, since placement's
mounted tier keys off it, steering shard placement at this disk). Remove
it on the way out, exactly as the Go loader already does; a volume that
already holds shards keeps them (the RPC's first-error-aborts contract).
Regression test covers both.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9

* rust: skip already mounted shards on a mount retry

Review follow-up: EcVolume::add_shard replaces self.shards[id] for a
shard the volume already holds, and the mount loop then bumps the
ec_shards gauge although the mounted count did not grow — gauge drift on
every mount retry, and a serving fd swapped for no reason. Skip shard
ids the volume already reports, mirroring Go's AddEcVolumeShard
added=false handling. Regression test pins the gauge across a duplicate
mount (unique collection label: the gauge is process-global and tests
run in parallel).

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9
2026-08-29 14:19:16 -07:00
Chris LuandGitHub 74b520113e ec: pin auto-selected shard placement to the disk that already owns the shard (#11029)
* ec: pin auto-selected shard placement to the disk that already owns the shard

A multi-disk server legitimately mounts one EC volume on several disks, so
FindEcShardTargetLocation's per-volume tiers tie at "mounted" and the
free-shard-count tie-break decides — pointing at whichever disk is emptier,
not at the disk that already holds the shard being placed. A re-copy of a
shard the server already has (a retried ec.balance / ec.rebuild move) then
lands on a sibling disk, and both disks register the same (volume, shard id):
the shard is reported to the master from two disk ids, and which claimant
serves reads or survives a later unmount/delete becomes an accident of
Locations order.

Add a tier above "mounted": a disk that already claims one of the shard ids
being placed wins, ahead of the space filters too — re-copying in place
needs no new shard slot, and a genuinely full disk should fail the write
rather than silently split the claim. Applied to the Go selector and the
VolumeEcShardsCopy auto-select (ReceiveFile refuses mounted EC volumes, so
no claim can exist there) and mirrored in the Rust volume server.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9

* ec: refuse a copy batch whose shards are already owned by different disks

Review follow-up: ownership-aware selection ranks a mixed-owner batch
(shard 0 on disk A, shard 2 on disk B — the legitimate multi-disk spread)
into one destination, so the copy would still duplicate the losing disk's
claim. No production caller sends such a batch (balance moves one shard,
rebuild and encode copy shards the target lacks), so fail closed: report
every owning disk via Store.EcShardOwnerDisks and refuse the copy with an
error naming them, telling the caller to split per shard or pass disk_id.
Go and Rust, with unit tests for the owner-reporting contract.

Claude-Session: https://claude.ai/code/session_01AWpefvdi4U3HLng18x5CJ9
2026-08-29 12:12:41 -07:00
Lorentz KindeandGitHub 83a929720a fix: Move cluster.LiveLock.generation to first field fix 32bit alignment (#11028)
On 32-bit systems the generation fields becomes unaligned. Not an issue
on 64bit.

Discovered on 32bit raspbian, validated by manual patch (in addition to
test).
2026-08-29 11:38:16 -07:00
Chris LuandGitHub 88c873ecd4 ec: uniform shard block layout (#10932)
* ec: uniform shard block layout

An EC volume is striped as 1GiB blocks until less than one row remains, then
1MiB blocks, and consecutive blocks land on different shards. With ec.encode's
-fullPercent 95 against the 30GiB default limit, ~30% of every volume sits in
that 1MiB tail, so a 4MB filer chunk there is five stripes on five servers.

New encodes now use one block per shard, sized ceil(datSize/dataShards) rounded
up to 1MiB and recorded in the .vif (EcShardConfig.block_size, also carried by
the .ecsum manifest). A needle now maps to one shard unless it is larger than
the block or straddles a boundary. The chosen size equals the legacy layout's
padded shard length for every input, so shard sizes, capacity math, and the
shard-size credibility checks are unchanged; only the byte placement moved.

Reads, decode, and scrub resolve the block sizes from the volume's .vif;
absence keeps the legacy interpretation, so existing EC volumes read exactly as
before. Rebuild is layout-agnostic. weed fix -ecx recovers the layout from the
.vif, else the .ecsum sidecar, and with neither de-stripes under both candidate
layouts and keeps the one that indexes more valid needles.

Same change in the Rust volume server, which now also streams the encode in
256KB sub-batches like Go instead of allocating whole blocks, and computes the
large-row count as shardSize/largeBlock to match Go on exact multiples. On a
26MB fixture both encoders produce byte-identical shards, and a Go-written .vif
parses in Rust with the block size intact.

* ec: resolve the rust ecx rebuild through the recorded layout

The Rust rebuild path regenerated a lost .ecx by scanning the logical .dat
through a hand-rolled pure-1MiB striping, which was already wrong for legacy
volumes with large-block rows and is wrong for any uniform volume with a block
past 1MiB. Route the scan through locate_data with the .vif-recorded block
size, the same mapping the read path uses. Also seed the new tests' random
data instead of the deprecated global math/rand.Read.

* ec: fail the Rust ecx rebuild on any shard read error

A read error mid-scan published the entries collected so far as a
successful .ecx, and read_at's byte count was ignored so a legal short
read passed as complete — a truncated or failing shard could produce a
silently incomplete recovery index. Exact-read semantics in
read_from_data_shards, error propagation in the needle walk, and a
truncated-shard regression test.

* ec: fail the mount on an unreadable or malformed vif

Both servers silently fell back to the legacy layout when an existing
.vif could not be read or parsed. Every new encode records a positive
uniform block size there, so the fallback mounted the same shards with
legacy offset math and could return wrong data. Absent stays legal
(legacy volumes predate the sidecar), and a zero-byte stub still reads
as absent (Go's MaybeLoadVolumeInfo convention, now mirrored in Rust);
a present-but-unreadable or malformed .vif fails the mount instead.

* ec: bound the reconstruct fan-out of one needle's intervals

A degraded interval fans out a read to every reachable shard location, each
with a buffer the size of the interval. Reading a needle's intervals in
parallel multiplied that by the interval concurrency: a needle spanning 8
blocks could hold 8 x MaxShardCount remote reads and buffers at once, where
the sequential version peaked at MaxShardCount. Give each needle a single
reconstruct budget its intervals share, held for the buffer's lifetime, so
separate reads stay independent but one read cannot multiply its own
fan-out.

* ec: drop the duplicated shard-size formula

calculateExpectedShardSize reimplemented the padding rule that
UniformBlockSize already owns — TestUniformBlockSizeMatchesLegacyShardSize
asserts the two agree for every input — so a change to the rule would have
had to be made in both. Defer to the helper, keeping the historic answer for
an empty .dat.

* ec: resolve the shard block layout from whatever records it

Four places still answered the layout question by inference when a record of
it was available, or accepted an answer that was not one:

- A mount with no .vif defaulted to the legacy layout; the bitrot sidecar
  records the same config at encode time, so take it when present, as
  weed fix -ecx already does. The vif itself is now parsed once per mount
  rather than twice.
- The Rust ecx rebuild derived its row count from the padded shard extent,
  which under the legacy layout reads a shard that is an exact large-block
  multiple as one row too many. Pass the encode-time .dat size from the .vif
  and keep the extent as the fallback.
- weed fix -ecx read the block size outside the EC-config guard (collapsing
  the unknown sentinel into a definitive legacy), only wrote the recovered
  layout back when the .vif was absent rather than unusable, and broke a
  scan tie by candidate order instead of the documented reach.
- The uniform layout tripped writeDatFile's large-block ambiguity guard,
  which cannot apply when the large and small blocks are the same size.

* ec: give the index-recovery tests a parseable vif

The fixtures wrote the literal bytes "volinfo" as the source .vif and the
recovery copies it verbatim, so the receiving server then mounted the volume
from a .vif it could not parse. That used to pass by silently defaulting to
the legacy layout; a mount now refuses a vif it cannot read, which is what
the tests were exercising all along without meaning to.

* ec: validate the layout a vif records, not just its syntax

Review follow-ups on the mount-strictness change:

- A .vif can parse and still record a block size no encoder could have
  produced (negative, or not a whole number of small blocks). Both servers
  took it and mapped every read through it. ValidateBlockSize / the Rust
  mirror now refuse the mount, the same way an unparseable vif does; 0 stays
  valid as the legacy two-tier layout.
- The bitrot-sidecar fallback accepted parity_shards == 0 and summed the
  counts in their own width, so values near the ceiling wrapped past the
  MaxShardCount bound. Require both counts and sum in a wider type.
- weed fix -ecx treated a config with only DataShards > 0 as usable, so a
  half-written .vif suppressed the recovery paths AND survived the rewrite.
  Require a complete, in-range config before trusting it.
- Returning the vif-load error left the .ecx and .ecj descriptors open;
  repeated mount attempts on malformed metadata could exhaust them.

* ec: refuse to act on a layout the metadata does not establish

- The worker encode only logged a failed .vif write and skipped it in the
  distribution set, and treated the .ecsum write as best-effort. A worker
  whose disk filled after the much larger shards landed could still
  distribute, mount, verify shard inventory, and delete the source replicas —
  leaving holders with shards whose geometry nothing records. Both writes and
  both inclusions are encode success conditions now.
- A generation-matching .ecsum that disagreed with the .vif geometry only
  disabled checksums in Go, and in Rust was not compared at all, so
  protection stayed On while reads used the other layout. Both files record
  the layout their generation was encoded with, so a disagreement now fails
  the mount.

* ec: reject an invalid recorded block size in weed fix -ecx

A .vif with valid shard counts but a negative or unaligned block size was
marked usable: a positive invalid value pinned the scan to a geometry that
de-stripes to garbage, and a negative one ran the dual scan but left the
invalid .vif in place afterwards. Validate it with the same rule the mount
applies, and when it fails leave the layout unknown so the scan recovers it
and the file is rewritten.

* ec: validate the sidecar layout weed fix -ecx recovers from

The .ecsum fallback was taken on DataShards > 0 alone, so a CRC-valid
sidecar carrying the wrong generation, an incomplete ratio, or an unaligned
block size would pin the reconstruction to one incorrect uniform-layout
candidate instead of letting the dual scan decide. Require generation 0, a
complete in-range ratio, and a valid block size; anything less leaves the
layout unknown, which is the answer that still recovers by scanning.

* ec: let only a genuinely absent sidecar choose the legacy layout

With no .vif the bitrot sidecar is the only record of a volume's layout, and
the mount fallback read a failed load, an unusable config, or a sidecar
stamped for another generation as "assume legacy". A uniform generation-0
volume could therefore mount with legacy or another generation's geometry and
answer reads with the wrong bytes. Present-but-unusable now fails the mount;
only actual absence keeps the legacy defaults. Shared as
EcShardConfigFromSidecar so every caller reads the sidecar the same way.

* ec: treat a recorded-but-impossible layout as corruption, not as legacy

- A .vif whose ecShardConfig is PRESENT but records an impossible ratio was
  answered with the default 10+4 and the legacy block layout, in both
  languages. That reads a uniform volume's shards at the wrong offsets and
  returns the wrong bytes. Only an entirely absent config still means "this
  predates the record"; a present one that cannot be true fails the mount.
- The shard-count bound summed two uint32 counts as int, which wraps on a
  32-bit build: 0x7fffffff + 0x7fffffff lands at -2 and slips under
  MaxShardCount. ValidEcShardCounts sums in uint64, and every EC call site
  that checked a recorded ratio now goes through it.

* ec: rebuild on the geometry the sidecar records, and flag it when it disagrees

The rebuild RPC passes BackgroundECContext, so RebuildEcFiles resolves the
layout itself — and it resolved a missing or invalid .vif to the default 10+4
with the legacy block size. Two consequences: a 12+4 volume was reconstructed
through a 10+4 matrix, which produces wrong bytes and never regenerates
shards 14-15; and the chosen geometry then contradicted a valid uniform
sidecar, which loadRebuildSidecar reported as BitrotOff — silently skipping
the input and regenerated-shard checksum checks precisely when the volume had
already lost its metadata.

The layout now resolves from the bitrot sidecar (found across the server's
disks, not just beside the base name) before falling back to the defaults,
and a present-but-impossible ratio fails instead of being replaced. A sidecar
that contradicts the chosen geometry is BitrotInvalid, which the existing
unsafeIgnoreSidecar override still lets an operator push past.

* ec: let the Rust rebuild read metadata off a sibling disk

read_ec_shard_config searches only the location the rebuild writes into, so a
volume whose .vif or generation-0 .ecsum sits on another of the server's
disks resolved to the default 10+4 with the legacy block layout — the Rust
half of the geometry-guessing the Go rebuild just stopped doing. It then
reconstructs a custom-ratio or uniform volume through the wrong
Reed-Solomon matrix and de-striping geometry.

The rebuild now looks for the .vif in its own location and then each sibling,
falls back to the generation-0 sidecar wherever that lives, and only defaults
when neither exists anywhere. The encode-time .dat size the ecx rebuild needs
is resolved the same way.

* ec: resolve a rebuild's vif from every directory that may hold it

RebuildEcFiles probed only <data-base>.vif. The caller knows the selected
location's index directory and the sibling locations, but passed neither for
metadata: additionalDirs carried shard directories only, and were searched
for shards and the checksum sidecar. A split -dir/-dir.idx layout, or a disk
holding only shards, therefore resolved a pre-sidecar custom-ratio volume to
10+4 and reconstructed through the wrong matrix — never regenerating shards
14-15.

The caller now hands over the index and sibling directories, and the resolver
probes the vif across all of them, matching what the Rust resolver already
does for both the vif and the sidecar.

* ec: make every rebuild consumer agree on the layout it resolved

- The post-rebuild bitrot backfill re-derived the geometry from this
  directory's .vif alone and dropped the block size entirely, so a rebuild
  that resolved its layout from a sibling, the sidecar, or a uniform vif wrote
  a manifest describing a DIFFERENT layout — one later mounts reject, or that
  covers only the default shard count. The layout is resolved once now,
  through an exported ResolveRebuildECContext, and the rebuild and the
  backfill share that answer.
- The Rust rebuild collected only each location's data directory, so a
  sibling's INDEX directory — where a split -dir/-dir.idx layout keeps
  .ecx/.ecj/.vif — was never probed, and a custom-ratio volume still resolved
  to 10+4 with the legacy layout. Both directories of every location are
  carried now, deduped against the rebuild's own.
- A shard delivery can bring the checksum manifest with it, but the receive
  path only writes the file: a server that already had the volume mounted kept
  its resolved protection state (off) until a remount. The mount RPC
  re-resolves it once the shards it describes have been added.

* ec: cover the rebuild's directory search with tests

Reviewers flagged the sibling index directory twice, and the fix that
closed it had no test of its own: the assembly sat inline in the rebuild
handler, reachable only through a gRPC call against a populated store.
Lifting it into rebuildSearchDirs / select_rebuild_location makes the
rule assertable — a sibling contributes BOTH its data and its index
directory, a shared index directory is listed once, and the rebuild's own
data directory never repeats.

Writing the Rust cases surfaced that the two implementations do not agree
on where the rebuild's own index directory belongs, and both are right:
Go's resolver takes a single directory list, so that directory has to be
inside it, while Rust's takes the rebuild's data and index directories as
their own arguments and would search them twice. The tests now state
which contract each side is holding to, so neither drifts into the
other's shape.

Pure refactor otherwise; no behaviour change.

* ec: search the index directory for the layout sidecar

The Rust resolver looked for the generation-0 .ecsum in the rebuild's
data directory and the sibling list, but not in the rebuild's own index
directory — while the .vif lookup directly above it did, and Go's
findBitrotSidecar has always checked both bases. On a split -dir/-dir.idx
location that directory is where the metadata lives, and callers leave it
out of the sibling list precisely because it is passed here separately,
so nothing searched it.

With no .vif anywhere the sidecar is the only surviving record of the
layout. Missing it resolved a 12+4 uniform volume to 10+4 with the legacy
striping — the test added here fails with (10, 4, 0) against the old
code — and the rebuild then reconstructs through the wrong matrix and
writes .ecx offsets that no reader can follow.

* ec: let the rebuild see its own index directory

The Rust rebuild takes a single flat directory list — the shape Go's
RebuildEcFiles uses — so it cannot be handed the rebuild location's index
directory separately the way the layout resolvers are, and the handler
was passing the sibling list, which deliberately omits exactly that
directory. On a split -dir/-dir.idx location that is where .ecx and .vif
live, so the shard and index lookups could not see them.

Go has always carried that directory in additionalDirs; this lines the
two call sites up.

* ec: let a config-free vif fall through to the layout sidecar

A .vif that carries no ecShardConfig answers nothing about the layout, so
it is no more informative than an absent one — but both trees treated its
mere existence as the end of the search. Go went straight to the 10+4
legacy defaults without consulting the sidecar at all; Rust returned
whatever ec_shard_config_from could make of a single directory. A 12+4
uniform volume with a legacy config-free vif therefore resolved as 10+4
legacy, and every read landed at the wrong shard offset.

The sidecar lookup was also single-directory on both sides, while a split
-dir/-dir.idx layout keeps .vif and .ecsum with the INDEX. Go's
findBitrotSidecar has always taken both bases; the callers here passed
only the data base, and the Rust bitrot resolver derived its path from
the data base alone. Rust's layout resolver now takes a candidate
directory list — data, index, then any siblings — and searches all of it,
which also removes the early return that made the vif's presence
decisive.

load_vif_info_across_dirs reported `dir` even when load_vif_info had
found the vif in `dir_idx`. Nothing reads that field today, so this
changes no behaviour; it stops the next caller that resolves the rest of
the volume's metadata against the answer from being sent to a disk
holding none of it.

Absence stays legal throughout: a volume with neither record is genuinely
legacy. Present-but-unusable still fails the mount, now in the
config-free-vif branch too.

* ec: activate a delivered sidecar on every per-disk runtime

A vid mounts as one EcVolume per disk, each with its own resolved
protection state, but the post-delivery reload used the first-match
lookup and so touched exactly one of them. The siblings kept reporting no
protection until a remount — and since shard distribution deduplicates
the metadata files onto the first target disk for a node, the runtime
that got the .ecsum is not necessarily the one the lookup returns.

Iterate every runtime instead, via a new FindAllEcVolumes and its Rust
mut equivalent. Combined with each runtime now resolving its sidecar
against its index directory as well as its data directory, a server
sharing one -dir.idx across its disks activates all of them from the
single delivered copy.

The Rust volume server had no post-mount reload at all; it gets one here,
matching Go.

* ec: resolve the delivered sidecar across every EC metadata directory

Reloading every per-disk runtime, added last round, did not by itself
make the delivered manifest reachable. Startup mirroring copies
.ecx/.ecj/.vif to every shard-bearing disk so each mounts
self-contained, but deliberately not .ecsum, and a repair delivers
exactly one copy. Each runtime was resolving against its own two
directories, so every sibling of the disk that received the file kept
reporting no protection however often it reloaded.

Resolve one authoritative copy across every EC metadata directory
instead of duplicating the file. Mirroring .ecsum would have to keep
pace with a file that is rewritten as shards are repaired, and would not
help the reported case at all: the delivery happens at runtime, and
mirroring only runs at startup.

The regression test pins both halves — a reload restricted to the
volume's own directories still finds nothing, and the same reload
given the server's metadata directories turns protection on.

* ec: ask every directory before writing a TOFU baseline

After a rebuild the opportunistic backfill asks whether this volume
already has a checksum manifest, and answered from the data base alone.
A split -dir/-dir.idx layout keeps the sidecar with the index, and a
multi-disk server may keep it on a sibling, so an existing manifest read
as absent.

The consequence is worse than a missed read. On a false "no" the backfill
writes a fresh sidecar at the data base from whatever the shards say right
now — and the data base is the first candidate every resolver checks, so
that TOFU baseline shadows the real manifest rather than sitting beside
it. A shard that was silently corrupt gets blessed, and the record that
would have caught it stops being consulted.

FindBitrotSidecar exports the search the package already used internally,
so the question is asked of the data base, the index base and the sibling
disks — the same candidates the rebuild resolves its layout from.

* ec: refuse a shard block size no encoder could have produced

weed fix -ecx derived one from the raw shard extent, so a truncated or
partially copied shard wrote a .vif that NewEcVolume then permanently
refuses — the volume the tool was run to rescue could never mount again.
An extent that is not a whole number of small blocks cannot have come
from a uniform encode, so it is no longer offered as a candidate, and
nothing unvalidated reaches the .vif.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: derive the .vif's dat size and block size from one measurement

VolumeEcShardsGenerate stat'ed the .dat before the encode while
WriteEcFiles stat'ed it again to size the blocks. A write landing
between the two produced a .vif whose own two fields describe different
files. WriteEcFiles now leaves both on the context, and fills a
placeholder context in place so the caller can read them back.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: keep the source volume until every holder serves its shard layout

The uniform layout rides in a .vif field older volume servers never
knew: they discard it, mount the shards as legacy and return wrong bytes
with nothing erroring, and the shard files are the same length either
way so no other check notices. The upgrade order lived only in the
release note. VolumeEcShardsInfo now reports the block size the holder
actually serves, in both the Go and Rust servers, and the pre-delete
verification refuses to drop the source unless every reachable holder
echoes the one the shards were encoded with — while a rollback still
exists. A server that predates the field answers 0, which is the
negative answer.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: drop the rebuild's dead block-size parameters

generateMissingEcFiles never reads largeBlockSize/smallBlockSize —
Reed-Solomon reconstruction is layout-agnostic — so passing the legacy
constants only advertised a layout the rebuild does not use. Also move
UniformBlockSize's doc off ValidateBlockSize.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: warn about EC defaults only when the mount used them

The "vif file not found, using defaults" warning fired even after the
bitrot sidecar supplied a non-default layout, sending anyone triaging
wrong bytes after the legacy layout the volume never mounted on.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: stat the distributed bitrot sidecar once

The strict check re-stat'ed the file immediately before the stat that
already gates inclusion, and a failed sidecar write now fails the encode
outright, so the first could only fire on a deletion between the two
lines.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* ec: say what the reconstruct budget actually bounds

A shard's buffer stays in bufs until its interval reconstructs, which is
after the read that filled it released its permit, so the semaphore
bounds round trips in flight and not retained bytes. Peak memory is the
intervals reconstructing at once times the shards each reaches times the
interval size.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7

* test: let the fake volume server report its delivered EC layout

The pre-delete verification now asks each holder which shard block
layout it serves, and a fake that always answered "unset" looked exactly
like a volume server too old to know the field. Distribution ships the
.vif to every holder alongside its shards, so read the layout back out
of it as a real holder does.

Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7
2026-08-28 20:46:59 -07:00
93666c90e9 filter by volume ids (#10983)
* filter by volume ids

* master: carry the volume ids VolumeList asks about in one repeated field

One id and a list of them ask the same question, so field 2 holds the list
rather than standing beside a second field that supersedes it.

Claude-Session: https://claude.ai/code/session_011qAmAdhrYvnzGkw7A9N4mP

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-28 20:43:29 -07:00
Chris LuandGitHub 3967ca23be rust: cover the READS scrub reconstruction path (#11027) 2026-08-28 17:17:47 -07:00
fcc2ea61d3 ec: scrub a volume through its parity data (#11006)
* Introduce a new `READS` scrub mode.

`READS` performs a full volume scrub but, unlike `FULL`, it will attempt to
reconstruct data for missing/damaged shard intervals from other shards in the cluster
when necessary.

The goal of this check is to ensure that EC volume contents _are readable by Seaweed_
even on a degraded storage state, by exercising parity data which is not read in `FULL`
mode. This is useful not only to validate data is user-readable, but also to detect potential
parity shard issues which may be difficult to pinpoint otherwise - particularly for older
volumes lacking sidecar data, and hence unaffected by `CHECKSUM` scrubs.

For regular volumes, this operation is equivalent to `FULL`.

Example:

```
> ec.shard.unmount --volumeId=1 --shardId=0,3,11 --delete --apply
Live shard topology for volume ID 1 (14 shards):
	0@10.200.18.89:9001
	1@10.200.18.89:9002
	2@10.200.18.89:9003
	3@10.200.18.89:9004
	4@10.200.18.89:9005
	5@10.200.18.89:9006
	6@10.200.18.89:9007
	7@10.200.18.89:9008
	8@10.200.18.89:9009
	9@10.200.18.89:9013
	10@10.200.18.89:9010
	11@10.200.18.89:9011
	12@10.200.18.89:9012
	13@10.200.18.89:9020

Will unmount + delete 3 shard(s):
	0@10.200.18.89:9001
	3@10.200.18.89:9004
	11@10.200.18.89:9011

Unmounting shard 0@10.200.18.89:9001 for volume ID 1...
Deleting shard 0@10.200.18.89:9001 for volume ID 1...
Unmounting shard 3@10.200.18.89:9004 for volume ID 1...
Deleting shard 3@10.200.18.89:9004 for volume ID 1...
Unmounting shard 11@10.200.18.89:9011 for volume ID 1...
Deleting shard 11@10.200.18.89:9011 for volume ID 1...

All done!

> ec.scrub --volumeId=1 --node=10.200.18.89:9002 --mode=full
using FULL mode
Scrubbing 10.200.18.89:9002 (1/1)...
Scrubbed 6 EC files and 1 volumes on 1 nodes

Got scrub failures on 1 EC volumes and 1 EC shards :(
Affected volumes: 10.200.18.89:9002:1
Affected shards:  10.200.18.89:9002:1:0

> ec.scrub --volumeId=1 --node=10.200.18.89:9002 --mode=reads
using READS mode
Scrubbing 10.200.18.89:9002 (1/1)...
Scrubbed 6 EC files and 1 volumes on 1 nodes
```

* ec: report the shards a READS scrub had to rebuild

A READS scrub that recovers an interval was recording nothing, so a volume
missing three shards came back clean and nobody repaired it. The unreadable
shard is now recorded before the rebuild is attempted: READS reports the same
broken shards as FULL and differs only in whether the needles themselves
failed, which is the signal worth having - shards are gone, data is still
there.

forceDeletedNeedlesCheck now applies to READS as well, in the shell and in the
RPC guard: it runs the same needle walk as FULL.

Regenerated the proto instead of hand-editing it, so the pancis typo (which
protoc-gen-go-grpc emits into eight other files here) and the header whitespace
stay as generated.

Mirrors into the Rust volume server, which also now honors
force_deleted_needles_check rather than hardcoding it off.

Claude-Session: https://claude.ai/code/session_014yMNebkUjSbx9sfUCWJJtq

* ec: answer a deleted needle from a READS rebuild as deleted

#11020 gave the Rust recovery a deleted flag alongside its bytes, and it
answers a deleted needle with no bytes at all. The READS scrub appended that
empty answer, which does not compile against the new signature and, once it
did, would leave the needle short and report the size mismatch as damage.

Zero-fill the interval instead, the way the direct read beside it already
does: the assembled needle then reaches read_bytes as the delete-state
mismatch the walk already tolerates. Go takes the same branch off the flag
its recovery returns, rather than discarding it.

Claude-Session: https://claude.ai/code/session_014yMNebkUjSbx9sfUCWJJtq

---------

Co-authored-by: Lisandro Pin <lisandro.pin@proton.ch>
2026-08-28 16:42:34 -07:00
Chris LuandGitHub ba5b14b457 master, filer, s3api: bound the collection deletes that strand a caller (#11026)
* master: bound each volume server DeleteCollection, and finish the fan-out

A collection delete fanned out to every volume server holding it with
context.Background(), so a server that accepted the connection and then
went quiet held the whole delete open with nothing to end it. Each RPC is
bounded now, on the same budget allocateVolumeTimeout gives the other
master-to-volume-server admin RPC. The volume server runs the delete to
completion regardless of the request context, so giving up costs the
confirmation and not the deletion.

The walk itself is the caller's, not a per-server one:

- It outlives the caller. A cancelled request must not abandon a
  destructive fan-out part-done, with volumes left behind and no request
  still running to come back for them.
- It no longer stops at the first server that refuses, which left the
  collection on every server after it in the list. The first failure is
  still what is reported, and the collection stays in the topology so a
  later delete comes back for the rest.
- It sends one RPC per server rather than one per replica.
  ListVolumeServers reports a node once for every replica it holds, while
  DeleteCollection removes the whole collection from the server it
  reaches, so a collection with thousands of volumes repeated the same
  whole-collection delete thousands of times over.

Both passes run too. Returning after a failed normal pass left the
collection's EC shards in place with nothing left to retry them.

Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP

* master: delete the EC shards behind /col/delete too

The HTTP handler carried its own copy of the volume-server walk and only
ever ran the normal pass, so a collection deleted through it kept its EC
shards. It shares the gRPC path now, which also gets it the bounded RPCs
and the one-per-server fan-out.

Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP

* filer: bound the collection delete a bucket delete leaves behind

Deleting a bucket entry deletes its collection afterwards, deliberately
detached from the request so a client that hangs up cannot strand the
bucket's volumes. Detached meant unbounded, though: with the master down
or mid-election the wait for a leader has nothing to end it, so the
handler parks, and the client retrying behind it parks another.

It keeps outliving the request and now carries a deadline of its own. The
budget bounds the wait, not the work: the master keeps deleting on its own
fan-out once asked, so giving up costs the confirmation.

Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP

* s3api: bound the collection RPCs a bucket creation and deletion issue

Neither carried a deadline, so a transient failure anywhere down the chain
held the S3 request open until the client gave up on it. Both budgets are
taken outside the filer failover walk, so one budget covers the whole walk
rather than granting each filer a fresh one.

The walk itself stops when that budget is spent, and stops without blaming
anyone: the caller's own expiry is not evidence against the filer that was
answering, and the next filer has no time left to answer in either.
Recorded as a filer failure, a slow master upstream would flag every filer
in the walk, and the three failures that open the circuit take unrelated
object reads down with them.

Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP

* s3api: a failed collection listing no longer fails a bucket creation

PutBucket lists collections to notice a leftover one it is about to reuse.
The result feeds a warning and nothing else -- s3a.exists is what decides
whether the bucket already exists -- yet a transient failure of that
listing returned 500 and refused the creation. It is advisory now, so a
failure is logged and the creation continues, exactly as it does when the
listing returns false.

Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
2026-08-28 16:32:30 -07:00
Chris LuandGitHub 7dc3835b02 s3: an abort answered mid-part no longer leaves the upload completable (#11025)
* s3: reject a part whose upload was aborted while its body was in flight

The upload-exists check runs before the part body is read. An abort answered
during the read deletes the upload directory, and the part write that follows
re-creates it, so the aborted upload is listed nowhere yet completes.

Re-check after the write: only createMultipartUpload stamps the destination
key on .uploads/<id>, so a directory without it is one the part write
resurrected. Drop it along with the part and answer NoSuchUpload.

Claude-Session: https://claude.ai/code/session_01ByZ49KQdUtHhmjG6SpNczT

* s3: reject a copied part whose upload was aborted mid-copy

UploadPartCopy has the same window as UploadPart: the upload-exists check
runs before the bytes are copied, and the part write that follows re-creates
the directory an abort removed. Both the re-encryption and the raw-copy path
re-check before answering.

Claude-Session: https://claude.ai/code/session_01ByZ49KQdUtHhmjG6SpNczT

* s3: do not complete an upload whose directory holds no upload record

A .uploads/<id> directory that a part write created rather than
createMultipartUpload carries no destination key, no owner and no
encryption settings. Completing one turned stray parts into an object;
answer NoSuchUpload instead.

Claude-Session: https://claude.ai/code/session_01ByZ49KQdUtHhmjG6SpNczT

* s3: log the part left behind when the resurrected directory survives

abortMultipartUpload can fail to remove what the part write re-created. The
client still hears NoSuchUpload, since the upload is gone either way and a
retry would only write another part, but the leftover is worth a line.

Claude-Session: https://claude.ai/code/session_01ByZ49KQdUtHhmjG6SpNczT
2026-08-28 16:12:09 -07:00
Chris LuandGitHub c858e01a09 ec: split the shard-interval recovery into a gather and a rebuild (#11005)
* ec: split the shard-interval recovery into a gather and a rebuild

Recovering an interval is now one function doing the local seeding, the waved
peer fetch, the shard accounting and the Reed-Solomon rebuild, under a memory
budget. Splitting the gather from the rebuild makes the rebuild a plain
function over a set of intervals, which is testable on its own and reusable by
the parity checks a full scrub wants.

The rebuild refuses a parity target, and the caller checks that before the
gather so a doomed target costs no fan-out. ReconstructData rebuilds data
shards only, so asking it for a parity shard returned no error and left the
slot nil, and the caller copied that out as a successful read of zeroes. Only
data shard ids reach here today, so this is a guard, not a live fix.

Claude-Session: https://claude.ai/code/session_014yMNebkUjSbx9sfUCWJJtq

* ec: rebuild only the EC shard the read asked for

ReconstructData rebuilds every missing data shard. The gather stops as soon as
DataShards intervals are in hand, so on a distributed volume it routinely
finishes holding parity where data is missing -- and each of those data shards
is then rebuilt into an interval-sized buffer, decoded, and never read. Ask for
the one shard the read needs.

The budget covers it now too: DataShards gathered plus the one the rebuild
allocates. It never covered the rebuild's output, and with ReconstructData that
output was up to ParityShards buffers.

The required mask is Total() long rather than DataShards. reedsolomon documents
both lengths, but its presence scan walks every shard and indexes the short
mask past its end, so the documented short form panics whenever a parity shard
is absent - which here it usually is.

Claude-Session: https://claude.ai/code/session_014yMNebkUjSbx9sfUCWJJtq
2026-08-28 15:55:12 -07:00
Chris LuandGitHub cd5013f116 Re-check an EC shard map a failed read has disproved (#11023)
* Re-check an EC shard map a failed read has disproved

A read that fails against a cached location drops that shard from the map,
which leaves it one short of complete -- and a map one short is trusted for
seven more minutes. So a moment's trouble between volume servers cost
minutes in which every read of that shard skipped the direct fetch and paid
for a Reed-Solomon recovery instead, at DataShards times the memory and the
peer load.

Mark the map when a read disproves it, and re-check a marked map on the
same eleven-second footing as one that never had enough shards to begin
with. The mark clears on refresh, so it buys one prompt re-check rather
than a master lookup per read. The tiers move into a helper; they were
three overlapping conditions in one expression, and the reading of them
was not obvious.

Rust keeps the entry rather than dropping it -- a dead peer fails fast on
the next attempt, and it was the freshness window, not the entry, hiding a
shard that had moved.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Invalidate the location of an EC shard whose own read failed

Recovery fans out to the other shards, so the one whose direct read just
failed is the only location nothing ever invalidates: a shard that moved to
another server was reconstructed on every read until the map's own window
expired, up to thirty-seven minutes for a map still complete. Mark the map
there too. The entry stays -- a moved shard's old holder fails fast, and
the next refresh is seconds away.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Consume the stale mark before the lookup, not after

A read that fails while the master is answering has disproved the very map
that answer is about to install, and clearing the mark on the refresh's
return swallowed it. Clear it where it is acted on instead. A lookup that
then fails loses the mark, which costs nothing: the refresh time is only
advanced on success, so the next read looks up regardless.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Judge the shard map and consume its mark in one critical section

Reading the mark and clearing it were two separate acquisitions, so a mark
raised between them was cleared by a refresh that had not seen it. In Go
that gap was a few instructions; in Rust the mark was read when the read
first snapshotted the volume and cleared at the decision point, with the
local interval reads in between. Take both under one hold. Rust needs a
mutex rather than an atomic to do it, and no longer carries the mark
through the snapshot.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Put the stale mark back when the lookup does not answer for it

Consuming the mark up front assumed the lookup would supersede it. A
lookup that fails, or comes back with fewer than DataShards holders,
supersedes nothing: the map is unchanged, its refresh time unadvanced, and
with the mark gone the map a read had disproved is trusted for its full
window again on the strength of a lookup that never landed. Put the mark
back on both branches.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN
2026-08-28 15:43:05 -07:00
Chris LuandGitHub 624deaf3a4 mount: implement fallocate (#11021)
* mount: implement fallocate instead of reporting it unsupported

Fallocate answered ENOSYS, so the kernel marked the mount as having no
fallocate and returned EOPNOTSUPP. glibc then fell back to its emulation,
which preads a byte from every block already inside the file to see if it
is allocated; on a write-only descriptor that pread is EBADF, and
posix_fallocate returned it.

Volume space is assigned when a write is flushed, so nothing can be
reserved up front: a range inside the file is answered OK untouched, and
one past the end grows the file the way a truncate would. A mode we
cannot honor is refused with ENOTSUP, not ENOSYS, so the kernel keeps
sending the ones we do.

Claude-Session: https://claude.ai/code/session_01XG6sAAdkqTwWffqK5h4rH9

* mount: let a fallocate that allocates nothing past the quota and worm guards

A range already inside the file, and any FALLOC_FL_KEEP_SIZE request,
reserve no space and rewrite no entry, but the preflight refused them
with ENOSPC on a full mount and EPERM on a worm-enforced file. Decide
the no-op first and guard only the growth.

Claude-Session: https://claude.ai/code/session_01XG6sAAdkqTwWffqK5h4rH9

* mount: charge a fallocate growth to the uncommitted byte counter

Write charges the counter by how much the file grew, so the writes that
fill a range fallocate already extended charge nothing and the real-time
quota check never sees that data — only the periodic filer refresh does.
Count the growth where it happens.

Claude-Session: https://claude.ai/code/session_01XG6sAAdkqTwWffqK5h4rH9

* mount: charge a truncate-up growth to the uncommitted byte counter

Same gap Fallocate had: Write charges the counter by how much the file
grew, so the writes that fill a range ftruncate already extended charge
nothing and the real-time quota check never sees that data. Count the
growth where it happens; a shrink still leaves the counter alone, since
it is only ever raised and then reset by the periodic filer refresh.

Claude-Session: https://claude.ai/code/session_01XG6sAAdkqTwWffqK5h4rH9
2026-08-28 14:56:50 -07:00
Chris LuandGitHub 7bb0a1c127 s3: replay a delete whose reply the transport dropped (#11022)
* s3: stop retrying a delete the filer refused for a non-empty folder

The filer looked and the children are there, so the answer will not change.
retryFilerOp spent six attempts and up to 3.1s of backoff on it before the
caller could act on the condition it was already holding.

Claude-Session: https://claude.ai/code/session_01XqaJrwgXQ5GSUpyzRbe5nD

* s3: thread the request context through the unversioned delete path

doDeleteEntry issued every DeleteEntry on context.Background(), so an S3
client that hung up left the gateway working on its behalf, out of reach of
both cancellation and the per-request retry allowance that
DeleteMultipleObjectsHandler installs.

Claude-Session: https://claude.ai/code/session_01XqaJrwgXQ5GSUpyzRbe5nD

* s3: treat a cancelled filer RPC as terminal, not transient

isRetryableFilerErr matched context.Canceled and DeadlineExceeded by
sentinel, which only holds while the error is still local. Once it has
crossed gRPC it is a status, so an abandoned request was retried six times
on behalf of a caller that had already gone.

Claude-Session: https://claude.ai/code/session_01XqaJrwgXQ5GSUpyzRbe5nD

* s3: replay a delete whose reply the transport dropped

A delete is idempotent at the filer, which answers an entry that is already
gone with an empty resp.Error, so a reply lost in transit can be reissued
rather than surfaced. Surfaced, it becomes a 500 on the bucket delete, which
boto3 resends and is then answered NoSuchBucket, or a per-key InternalError
inside the 200 of a multi-object delete, which no SDK retries at all.

The replay runs through retryFilerOp, so it draws on the allowance the
request already installs rather than paying a backoff per key, and stops for
a caller that has gone. rm and rmObject re-enter WithFilerClient per attempt,
so each one walks the failover list again on a connection the failed attempt
had invalidated; the multi-object loop holds one client for the batch, so
there the replay reuses it.

Classification stays structural. The filer reports its own refusals in
resp.Error, which carries no status and has the deleted path - and, for a
recursive delete, the children it stopped on - formatted into it, so no key
name can steer the decision either way.

rm and rmObject now take the caller's context. Cleanup and rollback paths
pass context.Background() deliberately: they have to run whether or not the
caller is still waiting.

Claude-Session: https://claude.ai/code/session_01XqaJrwgXQ5GSUpyzRbe5nD

* s3: share one retry allowance across multipart completion cleanup

The unused-entry loop deletes once per entry, and each delete now retries,
so a filer that stays unavailable held the response for 3.1s per entry after
the object was already committed.

Claude-Session: https://claude.ai/code/session_01XqaJrwgXQ5GSUpyzRbe5nD
2026-08-28 14:30:21 -07:00
Chris LuandGitHub af6f69740c Read metadata log chunks the way the mount reads every other chunk (#11018)
* Replay metadata log chunks the way the mount reads every other chunk

The subscription's log-chunk replay built its own lookup, which always
resolves volume server addresses. A mount started with
-volumeServerAccess=filerProxy cannot reach those, so every fresh
subscription failed on the previous minute's persisted segment and
resubscribed a second later, forever. Take the lookup from the caller
instead; the mount hands over the one it uses for file reads, which also
keeps publicUrl and the bounded location cache in play.

Claude-Session: https://claude.ai/code/session_01NGqrxYj7cHUpSrL249n3Z6

* Keep a log chunk read failure off the filer connection

A metadata subscriber reads persisted log chunks over HTTP from volume
servers and hands whatever went wrong back as the subscription's error.
"connection refused" from a volume server then matched the transport
patterns that decide a gRPC channel is dead, so every failed replay
closed the shared filer ClientConn and cancelled the assign and upload
RPCs riding on it with "the client connection is closing". Mark those
read failures so they are judged for what they are.

Claude-Session: https://claude.ai/code/session_01NGqrxYj7cHUpSrL249n3Z6
2026-08-28 14:15:39 -07:00
Chris LuandGitHub 95248f7492 Bound the memory an EC shard recovery holds (#11020)
* Reconstruct an EC shard from the shards already on this server

recoverOneRemoteEcShardInterval only ever fanned out to the cached shard
locations, so a server holding shards of the volume still fetched them
over gRPC from itself -- and when the peers were unreachable it could not
reconstruct at all, even holding the whole volume on local disk. Seed the
Reed-Solomon buffers from the locally mounted shards first; each one is a
peer round trip, and an interval-sized buffer, the fan-out no longer needs.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Fetch only the EC shards reconstruction still needs

The recovery fan-out read every surviving shard, so a 10+4 volume pulled
13 interval-sized buffers to feed Reed-Solomon 10 -- a third more memory
held, and a third more load asked of peers that were, by definition,
already having trouble. Fetch what is missing, and widen only when some of
those reads fail. A shard reporting the needle deleted ends the walk: the
rest would only answer the same.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Bound the bytes EC recovery holds in flight

Recovery is the one read path that multiplies the served bytes: it holds
an interval-sized buffer per shard until Reed-Solomon runs, and a peer that
is slow to fail keeps them all alive for the whole gRPC timeout. Nothing
bounded how many of those fan-outs ran at once, so a transient problem
between volume servers turned every read into a DataShards-fold allocation
and the server died of it -- 64 concurrent 4MB intervals pin 3.6GB, and
that is a small burst.

Charge each recovery against a process-wide budget, so a burst queues on
the semaphore instead of on the heap.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Answer a deleted EC needle as deleted, not as a failed recovery

A holder reporting the needle deleted is authoritative: deletes are never
invented and never undone. Recovery already collected that flag, then
dropped it on the branch where too few shards came back -- so a read of a
deleted needle that had to recover surfaced as "cannot recover shard", and
the volume server answered 500 where it owed a 404. Carry the flag out of
the shortfall, and let it decide ahead of the error it came with.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Check the encode run of a locally seeded EC shard in Rust

The Rust recovery seeded Reed-Solomon straight from the mounted shards,
without the encode-run check the remote reads and Go's
readLocalEcShardInterval both apply. A volume remounted from a newer
encode between the read's snapshot and its recovery would have fed
mixed-generation bytes into the reconstruction.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Say what the recovery budget actually guarantees

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN

* Seed Rust EC recovery from shards on every local disk

find_ec_volume returns the first disk's EcVolume, so a reconciled volume
whose shards are split across data dirs had the siblings ignored and could
report "cannot recover" while holding enough shards locally. Resolve each
shard together with the disk that owns it, the way Go's recovery already
does, and check that owner's encode run.

Claude-Session: https://claude.ai/code/session_01SM5ARdPvFcnvGWNpBPgNRN
2026-08-28 14:14:40 -07:00
Chris LuandGitHub 23241cf0f1 Let filer.sync move past a chunk the source cluster no longer has (#11019)
* Name the failure when the source cluster cannot locate a chunk's volume

LookupFileId formatted a nil err into the message it returned, so the only
thing a caller could do with "no locations for this volume" was match on the
text. Return a typed error instead.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Fail a source chunk read on a failure status instead of copying the error page

ReadPart never looked at the response status, so a volume server answering 404
for a needle vacuum had removed came back as a successful read whose body was
the error page. The caller counted those bytes as file content and reported a
size mismatch — a corruption claim about data the source had simply lost — and
a 404 from one replica ended the search instead of trying the next.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Stop retrying a chunk the source cluster can no longer produce

A chunk whose volume vacuum has removed fails the same way on every attempt, but
the retry loop had no way to say so and kept going forever. The sync job holding
it never finished, so it pinned the offset watermark at the event ahead of it and
filer.sync never checkpointed again — alive, quiet, and permanently behind.

Wait the source out for a grace period long enough to cover a volume server
restart or a master failover, then give up and mark the failure permanent.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Let replication continue past an entry whose source data is gone

An entry the source can no longer read holds the sync offset forever: the event
fails on every replay, so the checkpoint never moves past it and every later
event stays uncheckpointed, however long the sync keeps running. Nothing brings
those bytes back, so skip the entry with an error naming it and carry on.

Skip only while the source is demonstrably still serving other chunks. A volume
with no locations reads the same whether it was vacuumed away or every replica is
down, and during a cluster-wide outage that answer comes back for every chunk —
skipping then would drop live files wholesale.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Propagate a missing source chunk instead of waiting when supersession is unverifiable

An incremental sink's dated target keys cannot be mapped back to a source path,
so nothing here can tell a chunk the source lost from one a later version already
replaced. Waiting out the grace period would stall every vacuumed needle for half
an hour; hand the failure to the caller, which has the event's real source key.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Wait out a gone volume once, not once per file it held

A volume vacuum removed took every file it held with it, and each chunk was
timing its own grace period. With a bounded chunk executor those waits serialize,
so one gone volume holding many files stalls the sync for far longer than the
grace period — the wedge again, only slower.

Track the wait per source volume on the sink instead: the first chunk to find it
unlocatable starts the clock, every later chunk inherits it and gives up as soon
as it has run out, and a chunk the source does serve clears it.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK

* Probe the source with a read, not a lookup, before writing an entry off

A lookup only proves the source master still has the topology. If every volume
server is unreachable while the master still lists them, the probe passed and the
sink wrote off an entry whose data was merely out of reach. Read the probe chunk
instead, and say in the log that the entry stays unreplicated.

Claude-Session: https://claude.ai/code/session_01SRPEP4jRu29FbLSN6bjLaK
2026-08-28 14:09:56 -07:00
Chris LuandGitHub 60893c5ef3 Classify a filer error before a user-controlled path is wrapped into it (#11004)
* util, pb: classify a filer error by the status the server sent

DoSeaweedListWithSnapshot wrapped a failed ListEntries with %v, dropping the
gRPC status, so IsTransientError fell back to matching substrings against a
message that now held the caller's path. Keep the status with %w and let it
decide, reading the server's own text rather than the wrapper's.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* s3: keep the bucket and prefix out of the list retry decision

A bucket named transport, or a prefix under logs/unavailable/, made a
PermissionDenied listing look transient and got it retried; a key holding the
not-found sentence suppressed a retry that should have run. Both checks now
read the filer's status, and only fall back to the text when there is none.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* filer, s3: classify a delete failure before the path is wrapped into it

The filer put the non-empty-folder marker behind its own "delete directory %s"
wrapper and the gateway matched it as a substring, so a key named after the
marker turned a real delete failure into the demote-the-marker no-op and the
request answered 204. Keep the marker leading the message that crosses the
wire, turn it back into a sentinel where the response is read, and match that.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
2026-08-27 22:30:53 -07:00
Chris LuandGitHub 9e06e1d0f9 Report a delete the filer rejected instead of answering success (#11003)
* s3tables: report a delete the filer rejected

deleteDirectory discarded DeleteEntryResponse and checked only the
transport error, so DeleteTable, DeleteNamespace, DeleteView and
DeleteTableBucket answered 200 for a delete the filer refused. Call
filer_pb.DoRemove, which reads resp.Error and still treats a missing
entry as success.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* admin: report a delete the filer rejected

The bucket delete, the file browser handlers and the topic retention
purger all discarded DeleteEntryResponse, so a delete the filer refused
came back as success. Call filer_pb.DoRemove, which reads resp.Error.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* credential: report a delete the filer rejected

DeleteUser, DeletePolicy and the full-sync cleanup loops discarded
DeleteEntryResponse, so a rejected delete answered success and left the
credential file in place. The service account path in the same store
already read resp.Error; the rest now do too, via filer_pb.DoRemove
where not-found is already tolerated.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* shell: report a delete the filer rejected

remote.configure -delete, remote.cache and the remote metadata sync
discarded DeleteEntryResponse, so a rejected delete printed as removed.
Call filer_pb.DoRemove, which reads resp.Error.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* mq: report a delete the filer rejected

The consumer offset group purge and the coordinator assignment delete
discarded DeleteEntryResponse. Call filer_pb.DoRemove, which reads
resp.Error.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* iam: count only the revocation entries the filer actually deleted

The expiry sweep discarded DeleteEntryResponse, so a rejected delete was
counted as purged and the entry stayed. Call filer_pb.DoRemove, which
reads resp.Error, matching the role and provider stores beside it.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* mount: fail rmdir when the unary fallback delete was rejected

The streaming branch turns DeleteEntryResponse.Error into an error, the
unary fallback dropped it, so rmdir of a non-empty directory answered OK
off the stream and ENOTEMPTY on it. Surface it in both.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* s3tables: fail DeleteTableBucket when the directory delete is refused

The handler only failed when both the leaf entry and the directory
delete failed, so a refused bucket directory delete still answered 200
with the bucket in place. The directory is the bucket, so it decides;
the leaf entry stays best-effort.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
2026-08-27 22:29:48 -07:00
Chris LuandGitHub 742b2f5896 s3: share one retry allowance across a batch delete (#11001)
Every key in a multi-object delete drives its own retryFilerOp, so a filer
that is briefly unhealthy multiplied one op's ~3.1s of backoff by a key
count the client picks. The batch now carries a single allowance in its
context, sized to one op's worst case; once it is spent the remaining keys
fail fast with a per-key error instead of holding the request goroutine.
A single-object delete carries no allowance and keeps its full retries.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
2026-08-27 22:28:25 -07:00
Chris LuandGitHub 902a12fd6f wdclient: bound the wait for a master leader by the caller's context (#11002)
* wdclient: bound the wait for a master leader by the caller's context

WithClient waited on GetMaster with context.Background(), so a caller that
arrived while no master leader was known parked in a 200ms poll loop until one
appeared, whatever deadline it had already set on the RPC. Each retry above it
then left another goroutine in the same wait.

Take the context in WithClient and WithClientCustomGetMaster and hand it to
GetMaster, and stop the retry loop once it is done. The dial keeps
context.Background(): fn brings its own RPC context, so a cancellation seen
here cannot be attributed to the shared connection.

Call sites pass whatever they hold: the request context in the filer's
CollectionList, DeleteCollection and Statistics handlers and in the credential
store's propagation, the operation context in the shell's s3.bucket.delete and
the kafka gateway's broker and filer discovery, and context.Background() where
there is none - the shell commands, the admin dashboard wrapper, and the
exclusive locker's initial lease. The locker's release keeps its own
uncancelled context so a slow unlock cannot turn into a ghost lock.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* wdclient: test that WithClient gives up with the caller's context

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* wdclient: cut the master retry backoff short when the caller gives up

util.Retry sleeps unconditionally between attempts, so a transient error
arriving just before the caller's deadline still cost it a full backoff step.
Use the context-aware util.RetryWithBackoff, the same helper the volume lookup
in this file already uses.

Two call sites went with it: the shell's lock-holder lookup builds its three
second bound before WithClient so it also covers finding the leader, as its
comment already promised, and the filer's post-delete collection cleanup goes
back to an uncancelled context - the entry is already gone, so a caller that
hung up must not leave the collection behind.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU

* wdclient: test that a cancel during backoff ends the retry

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
2026-08-27 22:27:45 -07:00
Chris LuandGitHub d850f36513 s3: distinguish a failed bucket lookup from a missing bucket on HEAD (#11000)
HeadBucket treated any lookup error as ErrNoSuchBucket, so a transient
filer failure answered 404 instead of 500 and clients stopped retrying.
Split the two cases the way the bucket policy handlers already do.

Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
2026-08-27 22:26:15 -07:00
Chris LuandGitHub eed3c27d15 volume: cut the memory a server holding millions of volumes still uses (#10999)
* volume: stop the .vif guard depending on which entry the scan handed over

A volume has both an .idx and a .vif, and loadExistingVolume skipped a .vif
next to an .ecx as EC shard metadata. That was only ever correct because
os.ReadDir sorted .idx ahead of .vif: an interrupted encode, where the .idx is
still there, has to reach validateEcVolume to be reclaimed. Ask for the .idx
instead of trusting the order.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* volume: walk volume directories in batches instead of listing them whole

os.ReadDir builds, and sorts, a slice of every entry before the caller sees
the first one. A disk holding millions of volumes has a .dat, .idx and .vif
per volume, so each startup scan costs hundreds of MB of peak heap that the
runtime is slow to hand back -- and there are several of them before the
first volume loads.

Walk in batches instead, and keep only the entries each scan acts on:
loadAllEcShards now sorts and stats the shard and index files alone rather
than every file on the disk.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* volume: skip the sibling-.dat scan when no EC volume is loaded

pruneIncompleteEcWithSiblingDat only ever prunes EC volumes that are loaded,
but it first walks every disk and keys a map by every .dat on the server. On
a store with no EC volumes at all that is millions of map entries built to
answer no question.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* volume: stop keeping a departure message for every volume

The report state held a VolumeShortInformationMessage per volume copy so a
departure could be named, but almost no volume ever departs. Hold a handle to
the identity instead -- volumes share very few distinct ones -- and build the
message on the way out.

Measured over a populated report state: 195 -> 83 bytes per volume.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* rust volume: stop keeping a whole volume message per volume held

The send loop kept a VolumeInformationMessage for every volume just to notice
mounts and unmounts, and rebuilt the map from scratch on every beat. Keep the
identity a delta names, which is what the Go report state keeps for the same
reason.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* rust volume: keep only the EC files the shard scan acts on

load_all_ec_shards named every file on the disk twice -- once in the dedup set
and once in the sorted vector -- before deciding it only wanted .ec?? and .ecx.
Filter while reading instead. Mirrors the same change in loadAllEcShards.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy

* volume: share the strings every .vif repeats

A tiered volume's .vif names its replication and its backend, and every decode
allocates a fresh copy, so a server holding millions of them holds millions of
copies of the same handful of names. Route them through the interning table
the volume info decode already uses. The remote key names one volume and is
left alone.

Claude-Session: https://claude.ai/code/session_01NWpFUwAJcR2KUENrhLc9Sy
2026-08-27 22:25:15 -07:00
Chris LuandGitHub fdd8bd9478 s3: reject a request that names two operations (#10987)
The router matches bucket subresource routes in registration order while
the IAM action resolver matches its own list in a different order, so a
request carrying two operation subresources is authorized as one
operation and served as another. `PUT /bucket?policy&tagging` resolves to
s3:PutBucketTagging and runs PutBucketPolicy, letting an identity
delegated bucket tagging install an arbitrary bucket policy. The same
mismatch reaches PutBucketCors, PutBucketLifecycle, PutBucketVersioning,
PutObjectLockConfiguration, PutBucketRequestPayment and the policy and
cors deletes.

Reject the ambiguity where the other pre-routing checks live, so neither
list has to stay in step with the other. Keys that modify an operation
rather than select one -- versionId, partNumber, prefix -- still combine
freely.
2026-08-27 16:46:46 -07:00
Chris LuandGitHub 99cf7a66df shell: remove the directories emptied by volume.fsck's filer entry purge (#10992)
* shell: remove the directories emptied by volume.fsck's filer entry purge

volume.fsck -findMissingChunksInFiler -reallyDeleteFilerEntries deleted the
orphan entries but left their parent directories behind, so a namespace
accumulated empty directories that had to be cleaned up by hand.

Remember the parent of every purged entry and, once the purge is done, walk
up from each one deleting the directories that are now empty. The delete is
non-recursive, so the filer itself rejects a directory that still has
children; a bucket and a directory that is an S3 object of its own are left
alone.

Claude-Session: https://claude.ai/code/session_01BncsNo2RVANCDtdbw96Kfc

* shell: keep a directory volume.fsck saw change under it

The empty-directory sweep read the entry to spot an S3 directory key object
and then deleted unconditionally, so a directory promoted to an object in
between was removed anyway.

Delete with the mtime the lookup returned, leaving the filer to skip a
directory that has changed since.

Claude-Session: https://claude.ai/code/session_01BncsNo2RVANCDtdbw96Kfc

* shell: leave a directory volume.fsck just saw written for the next run

The mtime the delete is conditioned on has second resolution, so a write
landing in the same second as the one already on the directory is
indistinguishable from it and the directory would still be deleted.

Skip a directory modified within the last few seconds. A write after the
lookup then always carries a later second than the one the delete carries,
and the sweep picks the directory up on the next run.

Claude-Session: https://claude.ai/code/session_01BncsNo2RVANCDtdbw96Kfc

* shell: skip a directory volume.fsck cannot condition a delete on

A zero mtime disables the delete's condition at the filer, so a directory
whose entry carries none was removed unconditionally and a concurrent
promotion to an S3 object went with it.

Leave such a directory alone.

Claude-Session: https://claude.ai/code/session_01BncsNo2RVANCDtdbw96Kfc

* shell: hold volume.fsck's quiet period to the cutoff second itself

Mtime keeps whole seconds, so a directory whose mtime lands on the cutoff
second was written up to a second after it. Skip that directory too, so the
quiet period fails closed.

Claude-Session: https://claude.ai/code/session_01BncsNo2RVANCDtdbw96Kfc
2026-08-27 16:44:19 -07:00
Chris LuandGitHub 2a97e08caa s3: cover the directory marker key with object lock (#10988)
* s3: enforce object lock when deleting a directory marker

The key "dir/" is deleted the unversioned way, ahead of the branches
that enforce Object Lock, so a principal with plain delete permission
could remove a key the gateway was reporting as COMPLIANCE-retained --
retention set through PutObjectRetention is stored on the directory
entry and served back by GetObjectRetention, only the delete ignored it.

The same path also takes any key ending in "/" regardless of size, while
a PUT only makes a marker of one up to 1KiB. A larger one is a genuine
versioned object, and deleting it here dropped its whole history after
the versioned delete of the same key had been refused.

Enforce in the marker delete itself, so the single, versioned and
multi-object delete paths are all covered.

* s3: apply object lock headers on a directory marker PUT

The trailing-slash branch runs before the versioning and Object Lock
handling, so it accepted x-amz-object-lock-* headers and stored none of
them: a bucket owner could believe a key was retained while nothing
recorded it, and an invalid mode or a past retention date that a regular
key rejects came back 200 here.

Validate the headers the way the regular path does, store what they ask
for beside the owner the same callback already sets, and refuse to
replace a key that is already retained.

* s3: check every version a marker delete would remove

The marker delete clears any history under the key in one recursive
removal, while the lock check ahead of it resolves the latest version
only. A version retained under an unretained one was taken with the
rest, so enforce against each version the removal covers.

* test: pin the marker lock refusals to AccessDenied

A bare require.Error passes on any failure, including one that has
nothing to do with the lock. Assert the code, the key the batch delete
reports, and that the marker survives each refusal.

* s3: check the history entries a version list leaves out

The version list skips an entry without a version id, while the removal
takes it with the rest, so an entry an older build left unnamed escaped
the check. Walk the history directly instead, and refuse when an unnamed
entry is still under a retention or a legal hold of its own.

* s3: let a governance bypass reach an unnamed history entry

The unnamed branch refused every active retention, so a caller allowed
to bypass governance could not clear one, which the named path lets
through. Refuse a legal hold and compliance mode as before, and take the
bypass into account for governance.

* s3: keep the object lock decision in one place

The unnamed history entry had to repeat the retention and legal hold
rules inline because the enforcement helper only takes a key to look up.
Split the part that judges an entry out of it and call that from both.

* s3: guard a marker PUT on the entry it replaces

The overwrite check resolved the key's latest version, but mkdir builds
a fresh entry for the marker itself, dropping the lock metadata the old
one carried. Once the key had a history, an unlocked version answered
for a retained marker and a plain PUT replaced it. Judge the entry the
write is about to replace instead; a versioned write of the same key
still adds a version, which is its own to allow.

* s3: guard a marker delete on the entry it removes

The check ran against the key rather than the entry, so once the key had
a history it answered with a version and the retention recorded on the
marker itself went unseen. Judge the entry that is about to be removed,
the same way the PUT side now does; the versions under it are still
covered by the walk that follows.

* s3: take the object write lock for a marker PUT

The overwrite check read the entry that the mkdir after it replaces, so
two marker PUTs could both pass while one was still unlocked. The marker
delete already runs under this lock; hold it across the check and the
mkdir so the entry cannot change in between, and so the two paths are
serialized against each other.
2026-08-27 16:35:45 -07:00
Chris LuandGitHub ab8b34720a s3tables: delete only the location the dropped table owns (#10986)
DeleteTable authorizes the named table, then recursively purges the data
path derived from its stored MetadataLocation. That location is supplied
by the caller at create/register time and never bound to the table, so a
tenant allowed to drop one table could point it at a table in a sibling
namespace and have the delete destroy that table's catalog entry and
data files.

A legitimately decoupled location -- a rename source, or a leftover the
name was reused over -- has had its catalog attributes stripped, so a
surviving metadata marker identifies a path that belongs to another
entry. Refuse those, alongside the existing ancestor refusal.
2026-08-27 16:28:02 -07:00
Chris LuandGitHub 0b5fff2ccd filer, s3: reuse the volume server's guarded remote-storage client builder (#10990)
* volume: build the guarded remote storage client through a shared helper

Fold the endpoint validation, credential check and rebinding-safe dialer
that FetchAndWriteNeedle applies before dialing a caller-supplied remote
storage endpoint into a single BuildGuardedRemoteStorageClient helper, so
other callers that dial the same endpoints can reuse it. No behavior
change on this path.

Claude-Session: https://claude.ai/code/session_01AiH1FU3rmshSbFFTbJpaZN

* filer: build the remote-mount stream client through the guarded helper

streamFromRemote serves a cold remote-only entry straight from its mounted
origin. Build its client through BuildGuardedRemoteStorageClient so the
same endpoint checks the volume server applies cover this read path too.

Claude-Session: https://claude.ai/code/session_01AiH1FU3rmshSbFFTbJpaZN

* s3: build the remote-mount stream client through the guarded helper

openRemoteStream serves a remote-mounted object straight from its origin
when the local read cannot. Build its client through the same guarded
helper so the endpoint checks apply here as well.

Claude-Session: https://claude.ai/code/session_01AiH1FU3rmshSbFFTbJpaZN
2026-08-27 16:25:56 -07:00
Chris LuandGitHub 28862c866e Authorize an Iceberg table create before it writes (#10991)
* s3tables: share one CreateTable authorization gate

CreateTable and RegisterTable each carried their own copy of the name
validation, policy load and permission check. Fold them into
authorizeCreateTable, and expose it on the Manager for callers that write
into a table bucket before the table itself is registered.

Claude-Session: https://claude.ai/code/session_01QiJkka1T2NAWDWq4JQ8Vuy

* iceberg: authorize a table create before it writes

Stage-create returns before the S3Tables registration that authorizes a
create, and the plain create writes its metadata file before reaching it,
so a caller who may not create the table could still leave a staged
template, a marker and a v1.metadata.json in the target bucket - and get
vended credentials for a location of their choosing. Run the CreateTable
gate as soon as the table is known to be absent.

Claude-Session: https://claude.ai/code/session_01QiJkka1T2NAWDWq4JQ8Vuy

* iceberg: authorize a create-on-commit the same way

A commit against a table that does not exist creates it, writing the
metadata file first and only then reaching the registration that checks
the caller may create it. Denied callers saw a 500 for what is a 403.

Claude-Session: https://claude.ai/code/session_01QiJkka1T2NAWDWq4JQ8Vuy

* iceberg: pin that identity actions reach the create gate

The manager request is built from the caller's own context, so an identity
whose actions carry the permission still passes. Worth a test: a fresh
context here would silently deny every such caller.

Claude-Session: https://claude.ai/code/session_01QiJkka1T2NAWDWq4JQ8Vuy
2026-08-27 16:23:51 -07:00
Chris LuandGitHub bc06505b40 mount: keep metadata operations working on an unlinked open file (#10989)
* mount: serve metadata ops from the open handle of an unlinked file

ftruncate on a descriptor whose file was unlinked failed with ENOENT:
maybeReadEntry resolved the inode to a path first, and unlink had already
dropped it. GetAttr worked around that with its own handle fallback;
SetAttr and the xattr handlers had none.

Look the handle up first and let it answer whether or not a name still
points at the inode. GetAttr keeps reporting nlink 0 there, now off the
empty path.

Claude-Session: https://claude.ai/code/session_01U1R8BM4bVT46KwPDEj2Ega

* mount: read an open handle's attributes under the handle lock too

GetAttr held only the LockedEntry lock, which covers the async uploader's
chunk appends but not Write or the metadata flush: those rewrite size,
times and the whole chunk slice under the handle lock, so FileSize could
walk a slice mid-reassignment. The branch this replaced took both locks;
take both here, outer handle lock first, as Read and Lseek do.

Claude-Session: https://claude.ai/code/session_01U1R8BM4bVT46KwPDEj2Ega

* mount: report nlink 0 from SetAttr for an unlinked open file

The kernel caches the attributes a SETATTR reply carries, so an ftruncate
on an unlinked file left fstat reporting nlink 1 until the cache expired,
even though GetAttr had it right. Both replies go through the same rule.

Claude-Session: https://claude.ai/code/session_01U1R8BM4bVT46KwPDEj2Ega
2026-08-27 16:22:25 -07:00
Chris LuandGitHub e9a464840c webdav: describe a listed entry the way clients expect (#10993)
* webdav: name the entry, not its path, in a listing

DAV:displayname carried the full path of every entry. A client that
takes displayname for the child's name - Windows Explorer does - then
looks for /dir/name under /dir and finds nothing, so a folder shows up
empty while the root, where the two spellings differ only by a leading
slash, still lists.

Readdir now builds its entries with toFileInfo like stat does, so a
listing and a lookup describe a child the same way, and the wrapper that
was trimming the sub-folder back off a name goes away with it.

Claude-Session: https://claude.ai/code/session_01XCeuCWpF9xo9CfyHvCQE9c

* webdav: derive an ETag when nothing hashed the entry

Uploads through this gateway carry no content MD5, so filer.ETag comes
back empty and every file in a PROPFIND answered with an empty
DAV:getetag, which is not a valid entity-tag. Report it as unimplemented
instead, the way the sub-folder wrapper already did, and webdav falls
back to modification time and size. The wrapper's copy went with it - it
swallowed the stat error a caller was meant to see.

Claude-Session: https://claude.ai/code/session_01XCeuCWpF9xo9CfyHvCQE9c
2026-08-27 16:21:14 -07:00
Chris LuandGitHub d8a189f07f s3: keep a missing object a 404 under If-Match and If-Unmodified-Since (#10985)
* s3: keep a missing object a 404 under If-Match and If-Unmodified-Since

GET and HEAD resolved the target before evaluating the conditional headers, and
a missing target failed If-Match and If-Unmodified-Since outright, so absence
surfaced as 412 PreconditionFailed. AWS reports the missing object instead:
404 for HeadObject, NoSuchKey for GetObject, and 412 only when a live object
fails the condition. Clients cannot tell absence from a stale precondition
without an extra racy HEAD, so OpenDAL disabled its four conditional
stat/read capabilities against SeaweedFS.

A precondition now only fails against an object that exists; a missing one --
including a latest version that is a delete marker -- returns NoSuchKey.

Claude-Session: https://claude.ai/code/session_01X4kEbuwxd9DFsTnSXjfjgv

* s3: evaluate a conditional read against the version the request names

GET and HEAD resolved the latest version before evaluating the conditional
headers, so a request carrying versionId had its If-Match compared against a
different version than the one it was asking for: a live version whose ETag the
client held failed once a newer version -- or a delete marker -- became the
latest. resolveObjectEntry now resolves the named version on a versioned bucket,
the way DELETE already does.

A named version that resolves to nothing is left to the handler, which alone
knows whether the bucket is versioned and so whether it owes NoSuchVersion.

Claude-Session: https://claude.ai/code/session_01X4kEbuwxd9DFsTnSXjfjgv
2026-08-27 11:56:33 -07:00
Chris LuandGitHub 2d25c39da4 volume: resolve the disk IO slow-latency threshold per disk (#10976)
* volume: resolve the disk IO slow-latency threshold per disk

volume.toml keys [volume.disk.io.slow.latency] by disk type, but the
threshold was chosen once per server by switching on the raw -disk flag.
-disk is comma-separated, one entry per -dir, so a multi-disk server
matched no case and silently took the hdd threshold.

Carry the table on DiskIOProbeConfig and resolve it in CheckDiskSpace
from the location's own DiskType. A type with no entry keeps falling
back to the hdd threshold.

* volume: run the disk IO probe on multi-directory volume servers

The probe was disabled whenever more than one -dir was configured,
because a single server-wide slow-latency threshold could not describe
disks of different types. The threshold is per disk now, and the rest of
the probe already is: diskRegistry is keyed by directory, each
DiskLocation runs its own CheckDiskSpace, and Store consults
isDiskUnavailable per location.

* volume: reject duplicate -dir entries

Nothing deduplicated -dir, so the same directory listed twice produced two
DiskLocations that each loaded every volume in it, appending to the same .dat
under two independent locks. Compare directory identity with os.SameFile
rather than the path, so a symlink or bind mount aliasing an earlier entry is
rejected as well.

* volume: cover the per-disk slow-latency handoff

SlowLatencyFor has a test, but nothing asserted that CheckDiskSpace feeds it
the location's own disk type. Probe through a seam so the resolved threshold
is observable, and check hdd, ssd, nvme, the empty type, and an unlisted tag.
2026-08-27 10:01:05 -07:00
Chris LuandGitHub 8dcdb70594 mount: let a rename remove its source at the source's own version (#10973)
* mount: let a rename remove its source at the source's own version

A rename stamps the source and takes the name away at the same log position,
so the removal reaches the meta cache carrying exactly the version the source
already records. The version gate read that as a write already reflected and
dropped it, while the destination half of the same event still applied -- the
source stayed cached beside the destination, and readdir and stat went on
serving a name the filer no longer had:

    gate dropped removal of /winfsp-test-TestRenameOverExisting/src
      eventTs=1787761708173717200 record=1787761708173717200
      floor=1787761708173717200 tombstone=false

A removal asks a different question from a write. An entry still present at
exactly that version has the write reflected but not its removal, so only a
strictly newer record fences one out; a tombstone is the removal already
reflected and goes on fencing as before.

* mount: sweep a section's vanished name recorded at the snapshot

The refresh deletes the names its listing did not return, but asked the gate
whether a write at the snapshot was reflected. A name recorded at exactly that
version has the write reflected and not its removal, so it survived the sweep
and stayed cached until some later event happened to touch it.

Same reading as the rename source a commit earlier: the call site removes, so
it asks about a removal.
2026-08-26 10:20:38 -07:00
Chris LuandGitHub f5f1dcbd8c s3: keep verifying the request host when externalUrl is set (#10970)
* s3: keep verifying the request host when externalUrl is set

externalUrl was the only host candidate once set, so a client that dialed
the gateway directly instead of through the proxy always got
SignatureDoesNotMatch. Make it lead the candidate walk instead: every
candidate still needs a valid signature, and the request-derived hosts are
already trusted when the flag is unset, so a mixed proxy plus in-cluster
topology can now advertise a public endpoint and verify both planes.

* s3: cover virtual-hosted addressing behind externalUrl

The old pin also rejected an external client that signed
bucket.api.example.com, since only the bare externalUrl host was ever
tried. The candidate walk covers it; pin the case down.
2026-08-26 10:09:05 -07:00
Chris LuandGitHub 3431bdcb74 s3: fix UploadPartCopy with volume-data encryption (#10971)
* operation: give an encrypted chunk the plaintext ETag

With -encryptVolumeData the volume server stores ciphertext, so it cannot
echo a Content-MD5 back and the chunk lands with an empty ETag. Every ETag
derived from those chunks then comes out empty for a single chunk, or
d41d8cd98f00b204e9800998ecf8427e-N for several.

The caller already hashes the plaintext to send as Content-MD5, so keep that
digest as the chunk ETag instead of dropping it, and compute it for a
WantMd5 caller under cipher too.

* s3: re-encrypt a part copy from a volume-encrypted source

UploadPartCopy raw-copies source chunks when neither side uses SSE, which
also caught -encryptVolumeData sources. Those chunks are ciphertext a
whole-chunk cipher key decrypts, so copying a byte range out of one and
keeping the key leaves a destination that fails authentication on GET, and
the copied chunks carry no ETag for the part result to report.

Route them through the re-encrypting path already used for SSE: it reads the
source as plaintext, hashes the part, and writes the destination under the
gateway's own encryption.

* s3: fetch only the range a part copy asked for

The re-encrypting UploadPartCopy path opened the source at offset 0 and threw
the prefix away, so assembling an object part by part read the source once per
part. Now that volume-encrypted sources take this path too, that is the common
case rather than an SSE corner.

The chunk stream already seeks, so hand it the range.

* s3: reject an unsatisfiable copy-source-range

A part copy has no way to report a short part, so a range reaching past the
source cannot be clamped the way a GET clamps one. The fast path silently
produced a part shorter than asked for, or an empty one; the re-encrypting
path pads with zeros, so a 2 MiB source copied as bytes=1048576-9999999 came
back as 1 MiB of data followed by 7.5 MiB of nothing.

Answer InvalidRange instead, which is what s3-tests'
test_multipart_copy_invalid_range expects.
2026-08-26 10:05:49 -07:00
Chris LuandGitHub 12fd60f92e rust volume: stop racing the clock in torn_sdx_is_regenerated (#10966)
The test truncates a good .sdx and asserts the result still looks fresher
than its .idx, on the reasoning that truncation bumps the mtime. That holds
only at the filesystem's timestamp granularity: where both writes land in the
same tick the precondition fails and the run reports a failure that says
nothing about the code under test — as it did on CI. Backdate the .idx the
way the sibling stale_sdx_is_regenerated already does.
2026-08-26 08:58:34 -07:00
Chris LuandGitHub da087f77b3 mount: stop a replaced rename destination from flushing over the rename (#10965)
* mount: stop a replaced rename destination from flushing over the rename

Rename replaces whatever the destination held, which deletes that entry, but
only the source handle was told. A handle still open on the replaced entry
went on flushing its metadata under that name, and on Windows -- where the
close carrying the flush runs after the application's CloseHandle has already
returned -- the flush landed after the rename and put the destination's old
content back:

    dir Rename old_entry:{name:"src"} new_entry:{name:"dst" ... inode:...3416}
    doFlush /dst fh 1521468582993181449
    /dst saveToStorage 1,6872462993 [0,3)
    flushMetadataToFiler /dst inode 11939747521756968515
    InsertEntry /dst

The next read of the destination returned the content the rename was supposed
to replace. Unlink already handles this with markHandleDeleted, which raises
the flag under the handle's flush lock so a flush already writing finishes
first and any later one sees it; a rename that replaces an entry deletes it
just the same, so it now does likewise.

Verified on the Windows runner: TestRenameOverExisting 300/300, where the same
loop reproduced the corruption twice without this.

* test/winfsp: say which layer kept a renamed-away name

The failure only reported the stat. Which layer answered narrows the search a
lot: a listing reads no per-path cache, the mount's own forgets within a
second, and a name that survives both is still in the meta cache.

* mount: keep the destination barrier honest when the rename does not happen

Two gaps in the barrier the previous commit put in front of a replaced rename
destination:

The flag was raised before the filer rename, which can still fail. The
destination then stays exactly where it was, with its handle marked deleted
and its dirty metadata silently dropped from then on, so a rename that
returned an error has to put the flag back.

The handle was only found through the path mapping, which Forget drops while
the handle is still open. The source side already falls back to the inode the
entry carries; the destination now does the same, off the entry the sticky-bit
check had already loaded.

* mount: let only the caller that raised a delete mark lift it

Restoring the destination handle after a failed rename cleared isDeleted
outright, so an unlink that marked the same handle in between lost its mark and
a later flush could write the unlinked entry back.

Every raise of the flag already happens under the handle's flush lock, so
counting them there is enough to tell one caller's mark from another's: the
rename lifts only the mark it made itself.

* mount: drain the destination flush before marking it deleted

A flush already queued for the destination belongs to the entry as it stands.
Marking first meant the drain waited on a flush that then skipped its metadata
as deleted and released its handle, so a rename that failed afterwards had
nothing left to restore and the queued update was gone, its chunks orphaned.

Draining first lets that flush finish as itself, before the rename has taken
anything away.
2026-08-26 08:51:37 -07:00
Chris LuandGitHub eb3bbfeb1f filer: apply the path's storage rule TTL on every write path (#10963)
* filer: cover the storage rule TTL on the object transaction write path

An object written through ObjectTransaction used to land with ttlSec 0
even under an fs.configure TTL rule, while the same object written
through CreateEntry got the rule's TTL. Guard the shared stamping so the
two paths cannot drift apart again.

* filer: apply the path's storage rule to an appended entry

AppendToEntry resolved the storage option from the path - so its chunks
land on a TTL volume under an fs.configure TTL rule - but never stamped
the rule's TTL on the entry it creates, leaving an entry that outlives
its data. Route it through applyStorageDefaultsToEntry, which now feeds
the entry's own TTL into the option so the placement an existing entry's
appended chunks get is unchanged.

* filer: apply the path's storage rule to a completed TUS upload

The PATCH path resolves the storage option from the target, so a TUS
upload into an fs.configure TTL prefix writes its chunks to a TTL volume,
but completion built the final entry with ttlSec 0 - the entry outlived
the data it pointed at. Stamp it through applyStorageDefaultsToEntry,
which also subsumes the hand-rolled read-only check and supplies the
rule's name-length limit.

* filer: apply the destination's storage option TTL to a copied entry

The copy handler re-uploads the source's chunks under the destination's
storage option, so a copy into an fs.configure TTL prefix already lands
its data on a TTL volume. The entry, though, carried the source's ttlSec
- 0 for a source outside the prefix, or the source's own TTL where the
two rules differ - so it never expired with the data it pointed at. Take
the TTL from the same option the chunks were placed with, after the
data-only copy has restored the destination's metadata.
2026-08-26 08:49:25 -07:00
Chris LuandGitHub a02c0024e5 master: cap the reported capacity at what the disks hold (#10960)
* master: cap the reported capacity at what the disks hold

Statistics reported max volume count times the volume size limit, which is
how many volumes the cluster is allowed to place, not how much space it has.
A cluster given far more slots than its disks can fill reported a capacity it
could never reach -- 65536 slots at 30GB read as 1.9PB on a 460GB disk -- and
the number never moved, since writing data changes neither the slot count nor
the size limit.

The volume servers already report each filesystem's total and free bytes in
their heartbeats, so bound the answer by what they say is left.

* mount: keep the last known sizes when filer statistics fails

A failed Statistics call returned before df's answer was filled in, so a
mount whose filer or master was briefly unreachable reported an empty
filesystem rather than the sizes it already had.

* master: drop the disk ceiling when a volume server does not report

A cluster part way through an upgrade has volume servers that predate the disk
bytes in the heartbeat. Summing only the ones that answered left the quiet
server's free space out of the total, and the server holding the room is
exactly the one that could make the cluster read as full.

Answer with the disks only when every one of them reported.
2026-08-26 00:12:56 -07:00
Chris LuandGitHub b77d954f55 rust volume: fail closed on sorted-index failures and reconcile tier-up (#10956)
* rust volume: fail closed on sorted-index failures and reconcile tier-up

Follow-ups to the .sdx sorted needle map (#10951):

- get() folded open/read failures into None, so an EIO, a torn .sdx, or a
  failed pooled reopen answered reads with NotFound and let do_delete_request
  acknowledge the delete as Ok(0) without writing a tombstone. It now returns
  io::Result and every caller propagates; redb's get() had the same shape and
  is fixed with it. is_file_unchanged cannot propagate, so it reports unknown
  and logs rather than treating an unreadable index as proof of a change.
- A delete whose .idx append landed but whose .sdx mark failed left the map
  still resolving the old live entry, so deleted content stayed readable until
  a reload. The map now records the tombstone before touching .sdx and only
  clears it once the mark lands; lookups consult that first and report the
  needle deleted, which is what the next reload concludes anyway.
- Mode reconciliation ran one way. Entering remote mode made use_sorted_index()
  true, which returned early, so a volume tiered while the server runs kept its
  in-memory map and pinned .idx descriptor until restart — the RAM and fd win
  never applied. It now reconciles in both directions.
- Tier-down dropped the remote reference before the fallible refresh, so a
  failure left volume_info local, the remote backend attached, the .vif still
  remote, and a retry reporting "already on local disk". The transition is
  snapshotted and rolled back.
- The read-only fallback set no_write_or_delete but left no_write_can_delete,
  so metrics and mode checks called the volume delete-capable while every
  delete was refused.

* rust volume: count a sorted-map delete against the durable .idx append

The deletion counters sat after the in-place .sdx mark, so a mark that
failed left them at their pre-delete values while the tombstone was already
durable in .idx — and with retries now idempotent, nothing applied them
later either. Heartbeats, status responses, and the garbage calculation
would report the volume as free of that garbage until a reload.

Move them to the append that makes the delete durable, which is also what a
reload of .idx would count. Covered by a test that injects a mark failure
through a cfg(test) seam: no portable filesystem trick reproduces it, since
a read-only .sdx fails the borrow long before the mark.

* rust volume: hide a pending tombstone from the sorted-map scans too

The overlay that keeps a needle deleted after a failed .sdx mark was only
consulted by get(). visit_live_entries still read the stale valid record
straight off .sdx, so ascending_visit, iter_entries and save_to_idx all
reported the needle live — and compaction takes iter_entries for the
complete live set, so it would copy the deleted content forward and
save_to_idx would write it back into the rebuilt .idx as live.

Snapshot the overlay once per scan and skip its keys, which is the same
conclusion the next reload reaches from the .idx tombstone.

* rust volume: quarantine a durable write whose index lookup fails

The prior-mapping lookup that decides whether to index a fresh append runs
after the record is already down and flushed, so a failing lookup leaves
exactly the state a failing put leaves: a durable .dat record nothing
indexes. The put path marks the volume read only for it; this one returned
the error and kept taking writes, and the next append would bury the
orphan mid-file where the .dat tail check on reload cannot see it.

Give it the same treatment.
2026-08-25 23:14:30 -07:00
Chris LuandGitHub 7658305c76 mount: name the disk after the mounted path (#10958)
* mount: name the disk after the mounted path

Finder and Explorer labelled every mount with the filer address, so two
mounts from one filer were indistinguishable. Use the mounted path's last
segment, the way df already shows it, and keep the filer address only for
a whole-tree mount.

* mount: let a given mount option override the default

The options from -o were placed before the ones this mount derives, so
a volname or iosize given on the command line lost to the derived value.
Append them last, matching the Windows adapter.

* mount: document what labels the disk
2026-08-25 22:56:33 -07:00
Chris LuandGitHub 627b5e9d59 shell: parse every collection filter the same way (#10955)
* worker: move the collection filter parser into weed/util/wildcard

The parser sits beside the volume-list filtering it was written for, in
weed/plugin/worker, which imports weed/shell — so the shell commands that
parse the same filter three other ways can never call it. Move it down to
weed/util/wildcard, next to the comma-separated wildcard helper it already
replaced, leaving the behavior unchanged.

* shell: parse every collection filter the same way

The shell parsed a collection filter three ways: compileCollectionPattern
compiled one regex for ec.encode, ec.decode, volume.balance and the tier
commands; volume.list and volume.deleteEmpty matched a single wildcard; and
volume.tier.move, volume.fix.replication and volume.configure.replication
called filepath.Match on their own. None of them took a list, so
"ec.encode -collection=a,b" selected nothing, the same way the admin UI did.

They all go through the shared matcher now: a comma-separated list of names,
"*" and "?" wildcards, "_default" for the collection with no name, and regex
entries. The one thing that stays per-command is what an empty value means -
every collection for -collectionPattern, the unnamed collection for the ec
and tier -collection flag - so compileCollectionPattern keeps that mapping.

The matchers are compiled once per command instead of once per volume, and a
regex entry now has to match the whole name unless it anchors itself, so
-collection=bucket no longer picks up mybucket2.

* shell: keep dots in collection names, and commas inside a regex

A dot no longer marks an entry as a regex, so a collection named "my.bucket"
matches itself and not "my-bucket" - the difference decides which volumes
volume.deleteEmpty and volume.tier.move touch. A dot still counts when it is
quantified, so "bucket.*" stays a prefix regex.

The comma split also leaves alone the commas inside a character class or a
repetition count, so "bucket[0-9]{1,3}" stays one entry instead of becoming
two broken fragments.

* shell: let a regex entry match its own spelling

A collection named after regex syntax, say "logs(2024)", was unreachable:
the entry compiled to a pattern that matches "logs2024" instead. Match the
entry verbatim as well, so naming a collection always selects it, whatever
characters it holds.

* shell: reject a collection filter that names no collection

A value of "," parsed to no entries and then matched every collection, so a
typo widened ec.encode or volume.deleteEmpty to the whole cluster. Only a
genuinely empty filter means "all collections"; anything else has to name one.

* shell: keep commas inside a regex group out of the entry split

The split already left alone the commas inside a character class or a
repetition count, but not the ones inside a group, so "bucket(foo,bar)"
was cut into two fragments that no longer compile.

* shell: cover escaping a collection name that is not a regex

A name like "logs(2024" does not parse as a regex on its own; escaping it,
"logs\(2024", reaches it. Pin that so the escape hatch does not regress.

* shell: split entries only on commas inside a closed regex construct

An unmatched "{" or "[" made the splitter swallow every comma after it, so
"foo{bar,videos" became one entry that matches neither collection - the
silent no-op this filter work exists to remove. A construct now has to close
before its commas stop separating entries.

* shell: skip character classes while scanning a regex group

A ")" inside a class is a literal, so "(a[)],b)" ended its group early and
split into two fragments that no longer compile.

* shell: cover escaping a comma inside a collection name

A comma separates entries, so a name holding one is reached by escaping it.

* shell: follow the regexp parser when scanning a character class

A "]" leading a class is a member of it, and a POSIX class such as
"[:alpha:]" carries its own "]", so stopping at the first one cut a valid
filter like "(a[]),],b)" into fragments and rejected it.
2026-08-25 18:03:52 -07:00
Chris LuandGitHub 368b2035b2 s3: deny anonymous access when the identity config loads no identities (#10954)
* s3: deny anonymous requests when the identity config loads no identities

Naming a config file is the operator asking for authentication. A file that
yields no identity - an unpopulated secret mount, or a mistyped top-level key
the proto parser silently drops - left the gateway open to every anonymous
caller: ListBuckets returned 200, and anonymous PUT could create buckets and
write objects.

* s3: name the unknown top-level keys in an identity config

The proto parser discards what it does not recognise, so a mistyped
"identites" loads as an empty config. Naming the dropped keys at startup turns
the resulting lockout into a one-line diagnosis.

* s3: isolate the auth-enforcement tests from AWS environment credentials

* s3: use a singular "identity" as the unrecognised-key example

Codespell rejects the misspelling the example used.

* s3: cover the empty identity config alongside the unrecognised key

* s3: cover a config file whose body is an empty object
2026-08-25 15:31:32 -07:00
Chris LuandGitHub b58d52ac16 rust volume: search .sdx for read-only volumes instead of holding the index (#10951)
* rust volume: search .sdx for read-only volumes instead of holding the index

The Go volume server loads every read-only volume through SortedFileNeedleMap:
the index lives on disk as a sorted .sdx, a lookup is a binary search, and
since #10950 no descriptor is held between lookups. The Rust server had no
counterpart. Read-only volumes built a full in-memory CompactNeedleMap, and
cloud-tiered ones — noWriteCanDelete, so not the read-only branch — went
through the writable path and pinned an .idx append handle on top of it. At the
hundreds of thousands of tiered volumes a real server carries, that is an index
in RAM and a descriptor each, for volumes nobody reads.

Port the sorted map and the bounded handle pool. A tiered volume now costs zero
descriptors and zero index bytes when idle; the pool keeps the hot handles open
so a busy volume does not pay an open() per needle. Handles are Arc<File>, so
an eviction cannot close one a reader still holds.

The generated .sdx is byte-identical to Go's — same sort, same last-write-wins,
same dropped tombstones — so a volume moved between a Go and a Rust server reads
whichever copy is already on disk. A test pins the bytes against a Go-generated
fixture.

* rust volume: fail compaction on an unreadable .sdx, and rebuild the map on tier-down

Two ways the sorted map could lose data.

iter_entries swallowed read errors and returned however many entries it managed
to collect. Compaction takes that vector for the complete live set, so a
truncated .sdx or a mid-scan I/O fault would commit a volume missing every
needle past the failure. Return a Result instead and abort. redb's
collect_entries dropped errors the same way on the same path, so it goes with
it.

Tier-down clears the remote mode and publishes the volume as writable, but the
map it booted with is the read-only sorted one. Its put always fails, so the
first write would append to the local .dat and then fail to index it, leaving
bytes nothing references — and a non-fsync write repeats it. Fold the
reopen_idx_for_write swap into refresh_remote_write_mode so the map always
matches the mode it just published; a rebuild that fails pins the volume
read-only rather than letting it take writes it cannot record.

Go reaches neither: its tier-down leaves noWriteCanDelete set, so the volume
stays read-only until a reload or an explicit mark-writable, which already goes
through reopenIdxForWrite.

* rust volume: keep read-only volumes mountable on a read-only index dir, and batch the .sdx scan

Building .sdx writes to the index directory, and load_index_sorted_file also
created a missing .idx there. A volume whose index sits on a read-only mount
took both paths and failed to load, where before it mounted read-only off an
in-memory index and served reads. Create the .idx only where deletes are
allowed, and fall back to the in-memory map when the sorted one cannot be
built, so a directory nobody can write costs memory rather than availability.

The end-to-end scan behind iter_entries, ascending_visit and save_to_idx read
one entry per syscall. Read 1024 at a time instead, the batch size
idx::walk_index_file uses. Positional reads, not a cursor: the handle is shared
with any other borrower.

Also gate the Go byte-parity fixture on the 5bytes feature it describes, which
is otherwise dead code in a 4-byte-offset build.

* rust volume: roll back a failed writable mark, and rebuild a torn .sdx

set_writable clears the read-only flags before it can know the rest will
succeed, but only the map rebuild rolled them back. An .idx writer that fails to
attach left the volume advertising writable over a needle map with no writer, so
puts landed in memory and were gone after a restart — the exact failure the
function exists to prevent. The read-only-mount fallback made it reachable: that
path loads an in-memory map with no writer attached. All three steps now run
behind one rollback point.

A .sdx whose length is not a whole number of entries was accepted as long as it
looked fresh, and truncation is what makes it look fresh. The entry count then
floored, hiding the last needle from lookups and from compaction, which would
commit the shorter set. Treat a torn file like a stale one and rebuild it from
.idx. Go writes .sdx in place rather than through a temporary, so a crash
mid-generation is a real way to produce one.

Appends now start at the last whole .idx entry too, so a torn tail there is
overwritten by the next tombstone instead of misaligning every row after it.

* rust volume: trim a torn .idx before writing to it, keep delete-only volumes online, count sorted-map deletes

Three from review.

Flooring the sorted map's append offset only protected its own positional
writes. Every writable path appends at EOF instead, so a partial row left by a
short write pushed the next row off alignment and the following load parsed the
rest of the file as garbage. Drop the partial row before attaching any writable
index writer — it is unrecoverable anyway, and every loader already skips it.
Go refuses to load such a volume at all; trimming keeps it mountable with the
rows before the tear intact.

The unwritable-index-dir fallback stopped one step short for volumes that allow
deletes, which is every tiered one: the in-memory loader opens .idx read-write
there and fails on the same directory that just refused the .sdx, so the volume
stayed offline. Give up the deletes instead — without a writer no tombstone
could be recorded anyway — and a remount on a writable directory restores them.

Sorted-map deletes left the counters untouched, so a tiered volume reported
itself garbage-free until it restarted. They now land where a reload would put
them: the tombstone is another .idx row, and both it and the row it supersedes
count as deletions under the rule the load-time metric applies. Go skips this
too, and should not.
2026-08-25 15:21:37 -07:00
Chris LuandGitHub e482e67971 admin: accept a list of collections in the task collection filter (#10953)
The collection filter was parsed twice with two syntaxes: the master-side
volume listing compiled the whole string as one regex, while EC encode and
EC balance detection split it on commas and matched each entry as a
wildcard. A volume had to pass both, so "collection-a,collection-b" matched
nothing (no collection is named that), and the ALL_COLLECTIONS sentinel,
which the master side skips, dropped every volume at the task side.

Parse it once, in one place: a comma-separated list where an entry is a
name with optional * and ? wildcards, or a regex when it carries regex
syntax. A regex entry now has to match the whole name unless it anchors
itself, so listing a collection no longer picks up its longer namesakes.
2026-08-25 15:13:06 -07:00
ef4c9d9178 filter volume by local or remote storage name (#10946)
* filter volume by local or remote storage name

Signed-off-by: lou <alex1988@outlook.com>

* fix SelectsEverything

Signed-off-by: lou <alex1988@outlook.com>

* keep the proto sync out of this change

The branch copied weed/pb/*.proto over their seaweed-volume and Java
counterparts and regenerated every .pb.go with a different protoc and
protoc-gen-go-grpc. DiskStatus.error arriving that way broke the Rust
build, and the rest is toolchain churn in files this change has nothing
to say about.

---------

Signed-off-by: lou <alex1988@outlook.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-25 13:05:33 -07:00
Chris LuandGitHub 70c3adb983 volume: stop read-only volumes from pinning .idx and .sdx (#10950)
A read-only or cloud-tiered volume loads a SortedFileNeedleMap, which held
both its .idx and its .sdx open for the life of the process. On a server with
~600K tiered volumes that is 1.2M descriptors before a single read, enough to
exhaust the fd limit and take the listeners down. The .dat is not the problem:
a tiered volume serves it from the remote backend.

Neither index file is needed except while a lookup is in flight, so borrow them
from a bounded process-wide pool instead. An idle volume now holds zero
descriptors; a busy one keeps its handles hot rather than paying an open() per
needle. Reads borrow O_RDONLY, so a volume on a read-only mount answers lookups
that previously failed at load. Sync tracks whether a tombstone was appended,
which also drops the fsync-per-volume storm at shutdown.
2026-08-25 10:23:00 -07:00
Chris LuandGitHub b77431c142 master: stop hintless small-file assigns from marking volumes full (#10944)
* master: estimate a hintless assign's size from the volume's average file size

An assign that carries no dataSize hint charged a flat 1MB per file id
against the volume's effective size. A small-file workload overpays by
orders of magnitude: bulk-writing 4KB files marks volumes holding a few
hundred MB of real data as crowded and then full, so the master grows
unnecessary volumes and, once every volume is spuriously full, fails all
assigns. Estimate from the volume's own average file size instead, and
keep the 1MB fallback only for volumes with no history.

* master: decay pending assign sizes for volumes gone quiet

The decay that corrects pending assign estimates runs only when a
heartbeat reports the volume, and a heartbeat only reports a volume
whose content changed. A volume held out of the writable list takes no
writes, so once inflated estimates mark every volume full, nothing is
ever reported again, nothing decays, and the cluster refuses all writes
until a restart. Run the decay from the master's periodic loop for
volumes no heartbeat has reported within two pulses, feeding the last
reported size back through the same path an unchanged heartbeat would
take.

* master: trim the comments on the assign size estimate

* master: keep the periodic decay out of the replica-dedup window

UpdateVolumeSize ignores a report arriving within two seconds of the last
one, so replicas of the same volume do not each halve the pending
estimate. The periodic decay went through the same path and stamped that
window, so a real heartbeat landing right behind it was dropped along
with its reported size and compact revision. Only a volume whose content
changed is reported at all, so nothing would send that size again and
the master kept a stale one. Let the dedup window belong to volume
server reports alone.

* master: let the decay read the size record under the lock it mutates

The periodic decay picked its volumes under a read lock and replayed
them under a write one, carrying the size it had read across the gap. A
heartbeat landing in between was rolled back: the replay wrote the older
size and compact revision over the fresh ones, and a compaction report
lost that way is never resent, since only a volume whose content changed
is reported. The decay has no size of its own to contribute, so it now
reads the record under the same lock it mutates.

* master: let a heartbeat that beat the decay stand for the cycle

The decay chooses its volumes under a read lock and applies them under a
write one. A heartbeat landing in that gap already did the halving the
cycle owed, so applying the decay on top of it halved twice and forgot
pending bytes the volume has not written yet - the double-halving the
replica-dedup window exists to prevent. Both callers now give way to a
report already handled for this cycle; only a real report still advances
lastUpdateTime, so a quiet volume keeps decaying every pulse.

* master: keep genuinely full volumes out of the decay pass

A volume the disk really did fill keeps its fullSince set for good, so it
was selected every pulse for a decay that cannot help it: UpdateVolumeSize
refuses to recover a volume whose reported size is at the limit, and
replaying a size that cannot move leaves the record as it found it. Full
and quiet is the ordinary resting state of a cluster, so this was most of
the pass, taking the layout write lock away from the heartbeats to do
nothing. On a million tracked volumes with a hundredth of them phantom-full
it costs ten thousand write locks a pulse instead of a million.

* master: put the stale-replay test back on the path it guards

Giving the decay the dedup window left this test short-circuiting there,
so it no longer reached the locked read it was written for and passed
with that read removed. Age the record past the window, which is the only
case where reading it under the lock is what saves the report.
2026-08-25 10:16:48 -07:00
Chris LuandGitHub 50b388771a s3: stop one abandoned request from cancelling every concurrent upload (#10948)
* grpc: a non-cancellable context is no evidence of a stale channel

shouldInvalidateConnection only invalidates on Canceled/DeadlineExceeded
while the context handed to WithGrpcClient is still live, so that an RPC
timing out on its own does not close the shared cached ClientConn and
cancel every other in-flight RPC on it. context.Background()/TODO never
expire, so Err() stays nil forever and that guard always answered
"invalidate" - and Background is what almost every caller passes, the S3
gateway included.

One S3 request whose RPC rode an abandoned HTTP request context therefore
closed the shared filer connection, and every multipart part in flight
died with "the client connection is closing", surfacing to the client as
400 InvalidRequest.

Only a cancellable context bounds an RPC attempt, so require one before
reading it. A genuinely stale channel (a peer restart behind a stable L4
endpoint) surfaces as Unavailable, which invalidates on its own branch.

* grpc: a bystander of a connection teardown is not a stale-channel witness

gRPC raises ErrClientConnClosing locally, before an RPC reaches the wire,
when this process has already closed the ClientConn. Every caller that
touches a channel during another goroutine's teardown gets it, so reading
it as a stale-channel signal lets one teardown re-arm itself across the
whole herd of callers it just cancelled.

The cached-connection version check keeps those callers from closing a
replacement channel, but the streaming path invalidates by address alone
and has no such guard.

* grpc: end a stream without dropping the peer connection under it

A streaming caller gets its own ClientConn, but on any error it also drops
the cached non-streaming ClientConn every request handler shares with that
peer, to recover a peer restart hidden behind a stable L4 endpoint. Any
error includes the ordinary ones: a metadata subscription that reached its
stop point, a follow callback that refused an event, a caller that gave up.

The S3 gateway follows filer metadata on such a stream and reconnects
forever, so each ordinary end of it cancelled every S3 request in flight
against the filer. Drop the shared channel only for errors that say the
peer went away, which is what invalidation is for.

* test: close the connections the cascade tests leave cached

Each test swaps in a fresh connection cache and restores the previous one,
dropping its own entries without closing them, so the ClientConn's
transport and reconnect goroutines outlive the fake filer they dialed.

* grpc: say why ErrClientConnClosing's deprecation notice does not apply

It points at codes.Canceled, which is the code this function exists to
disambiguate. Only the message distinguishes a teardown a caller merely
walked into, so the sentinel stays.
2026-08-25 10:15:47 -07:00
Chris LuandGitHub 44115c1051 filer: stop TUS uploads from turning into garbage (#10945)
* filer: store TUS sub-chunks through the regular chunk writer

A TUS sub-chunk was written with one assigned file id, retried up to
three times against that same id, and abandoned on failure: an attempt
that had landed on some replicas left a needle no session record and no
entry ever references, unreclaimable by vacuum.

dataToChunkWithSSE, which the regular write path uses per chunk, assigns
a fresh file id per attempt and hands back the file ids of failed
attempts, which are now freed the way the regular write path frees them.

* filer: retry a chunk write on a fresh volume when the server 5xxs

The filer's chunk writer assigns a fresh file id per attempt but only
retried transient network errors, so a volume filling up and turning
read-only mid-write failed the whole request even though the very next
assignment would have landed elsewhere. Every other write client already
routes this through ShouldReassignUpload; the filer's own write path now
does the same, for regular uploads and TUS sub-chunks alike.

* filer: export the chunk deletion queue

The filer test harness in weed/server builds filer.Filer as a struct
literal, so any code path reaching DeleteChunks dereferenced a nil
queue. Exported like the neighboring DeletionRetryQueue so the harness
can arm it.

* filer: complete a TUS upload whose chunk records overlap

A PATCH retried while its predecessor was still storing a sub-chunk -
a proxy timeout with an immediate retry is enough - records the same
range twice. HEAD computes Upload-Offset as the covered watermark and
reported the upload fully received, but completion demanded exactly
adjacent records and failed every attempt: the client concluded success
from offset == length, no entry was created, and the session eventually
expired, turning the entire upload into deleted needles for the vacuum
to chew through.

Completion now validates gapless coverage with the same watermark HEAD
uses. A record extending coverage joins the entry - the read path
resolves partial overlaps by ModifiedTsNs, and the raced copies carry
identical bytes - while a fully covered duplicate is freed once the
entry lands.

* filer: allow one mutating TUS request per session at a time

Nothing stopped two PATCHes from writing the same range concurrently:
both loaded the same offset, both passed the conflict check, and both
recorded their sub-chunks. A client whose request timed out in a proxy
retries immediately while the server side is still storing the buffered
sub-chunk, which is exactly that race.

A session now accepts one PATCH or DELETE at a time, the way tusd locks
uploads; a concurrent one is refused with 423 Locked, which TUS clients
retry, and HEAD keeps answering so progress polling is unaffected. The
chunk state is loaded under the claim, so a retried PATCH sees every
record its predecessor left and conflicts cleanly instead of duplicating
data.

* test: cover a TUS PATCH raced by its own retry

Stalls a PATCH mid-body over a raw connection, retries the same range
while it is in flight, and expects the retry refused with 423 Locked;
the upload then resumes from the reported offset and the final content
must be intact.

* filer: never free a TUS duplicate the entry still references

Coverage is computed from ranges, so a record fully covered by another
is treated as a duplicate no matter which needle it names. A malformed
record naming a file id the entry keeps would have had that needle freed
right after the entry landed - the corruption this change set exists to
stop. The duplicates are now freed in one batch, skipping any file id
the entry references; their records go with the session directory.

* test: bound the raw TUS connection reads

http.ReadResponse on the stalled PATCH's connection blocked until the
whole go test timeout if the filer never answered.

* filer: free the needles of chunk write attempts a retry replaced

A volume server stores the needle locally and only then fans out to the
replicas, so a replication failure 5xxs with the data already written.
Each attempt assigns its own file id, so once a later attempt lands
elsewhere nothing references the earlier ones: the caller only sees the
chunk that succeeded, and the failed ids were dropped.

They are now freed the way the caller frees them when the whole write
fails. Retrying on a 5xx makes this reachable on every read-only or full
volume, which is exactly the condition that filled the reporter's
volumes.
2026-08-25 09:24:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
c69bb10407 build(deps): bump github.com/getsentry/sentry-go from 0.44.1 to 0.48.0 (#10921)
* build(deps): bump github.com/getsentry/sentry-go from 0.44.1 to 0.48.0

Bumps [github.com/getsentry/sentry-go](https://github.com/getsentry/sentry-go) from 0.44.1 to 0.48.0.
- [Release notes](https://github.com/getsentry/sentry-go/releases)
- [Changelog](https://github.com/getsentry/sentry-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/getsentry/sentry-go/compare/v0.44.1...v0.48.0)

---
updated-dependencies:
- dependency-name: github.com/getsentry/sentry-go
  dependency-version: 0.48.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* bump cockroachdb/errors to v1.14.0 for sentry-go 0.48.0

sentry-go 0.48.0 removed Event.Extra, which cockroachdb/errors v1.11.3
still references; v1.14.0 builds against the new API.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-25 02:00:52 -07:00
Chris LuandGitHub 68f0793b6f mount: register UNC mount points as WinFsp network file systems (#10943)
A \\server\share -dir was passed to WinFsp as a plain mount point, which
treats it as a directory path on an actual remote server and fails. Turn it
into the VolumePrefix option instead, so the mount registers with the WinFsp
network provider: the UNC path is then reachable from every logon session,
which a drive letter mounted from a service is not, and each user can map
their own drive letter to it.
2026-08-25 01:28:50 -07:00
Chris LuandGitHub 40f77503d0 helm: trim the Lance chart comments (#10940)
Comments only, no rendering change: the values paragraphs compress to
the density of the file around them, the env-var note becomes a
template comment instead of leaking into the rendered manifest, and the
two spots that invite a wrong simplification - the unconditionally
rendered -port.lance and the empty-placeholder platform guard - each
get their one-line why.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-24 23:51:35 -07:00
Chris LuandGitHub d9d7d0be74 helm: serve the Lance catalog and deploy the Rust worker (#10936)
* helm: serve the S3 gateway's Lance Namespace, on by default

Standalone `weed s3` serves the Lance Namespace API on 9101 unless told
not to, so the chart defaulting s3.lancePort to 9101 matches weed's own
posture instead of hiding the port behind a null. The flag is always
rendered, so lancePort: 0 reaches weed as -port.lance=0 and genuinely
disables the namespace rather than silently falling back to the binary
default; 0 also drops the service port and the optional lanceIngress,
which otherwise mirror the iceberg wiring. The NetworkPolicy admits the
port the same way it admits icebergPort.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* helm: run the Lance maintenance worker beside the Go worker

The Go and Rust workers have no overlapping jobs - Go serves vacuum,
balance, EC and iceberg_maintenance, only /usr/bin/weed-worker serves
the lance_* family - so a cluster serving Lance tables needs both, not
an either/or switch. The worker deployment now adds a worker-lance
container whenever the namespace is reachable: worker.namespaceUrl, or
derived from the release's S3 service and s3.lancePort. Untouched Go
container; admin address derived the same way; mTLS flags point at the
already-mounted worker cert when security is on; metrics on their own
worker.lanceMetricsPort (9328, next in the 932x convention) with the
same health probes, service port and scrape endpoint the Go container
gets, and the worker NetworkPolicy admits that port exactly when the
container renders. The image carries an empty placeholder on armv7/386
where exec falls back to the shell and exits 0, so the command refuses
those platforms by name; s3.lancePort: 0 is the escape hatch there.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-24 21:02:37 -07:00
Chris LuandGitHub b3be2f5449 filer.backup, filer.sync: stop sharing resume checkpoints across destinations (#10934)
* filer.backup: key the checkpoint by source path and sink destination

The checkpoint id hashed only sink name + directory, so two backups to
different buckets or endpoints sharing a directory layout advanced one
checkpoint: whichever job was running pushed the shared offset forward,
and a stopped or failing job later resumed from the other's position,
silently skipping changes. Backups of different source paths to the same
destination shared a checkpoint the same way.

Each sink now reports a destination identity (endpoint or account,
bucket or container, directory) and the checkpoint is keyed by the
source path plus that identity. Reads fall back to the historical
name+directory key when the new key has no value, so existing backups
resume where they left off; writes go only to the new key.

* filer.sync: include the target path in the offset key

The offset stored on the target filer was keyed by source path and
source filer signature only, so two syncs from the same source cluster
and path to different directories on the same target cluster advanced
one shared checkpoint, and the slower one could resume past events it
never applied. The target path now participates in the key; "/" keeps
the historical form, and a sync with a non-root target path falls back
to the historical key once when its own key has no value yet.

* join checkpoint key fields with NUL so they cannot alias

A path or configuration value spelling out the separator could
concatenate two different field tuples to the same checkpoint key.
NUL cannot appear in a CLI path argument or any sane configuration
value, making the encoding injective.
2026-08-24 19:30:20 -07:00
Chris LuandGitHub 4a2879abad admin: show a copyable S3 object URL in the bucket file browser (#10933)
* admin: offer copyable S3 object URLs in the bucket file browser

* admin: hide object urls when the bucket type lookup fails

* admin: ignore an s3.public_endpoint that is not an absolute http url

* mini: build the seeded s3 endpoint with JoinHostPort for ipv6

* admin: reject a query or fragment in s3.public_endpoint

* mini: drop the seeded s3 endpoint when a later run disables s3

* admin: reject userinfo and bare delimiters in s3.public_endpoint, redact the warning

* mini: pass its s3 endpoint as an admin option instead of mutating viper

* admin: keep the rejected s3.public_endpoint value out of the log
2026-08-24 19:29:01 -07:00
Chris LuandGitHub 2a70532d0d s3: log each request at -v=2 (#10931)
* s3: log each request at -v=2

* s3: quote requester and path in the access log line

* s3: record the post-policy signing identity as the requester
2026-08-24 18:39:30 -07:00
Chris LuandGitHub d2c470af1b S3: commit SSE GET status only after the first read succeeds (#10935)
The SSE streaming path kept writing 200/206 from filer metadata before
fetching or decrypting anything, so a missing needle or failed decrypt
setup surfaced as a broken 200 body. Same deferral as the plain path:
the status commits on the first body write, and every failure before
that returns to the handler for a clean S3 error response.
2026-08-24 18:36:46 -07:00
Chris LuandGitHub 115756dd41 helm: expose loadBalancerClass, loadBalancerIP, loadBalancerSourceRanges on services (#10929) 2026-08-24 16:31:47 -07:00
Chris LuandGitHub d9d5fab35b S3: commit GET status only after the first read succeeds (#10930)
streamFromVolumeServers wrote the 200/206 status from filer metadata
before any byte had been fetched from a volume server, so a missing or
corrupted needle surfaced as a broken 200 body and the request metrics
recorded a success. Defer the status commit to the first body write: a
failed first read now returns a clean 500 before headers, while the
wire timing of successful responses is unchanged since net/http buffers
the status line until body bytes arrive anyway.
2026-08-24 16:14:52 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d8545997c2 build(deps): bump github.com/seaweedfs/goexif from 1.0.3 to 2.0.0+incompatible (#10917)
build(deps): bump github.com/seaweedfs/goexif

Bumps [github.com/seaweedfs/goexif](https://github.com/seaweedfs/goexif) from 1.0.3 to 2.0.0+incompatible.
- [Release notes](https://github.com/seaweedfs/goexif/releases)
- [Commits](https://github.com/seaweedfs/goexif/commits)

---
updated-dependencies:
- dependency-name: github.com/seaweedfs/goexif
  dependency-version: 2.0.0+incompatible
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:24:08 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
9121177f48 build(deps): bump github.com/shirou/gopsutil/v4 from 4.26.6 to 4.26.7 (#10918)
Bumps [github.com/shirou/gopsutil/v4](https://github.com/shirou/gopsutil) from 4.26.6 to 4.26.7.
- [Release notes](https://github.com/shirou/gopsutil/releases)
- [Commits](https://github.com/shirou/gopsutil/compare/v4.26.6...v4.26.7)

---
updated-dependencies:
- dependency-name: github.com/shirou/gopsutil/v4
  dependency-version: 4.26.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:23:58 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
9ba3d473d7 build(deps): bump cloud.google.com/go/pubsub from 1.51.0 to 1.51.1 (#10919)
Bumps [cloud.google.com/go/pubsub](https://github.com/googleapis/google-cloud-go) from 1.51.0 to 1.51.1.
- [Release notes](https://github.com/googleapis/google-cloud-go/releases)
- [Changelog](https://github.com/googleapis/google-cloud-go/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-cloud-go/compare/pubsub/v1.51.0...pubsub/v1.51.1)

---
updated-dependencies:
- dependency-name: cloud.google.com/go/pubsub
  dependency-version: 1.51.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:23:49 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e36b01c18f build(deps): bump github.com/rabbitmq/amqp091-go from 1.13.0 to 1.14.0 (#10920)
Bumps [github.com/rabbitmq/amqp091-go](https://github.com/rabbitmq/amqp091-go) from 1.13.0 to 1.14.0.
- [Release notes](https://github.com/rabbitmq/amqp091-go/releases)
- [Changelog](https://github.com/rabbitmq/amqp091-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/rabbitmq/amqp091-go/compare/v1.13.0...v1.14.0)

---
updated-dependencies:
- dependency-name: github.com/rabbitmq/amqp091-go
  dependency-version: 1.14.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:23:40 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
dfa75ce231 build(deps): bump actions/upload-artifact from 4 to 7 (#10922)
Bumps [actions/upload-artifact](https://github.com/actions/upload-artifact) from 4 to 7.
- [Release notes](https://github.com/actions/upload-artifact/releases)
- [Commits](https://github.com/actions/upload-artifact/compare/v4...v7)

---
updated-dependencies:
- dependency-name: actions/upload-artifact
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:17:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
61d588e455 build(deps): bump github/codeql-action from 4.37.6 to 4.37.8 (#10923)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.6 to 4.37.8.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.6...v4.37.8)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:17:44 -07:00
a3afe4460b tarantool: fix upsert data corruption and missing context propagation (#10926)
Co-authored-by: Marat Karimov <karimov_m@inbox.ru>
2026-08-24 15:17:08 -07:00
a0ddf22f17 Bump Tarantool client library from 3.0.0 to 3.0.1 (#10925)
Co-authored-by: Marat Karimov <karimov_m@inbox.ru>
2026-08-24 15:10:48 -07:00
Chris LuandGitHub 863fec6c3f S3: let a key that is a prefix of other keys be an object (#10912)
* filer: keep the sentinel when CreateEntry reports an update failure

CreateEntry flattened the error UpdateEntry wraps, so errors.Is stopped
matching and ErrExistingIsDirectory and ErrExistingIsFile never reached
the S3 mapper, which answered a retryable 500 instead.

* s3: let a key that is a prefix of other keys be an object

S3 keys are flat, so "a/b" and "a/b/c" are independent objects that
coexist in either write order. The filer stores a key as a path, so one
of them has to live on the directory the other is nested under.

Writing the nested key first refused the prefix key outright. Writing it
second promoted the file to a directory, which kept its data but lost the
key: an empty object left nothing to recognise it by and disappeared, and
one with data listed under a trailing slash it never had.

Mark the directory that carries such a key, and write the object onto it
when the path is already a directory. The mark makes an empty prefix
object visible to listings and readable by GET and HEAD, keeps the empty
folder cleaner off it, and lists it under the key it was written with.
Deleting the key strips the mark back off along with the data.

* filer: keep a TTL off a directory that stands for an object

An expired entry is deleted a row at a time, so expiring a directory
removes it and leaves everything under it unreachable. Promoting a file
to a directory carried its TTL across, and a promoted file is exactly the
one that has keys nested under it.

Drop the TTL on promotion, and leave one an older build wrote alone. The
lifecycle worker still expires the object, through the delete that leaves
the directory behind.

* s3: delete the null version of a key other keys are nested under

The routed delete cannot remove an entry that other keys live under, and
answered a retryable 500 rather than falling back to the lock path the
unversioned delete already falls back to. That path then looked the entry
up under the bucket with the whole key as its name, so the demote wrote it
back one directory too high and failed as not found.

Fall back on any non-precondition error, and split the key before deleting
it. Trailing-slash directory markers with children reach the same delete.

* filer: keep the sentinel when MkFile and Mkdir report a create failure

Same flattening one layer out: every mkFile caller lost the sentinel, so
a CopyObject onto a key that other keys are nested under answered a
retryable 500 where a PutObject of the same key answers 409.

* s3: copy and rename a key that other keys are nested under

Such a key is stored on the directory those keys live in, and copy and
rename both refused it: the source lookup maps every directory entry to
NoSuchKey, so a key a plain GET serves could not be copied or moved, and
the destination side refused it as a directory conflict.

The source is read through a view of the entry as the object it names.
The destination is written the way a PutObject of that key writes it. A
rename at either end copies the object's own data across and strips it off
the source key rather than going through AtomicRenameEntry, which moves a
directory by moving everything under it - the nested keys are not part of
what is being renamed.
2026-08-24 15:10:34 -07:00
Chris LuandGitHub 46ce2c45a2 mini: reserve the admin gRPC port instead of binding it late (#10928)
* mini: reserve the admin gRPC port instead of binding it late

Port selection probes every port with a throwaway listener and closes it.
Master, filer, volume and S3 bind a moment later, but the admin waits for
all of them first and only then binds its worker gRPC port, roughly two
seconds in. That port defaults to the admin http port + 10000, which lands
inside the Linux ephemeral range, so one of the cluster's own outgoing gRPC
dials can take it during the gap and the admin dies on bind, taking the
worker with it.

Keep the listener from the availability check and hand it to the admin.

* mini: clear the admin gRPC reservation before retaking it

A rerun inside one process would otherwise inherit the closed listener of
the previous run whenever the reservation fails, and the admin would accept
it and only find out inside Serve.

* mini: snapshot the admin options for the startup goroutine

The cleanup path read the package-level options long after the goroutine
started, so a later in-process run could have its reserved listener closed
by the previous run.
2026-08-24 14:55:53 -07:00
Chris LuandGitHub 51eb5333d3 ec: read a needle's intervals in parallel (#10911)
* ec: read a needle's intervals in parallel

A needle spanning more than one EC block gets one interval per block, and
consecutive blocks live on different shards. We read those intervals in
sequence, so a 4MB chunk landing in a volume's 1MB small-block region cost
five round trips to five different servers.

Read them concurrently into disjoint slices of a single buffer, at most 8 in
flight. Same change in the Rust volume server's phase C.

* ec test: seed the random payload instead of the deprecated rand.Read
2026-08-24 14:03:44 -07:00
Chris LuandGitHub 69cc2869ad Fixes from the review of the admin bucket policy UI (#10907)
* admin: treat a missing S3 Tables policy as an empty load, not an error

The bucket/table policy GET relayed the backend's 404 NoSuchPolicy to the
dialog, whose loader treats any non-OK response as a load failure and
keeps Save and Delete blocked. A bucket or table without a policy could
never be given one. Return policy null instead, the same contract
ShowBucketPolicy uses for classic buckets.

* admin: reject policy documents the structured editor would misread

A top-level JSON array passed the object guard (typeof [] is 'object')
and loaded as a zero-statement policy, which the next commit would
rewrite to an empty document. Object elements in Action/Resource were
coerced to '[object Object]' and saved that way on the s3tables surface,
which stores policies verbatim. Both now throw, which routes the
document to the JSON tab like other unrepresentable shapes.

* admin: let the JSON tab save documents the structured editor can't model

Save with the JSON tab active required a round-trip through
policyDocToEditorState, so exactly the documents the dialogs shunt to
'JSON tab only' mode (unrepresentable Effect, Resource+NotResource, and
the like) could never be saved - Delete was the only mutation left.
Invalid JSON still blocks; an unrepresentable document now saves and the
editor state stays marked unparsed.

* admin: pin the policy editor to what each consumer's backend supports

The s3tables evaluator has no NotResource/NotPrincipal fields - it
silently drops them, turning Allow+NotResource into allow-everything and
making Deny+NotPrincipal inert - and it only matches s3tables: actions
against s3tables ARNs, while the editor suggested s3: actions and
arn:aws:s3::: resources. New registerPolicyEditor knobs: allowNegation
hides the Not* modes and routes documents using them to the JSON tab;
resourceSuggestions pins the Resource autocomplete to the open
resource's ARN; the S3 Tables dialogs get an s3tables-only action
datalist. requirePrincipal now also hides NotPrincipal, which
policy_engine.ValidateBucketPolicy always rejects, and the client-side
check requires Principal specifically to match that server rule.

* admin: save S3 Tables policies from a button, not form submission

The multi-input structured editor sits inside a form whose Save button
was type=submit, so Enter in any single-line editor input - accepting an
autocomplete suggestion, say - implicitly submitted whatever half-built
statement the editor held, and the backend stores the document verbatim.
A lone statement with no Principal matches nobody, locking out every
non-owner. Save is now an ordinary button and the form ignores
submission.

* admin: block zero-statement policy saves

Committing the active tab before the emptiness check made 'Policy JSON
is required' dead code: an empty editor serializes to {"Statement":[]},
which the s3tables backend stores verbatim - evaluated default-deny for
every non-owner, while the statement-count column keeps showing 'Not
configured'. All three policy dialogs now refuse a save with no
statements and point at Delete instead. The classic bucket modal only
gained a clearer message; the server already rejected the document.

* admin: guard S3 Tables policy mutations against stale and overlapping requests

The save/delete completions ran against whatever resource the shared
modal happened to show by then: a slow PUT for one bucket would hide the
modal mid-edit of another and misattribute its alerts, a late DELETE
cleared the shared textarea over the newly opened resource with its
loaded flag set, and nothing stopped a double-click from firing two
overlapping mutations. Ported the classic modal's pattern: capture the
target on start, flag the mutation in flight with the buttons disabled,
and only touch the UI when the completion still matches the open
resource. Success now reloads the page, which also keeps the Policy
column's statement count honest.

* admin: confirm before deleting an S3 Tables policy

Delete Policy sat next to Save and fired on a single click; with
default-allow enabled one stray click silently dropped the resource
policy and left the bucket open to every principal. Same confirmation
the classic bucket modal already has.

* admin: let a corrupt stored bucket policy be shown, fixed, and deleted

A stored document the decoder rejects made the policy GET 500, and with
the loaded flag never set the modal blocked both Save and Delete - the
one policy an operator most needs to remove was the one they couldn't,
even though the delete path never reads the document. The GET now
returns the raw bytes alongside a null policy; the dialog hands them to
the JSON tab and unblocks the buttons.

* admin: url-encode the bucket name in the policy API calls

The filer lists any directory under the buckets path, names S3 would
never allow included; one carrying '#' or '%' broke the fetch URL or
addressed a different name than the modal shows.

* admin: drop stale edit-policy responses on the IAM policies page

The same race the bucket and S3 Tables dialogs already guard against:
open one policy's editor while its GET stalls, open another, and the
late response populates the editor under the second policy's name -
Update then saves the first policy's statements over the second.

* admin: warn before a bucket policy save drops unsupported fields

The editor tracks unmodeled top-level keys precisely so
confirmPolicyFieldDiscard can warn before the server's Version+Statement
decode discards them, but only the IAM page called it; the bucket modal
saved a pasted document with e.g. a console-generated Id without a word
while the editor kept displaying the field.

* s3: enforce the bucket policy size cap on both surfaces

The 20KB cap lived only in the admin UI, so a larger policy stored via
the S3 API displayed there but could never be re-saved, desyncing the
two writers the cap comment claimed could not desync. The constant now
lives in policy_engine next to the shared validator and PutBucketPolicy
rejects oversized documents with PolicyTooLarge, matching AWS.

* admin: ship the policy editor's fieldset styles with the editor

The .policy-stmt-* rules that undo Bootstrap's full-width legend reset
stayed behind in policies.templ when the editor markup moved to the
shared script, so the bucket and S3 Tables dialogs rendered Actions/
Resource/Principal as full-width jumbo headings. PolicyDatalists is the
component every consumer already renders once; the styles live there
now.

* s3: mirror bucket policy changes into the IAM store from the metadata subscription

The advanced-IAM path appends the bucket-policy:<bucket> document to
every STS/session evaluation, but only this gateway's own PutBucketPolicy
maintained that mirror - a policy tightened or created through the admin
UI (or another gateway) never reached it, so revoked access stayed live
indefinitely, and the delete side was an unimplemented TODO in any case.
The metadata subscription now diffs the stored policy on every bucket
entry change and updates or removes the mirror, covering all writers and
deletion with one mechanism; IAMManager gains the missing
RemoveBucketPolicy.

* admin: deduplicate the bucket policy write path

Set and Delete carried line-for-line identical filer closures;
bucketPolicyMutation already treats nil as clear-the-key. The shared
helper sits below Set's validation, since ValidatePolicy cannot take the
nil document Delete passes.

* s3: drop ValidateBucketPolicy's re-checks of ValidatePolicy rules

Both callers run ValidatePolicy first, which already enforces the
version and at-least-one-statement rules; the duplicates were dead code
with drifted error text.

* admin: seed a new statement's Resource from the pinned suggestions

A fresh statement on the S3 Tables dialogs started with no resource row
at all; seed it with the broadest pinned ARN the same way cfg.bucket
already seeds the classic modal.

* admin: refuse to save Not* fields the backend would silently drop

Hiding the NotResource/NotPrincipal modes was not enough where negation
is disallowed: the JSON tab accepts any valid document (that is its
job), and a statement's Advanced-fields box can reintroduce the keys, so
an s3tables save could still store fields the evaluator drops - turning
Allow+NotResource into allow-everything. commitPolicyActiveTab now runs
a final document-level check over what would actually be saved; Delete
stays available for cleanup.

* s3: move the IAM bucket policy mirror on a bucket rename

A same-directory rename delivers one event carrying both entries, and
the byte-equality short-circuit skipped the new name's mirror when the
policy was unchanged - while the replayed delete for the old name
removed its mirror, leaving the renamed bucket unmirrored. The mirror
decision is now a pure function that removes the old name and writes the
new one regardless of byte equality, with the rename cases unit tested.

* s3: backfill the IAM bucket policy mirror on lazy bucket loads

The metadata subscription only mirrors changes, so a policy that
predates the IAM integration never reached the bucket-policy:<bucket>
mirror and its grants did not bind on the IAM path until the policy was
next modified. The gateway is deliberately lazy at startup (nothing
lists all buckets), so the backfill hooks the same place a bucket's
policy first becomes known: the cold bucket-config load. EnsureBucketPolicy
writes only when no mirror is stored, so repeat loads cost one cached
read.

* s3: reconcile the bucket policy backfill against concurrent changes

The backfill's check-then-write could race an event-driven mirror update
or removal and re-store bytes that were already stale, with no later
event to heal it. EnsureBucketPolicy now reports whether it wrote, and a
write is reconciled against a fresh authoritative entry read: a changed
policy is re-mirrored, a removed one is removed. Anything changing after
that read fires its own event, which finds the backfill's write already
present and supersedes it. The backfill also carries the entry's raw
bytes rather than a re-marshaled document, so the reconcile can
byte-compare.

* s3: prime the bucket policy mirror before advanced-IAM authorization

The backfill ran from the lazy bucket-config load, but IAM authorization
evaluates the bucket-policy:<bucket> mirror before any handler runs - a
grant carried only by a not-yet-mirrored policy denied forever, and the
denied request never reached the code that would have loaded the bucket.
authorizeWithIAM now primes the bucket config first (an in-memory cache
hit once warm), and the backfill runs synchronously on the cold load so
the very first authorization already sees the mirror.
2026-08-24 00:52:01 -07:00
Chris LuandGitHub 68ec8ca655 admin: honor a persisted or admin.toml maintenance enabled=false (#10909)
* admin: honor a persisted or admin.toml maintenance enabled=false

The startup path discarded an operator's enabled=false twice over:
ApplyDefaultsToProtobuf treated the bool zero value as unset and applied
the schema default of true, and a force-enable migration block flipped
any survivor. With the legacy /maintenance UI routes gone, nothing could
write the config either, so the maintenance system ran unconditionally.

Keep the persisted enabled flag across schema-default application in
LoadMaintenanceConfig, drop the force-enable block, and add a top-level
[maintenance] enabled key to admin.toml as the config surface, persisted
through SaveMaintenanceConfig like the per-task settings. Absent config
still defaults to enabled.

* admin: track presence on the maintenance enabled flag

A plain proto3 bool cannot distinguish an operator's persisted false
from a legacy file that simply omits the field, so honoring false would
have silently switched maintenance off for configs written before the
toggle could be persisted. Make the field optional: files that predate
presence tracking keep the enabled default, while a file that explicitly
persists the toggle is honored either way.
2026-08-24 00:01:48 -07:00
Mathieu ArnoldandGitHub e931cccc7b Manage bucket policies via the admin ui (#10895)
* admin: manage S3 bucket policies from the admin UI

Bucket policies were only manageable through the S3 PutBucketPolicy API;
the admin UI had no equivalent to the quota/owner/lifecycle editors it
already offers. Add GET/PUT/DELETE for a bucket's policy, sharing the
exact validation the S3 gateway uses.

- Extract validateBucketPolicy/validateResourceForBucket out of
  s3api_bucket_policy_handlers.go into policy_engine.ValidateBucketPolicy /
  ResourceMatchesBucket so both the S3 API and the admin UI enforce
  identical rules.
- weed/admin/dash/bucket_policy.go: Get/Set/DeleteBucketPolicy, writing
  through ObjectTransaction + PATCH_EXTENDED (the lifecycle pattern) so a
  concurrent owner/quota/lifecycle change on the same bucket entry isn't
  clobbered. Propagation to every S3 gateway is automatic via the existing
  filer metadata log subscription. The S3 gateway's IAM policy mirror is
  deliberately not replicated here (its delete path is already an
  unimplemented TODO on the S3 side).
- New GET/PUT/DELETE /api/s3/buckets/{bucket}/policy routes, CSRF-guarded
  on writes.
- Bucket list and details modal now show a statement-count badge, read
  from the entry already fetched (no extra RPC).
- UI: a JSON-textarea policy editor modal, matching the lifecycle modal's
  structure.

* admin: reuse the visual policy editor for bucket policies

Extract the structured policy editor (add/remove statement, action/
resource/principal rows with autocomplete, JSON tab kept in sync) out of
policies.templ's inline script into a shared
weed/admin/static/js/policy_editor.js, and wire the bucket policy modal
in s3_buckets.templ up to it instead of a bare JSON textarea.

- registerPolicyEditor(which, config) replaces the hardcoded create/edit
  id derivation with a per-instance config (textarea/tab/body ids,
  datalist ids, requirePrincipal, bucket). The IAM policies page keeps its
  exact pre-extraction ids via two registerPolicyEditor calls, so its
  markup is unchanged.
- New policy_datalists.templ exposes the three shared <datalist>s
  (actions/resources/principals) as @PolicyDatalists(), now rendered by
  both policies.templ and s3_buckets.templ.
- requirePrincipal seeds new bucket-policy statements with Principal: "*"
  and adds a client-side check before save (the server, via
  policy_engine.ValidateBucketPolicy, remains the actual authority); the
  bucket config pins the Resource autocomplete to the open bucket instead
  of fetching every bucket in the cluster.
- layout.templ loads policy_editor.js globally, after admin.js/
  modal-alerts.js (basePath/escapeHtml/showAlert) which it depends on.

3a (the extraction) is a byte-preserving move verified against the
unchanged policies.templ behavior before layering 3b's parameterization
and the bucket-policy wiring on top.

* admin: migrate S3 Tables bucket/table policy editors to the shared editor

Third consumer of the shared visual policy editor: the S3 Tables bucket
and table policy modals (a bare JSON textarea each) now get the same
structured Editor/JSON tabs as the bucket policy and IAM policy pages,
via registerPolicyEditor('s3tablesBucketPolicy'/'s3tablesTablePolicy',
{ textareaId: ... }). Storage and validation are untouched - S3 Tables
policies still go through their own s3tables.PolicyDocument type and the
s3tables.policy extended attribute, unrelated to policy_engine and
s3-bucket-policy; only the editor UI is shared.

Fix a real bug surfaced by adding this second load path: the bucket
policy modal (and the naive first draft of this s3tables port) called
commitPolicyTextareaToEditor() right after a GET and then force-switched
to the Editor tab. commitPolicyTextareaToEditor() is designed to leave
the current tab in place and the editor state untouched when a document
fails to parse (so an in-progress edit survives a bad tab switch), so
forcing the Editor tab afterwards could show empty/stale editor state
that a careless Save would then serialize over a perfectly valid but
structurally-unusual stored policy. Add
loadPolicyTextareaIntoEditor(which) to policy_editor.js, which has no
"current tab" to defer to and instead falls back to the JSON tab with an
alert on a document the structured editor can't represent - the same
safety editPolicy already had in policies.templ - and use it at all three
"populate the editor right after a GET" call sites (bucket policy,
S3 Tables bucket policy, S3 Tables table policy).

* admin: show policy statement count on the S3 Tables buckets page

Mirrors the "Policy" column already added to the classic S3 buckets
list: a clickable badge with the statement count when the table bucket
has a resource policy, "Not configured" otherwise. S3 Tables policies
are a separate mechanism (s3tables.PolicyDocument under the
s3tables.policy extended attribute) from the S3 bucket policy work
elsewhere in this branch (policy_engine.PolicyDocument /
s3-bucket-policy), so this is a parallel implementation of the same
pattern rather than shared code.

- S3TablesBucketSummary gains PolicyStatementCount, populated in
  GetS3TablesBucketsData from entry.Entry.Extended[s3tables.ExtendedKeyPolicy]
  via the new extractS3TablesPolicyStatementCountFromEntry - no extra RPC,
  the entry is already fetched for ExtendedKeyMetadata.
- The badge reuses the existing .s3tables-bucket-policy-btn class, so it
  opens the same policy modal as the row's action button with no JS
  changes.

* admin: don't let a failed policy GET open the door to an empty overwrite

loadS3TablesBucketPolicy/loadS3TablesTablePolicy cleared the textarea,
then unconditionally called loadPolicyTextareaIntoEditor() regardless of
whether the GET actually succeeded - including when fetch() rejected or
the response was not ok, silently logged to console only. That leaves
the structured editor holding a legitimate-looking empty policy
({version, statements: []}), with the Editor tab active by default.

If Save is then clicked, commitPolicyActiveTab() serializes that empty
state into the textarea as `{"Version":"2012-10-17","Statement":[]}` -
a non-empty string - before the "Policy JSON is required" guard ever
sees it, so the guard passes and the transient load failure gets
written over whatever policy was actually stored.

Add s3tablesBucketPolicyLoaded/s3tablesTablePolicyLoaded, set true only
once a GET has actually completed (ok, including a genuinely empty
policy) and false on any failure path (fetch rejection or a non-ok
response, which previously fell through silently). Both submit handlers
now check the flag before touching the editor at all, and a failed load
surfaces via alert() instead of only a console.error - the user
previously had no visible indication the load had failed.

Verified with a jsdom simulation driving the real rendered page against
a stubbed fetch: a failed GET followed by Save now sends no PUT at all
(previously it sent Statement: []); a successful GET followed by Save
still PUTs the loaded policy unchanged.

* admin: address code review findings on the policy editor

1. policy_editor.js: policyEditors is only pre-populated for 'create'/
   'edit'; every other `which` (bucket, s3tablesBucket, s3tablesTable)
   stays undefined until its first successful async load. Nothing in
   this file enforces that a page hide its Editor/JSON tabs and
   Add-statement button until that load completes - the S3 Tables policy
   modals don't - so a click in that window (e.g. Add statement, or
   switching to the JSON tab) threw "Cannot read properties of undefined
   (reading 'unparsed')". Add policyEditorState(which), which lazily
   initializes a default state, and route addPolicyStatement, the
   jsonTabBtn 'show.bs.tab' handler, commitPolicyActiveTab, and
   renderPolicyEditor through it. Verified with a jsdom simulation
   against a never-resolving fetch: the exact click threw on the
   pre-fix code and no longer does.

2. s3_buckets.templ: the bucket-policy Save handler checked the
   textarea for emptiness before calling commitPolicyActiveTab(), which
   is what actually serializes the structured Editor tab's fields into
   that textarea. A policy entered entirely through the Editor tab (the
   primary path - never touching the JSON tab) left the textarea at
   whatever it was at load time, so creating a new policy this way hit
   "Enter a policy document" and Save silently did nothing. Move the
   commit before the emptiness check, preserving the existing alert and
   early-return. Verified with a jsdom simulation: Add-statement then
   Save (no tab switch) now PUTs the entered statement; before the fix
   the same sequence never reached fetch().

3. s3tables_buckets.templ / s3tables_tables.templ: the policy Editor/
   JSON nav-tabs were missing the ARIA roles Bootstrap's own tab pattern
   expects (role="tab"/"tabpanel", aria-selected, aria-controls,
   aria-labelledby) - screen readers had no way to tell these were tabs
   or which pane went with which button. Added the standard Bootstrap 5
   tab markup to both.

* admin: guard policy load/save flows against overlapping requests

1. s3tables.js: loadS3TablesBucketPolicy/loadS3TablesTablePolicy had no
   protection against overlapping loads. Opening one bucket's (or
   table's) policy dialog and then another's before the first GET
   resolved let the late response write its document into the shared
   textarea and mark the dialog "loaded" while it was now targeting the
   second resource - a subsequent Save would then push the first
   resource's policy onto the second. Add a per-load monotonic sequence
   number (s3tablesBucketPolicyRequestSeq / s3tablesTablePolicyRequestSeq,
   the same pattern already used for the classic bucket-policy load in
   s3_buckets.templ); a response is only applied - textarea, loaded flag,
   editor state - if its captured sequence still matches the latest one
   issued.

   Verified with a jsdom simulation: bucket A's policy load (artificially
   slow) followed immediately by bucket B's (fast) previously left A's
   policy in the textarea once A's late response landed; it now correctly
   keeps B's.

2. s3_buckets.templ: the bucket-policy Save button lives outside the
   (initially hidden) editor wrapper, so it stays clickable while a load
   is still in flight - the existing policyRequestSeq guard only protects
   the *load* from a stale response, not Save from firing before any
   load for the current bucket has completed. Add bucketPolicyLoaded,
   reset before each GET and set only once the matching response lands,
   and check it at the top of the Save handler.

   Verified with a jsdom simulation: clicking Save immediately after
   opening the dialog, before a (deliberately never-resolving) GET
   settles, now sends no PUT; a normal load-then-save sequence still
   PUTs the loaded policy unchanged.

* admin: address further code review findings on the policy editor

1. s3tables.js: loadS3TablesBucketPolicy/loadS3TablesTablePolicy only
   reset the JSON textarea when a new load starts; the structured editor
   kept showing the previously loaded resource's statements (Editor tab
   is the default active one) until the new fetch resolved. Call
   loadPolicyTextareaIntoEditor() against the now-cleared textarea
   immediately, so switching resources visibly resets the editor right
   away instead of only once its own load completes. Verified with jsdom:
   opening bucket A (loads fully) then bucket B (GET never resolves) no
   longer leaves A's statements visible in B's editor.

2. s3tables.js: deleteS3TablesBucketPolicy/deleteS3TablesTablePolicy had
   no loaded-state check, so a failed GET (which already blocks Save)
   left Delete fully able to remove the resource's stored policy sight
   unseen. Add the same s3tablesBucketPolicyLoaded/s3tablesTablePolicyLoaded
   guard Save already uses. Verified with jsdom: delete after a failed
   load now sends no DELETE; delete after a successful load is unaffected.

3. s3_buckets.templ: the bucket-policy Editor/JSON nav-tabs were missing
   the same ARIA roles already added to the S3 Tables policy tabs in an
   earlier round (role="tab"/"tabpanel", aria-selected, aria-controls,
   aria-labelledby) - this instance was out of scope for that review
   comment but is the same gap. Bootstrap's own tab.js already manages
   aria-selected on tab switch once the attribute exists, so no extra JS
   was needed.

4. s3_buckets.templ: neither the bucket-policy Save nor Delete handler
   guarded against a double-click, or against firing while the other was
   still in flight - two overlapping PUT/DELETE requests for the same
   bucket could land in either order. Add a shared
   bucketPolicyMutationInFlight flag: set (and both buttons disabled)
   before each fetch, cleared (and buttons re-enabled) on failure so the
   user can retry, left set through the existing success hide-and-reload
   path, and also reset when a new bucket's dialog opens so an abandoned
   in-flight request from a closed dialog can't leave the buttons stuck
   disabled. Verified with jsdom: double-clicking Save now sends exactly
   one PUT, and a Delete click while that PUT is still pending sends no
   DELETE.

* admin: scope bucket-policy mutation completions to the bucket that started them

1. The previous round's fix reset bucketPolicyMutationInFlight whenever a
   new bucket's policy dialog opened, to avoid leaving Save/Delete stuck
   disabled if the modal was closed mid-request. That traded one bug for
   a worse one: if bucket A's PUT/DELETE was still in flight when the
   user opened bucket B's dialog, the reset let B's Save/Delete fire
   immediately, and A's completion handler - unaware anything had
   changed - would still hide the (now B's) modal and reload the page
   out from under whatever the user was doing with B, on success, or
   alert a message with no bucket context, on failure.

   Stop resetting on reopen, so a pending mutation for a previous bucket
   keeps this bucket's Save/Delete blocked until it settles (matches the
   "preventing overlapping mutations" the review comment describes).
   Instead, capture policyEditorBucket as targetBucket right before each
   fetch and compare it against policyEditorBucket again in the
   completion handler: the in-flight flag is always released so the
   buttons never get stuck, but the modal-hide/reload/alert only fire if
   this bucket is still the one showing; a stale completion for an
   abandoned bucket just logs to the console instead.

   Verified with a jsdom simulation: opening bucket B while bucket A's
   Save is still pending leaves B's Save button disabled and a click on
   it a no-op; once A's PUT resolves, B's button re-enables but no
   modal.hide()/reload() fires (previously both fired unconditionally).

2. bucketPolicyDeleteBtn had no bucketPolicyLoaded check, unlike Save -
   a failed GET blocked Save but left Delete free to remove a policy the
   client never actually saw (the same gap already fixed for the S3
   Tables policy modals in an earlier round). Added the same guard,
   ahead of the confirm() dialog. Verified with jsdom: Delete after a
   failed load now sends no DELETE request.

* admin: fix spelling mistake
2026-08-23 22:11:18 -07:00
孙超andGitHub c80664ec21 s3: propagate storage rule fsync to volume server uploads (#10906)
The storage rule's fsync decision was computed by the filer
(detectStorageOption -> rule.Fsync) and applied on the filer's own HTTP
write path, but was never carried onto the chunk uploads S3 issues: the
AssignVolumeResponse had no fsync field, so the s3api client could not
learn the decision, and the chunked upload URL was hardcoded without it.
Every S3 write to a path with fsync configured went to the volume server
as a non-fsync write.

Carry the decision through the assign response:

- filer.proto: AssignVolumeResponse gains bool fsync, filled from the
  storage option the assign resolved.
- operation.AssignResult gains Fsync, so uploadChunk can append
  ?fsync=true to the volume server upload URL (single and replica
  fan-out paths).
- The S3 PUT/UploadPart assignFunc, the S3 copy path, the admin file
  browser upload, and the Iceberg worker assign functions all forward
  the response field.

Adds TestUploadReaderInChunksAppendsFsyncWhenAssigned.
2026-08-23 22:11:08 -07:00
Chris LuandGitHub 71a8c77a36 telemetry: let the dashboard pick the confirmation window (#10904)
* telemetry: let the dashboard pick the confirmation window

* telemetry: cover the serialized threshold map through the stats handler
2026-08-23 21:48:33 -07:00
Chris LuandGitHub 9c8d3b6a81 ec: refund the cleared leftover shards' slots in the encode source health check (#10903)
* erasure_coding: one home for the shard-count to volume-slots conversion

* ec: refund the cleared leftover shards' slots in the encode source health check
2026-08-23 21:48:20 -07:00
Chris LuandGitHub 36c97344ef s3: confine a Lance catalog table location to the caller's own bucket (#10901)
The Lance namespace gateway took the request-body location field, trimmed a
trailing slash, and passed it straight to the marker sink. That location feeds
TableDataDirFromMetadataLocation, which joins it under /buckets and collapses
any ../ segments, and writeMarker's CreateEntry then auto-creates every missing
parent. A caller could point the location at another tenant's bucket, or escape
/buckets entirely, and plant a fixed-name marker (recursively creating the
parents) or hide a victim's live table with .lance-deregistered.

Confine the declared location the way the Iceberg gateway already does: require
an s3:// URI whose bucket is the caller's own and whose path carries no
traversal segment, on both the declare and register handlers.
2026-08-23 11:49:52 -07:00
Chris LuandGitHub 74038e1b14 master: don't let a dead KeepConnected handler close its successor's channel (#10900)
A client that reconnects before the old handler exits re-registers the
same client name, and addClient overwrites the map entry. The old
handler's deferred deleteClient then closed whatever channel the map
held under that name: the new, live stream's. Receiving from a closed
channel returns nil immediately and forever, so the new handler's send
loop degenerated into sending empty responses at wire speed, pinning a
core on each side until the client killed the connection.

deleteClient now closes the channel its own handler registered and
leaves the map entry alone unless it still points to that channel. This
also closes the previously orphaned old channel, whose drain goroutine
used to leak. The send loop treats a closed channel as an exit instead
of a message stream.
2026-08-23 11:36:00 -07:00
Chris LuandGitHub cf0dba334c s3api: no filer failover after the callback has consumed part of a response (#10902)
s3api: no filer failover after fn has consumed part of a response

withFilerClientFailover replays fn verbatim on the next filer, so a filer
that died mid-stream followed by a healthy peer returned success with the
callback's closure-captured accumulator holding the dead filer's prefix
twice; the per-attempt accumulator in listWithRetry could not close this,
because the replay happens inside a single attempt. Track delivery on the
connection handed to fn: once a unary reply or streamed message has reached
the callback, surface the transport error unwrapped instead of failing
over, and let callers replay from a clean slate. A filer that fails before
delivering anything fails over exactly as before.
2026-08-23 11:30:43 -07:00
Chris LuandGitHub c167af541e telemetry: confirm a cluster after a week of reports, not two days (#10899)
* telemetry: sync the server module to go 1.26

The root module moved to go 1.26 but the telemetry server module, which
replaces seaweedfs with the repo root, stayed on 1.25.8, so go refuses
to build or test it until the directive catches up.

* telemetry: confirm a cluster after a week of reports, not two days

Two days of history still lets recurring CI and demo clusters into the
confirmed fleet: anything torn down and rebuilt across a UTC midnight
counts. Requiring seven distinct UTC days keeps the fleet charts and the
version/OS distributions to clusters that actually stay up; real
clusters qualify after their first week, and the fallback to all active
clusters while none is confirmed is unchanged.
2026-08-23 11:14:17 -07:00
Chris LuandGitHub 0f85d005ad server: 416 only when no requested range overlaps, with Content-Range, and the Rust mirror (#10889)
* filer, volume server: return 416 when no requested range overlaps the content

* seaweed-volume: return 416 when no requested range overlaps the content

* server: check the range test error, use the request context, fix the no-overlap comment boundary
2026-08-23 11:13:36 -07:00
Chris LuandGitHub 173adbc291 master: never re-seed a raft cluster over committed state under -raftBootstrap (#10883)
* master: never re-seed a raft cluster over committed state

-raftBootstrap deleted logs.dat, stable.dat and snapshots on every start and
then bootstrapped a fresh cluster. Since hashicorp raft only snapshots after
8192 log entries, the TopologyId lives in the log, not in a snapshot, so the
pre-wipe snapshot recovery found nothing and each restart minted a new cluster
identity. A master that came up while it could not reach its peers seeded a
rival cluster; when the two logs met, SetTopologyId's split-brain guard fatally
stopped every master holding the other id, and the master layer crash-looped
with no quorum.

Bootstrapping is genesis. Drop the wipe and the inline bootstrap. The first
master in -peers already mints a cluster once it has confirmed no peer has a
leader, so the flag has nothing left to do and is now ignored; keeping that one
master the sole bootstrap authority is what stops a partition from minting two
clusters, so the flag must not widen it either. A master with state rejoins its
peers, and one whose data dir was reset is admitted by the sitting leader
instead of forking again.

* test: cover -raftBootstrap restarts in the multi-master suite

Three masters start with -raftBootstrap, the way the helm chart renders it on
every master on every roll, and the cluster has to hold one TopologyId after
they all restart. /dir/status is proxied to the leader, so each master's own
view of the identity is read out of its log, which is where a fork shows up.
Before the fix the hashicorp case minted a new id on each restart.
2026-08-23 11:10:20 -07:00
Junker der ProvinzandGitHub fa3bd5b5a7 mount: use the kernel-resolved node id in Link, not the persisted attribute (#10885)
* fix(mount): reply to LINK with the kernel node id, not the stored inode

Link() answered the kernel with out.NodeId = oldEntry.Attributes.Inode.
That attribute is a mount-runtime number and only entries created through a
mount carry one. An entry written by the S3 API, WebDAV or a direct filer
call persists inode 0, so the LINK reply named node id 0, which the kernel
rejects as invalid_nodeid and reports as EIO. The hard link itself had
already been written to the filer, which is why it looked correct again
after a mount restart.

The same stale number was also used as an inodeToPath key. AddPath(0, path)
filed the new link under inode 0, so a later Lookup on that name handed the
kernel node id 0 as well, and a LOOKUP reply carrying node id 0 means no
such entry.

in.Oldnodeid is the node id the kernel already holds for the source, and it
is the key inodeToPath is indexed by, so use it for the reply, for AddPath
and for the sibling sync.

Fixes #8404

* test(mount): cover the sibling sync in Link with a third hard link

The two existing cases never reach the body of syncHardLinkSiblings: with
two links the source alias and the name just created are both in skipPaths,
so the loop iterates over nothing and a change to that site goes unnoticed.
A third link leaves one name that no other part of Link() writes.

The new case drives three links off one source. It guards against covering
nothing (it fails if every path turns out to be a skipPath), checks that
every name of the file reports nlink 3, and then drives the sync with both
candidate keys to pin down which one it has to be: keyed by the source's
persisted Attributes.Inode, which is 0 for an entry written outside a mount,
GetAllPaths has no path to walk, while the kernel node id reaches the
sibling.

That second half is driven directly because Link() alone cannot tell the two
keys apart. The meta cache keeps one blob per hard link id (FilerStoreWrapper
setHardLink/maybeReadHardLink), so a read of any sibling returns the
attributes of the last write to any of them whether or not the sync ran.
2026-08-23 10:43:22 -07:00
Junker der ProvinzandGitHub 5ebc9c9f4b server: reject a Range start offset equal to the file size (#10898) 2026-08-23 08:20:25 -07:00
Chris LuandGitHub 3b10e43d5d test: wait for volume server registration in the FUSE p2p harness (#10897) 2026-08-23 02:14:50 -07:00
Chris LuandGitHub 9d06f2c378 test: keep per-test log directories in the FUSE DLM harness (#10893) 2026-08-23 01:19:22 -07:00
Chris LuandGitHub 9d4270f118 test: wait for volume server registration in the FUSE DLM harness (#10891) 2026-08-23 00:59:47 -07:00
Chris Lu 3b8931c2f6 admin: address review feedback on the maintenance scanner fix 2026-08-23 00:26:38 -07:00
Chris Lu 8d8a25b1cf s3api: remove the duplicated listing retry helpers left by overlapping merges 2026-08-23 00:26:02 -07:00
c58795354a s3api: retry a transient filer failure on metadata listings (#10890)
* s3api: retry a transient failure when listing multipart uploads/parts

A blip on the way to the filer failed the whole ListMultipartUploads or
ListParts request. Both reported failure points sit inside one streaming
listing: the ListEntries call that opens the stream, and the stream.Recv
calls that drain it. Neither retried, so a single Unavailable answer from
a filer that was restarting turned into a 500 for the S3 client.

Replay the listing instead, bounded to three attempts with a 100ms
backoff that doubles. Only a transient failure is replayed. A not-found
answer stays authoritative so the empty-list branch still works, and
every other error still reaches the client on the first attempt.

This is scoped to (*S3ApiServer).list rather than added inside
DoSeaweedListWithSnapshot, which mount, the shell and the other object
listings share, and where a retry after a partial stream would
re-deliver entries the callback had already seen. Within one call to
list, a replay is safe: it collects into a fresh slice each time, so it
can neither duplicate nor drop entries.

That guarantee does not extend past this function. withFilerClientFailover
already re-runs its callback against the next filer on any non-NotFound
error without resetting the caller's accumulator, so on a multi-filer
gateway a mid-listing failover can itself produce a duplicated result
with err == nil, independent of this change and not fixed by it. Noted
in the PR rather than silently left for someone to rediscover.

Fixes #7221
References #7235

* s3api: move the listing retry inside list itself

---------

Co-authored-by: Junker der Provinz <jdp@braethoria.com>
2026-08-22 23:42:33 -07:00
Junker der ProvinzandGitHub 8d2c0273bd admin: stop the maintenance scanner pinning itself to one scan per second after a transient failure (#10887)
* admin: honour persisted task configs when building the maintenance policy

buildPolicyFromTaskConfigs passed a literal nil to vacuum, erasure_coding
and balance LoadConfigFromPersistence. Those functions look for their
LoadXTaskPolicy() accessor via a type assertion, which a nil interface can
never satisfy, so every call fell through to NewDefaultConfig() and the
policy came back with the compiled-in defaults - Enabled: true among them.
A task disabled on disk was therefore still scheduled, and the only trace
was a glog.V(1) "Using default ... configuration" line.

Thread the real ConfigPersistence through instead. There are two copies of
this function: the one in weed/admin/dash builds config.Policy on the
normal admin startup path and can simply take cp as its receiver, and the
one in weed/admin/maintenance is the fallback used when the config carries
no policy yet, which now receives the store from NewMaintenanceManager.
weed/admin/dash already imports weed/admin/maintenance, so the maintenance
side has to keep the duck-typed interface{} parameter that the task
loaders already use rather than importing the concrete type back.

The store is only handed over when a data directory is configured: an
unconfigured one has nothing to read, and a typed nil pointer would pass
the loaders' type assertion and then panic on first use.

Fixes #10874

* admin: restore the maintenance scan cadence after an error backoff

scanLoop shortens its ticker to the error backoff delay after a failed
scan, but it decided whether to replace the ticker by comparing the
target interval against the configured scan interval instead of against
the interval the ticker was actually running at. Once the errors stopped,
getScanInterval returned the configured interval again, the comparison
came out false, and the ticker was left at the backoff delay - so a
single transient scan failure pinned the scanner to one scan per second
for the rest of the process lifetime. That is the ~1/second cadence in
issue #10874: 658 KB/s of "Cancelled N stale pending balance tasks
before re-detection" and 193k orphaned task files over two days.

Track the interval the ticker is running at and compare against that, so
both entering the backoff and returning to the normal cadence replace the
ticker.

While in here:

- defer ticker.Stop() bound the ticker that was current when the defer
  was registered, so every replacement ticker leaked on return. Wrap it
  in a closure.
- running was written by Start/Stop and read by all three background
  loops without synchronisation. Guard it with the existing mutex, fold
  the running check in triggerScanInternal into the lock it already
  takes, and make Stop a no-op when not running so a second call cannot
  close the stop channel twice.

Refs #10874

* admin: make the maintenance policy actually reach the task detectors

Loading the persisted task configs into the maintenance policy only
matters if something reads that policy, and nothing did.

MaintenanceIntegration pushes the policy into every registered detector
and scheduler through interface{ SetEnabled(bool) } and
interface{ SetMaxConcurrent(int) } type assertions. Every task registered
through base.RegisterTask is backed by base.GenericDetector and
base.GenericScheduler, and neither implemented either method, so all four
assertions failed silently for every task on every startup. The policy's
enabled flag reached nothing: ScanWithTaskDetectors gates on
detector.IsEnabled(), and the queue's policy lookups for max concurrent
and repeat interval are fallbacks that only fire when the scheduler
reports zero, which the generic scheduler never does.

Add the setters, delegating to the TaskConfig.SetEnabled the interface
already declares and to TaskDefinition.MaxConcurrent, which is what
GetMaxConcurrent returns.

Applying the policy required three more fixes, because with the
assertions working the policy could now do damage as well as good:

- IsTaskEnabled reports false for a task type the policy has no entry
  for, so applying it unconditionally would have disabled every task the
  policy does not list. Skip task types with no policy entry: no entry
  means no opinion, not disabled.

- ec_balance was exactly such a task. It is registered like the other
  three but had no entry in the policy builder and no accessor on
  ConfigPersistence at all, so its configuration could never be
  persisted. Add SaveEcBalanceTaskPolicy/LoadEcBalanceTaskPolicy, the
  task_ec_balance.pb file, the SaveTaskPolicy dispatcher case, and the
  policy entry.

- InitMaintenanceManager ran before loadTaskConfigurationsFromPersistence,
  which replaces each task's whole config object, so the policy was
  applied and then immediately thrown away. Swap the order. Both read the
  same files, so the policy is now the last writer and stays
  authoritative.

MaintenanceManager.UpdateConfig also updated the queue's and the
scanner's policy but not the integration's, so a policy changed at
runtime never reached the detectors. Add MaintenanceIntegration.SetPolicy
and call it.

While building the policy, stop hand-copying each task's fields and use
the task's own ToTaskPolicy(). The hand-written version was a second
definition of every task's policy and had already lost the erasure coding
preferred tags and replica placement and the balance IO rate limit. For
the same reason, the "nothing persisted yet" branches of
LoadVacuumTaskPolicy, LoadErasureCodingTaskPolicy and
LoadBalanceTaskPolicy now derive from each task's NewDefaultConfig()
instead of a third hand-written copy. Those copies had drifted, so with a
data directory but no config file on disk the effective defaults differed
from what the task and the admin UI schema both advertise:

  vacuum          scan interval  24h  -> 2h
  balance         scan interval   6h  -> 30m
  balance         imbalance      0.1  -> 0.2
  erasure coding  scan interval 168h  -> 1h
  erasure coding  fullness      0.90  -> 0.95
  erasure coding  min volume   1024MB -> 30MB

Finally, weed/admin/dash and weed/admin/maintenance each carried a copy
of the policy builder and they had already diverged. Export the
maintenance one as BuildPolicyFromTaskConfigs and have dash call it.

Refs #10874

* worker: warn when a config store cannot supply a task's persisted config

LoadConfigFromPersistence logged a single glog.V(1) "Using default X
configuration" for every way of not loading anything, so the bug in
issue #10874 - a store handed in that the type assertion rejects, leaving
a task running on compiled-in defaults - looked exactly like the normal
"no data directory configured" case. The reporter had to read the source
to work out why their disabled task kept running, and asked for this
specifically.

Separate the cases. A non-nil store that does not provide the accessor is
always a wiring bug and is now logged at warning level, naming the type
and the missing method. A read error or a policy that will not apply is
also a warning. No persistence configured, and a store with nothing saved
yet, stay at V(1): those are normal.

Refs #10874

* admin: stop GetTaskPolicy panicking on a maintenance policy that is nil

GetTaskPolicy dereferenced its MaintenancePolicy argument to look at
TaskPolicies, so IsTaskEnabled, GetMaxConcurrent and GetRepeatInterval
all took the admin process down when handed a nil policy. A nil policy is
not a programming error here: MaintenanceConfig.Policy is unset until
something builds one, DefaultMaintenanceConfig returns a config with no
policy at all, and UpdateConfig installs whatever config it is given.
Found by calling IsTaskEnabled with the policy from a freshly defaulted
MaintenanceConfig.

Treat a nil policy as "no entry": no task enabled, the safe concurrency
default of 1, and a repeat interval of 0 so callers fall back to their
own default instead of reading DefaultRepeatIntervalSeconds off nil.

Also add the startup test this was found with. It walks the admin
server's startup sequence over a data directory that has balance saved as
disabled and checks the state that decides whether issue #10874 happens:
the balance detector reports disabled, vacuum stays enabled, and tasks
whose config was never saved keep their compiled-in default.

Refs #10874

* admin: document the synchronisation SetPolicy would need beyond startup

ConfigureTasksFromPolicy now really writes TaskDefinition.Config and
TaskDefinition.MaxConcurrent, which the scan loop reads through
detector.IsEnabled() with nothing synchronising the two. Every caller
runs during admin server startup today, before the scan loop exists, so
there is no live race - but the next caller has to add the locking, and
the same already applies to UpdateAllConfigs replacing the whole config
object. Write it down at the seam instead of leaving it to be
rediscovered.

Refs #10874
2026-08-22 23:41:45 -07:00
Junker der ProvinzandGitHub f710b6003a s3api: retry a transient failure when listing multipart uploads/parts (RFC on layering) (#10886)
s3api: retry a transient failure when listing multipart uploads/parts

A blip on the way to the filer failed the whole ListMultipartUploads or
ListParts request. Both reported failure points sit inside one streaming
listing: the ListEntries call that opens the stream, and the stream.Recv
calls that drain it. Neither retried, so a single Unavailable answer from
a filer that was restarting turned into a 500 for the S3 client.

Replay the listing instead, bounded to three attempts with a 100ms
backoff that doubles. Only a transient failure is replayed. A not-found
answer stays authoritative so the empty-list branch still works, and
every other error still reaches the client on the first attempt.

This is scoped to (*S3ApiServer).list rather than added inside
DoSeaweedListWithSnapshot, which mount, the shell and the other object
listings share, and where a retry after a partial stream would
re-deliver entries the callback had already seen. Within one call to
list, a replay is safe: it collects into a fresh slice each time, so it
can neither duplicate nor drop entries.

That guarantee does not extend past this function. withFilerClientFailover
already re-runs its callback against the next filer on any non-NotFound
error without resetting the caller's accumulator, so on a multi-filer
gateway a mid-listing failover can itself produce a duplicated result
with err == nil, independent of this change and not fixed by it. Noted
in the PR rather than silently left for someone to rediscover.

Fixes #7221
References #7235
2026-08-22 23:00:50 -07:00
Junker der ProvinzandGitHub f3caf6e7da admin: count plugin-runtime workers in worker metrics (#10884)
* admin: count plugin-runtime workers in worker metrics

The admin server keeps two worker registries: the legacy maintenance-worker
map, filled by workers registering over the worker gRPC stream, and the plugin
worker registry, filled by workers started as `weed worker`. Both the
SeaweedFS_admin_workers_connected / SeaweedFS_admin_worker_slots gauges and the
dashboard's Workers card read only the legacy map, so a cluster that runs the
admin and its workers as separate components reported 0 workers even while its
workers showed up on the plugin pages and ran scheduled jobs.

Aggregate both registries instead. The two are merged by worker ID: `weed mini`
starts both runtimes out of one working directory, so they share the persisted
worker ID and must not be counted twice. For such a worker the slot numbers
still come from the legacy registry, which keeps mini's existing readings.
Plugin workers report their slots in the heartbeat, so detection and execution
slots are summed from there; a worker that has connected but not yet sent a
heartbeat counts as connected with zero slots.

Fixes #10525

* admin: clamp negative worker-reported slot values in metrics merge

A plugin worker's self-reported heartbeat slot counts are untrusted
input; clamp them to 0 before summing so a stale or misbehaving
worker can't drive the aggregate gauge negative, matching the same
defensiveness already used in registry.go's own slot arithmetic.
2026-08-22 22:54:39 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
004fc32503 build(deps): bump github.com/moby/go-archive from 0.2.0 to 0.3.0 (#10826)
Bumps [github.com/moby/go-archive](https://github.com/moby/go-archive) from 0.2.0 to 0.3.0.
- [Release notes](https://github.com/moby/go-archive/releases)
- [Changelog](https://github.com/moby/go-archive/blob/main/changes_test.go)
- [Commits](https://github.com/moby/go-archive/compare/v0.2.0...v0.3.0)

---
updated-dependencies:
- dependency-name: github.com/moby/go-archive
  dependency-version: 0.3.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-08-22 13:01:07 -07:00
641fc8b031 admin: add visual iam policy editor (#10878)
* admin: add visual iam policy editor

Add a structured, tabbed editor (Editor / JSON) for creating and editing
IAM policies in the admin dashboard, alongside the existing raw-JSON
textarea:

- policies.templ: per-statement cards for Sid, Effect, Action, and
  Resource, with unmanaged fields (Principal, NotPrincipal, NotResource,
  Condition, or anything else) preserved verbatim in a per-statement
  "advanced fields" JSON box so nothing is lost on round-trip. Switching
  tabs commits and reparses in both directions. Restored the "Use Sample
  Policy" button, now filling both the structured editor and the JSON
  tab. The "Validate" button now calls the existing but previously
  unused POST /api/object-store/policies/validate endpoint instead of
  doing JS-only checks.
- Progressive Resource ARN autocomplete: suggests bucket names first,
  then once "bucket/" is typed, suggests bucket/* plus the bucket's
  direct subfolders, drilling down one path segment at a time as the
  user types further "/" characters.
- New GET /api/files/list-folders endpoint (file_browser_handlers.go)
  backing the folder autocomplete: wraps the existing file browser data
  function and returns just the subdirectory names as JSON, scoped to
  paths under /buckets.
- Action-name suggestions (datalist) for the Action field, sourced from
  the existing s3_constants.S3_ACTION_* constants plus new
  s3_constants.S3TABLES_ACTION_* constants (extracted from the s3tables
  operation dispatch switch) so the suggestion list can't drift from the
  strings the engines actually understand.
- policy_handlers.go: ValidatePolicy now accepts a statement with only
  NotResource set (previously required Resource), matching
  policy_engine.validateStatement and the fact the new editor makes such
  statements reachable from the UI.
- Tests: ValidatePolicy behavior, route registration for the policy API
  and the new list-folders endpoint, list-folders path scoping, and the
  action-suggestion list's shape.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: fix XSS, cache poisoning, and cap overshoot in policy editor

Address code review findings on the IAM policy editor added in the
previous commit:

- policies.templ (displayPolicyDetails): escape every interpolated
  policy value (Sid, Effect, Action, Resource, policy name, and the raw
  JSON document) before assigning to innerHTML. Policy documents can
  come from other admins or an import, so an unescaped field could
  execute script when the "View" modal renders it.
- policies.templ (policyEditorStateToDoc): reject JSON arrays in a
  statement's "advanced fields" box, not just invalid JSON. `typeof []
  === 'object'` was true, so a JSON array was assigned to the statement;
  subsequent property assignments (Sid, Effect, ...) landed on the array
  object but JSON.stringify of an array only serializes numeric indices,
  silently dropping them.
- policies.templ (loadPolicyFolderNames): on a failed folder lookup,
  remove the cache entry instead of permanently caching the empty
  fallback, so a transient network/server error doesn't block retries
  for the rest of the page's lifetime.
- file_browser_handlers.go (ListFolders): stop appending directory
  names as soon as the running count reaches maxListFoldersEntries,
  instead of only checking the cap after a full page is processed,
  so the returned list never exceeds the configured cap.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001
to keep the diff scoped to this file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: stop policy editor from clobbering the active tab and dropping malformed advanced fields

Address two review findings on the IAM policy editor (Issue 3, stored-XSS
in displayPolicyDetails, was already fixed by the previous commit and is
unchanged here):

- createPolicy, updatePolicy, and validatePolicyDocument always committed
  the structured editor's (possibly stale) state into the JSON textarea
  before submitting, even when the user had just edited the JSON tab
  directly. That silently discarded the user's JSON edits and
  validated/saved the old structured-editor state instead, which could
  leave broader permissions in force than intended.

  Added commitPolicyActiveTab(which), which commits whichever tab is
  currently visible into the other side instead of unconditionally
  overwriting the JSON tab from the editor: if the JSON tab is active it
  parses that JSON back into the structured editor (without touching the
  textarea itself), otherwise it serializes the structured editor into
  the textarea as before. All three call sites, plus the JSON-tab
  "show.bs.tab" handler, now use this and abort with an alert if the
  currently active tab's content can't be committed.

- policyEditorStateToDoc silently continued with an empty object when a
  statement's "advanced fields" box held invalid JSON, so switching
  tabs, validating, or saving would drop Principal/NotResource/Condition
  from that statement without telling the user. It now throws (with the
  statement number and parse error) on invalid or non-object JSON there,
  and callers surface that via showAlert and abort instead of proceeding.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: keep unmanaged top-level policy fields across editor tab switches

policyDocToEditorState only carried Version and Statement into editor
state, so any other top-level key (e.g. Id) present in the JSON tab was
silently rewritten away as soon as the user switched to the Editor tab
and back. Capture those keys in state.otherFields and merge them back in
policyEditorStateToDoc before Version and Statement are written, so the
two tabs stay faithful to each other and the editor never rewrites text
the user typed.

Note this is editor fidelity only: the admin API's
policy_engine.PolicyDocument carries just Version and Statement, and
DocumentJSON is never populated, so such fields are still discarded by
the server once a policy is saved. Making them survive a save would
require a backend change, which is out of scope here.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: warn before a policy save discards unsupported top-level fields

The editor round-trips unmanaged top-level keys (e.g. Id) between the
Editor and JSON tabs, but the admin API's policy_engine.PolicyDocument
carries only Version and Statement, so the server drops them on save and
the user saw no indication.

Added confirmPolicyFieldDiscard(), called from createPolicy and
updatePolicy after the active tab is committed (so the field list is
accurate whichever tab is showing). It names the fields that will be
lost and lets the user confirm or cancel. Not wired into
validatePolicyDocument, which doesn't persist anything.

Chose the warning over the alternative of persisting these fields
through the backend: policy_engine.PolicyDocument is shared by the S3
bucket-policy engine and IAM evaluation, so extending it would change
the stored document shape for every policy in the codebase - far beyond
the scope of this editor.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: reject malformed Effect and Resource/NotResource conflicts in policy editor

Two review findings on the IAM policy editor:

- policyDocToEditorState defaulted any non-"Deny" Effect (missing,
  misspelled, wrong case) to "Allow". A statement meant to be "Deny" with
  a typo like "deny" would silently become a permissive "Allow" instead
  of being rejected. It now throws on anything but an exact "Allow" or
  "Deny", naming the offending statement and value.
  commitPolicyTextareaToEditor catches this the same way it already
  catches invalid JSON: alert the user and keep the JSON tab active
  instead of switching to the Editor tab with wrong data.

- policyEditorStateToDoc could save a statement with both Resource (from
  the structured field) and NotResource (surviving in the "advanced
  fields" extras from before the user switched to using Resource) set at
  once - a contradictory combination neither the admin's ValidatePolicy
  handler nor policy_engine's evaluator rejected. When the structured
  Resource field is non-empty it now deletes any leftover NotResource
  from extras, consistent with the file's existing rule that structured
  fields take precedence over extras. Mirrored the existing
  Principal/NotPrincipal exclusivity check in
  weed/admin/handlers/policy_handlers.go's ValidatePolicy to reject the
  same combination server-side, since create/update perform no
  validation at all. Deliberately left policy_engine.validateStatement
  (used by the S3 bucket-policy PUT handler for every bucket policy in
  the product) unchanged - extending that shared validator is a larger,
  separate change outside this admin-editor fix's scope.

Added a handler test for the new Resource+NotResource rejection.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* admin: add NotResource support to the visual policy editor

Since Resource and NotResource are mutually exclusive (enforced by a
previous fix), NotResource could previously only be set through the raw
JSON in a statement's "advanced fields" box. Promote it to a first-class
mode of the structured editor:

- The static "Resources" label is now a Resource/NotResource dropdown;
  the same list of values underneath is reused for either key depending
  on the selected mode, with a short form-text explaining the semantics.
- NotResource is added to POLICY_STATEMENT_KNOWN_KEYS, since it's now a
  managed field like Resource rather than something that falls through
  to extras.
- policyDocToEditorState derives resourceMode from which key is present
  on load, and throws (same handling as the existing malformed-Effect
  case: alert, keep the JSON tab active) if a hand-edited document has
  both Resource and NotResource on one statement, since that can't be
  represented by the dropdown.
- policyEditorStateToDoc writes only the key matching the selected mode,
  replacing the previous one-directional "delete NotResource whenever
  Resource is set" fix with mode-driven logic that also deletes Resource
  when NotResource is selected.
- displayPolicyDetails (the read-only View modal) now shows the actual
  NotResource values with a distinct label instead of a static
  "(NotResource used instead)" placeholder.

No backend changes: the server-side "cannot specify both" check added
previously in policy_handlers.go's ValidatePolicy already covers this.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: reject non-object policy documents; catch Principal/NotPrincipal conflicts server-side

Two review findings:

- policyDocToEditorState treated a top-level JSON value that wasn't an
  object (null, or a bare string/number/boolean) as an empty statement
  list instead of failing explicitly. If the user typed e.g. "hello" or
  42 in the JSON tab and switched to the Editor tab, their input was
  silently discarded and replaced with an empty policy - the same class
  of "guess instead of reject" bug fixed for malformed Effect and
  Resource/NotResource conflicts previously. Added an explicit check
  that throws for null/scalar input, while leaving array and object
  document shapes accepted exactly as before.

- weed/admin/handlers/policy_handlers.go's ValidatePolicy checked the
  Resource/NotResource conflict by non-empty length
  (len(...Strings()) > 0), which misses a statement where Resource is
  explicitly present but an empty list (e.g. "Resource": []) alongside a
  non-empty NotResource. Switched that check to field presence (!= nil),
  matching how policy_engine's own validateStatement already treats
  Principal/NotPrincipal exclusivity. Also added the equivalent
  Principal/NotPrincipal presence check to this handler, which had none
  before - the advanced-fields box in the visual editor lets a user set
  both today, and nothing server-side caught it. The existing
  non-empty "Resource or NotResource is required" check is left as a
  length check, since an empty array shouldn't count as "provided".

Added test cases for both conflict checks in policy_handlers_test.go.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: add Principal/NotPrincipal support to the visual policy editor (v1, AWS-only)

Adds a first, deliberately narrow structured editor for a statement's
Principal/NotPrincipal, left out when NotResource support was added:

- A Principal/NotPrincipal mode dropdown mirrors the existing
  Resource/NotResource one (same mutual-exclusivity handling: the two
  fields can't be set at once, and switching modes reuses the same
  value list).
- A simple repeatable text-value list feeds a single {"AWS": [...]}
  object on save - always the AWS type, never the "bare" (untyped)
  SeaweedFS-extension shape. Per policy_engine's allowedPrincipalKeys,
  Service/Federated/CanonicalUser also parse successfully, but nothing
  in the S3 bucket-policy evaluation path ever sets a real caller's
  principal to a service name, an OIDC provider ARN, or a canonical
  user ID, so only AWS is functionally meaningful today - out of scope
  for this v1.
- On load, only the exact {"AWS": ...} single-key shape is unwrapped
  into the structured field and removed from "extras". Anything else
  (bare string/array, a different single type key, or several type
  keys at once) is left untouched in "extras" exactly as before, with a
  visible warning under the dropdown so the user knows a
  Principal/NotPrincipal exists but isn't shown there. Saving with the
  structured field left empty never touches whatever's already in
  extras, so a preserved complex form isn't silently dropped just
  because the user didn't touch this field.
- The read-only View modal now displays Principal/NotPrincipal for any
  shape (via a small generic summarizer), not just the AWS-simple one.
- Generalized the action/resource field-to-state-key mapping (used by
  commitPolicyEditorForm and the add/remove-item click handler) into a
  shared lookup table instead of stacking another ternary, now that a
  third field (principal) exists.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: support the bare "*" wildcard Principal in the visual editor

"Principal": "*" (and NotPrincipal: "*") is the standard AWS shorthand
for "everyone" and is common in real bucket policies, but the v1
Principal/NotPrincipal editor only recognized the {"AWS": ...} object
form, leaving a bare "*" statement's principal hidden in Advanced
fields.

parseSimpleAwsPrincipal now also accepts the bare string "*" as a
simple, structurally-editable value. On save, a principal value list
containing exactly ["*"] is written back as the bare "*" string
(matching the common convention) rather than wrapped as {"AWS": "*"};
anything else still wraps under AWS as before. Updated the field's
form-text hint accordingly.

Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: add Principal field autocomplete backed by users + IAM roles

Adds a datalist-backed autocomplete for the policy editor's Principal/
NotPrincipal text fields, sourced from a new API listing existing
identities:

- weed/admin/dash/principal_suggestions.go: AdminServer.GetPrincipalSuggestions
  combines S3 user ARNs (via the existing GetObjectStoreUsers +
  iam.UserArn) with IAM role ARNs (via integration.NewFilerRoleStore /
  ListRoles, reusing the exact same construction already used in
  iam_manager.go - no new dependency risk introduced). Role ARNs are
  reconstructed from the role name using SeaweedFS's default
  arn:aws:iam::role/<name> convention rather than fetching each role's
  stored definition, since this only backs a suggestion list. Role
  listing failures are logged and swallowed rather than failing the
  whole request - an incomplete suggestion list is fine, blocking
  policy editing over it is not. Service accounts are deliberately not
  listed separately: a service account's ARN is identical to its parent
  user's, already covered by the user list.
- weed/admin/handlers/policy_handlers.go: GetPrincipalSuggestions handler
  exposing this as {"principals": [...]}.
- Route registered at the API root (GET /api/principals) rather than
  under policyApi's "/object-store/policies" prefix, since that
  subrouter's existing "/{name}" GET route would shadow any
  single-segment GET route registered after it (the same class of
  gotcha previously seen with "/validate").
- weed/admin/view/app/policies.templ: a shared, lazily-fetched-once
  policyPrincipalSuggestions datalist (flat list - unlike the
  progressive per-folder Resource ARN autocomplete, users/roles aren't
  hierarchical), wired into policyListRowHtml for field:"principal" and
  populated on input/focus, with "*" always offered first.

Added tests for the new ARN-construction helper and route registration.
Regenerated policies_templ.go with the already-stamped templ v0.3.1001.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: fix fieldset/legend styling in the structured policy editor

Bootstrap's form reset stretches <legend> to the fieldset's full width
(float: left; width: 100%), which loses the native "notch in the
border" look and makes each section's label bar as wide as the card.

Add two scoped classes: .policy-stmt-fieldset (border, rounded
corners, spacing between sections) and .policy-stmt-legend (undoes the
float/width so the legend hugs its content, with a little padding).
Applied to the three per-statement sections (Actions,
Resource/NotResource, Principal/NotPrincipal), replacing the ad hoc
"border rounded" utility classes that were doubling up with the
fieldset's own border. Also gave the "Advanced fields" <details> a
small top margin to match the new spacing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: suggest bucket/* alongside the bucket itself in Resource autocomplete

At the bucket-name stage of the Resource field's progressive
autocomplete, only "arn:aws:s3:::bucket" was offered. Add
"arn:aws:s3:::bucket/*" right alongside it, since granting access to
everything in a bucket is the more common case and previously required
typing a "/" first to reach the folder-level "*" suggestion.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eb2a51LciCsyY35sqDoGNe

* admin: keep an unparseable policy in the JSON tab instead of wiping it

editPolicy() built the structured state inside the fetch .then, so a
policy the editor cannot model threw into the sibling .catch, which
alerted and called hide(). Showing the alert at that moment left the
modal on screen with an empty editor and the document only in the JSON
tab, and Save Changes then serialized the empty state over the policy.

Reachable two ways, since neither create path rejects these: the admin
API never validates on create, so "Effect":"allow" is stored as typed,
and policy_engine.validateStatement lets Resource and NotResource sit
in the same statement.

Hand the document to the JSON tab instead, which is what that tab is
for, and mark the state so nothing serializes the placeholder over it.

* admin: validate a policy document before saving it

Validation was wired only to the Validate button, so nothing stopped a
document the server's own validator rejects from being stored. With the
structured editor supplying the boilerplate and required dropped from
the textarea, opening the modal, typing a name and clicking Create
Policy was enough to save a statement-less policy.

Share validatePolicyJSON with the two save paths and abort on failure.

* admin: bound the folder autocomplete listing

maxListFoldersEntries caps the folders collected, but nothing capped the
entries paged through to find them, so a bucket holding only flat object
keys - no subfolders to count - was walked to the end, 200 entries per
round trip, behind one keystroke. Measured against an in-process filer:
6 entries 0.5ms, 3k entries 7.7ms, 30k entries 53ms, all of it linear in
the directory rather than in the answer.

Cap the scan as well, and let GetFileBrowser take a prefix so the segment
the user is still typing is filtered by the filer instead of by paging.
The same 30k directory now answers in 0.6ms once a prefix is typed.

* admin: clean the path before scoping list-folders to /buckets

util.CleanWindowsPath only rewrites backslashes, so "/buckets/../etc"
walked straight past the prefix check the endpoint relies on for its
scope. Nothing leaked - filer paths are literal keys, so the traversal
resolved to nothing - but the check reads as a boundary and wasn't one,
and the test asserting it didn't cover the one input that would try.

validateAndCleanFilePath in the same file already does this.

* admin: stringify policy values before escaping them

escapeHtml calls text.replace directly, and the Sid, the per-item action
and resource inputs, and the View modal's Resource/NotResource all pass
values straight out of JSON.parse. A policy carrying "Sid": 5 or
"Action": [1] threw "text.replace is not a function" and took the render
with it. escapedJoin already coerced; use it everywhere and coerce the
editor state at the point it's built.

* admin: only show the NotResource hint in NotResource mode

The hint rendered unconditionally, so it sat under a selector reading
"Resource" telling the user the statement applies to everything except
what they'd listed. Redraw the card when the selector changes so it
follows the mode.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-22 12:33:08 -07:00
Chris LuandGitHub 228500fe37 install.sh: install the Rust maintenance worker (#10882)
* install.sh: install the Rust maintenance worker

The release publishes weed-worker but the installer only knew weed and the
Rust volume server, so the one binary that cannot be built without a Rust
toolchain was the one you had to download by hand.

--component all skips it on a platform it has no build for rather than failing
an install that already put two binaries in place; asking for it by name there
still says so.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* install.sh: clean each component's temp directory as it finishes

The EXIT trap is per-process, so installing more than one component left every
extraction but the last behind. Cleaning at the end of the function keeps the
trap for the paths that exit early.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-22 12:16:28 -07:00
Chris LuandGitHub 7ebf2ebac3 Build the Rust worker against the protoc that ships with the build (#10881)
* worker: compile plugin.proto with the protoc that ships with the build

seaweed-volume already does this: protoc-bin-vendored carries the binary, so
the build needs no package manager and every build sees the same version. An
explicit PROTOC still wins, which is what lets the lance crates - whose own
build scripts read the same variable - share it.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: point the worker builds at the vendored protoc

The jobs installed protobuf-compiler for lance's build scripts. They read
PROTOC, so pointing it at the binary protoc-bin-vendored already puts in the
registry serves them without a system package - one less apt call on the way
to a release, and the same protoc a developer's build uses.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* docs: say what the worker build needs from protoc

The lance crates' build scripts are the ones that need it, not ours, and they
take the same vendored binary.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-22 11:34:33 -07:00
Chris LuandGitHub cc8364a03e Ship the Rust maintenance worker with the release (#10879)
* worker: name the binary weed-worker

It is the Rust side of `weed worker`, the way weed-volume is the Rust side of
`weed volume`, and lance is the first family of jobs it carries rather than
the only one it ever will. The crate keeps its own name: when a second family
arrives the bin target moves to a crate of its own, under this name.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* docker: ship the Rust maintenance worker in the image

Lance table buckets need a worker that can read the format, and until now the
only way to get one was a Rust toolchain and a cargo build. It now sits at
/usr/bin/weed-worker beside the Rust volume server, reached as
`docker run chrislusf/seaweedfs worker-rust --admin host:23646` — the verb
mirrors volume-rust, so plain `worker` still runs the Go one.

Taken pre-built or not at all: the lance jobs pull in arrow and datafusion, far
too large a tree to compile inside the image build, so an architecture CI did
not build for gets the empty placeholder the entrypoint refuses to exec, the
way the Rust volume server already does.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: build the Rust worker for the container images

The same native cross-compile the volume server uses, so the release, latest
and dev images all carry it on amd64 and arm64. The artifact holds both
binaries now, so it is named for that rather than for the volume server.

Only the release directory each job builds is cached: with a debug profile
beside it the worker's target/ reaches 24GB, against a 10GB cache budget.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: publish Rust worker binaries with the release

Linux amd64 and arm64 only: the worker runs beside the cluster it maintains,
and its dependency tree makes every extra target an expensive build.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: build and test the Rust workers on change

Nothing built seaweed-worker in CI, so the release and the container images
would have been the first place a break showed up. Tests run in release too,
rather than compiling lance, arrow and datafusion again in another profile.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* docs: say how to get a released worker

Neither the image nor the release tarballs were mentioned; a toolchain and a
cargo build read as the only way in.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: install protoc for the Rust worker builds

lance's crates compile their own protos, and unlike seaweed-volume they do not
vendor a protoc to do it with, so every job that builds the worker failed at
lance-encoding's build script.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* ci: do not persist credentials in the worker release checkout

The upload step is handed a token explicitly; a cargo build script should not
find another one sitting in the checkout's git config.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* docker: keep the worker's argument boundaries

Unquoted $@ splits on whitespace and expands globs, so an argument carrying
either arrived as something else.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-22 10:22:33 -07:00
Chris LuandGitHub c3f4d799b5 helm: install chart CI against an image tag that exists (#10877)
* helm: install chart CI against an image tag that exists

The release bumps appVersion on master well before the container build
publishes that tag, and the chart CI runs on the bump commit, so every
release turns it red with ImagePullBackOff. Resolve the tag first and
fall back to latest while the new one is still building.

* helm: run the chart CI when the workflow itself changes

* helm: bound the registry lookup in the chart CI

An unbounded curl can hold the job, and the log did not say why the tag
was rejected. Cap it and print the status.
2026-08-22 09:42:28 -07:00
495 changed files with 28895 additions and 4103 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.6
uses: github/codeql-action/init@v4.37.9
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.6
uses: github/codeql-action/autobuild@v4.37.9
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.6
uses: github/codeql-action/analyze@v4.37.9
+39 -5
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'weed/**'
- 'seaweed-volume/**'
- 'seaweed-worker/**'
- 'docker/**'
- 'go.mod'
- 'go.sum'
@@ -16,7 +17,7 @@ permissions:
jobs:
# ── Pre-build Rust volume server binaries natively ──────────────────
# ── Pre-build the Rust binaries natively ────────────────────────────
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -55,10 +56,26 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-docker-dev-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-dev-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-docker-dev-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build normal variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -67,11 +84,19 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-volume-${{ matrix.arch }}
path: weed-volume-normal-${{ matrix.arch }}
name: rust-bins-${{ matrix.arch }}
path: |
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
build-dev-containers:
needs: [build-rust-binaries]
@@ -84,7 +109,7 @@ jobs:
- name: Download pre-built Rust binaries
uses: actions/download-artifact@v8
with:
pattern: rust-volume-*
pattern: rust-bins-*
merge-multiple: true
path: ./rust-bins
@@ -98,7 +123,16 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Docker meta
id: docker_meta
+47 -6
View File
@@ -59,7 +59,7 @@ jobs:
echo "publish=true" >> "$GITHUB_OUTPUT"
fi
# ── Pre-build Rust volume server binaries natively ──────────────────
# ── Pre-build the Rust binaries natively ────────────────────────────
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -100,10 +100,26 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-docker-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build large-disk variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -120,13 +136,20 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-volume-${{ matrix.arch }}
name: rust-bins-${{ matrix.arch }}
path: |
weed-volume-large-disk-${{ matrix.arch }}
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
build:
needs: [setup, build-rust-binaries]
@@ -174,7 +197,7 @@ jobs:
- name: Download pre-built Rust binaries
uses: actions/download-artifact@v8
with:
pattern: rust-volume-*
pattern: rust-bins-*
merge-multiple: true
path: ./rust-bins
@@ -188,7 +211,16 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Docker meta
id: docker_meta
@@ -286,7 +318,7 @@ jobs:
if: needs.setup.outputs.publish != 'true'
uses: actions/download-artifact@v8
with:
pattern: rust-volume-*
pattern: rust-bins-*
merge-multiple: true
path: ./rust-bins
- name: Place Rust binaries in Docker context for local scan
@@ -304,7 +336,16 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Create BuildKit config for local scan build
if: needs.setup.outputs.publish != 'true'
run: |
@@ -364,7 +405,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.6
uses: github/codeql-action/upload-sarif@v4.37.9
if: always()
with:
sarif_file: trivy-results.sarif
@@ -42,9 +42,10 @@ concurrency:
jobs:
# ── Pre-build Rust volume server binaries natively ──────────────────
# Cross-compiles for amd64 and arm64 without QEMU, turning a 5-hour
# emulated cargo build into ~15 minutes of native compilation.
# ── Pre-build the Rust binaries natively ────────────────────────────
# The volume server and the Rust maintenance worker, cross-compiled for
# amd64 and arm64 without QEMU, turning a 5-hour emulated cargo build into
# ~15 minutes of native compilation.
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -83,10 +84,26 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-docker-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build large-disk variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -103,13 +120,20 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-volume-${{ matrix.arch }}
name: rust-bins-${{ matrix.arch }}
path: |
weed-volume-large-disk-${{ matrix.arch }}
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
# One job per (variant, platform) on a native runner, pushed by digest;
# the merge job stitches the digests into one multi-arch tag.
@@ -155,7 +179,7 @@ jobs:
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
uses: actions/download-artifact@v8
with:
pattern: rust-volume-*
pattern: rust-bins-*
merge-multiple: true
path: ./rust-bins
@@ -170,7 +194,16 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Free Disk Space
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -397,7 +430,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.6
uses: github/codeql-action/upload-sarif@v4.37.9
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
+31 -4
View File
@@ -3,10 +3,10 @@ name: "helm: lint and test charts"
on:
push:
branches: [ master ]
paths: ['k8s/**']
paths: ['k8s/**', '.github/workflows/helm_ci.yml']
pull_request:
branches: [ master ]
paths: ['k8s/**']
paths: ['k8s/**', '.github/workflows/helm_ci.yml']
permissions:
contents: read
@@ -1549,11 +1549,37 @@ jobs:
echo "All template rendering tests passed!"
- name: Resolve an image tag that is published
run: |
set -e
# A release bumps appVersion on master ~40 minutes before the container
# build publishes that tag, and this workflow runs on the bump commit.
# Install the last released image for the length of that window instead
# of failing on ImagePullBackOff.
IMAGE=$(helm template test k8s/charts/seaweedfs \
-s templates/master/master-statefulset.yaml | awk '$1 == "image:" {print $2; exit}')
REPO=${IMAGE%:*}
TAG=${IMAGE##*:}
# Anything but a published tag installs latest, which between releases
# is the same digest as the chart's own appVersion, so a registry blip
# costs nothing while failing the job on one would cost a red build.
STATUS=$(curl -sSL --connect-timeout 5 --max-time 15 -o /dev/null \
-w '%{http_code}' "https://hub.docker.com/v2/repositories/$REPO/tags/$TAG" || true)
if [ "$STATUS" = 200 ]; then
echo "installing $IMAGE"
else
echo "$IMAGE is unavailable (HTTP ${STATUS:-none}), installing $REPO:latest"
TAG=latest
fi
echo "IMAGE_TAG=$TAG" >> $GITHUB_ENV
- name: Create kind cluster
uses: helm/kind-action@v1.14.0
- name: Run chart-testing (install)
run: ct install --target-branch ${{ github.event.repository.default_branch }} --all --chart-dirs k8s/charts
run: |
ct install --target-branch ${{ github.event.repository.default_branch }} --all --chart-dirs k8s/charts \
--helm-extra-set-args "--set=image.tag=$IMAGE_TAG"
- name: Verify SFTP host key secret lifecycle
run: |
@@ -1561,7 +1587,7 @@ jobs:
CHART_DIR="k8s/charts/seaweedfs"
NS="sftp-hostkey"
SECRET="hk-seaweedfs-sftp-ssh-secret"
SFTP_ARGS="--set sftp.enabled=true --set master.enabled=false --set volume.enabled=false --set filer.enabled=false"
SFTP_ARGS="--set image.tag=$IMAGE_TAG --set sftp.enabled=true --set master.enabled=false --set volume.enabled=false --set filer.enabled=false"
kubectl create namespace "$NS"
echo "=== install generates a host key, upgrade keeps it ==="
@@ -1638,6 +1664,7 @@ jobs:
# release if the hook Job does not finish, so a clean install is the
# assertion.
helm install np $CHART_DIR -n "$NS" --wait --timeout 8m \
--set image.tag=$IMAGE_TAG \
--set s3.enabled=true \
--set s3.createBuckets[0].name=testbucket \
--set networkPolicy.enabled=true \
+1 -1
View File
@@ -34,7 +34,7 @@ jobs:
id: go
- name: Set up Java
uses: actions/setup-java@v5
uses: actions/setup-java@v6
with:
java-version: ${{ matrix.java }}
distribution: 'temurin'
+1 -1
View File
@@ -73,7 +73,7 @@ jobs:
echo "version=${VERSION}" >> "$GITHUB_OUTPUT"
- name: Set up JDK 17
uses: actions/setup-java@v5
uses: actions/setup-java@v6
with:
java-version: '17'
distribution: 'temurin'
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
uses: actions/checkout@v7
- name: Set up Java
uses: actions/setup-java@v5
uses: actions/setup-java@v6
with:
java-version: ${{ matrix.java }}
distribution: 'temurin'
+2 -2
View File
@@ -131,7 +131,7 @@ jobs:
- name: Results
if: always()
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@v7
with:
name: results-windows
path: C:\results-*.json
@@ -196,7 +196,7 @@ jobs:
- name: Results
if: always()
uses: actions/upload-artifact@v4
uses: actions/upload-artifact@v7
with:
name: results-linux
path: /tmp/results-linux.json
+82
View File
@@ -0,0 +1,82 @@
name: "Rust Plugin Worker Tests"
on:
pull_request:
branches: [ master ]
paths:
- 'seaweed-worker/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
push:
branches: [ master, main ]
paths:
- 'seaweed-worker/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
concurrency:
group: ${{ github.workflow }}-${{ github.head_ref || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
rust-worker-build:
name: Rust Plugin Worker Build and Unit Tests
runs-on: ubuntu-22.04
timeout-minutes: 45
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# cargo tracks its own inputs but not the runner's C toolchain, so a cached
# target/ can carry C objects built against a different glibc than we link against.
- name: Fingerprint build toolchain
id: toolchain
run: echo "fingerprint=$(getconf GNU_LIBC_VERSION | tr ' ' '-')-rustc-$(rustc -V | awk '{print $2}')" >> "$GITHUB_OUTPUT"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/release
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-worker-${{ steps.toolchain.outputs.fingerprint }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
# The release profile is what ships, and it is where the release and the
# container builds would otherwise discover a break for the first time.
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
# Release, so this reuses the build above rather than compiling lance,
# arrow and datafusion a second time in another profile.
- name: Run unit tests
run: cd seaweed-worker && cargo test --release --workspace
+97 -1
View File
@@ -1,4 +1,4 @@
name: "rust: build versioned volume server binaries"
name: "rust: build versioned binaries"
on:
push:
@@ -113,6 +113,102 @@ jobs:
weed-volume_${{ matrix.asset_suffix }}.tar.gz
weed-volume_${{ matrix.asset_suffix }}.tar.gz.md5
# The Rust maintenance worker: Linux only, because it runs beside the cluster
# it maintains rather than on a laptop, and its dependency tree (lance, arrow,
# datafusion) makes every extra target an expensive build.
build-rust-worker-linux:
permissions:
contents: write
runs-on: ubuntu-22.04
strategy:
matrix:
include:
- target: x86_64-unknown-linux-gnu
asset_suffix: linux_amd64
- target: aarch64-unknown-linux-gnu
asset_suffix: linux_arm64
cross: true
steps:
- uses: actions/checkout@v7
with:
# The upload step is handed a token explicitly; a cargo build script
# should not find another one sitting in the checkout's git config.
persist-credentials: false
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
targets: ${{ matrix.target }}
- name: Install cross-compilation tools
if: matrix.cross
run: |
sudo dpkg --add-architecture arm64
sudo sed -i 's/^deb /deb [arch=amd64] /' /etc/apt/sources.list
echo "deb [arch=arm64] http://ports.ubuntu.com/ jammy main restricted universe multiverse" | sudo tee /etc/apt/sources.list.d/arm64.list
echo "deb [arch=arm64] http://ports.ubuntu.com/ jammy-updates main restricted universe multiverse" | sudo tee -a /etc/apt/sources.list.d/arm64.list
sudo apt-get update
sudo apt-get install -y gcc-aarch64-linux-gnu
echo "CARGO_TARGET_AARCH64_UNKNOWN_LINUX_GNU_LINKER=aarch64-linux-gnu-gcc" >> "$GITHUB_ENV"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/${{ matrix.target }}/release
key: rust-worker-release-${{ matrix.target }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-worker-release-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
- name: Package binary
run: |
cp seaweed-worker/target/${{ matrix.target }}/release/weed-worker weed-worker
tar czf weed-worker_${{ matrix.asset_suffix }}.tar.gz weed-worker
rm weed-worker
md5sum weed-worker_${{ matrix.asset_suffix }}.tar.gz > weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
- name: Upload release assets
if: startsWith(github.ref, 'refs/tags/')
uses: softprops/action-gh-release@v3
with:
files: |
weed-worker_${{ matrix.asset_suffix }}.tar.gz
weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- name: Upload artifacts
if: ${{ !startsWith(github.ref, 'refs/tags/') }}
uses: actions/upload-artifact@v7
with:
name: rust-worker-${{ matrix.asset_suffix }}
path: |
weed-worker_${{ matrix.asset_suffix }}.tar.gz
weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
build-rust-volume-darwin:
permissions:
contents: write
@@ -64,6 +64,14 @@ jobs:
echo "=== Running S3 Empty Directory Marker Tests ==="
go test -v -timeout=180s -run TestS3ListObjectsEmptyDirectoryMarkers ./...
- name: Run S3 Prefix Object Tests
timeout-minutes: 15
working-directory: test/s3/normal
run: |
set -x
echo "=== Running S3 Prefix Object Tests ==="
go test -v -timeout=180s -run TestS3PrefixObjectKeys ./...
- name: Run IAM Integration Tests
timeout-minutes: 15
working-directory: test/s3/normal
@@ -34,7 +34,7 @@ jobs:
uses: actions/checkout@v7
- name: Set up JDK 11
uses: actions/setup-java@v5
uses: actions/setup-java@v6
with:
java-version: '11'
distribution: 'temurin'
+17 -2
View File
@@ -29,6 +29,7 @@ FROM alpine:3.23 as rust_builder
ARG TARGETARCH
ARG TAGS
COPY weed-volume-prebuilt/ /prebuilt/
COPY weed-worker-prebuilt/ /prebuilt-worker/
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-volume /build/seaweed-volume
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/weed /build/weed
WORKDIR /build/seaweed-volume
@@ -47,16 +48,30 @@ RUN if [ -f "/prebuilt/weed-volume-${TARGETARCH}" ]; then \
echo "Skipping Rust build for $TARGETARCH (unsupported)" && \
touch /weed-volume; \
fi
# The Rust maintenance worker is taken pre-built or not at all: the lance jobs
# it carries pull in arrow and datafusion, a far larger dependency tree than
# the image build can carry, so an architecture CI did not build for gets the
# same empty placeholder the entrypoint refuses to exec.
RUN if [ -f "/prebuilt-worker/weed-worker-${TARGETARCH}" ]; then \
echo "Using pre-built Rust worker for ${TARGETARCH}" && \
cp "/prebuilt-worker/weed-worker-${TARGETARCH}" /weed-worker; \
else \
echo "No pre-built Rust worker for ${TARGETARCH}" && \
touch /weed-worker; \
fi
# Pre-built binaries arrive via GitHub Actions artifacts, which drop the
# executable bit, so the copied file is 0644 and exec fails with "Permission
# denied". Restore it (no-op for the empty placeholder, which stays size 0).
RUN chmod 0755 /weed-volume
# denied". Restore it (no-op for the empty placeholders, which stay size 0).
RUN chmod 0755 /weed-volume /weed-worker
FROM alpine AS final
LABEL author="Chris Lu"
COPY --from=builder /go/bin/weed /usr/bin/
# Copy Rust volume server binary (real binary on amd64/arm64, empty placeholder on other platforms)
COPY --from=rust_builder /weed-volume /usr/bin/weed-volume
# Same for the Rust maintenance worker, which serves Lance table buckets
COPY --from=rust_builder /weed-worker /usr/bin/weed-worker
RUN mkdir -p /etc/seaweedfs
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/docker/filer.toml /etc/seaweedfs/filer.toml
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/docker/entrypoint.sh /entrypoint.sh
+10
View File
@@ -90,6 +90,16 @@ case "$1" in
exec /usr/bin/weed-volume $ARGS $@
;;
'worker-rust')
shift
if [ ! -s /usr/bin/weed-worker ]; then
echo "Error: Rust maintenance worker is not available on this platform ($(uname -m))." >&2
echo "Use 'worker' for the Go maintenance worker instead." >&2
exit 1
fi
exec /usr/bin/weed-worker "$@"
;;
'server')
ARGS="-dir=/data -volume.max=0 -master.volumeSizeLimitMB=1024"
if isArgPassed "-volume.max" "$@"; then
+27 -26
View File
@@ -4,7 +4,7 @@ go 1.26
require (
cloud.google.com/go v0.123.0 // indirect
cloud.google.com/go/pubsub v1.51.0
cloud.google.com/go/pubsub v1.51.1
cloud.google.com/go/storage v1.64.0
github.com/Shopify/sarama v1.38.1
github.com/aws/aws-sdk-go v1.55.8
@@ -32,7 +32,7 @@ require (
github.com/google/btree v1.1.3
github.com/google/uuid v1.6.0
github.com/google/wire v0.7.0 // indirect
github.com/googleapis/gax-go/v2 v2.23.0 // indirect
github.com/googleapis/gax-go/v2 v2.24.0 // indirect
github.com/gorilla/mux v1.8.1
github.com/hashicorp/errwrap v1.1.0 // indirect
github.com/hashicorp/go-multierror v1.1.1 // indirect
@@ -64,7 +64,7 @@ require (
github.com/prometheus/procfs v0.21.1
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/seaweedfs/goexif v1.0.3
github.com/seaweedfs/goexif v2.0.0+incompatible
github.com/seaweedfs/raft v1.2.0
github.com/sirupsen/logrus v1.9.4 // indirect
github.com/spf13/afero v1.15.0 // indirect
@@ -99,15 +99,15 @@ require (
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.293.0
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 // indirect
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6
google.golang.org/protobuf v1.36.11
google.golang.org/api v0.294.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/protobuf v1.36.12
gopkg.in/inf.v0 v0.9.1 // indirect
modernc.org/b v1.0.0 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
modernc.org/sqlite v1.56.0
modernc.org/sqlite v1.57.0
)
require (
@@ -122,14 +122,14 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.43.5
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2/config v1.32.33
github.com/aws/aws-sdk-go-v2/credentials v1.19.34
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
github.com/fluent/fluent-logger-golang v1.10.1
github.com/getsentry/sentry-go v0.44.1
github.com/getsentry/sentry-go v0.48.0
github.com/go-ldap/ldap/v3 v3.4.13
github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/flatbuffers/go v0.0.0-20230108230133-3b8644d32c50
@@ -143,20 +143,20 @@ require (
github.com/orcaman/concurrent-map/v2 v2.0.1
github.com/parquet-go/parquet-go v0.32.0
github.com/pkg/sftp v1.13.11
github.com/rabbitmq/amqp091-go v1.13.0
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rclone/rclone v1.75.0
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.6
github.com/shirou/gopsutil/v4 v4.26.7
github.com/tarantool/go-option v1.1.0
github.com/tarantool/go-tarantool/v3 v3.0.0
github.com/tarantool/go-tarantool/v3 v3.0.1
github.com/testcontainers/testcontainers-go v0.43.0
github.com/tikv/client-go/v2 v2.0.7
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
@@ -193,7 +193,7 @@ require (
github.com/cenkalti/backoff/v5 v5.0.3 // indirect
github.com/clipperhouse/uax29/v2 v2.7.0 // indirect
github.com/cockroachdb/apd/v3 v3.2.1 // indirect
github.com/cockroachdb/errors v1.11.3 // indirect
github.com/cockroachdb/errors v1.14.0 // indirect
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506 // indirect
github.com/cockroachdb/redact v1.1.5 // indirect
github.com/cockroachdb/version v0.0.0-20250314144055-3860cd14adf2 // indirect
@@ -238,12 +238,12 @@ require (
github.com/lpar/calendar v0.2.0 // indirect
github.com/magiconair/properties v1.8.10 // indirect
github.com/moby/docker-image-spec v1.3.1 // indirect
github.com/moby/go-archive v0.2.0 // indirect
github.com/moby/go-archive v0.3.0 // indirect
github.com/moby/moby/api v1.54.2 // indirect
github.com/moby/moby/client v0.4.0 // indirect
github.com/moby/patternmatcher v0.6.1 // indirect
github.com/moby/sys/sequential v0.6.0 // indirect
github.com/moby/sys/user v0.4.0 // indirect
github.com/moby/sys/sequential v0.7.0 // indirect
github.com/moby/sys/user v0.4.1 // indirect
github.com/moby/sys/userns v0.1.0 // indirect
github.com/moby/term v0.5.2 // indirect
github.com/oklog/ulid/v2 v2.1.1 // indirect
@@ -260,6 +260,7 @@ require (
github.com/rclone/Proton-API-Bridge v1.0.4 // indirect
github.com/rclone/go-proton-api v1.0.3 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd // indirect
github.com/ryanuber/go-glob v1.0.0 // indirect
github.com/sasha-s/go-deadlock v0.3.1 // indirect
github.com/smarty/assertions v1.15.0 // indirect
@@ -291,11 +292,11 @@ require (
require (
cel.dev/expr v0.25.2 // indirect
cloud.google.com/go/auth v0.23.0 // indirect
cloud.google.com/go/auth v0.23.2 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
cloud.google.com/go/iam v1.11.0 // indirect
cloud.google.com/go/monitoring v1.29.0 // indirect
cloud.google.com/go/iam v1.12.0 // indirect
cloud.google.com/go/monitoring v1.30.0 // indirect
filippo.io/edwards25519 v1.2.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0
@@ -336,7 +337,7 @@ require (
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4
github.com/aws/smithy-go v1.27.7
github.com/aws/smithy-go v1.28.1
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
github.com/buengese/sgzip v0.1.1 // indirect
@@ -354,7 +355,7 @@ require (
github.com/d4l3k/messagediff v1.2.1 // indirect
github.com/dgryski/go-farm v0.0.0-20200201041132-a6ae2369ad13 // indirect
github.com/dropbox/dropbox-sdk-go-unofficial/v6 v6.4.0 // indirect
github.com/ebitengine/purego v0.10.1 // indirect
github.com/ebitengine/purego v0.10.2 // indirect
github.com/elastic/gosigar v0.14.3 // indirect
github.com/emersion/go-message v0.18.2 // indirect
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 // indirect
@@ -474,7 +475,7 @@ require (
github.com/winfsp/cgofuse v1.6.1-0.20260126094232-f2c4fccdb286
github.com/xanzy/ssh-agent v0.3.3 // indirect
github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e // indirect
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b // indirect
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705 // indirect
github.com/ydb-platform/ydb-go-yc v0.12.1 // indirect
github.com/ydb-platform/ydb-go-yc-metadata v0.6.1 // indirect
github.com/yunify/qingstor-sdk-go/v3 v3.2.0 // indirect
@@ -496,8 +497,8 @@ require (
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
gopkg.in/yaml.v2 v2.4.0 // indirect
+56 -54
View File
@@ -94,8 +94,8 @@ cloud.google.com/go/assuredworkloads v1.7.0/go.mod h1:z/736/oNmtGAyU47reJgGN+KVo
cloud.google.com/go/assuredworkloads v1.8.0/go.mod h1:AsX2cqyNCOvEQC8RMPnoc0yEarXQk6WEKkxYfL6kGIo=
cloud.google.com/go/assuredworkloads v1.9.0/go.mod h1:kFuI1P78bplYtT77Tb1hi0FMxM0vVpRC7VVoJC3ZoT0=
cloud.google.com/go/assuredworkloads v1.10.0/go.mod h1:kwdUQuXcedVdsIaKgKTp9t0UJkE5+PAVNhdQm4ZVq2E=
cloud.google.com/go/auth v0.23.0 h1:6Gg1CMgpgubRG7DGz5Vf1pcoNo8RfiRiRAPS4crTp54=
cloud.google.com/go/auth v0.23.0/go.mod h1:4DhBRcqvtljQN3dJ57qtqbib5ZGCYE5f2crfiiC2EM0=
cloud.google.com/go/auth v0.23.2 h1:pxSCpfiji41hpzpPdMCftEUCezpgpqmmDdYiAjCKXxo=
cloud.google.com/go/auth v0.23.2/go.mod h1:4DhBRcqvtljQN3dJ57qtqbib5ZGCYE5f2crfiiC2EM0=
cloud.google.com/go/auth/oauth2adapt v0.2.8 h1:keo8NaayQZ6wimpNSmW5OPc283g65QNIiLpZnkHRbnc=
cloud.google.com/go/auth/oauth2adapt v0.2.8/go.mod h1:XQ9y31RkqZCcwJWNSx2Xvric3RrU88hAYYbjDWYDL+c=
cloud.google.com/go/automl v1.5.0/go.mod h1:34EjfoFGMZ5sgJ9EoLsRtdPSNZLcfflJR39VbVNS2M0=
@@ -283,8 +283,8 @@ cloud.google.com/go/iam v0.7.0/go.mod h1:H5Br8wRaDGNc8XP3keLc4unfUUZeyH3Sfl9XpQE
cloud.google.com/go/iam v0.8.0/go.mod h1:lga0/y3iH6CX7sYqypWJ33hf7kkfXJag67naqGESjkE=
cloud.google.com/go/iam v0.11.0/go.mod h1:9PiLDanza5D+oWFZiH1uG+RnRCfEGKoyl6yo4cgWZGY=
cloud.google.com/go/iam v0.12.0/go.mod h1:knyHGviacl11zrtZUoDuYpDgLjvr28sLQaG0YB2GYAY=
cloud.google.com/go/iam v1.11.0 h1:KieQ9Pb+LLPak1O3Rv3GgCxhnmkYf7Xyh0P5HfF1jFM=
cloud.google.com/go/iam v1.11.0/go.mod h1:KP+nKGugNJW4LcLx1uEZcq1ok5sQHFaQehQNl4QDgV4=
cloud.google.com/go/iam v1.12.0 h1:Aki3bX9aHUDKPHfnRJfDcTdVedvy6quGBQcTqx3DRXk=
cloud.google.com/go/iam v1.12.0/go.mod h1:FEZ4lXpADAC2AIpQY7LANNjjwyQ2jK439CI2VaD+sLY=
cloud.google.com/go/iap v1.4.0/go.mod h1:RGFwRJdihTINIe4wZ2iCP0zF/qu18ZwyKxrhMhygBEc=
cloud.google.com/go/iap v1.5.0/go.mod h1:UH/CGgKd4KyohZL5Pt0jSKE4m3FR51qg6FKQ/z/Ix9A=
cloud.google.com/go/iap v1.6.0/go.mod h1:NSuvI9C/j7UdjGjIde7t7HBz+QTwBcapPE07+sSRcLk=
@@ -310,8 +310,8 @@ cloud.google.com/go/lifesciences v0.6.0/go.mod h1:ddj6tSX/7BOnhxCSd3ZcETvtNr8NZ6
cloud.google.com/go/lifesciences v0.8.0/go.mod h1:lFxiEOMqII6XggGbOnKiyZ7IBwoIqA84ClvoezaA/bo=
cloud.google.com/go/logging v1.6.1/go.mod h1:5ZO0mHHbvm8gEmeEUHrmDlTDSu5imF6MUP9OfilNXBw=
cloud.google.com/go/logging v1.7.0/go.mod h1:3xjP2CjkM3ZkO73aj4ASA5wRPGGCRrPIAeNqVNkzY8M=
cloud.google.com/go/logging v1.18.0 h1:KhzZq+1cSkPH9YUaKLLhLtQxIHitVayBmk0sGfoM9+k=
cloud.google.com/go/logging v1.18.0/go.mod h1:ZGKnpBaURITh+g/uom2VhbiFoFWvejcrHPDhxFtU/gI=
cloud.google.com/go/logging v1.19.0 h1:NCqhdVUg3wQ8Cobdf16FDSuTGi3+6+hdSBHrY5TsR6Q=
cloud.google.com/go/logging v1.19.0/go.mod h1:i40NZCHC9Gqvod4yE+yQfDWwlgwW/SrshkkGibCHxcA=
cloud.google.com/go/longrunning v0.1.1/go.mod h1:UUFxuDWkv22EuY93jjmDMFT5GPQKeFVJBIF6QlTqdsE=
cloud.google.com/go/longrunning v0.3.0/go.mod h1:qth9Y41RRSUE69rDcOn6DdK3HfQfsUI0YSmW3iIlLJc=
cloud.google.com/go/longrunning v0.4.1/go.mod h1:4iWDqhBZ70CvZ6BfETbvam3T8FMvLK+eFj0E6AaRQTo=
@@ -338,8 +338,8 @@ cloud.google.com/go/metastore v1.10.0/go.mod h1:fPEnH3g4JJAk+gMRnrAnoqyv2lpUCqJP
cloud.google.com/go/monitoring v1.7.0/go.mod h1:HpYse6kkGo//7p6sT0wsIC6IBDET0RhIsnmlA53dvEk=
cloud.google.com/go/monitoring v1.8.0/go.mod h1:E7PtoMJ1kQXWxPjB6mv2fhC5/15jInuulFdYYtlcvT4=
cloud.google.com/go/monitoring v1.12.0/go.mod h1:yx8Jj2fZNEkL/GYZyTLS4ZtZEZN8WtDEiEqG4kLK50w=
cloud.google.com/go/monitoring v1.29.0 h1:AHhDsFaSax1/4k+qlIDX/SDGe6hggnfXJ9dkgD9qBPY=
cloud.google.com/go/monitoring v1.29.0/go.mod h1:72NOVjJXHY/HBfoLT0+qlCZBT059+9VXLeAnL2PeeVM=
cloud.google.com/go/monitoring v1.30.0 h1:r/d+JUbyKmJ8b07iznuKfzVzrIXTWxHQ3lBRm3x2LlY=
cloud.google.com/go/monitoring v1.30.0/go.mod h1:htlUR0QWVMrjFzZmN4LGnMAve9xB/eduwjmINxVZ8RM=
cloud.google.com/go/networkconnectivity v1.4.0/go.mod h1:nOl7YL8odKyAOtzNX73/M5/mGZgqqMeryi6UPZTk/rA=
cloud.google.com/go/networkconnectivity v1.5.0/go.mod h1:3GzqJx7uhtlM3kln0+x5wyFvuVH1pIBJjhCpjzSt75o=
cloud.google.com/go/networkconnectivity v1.6.0/go.mod h1:OJOoEXW+0LAxHh89nXd64uGG+FbQoeH8DtxCHVOMlaM=
@@ -391,8 +391,8 @@ cloud.google.com/go/pubsub v1.3.1/go.mod h1:i+ucay31+CNRpDW4Lu78I4xXG+O1r/MAHgjp
cloud.google.com/go/pubsub v1.26.0/go.mod h1:QgBH3U/jdJy/ftjPhTkyXNj543Tin1pRYcdcPRnFIRI=
cloud.google.com/go/pubsub v1.27.1/go.mod h1:hQN39ymbV9geqBnfQq6Xf63yNhUAhv9CZhzp5O6qsW0=
cloud.google.com/go/pubsub v1.28.0/go.mod h1:vuXFpwaVoIPQMGXqRyUQigu/AX1S3IWugR9xznmcXX8=
cloud.google.com/go/pubsub v1.51.0 h1:XOaCejsqX7EEtUdQz+WPag66wWsUUGliyCOfGPKfo90=
cloud.google.com/go/pubsub v1.51.0/go.mod h1:NERXf11sd82UV3VnflcUj8POIyQUXT/QwrKlxD8di/I=
cloud.google.com/go/pubsub v1.51.1 h1:R3G1wCOxBO7jRpL8x2pdZMv1GAJDF6ax/m2zPOtvTNE=
cloud.google.com/go/pubsub v1.51.1/go.mod h1:y2T0IKtW1iWwVvazYaRpqOAFO4gy2+O7dTDt9TWY/5U=
cloud.google.com/go/pubsub/v2 v2.6.0 h1:8pjR0id+GTB+krKx5G6AGJoYrHog58w2Q89PCOrfM64=
cloud.google.com/go/pubsub/v2 v2.6.0/go.mod h1:4anqvV/w8Pcgu2tO0qr2XgsF3GXHowzryfQ5gOnVmWY=
cloud.google.com/go/pubsublite v1.5.0/go.mod h1:xapqNQ1CuLfGi23Yda/9l4bBCKz/wC3KIJ5gKcxveZg=
@@ -707,8 +707,8 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.43.5 h1:yKT5GYnFWhuDo+DqKvE5ZPwVn3RjC4MAeBtZGlh6AVM=
github.com/aws/aws-sdk-go-v2 v1.43.5/go.mod h1:wZjAJppCntyOGgVSmgVTfDyRJK5PHOasO6Wsy8U7Axk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 h1:3IZY0XAJquT3aHzbkHfPzy4ACPcEjVG0x87KOwtpqGY=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14/go.mod h1:zwM6veDkhGgQFqkBy+uT28AAYpLu+uFMlPl+rCg/73E=
github.com/aws/aws-sdk-go-v2/config v1.32.33 h1:M1m/Q6f0OKDEDGwhiNOqx1OjTdrewe3v+GDbHmKczWk=
@@ -749,8 +749,8 @@ github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 h1:AsbZcJAQPRmHDJG8K1N0pof/
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4/go.mod h1:6imqztH0//t0mKbl6yWl7swSEl7F/w32oAmqB3vP1ag=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4 h1:w/AryDYMjSUANSQ2uoZxJovUsMTwWJNTv3IMex30Y+4=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4/go.mod h1:WeBiAa67azG7Su9Vf+ChGDBLiAozJCXzdjXiPBUwtbc=
github.com/aws/smithy-go v1.27.7 h1:Zgj5z4LfcDYoQIVk+n/yGdTkP/2y6ZT5vYxe0fp7bqE=
github.com/aws/smithy-go v1.27.7/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
github.com/bahlo/generic-list-go v0.2.0/go.mod h1:2KvAjgMlE5NNynlg/5iLrrCCZ2+5xWbdbCW3pNTGyYg=
github.com/bazelbuild/rules_go v0.46.0 h1:CTefzjN/D3Cdn3rkrM6qMWuQj59OBcuOjyIp3m4hZ7s=
@@ -848,8 +848,8 @@ github.com/cncf/xds/go v0.0.0-20260202195803-dba9d589def2 h1:aBangftG7EVZoUb69Os
github.com/cncf/xds/go v0.0.0-20260202195803-dba9d589def2/go.mod h1:qwXFYgsP6T7XnJtbKlf1HP8AjxZZyzxMmc+Lq5GjlU4=
github.com/cockroachdb/apd/v3 v3.2.1 h1:U+8j7t0axsIgvQUqthuNm82HIrYXodOV2iWLWtEaIwg=
github.com/cockroachdb/apd/v3 v3.2.1/go.mod h1:klXJcjp+FffLTHlhIG69tezTDvdP065naDsHzKhYSqc=
github.com/cockroachdb/errors v1.11.3 h1:5bA+k2Y6r+oz/6Z/RFlNeVCesGARKuC6YymtcDrbC/I=
github.com/cockroachdb/errors v1.11.3/go.mod h1:m4UIW4CDjx+R5cybPsNrRbreomiFqt8o1h1wUVazSd8=
github.com/cockroachdb/errors v1.14.0 h1:EfdVEJpN3z8rPMo43Yit59LxoiIa470fSXpZXuEs+ZI=
github.com/cockroachdb/errors v1.14.0/go.mod h1:xRa70jZ9sNBQmISt5KmJmAD++E4dQHm89oCRiZGEdq0=
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506 h1:ASDL+UJcILMqgNeV5jiqR4j+sTuvQNHdf2chuKj1M5k=
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506/go.mod h1:Mw7HqKr2kdtu6aYGn3tPmAftiP3QPX63LdK/zcariIo=
github.com/cockroachdb/redact v1.1.5 h1:u1PMllDkdFfPWaNGMyLD1+so+aq3uUItthCFqzwPJ30=
@@ -956,8 +956,8 @@ github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 h1:Oy0F4A
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3/go.mod h1:YvSRo5mw33fLEx1+DlK6L2VV43tJt5Eyel9n9XBcR+0=
github.com/eapache/queue v1.1.0 h1:YOEu7KNc61ntiQlcEeUIoDTJ2o8mQznoNvUhiigpIqc=
github.com/eapache/queue v1.1.0/go.mod h1:6eCeP0CKFpHLu8blIFXhExK/dRa7WDZfr6jVFPTqq+I=
github.com/ebitengine/purego v0.10.1 h1:dewVBCBT2GaMu1SrNTYxQhgQBethzfhiwvZiLGP/qyY=
github.com/ebitengine/purego v0.10.1/go.mod h1:iIjxzd6CiRiOG0UyXP+V1+jWqUXVjPKLAI0mRfJZTmQ=
github.com/ebitengine/purego v0.10.2 h1:W809HbnvzAxgdm+aOvlSekrM16wGCdT/e76+9tS7gzE=
github.com/ebitengine/purego v0.10.2/go.mod h1:iIjxzd6CiRiOG0UyXP+V1+jWqUXVjPKLAI0mRfJZTmQ=
github.com/eiannone/keyboard v0.0.0-20220611211555-0d226195f203 h1:XBBHcIb256gUJtLmY22n99HaZTz+r2Z51xUPi01m3wg=
github.com/eiannone/keyboard v0.0.0-20220611211555-0d226195f203/go.mod h1:E1jcSv8FaEny+OP/5k9UxZVw9YFWGj7eI4KR/iOBqCg=
github.com/elastic/gosigar v0.14.3 h1:xwkKwPia+hSfg9GqrCUKYdId102m9qTJIIr7egmK/uo=
@@ -1031,8 +1031,8 @@ github.com/gabriel-vasile/mimetype v1.4.13 h1:46nXokslUBsAJE/wMsp5gtO500a4F3Nkz9
github.com/gabriel-vasile/mimetype v1.4.13/go.mod h1:d+9Oxyo1wTzWdyVUPMmXFvp4F9tea18J8ufA774AB3s=
github.com/geoffgarside/ber v1.2.0 h1:/loowoRcs/MWLYmGX9QtIAbA+V/FrnVLsMMPhwiRm64=
github.com/geoffgarside/ber v1.2.0/go.mod h1:jVPKeCbj6MvQZhwLYsGwaGI52oUorHoHKNecGT85ZCc=
github.com/getsentry/sentry-go v0.44.1 h1:/cPtrA5qB7uMRrhgSn9TYtcEF36auGP3Y6+ThvD/yaI=
github.com/getsentry/sentry-go v0.44.1/go.mod h1:XDotiNZbgf5U8bPDUAfvcFmOnMQQceESxyKaObSssW0=
github.com/getsentry/sentry-go v0.48.0 h1:FRZNr7Uk1C86ev1bSJmYlUkL9oyivQA6YOcdYfaaMmY=
github.com/getsentry/sentry-go v0.48.0/go.mod h1:E5UkA5wp1qR2+MDydNYlVeUiNN2xEdjYMidkgf0Qoss=
github.com/ghodss/yaml v1.0.0/go.mod h1:4dBDuWmgqj2HViK6kFavaiC9ZROes6MMH2rRYeMEF04=
github.com/gin-contrib/sse v1.1.0 h1:n0w2GMuUpWDVp7qSpvze6fAu9iRxJY4Hmj6AmBOU05w=
github.com/gin-contrib/sse v1.1.0/go.mod h1:hxRZ5gVpWMT7Z0B0gSNYqqsSCNIJMjzvm6fqCz9vjwM=
@@ -1267,8 +1267,8 @@ github.com/googleapis/gax-go/v2 v2.4.0/go.mod h1:XOTVJ59hdnfJLIP/dh8n5CGryZR2LxK
github.com/googleapis/gax-go/v2 v2.5.1/go.mod h1:h6B0KMMFNtI2ddbGJn3T3ZbwkeT6yqEF02fYlzkUCyo=
github.com/googleapis/gax-go/v2 v2.6.0/go.mod h1:1mjbznJAPHFpesgE5ucqfYEscaz5kMdcIDwU/6+DDoY=
github.com/googleapis/gax-go/v2 v2.7.0/go.mod h1:TEop28CZZQ2y+c0VxMUmu1lV+fQx57QpBWsYpwqHJx8=
github.com/googleapis/gax-go/v2 v2.23.0 h1:Tchl7qkvE7Ip3y+ztvNufYFvkfqTe7NfLTYGIdJRLuE=
github.com/googleapis/gax-go/v2 v2.23.0/go.mod h1:rBQKOVJCdb8IFEzg+FCwlt1LP/xMDGuqUXhUG+XMXEg=
github.com/googleapis/gax-go/v2 v2.24.0 h1:myMaPYyF9MecEmvQqMqomIwn9t/4KCZN9qnwsS76wlg=
github.com/googleapis/gax-go/v2 v2.24.0/go.mod h1:IaTHBDd7NHxSCiu0vEs8pQZu4dGZrWwuSoxCnk16OFM=
github.com/googleapis/go-type-adapters v1.0.0/go.mod h1:zHW75FOG2aur7gAO2B+MLby+cLsWGBF62rFAi7WjWO4=
github.com/googleapis/google-cloud-go-testing v0.0.0-20200911160855-bcd43fbb19e8/go.mod h1:dvDLG8qkwmyD9a/MJJN3XJcT3xFxOKAvTZGvuZmac9g=
github.com/gookit/assert v0.1.1 h1:lh3GcawXe/p+cU7ESTZ5Ui3Sm/x8JWpIis4/1aF0mY0=
@@ -1543,8 +1543,8 @@ github.com/moby/buildkit v0.29.0 h1:wxLEFbCOJntEDjSNNN2YWd8zxltZxT5muDQ0LzpbtpU=
github.com/moby/buildkit v0.29.0/go.mod h1:Dmv2FeDe34t75QuzeU87rBoZpAAkcpT5zeu4hXzmASc=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
github.com/moby/docker-image-spec v1.3.1/go.mod h1:eKmb5VW8vQEh/BAr2yvVNvuiJuY6UIocYsFu/DxxRpo=
github.com/moby/go-archive v0.2.0 h1:zg5QDUM2mi0JIM9fdQZWC7U8+2ZfixfTYoHL7rWUcP8=
github.com/moby/go-archive v0.2.0/go.mod h1:mNeivT14o8xU+5q1YnNrkQVpK+dnNe/K6fHqnTg4qPU=
github.com/moby/go-archive v0.3.0 h1:nos4BtzzUIqB406BgQnWGMI4qib9BZ8XUHU+ucv/n1c=
github.com/moby/go-archive v0.3.0/go.mod h1:Npdv43fFqlhZW7Xo8fbm3ZMYFvAGNviUPqX21VERbcE=
github.com/moby/locker v1.0.1 h1:fOXqR41zeveg4fFODix+1Ch4mj/gT0NE1XJbp/epuBg=
github.com/moby/locker v1.0.1/go.mod h1:S7SDdo5zpBK84bzzVlKr2V0hz+7x9hWbYC/kq7oQppc=
github.com/moby/moby/api v1.54.2 h1:wiat9QAhnDQjA7wk1kh/TqHz2I1uUA7M7t9SAl/JNXg=
@@ -1559,14 +1559,14 @@ github.com/moby/sys/capability v0.4.0 h1:4D4mI6KlNtWMCM1Z/K0i7RV1FkX+DBDHKVJpCnd
github.com/moby/sys/capability v0.4.0/go.mod h1:4g9IK291rVkms3LKCDOoYlnV8xKwoDTpIrNEE35Wq0I=
github.com/moby/sys/mountinfo v0.7.2 h1:1shs6aH5s4o5H2zQLn796ADW1wMrIwHsyJ2v9KouLrg=
github.com/moby/sys/mountinfo v0.7.2/go.mod h1:1YOa8w8Ih7uW0wALDUgT1dTTSBrZ+HiBLGws92L2RU4=
github.com/moby/sys/sequential v0.6.0 h1:qrx7XFUd/5DxtqcoH1h438hF5TmOvzC/lspjy7zgvCU=
github.com/moby/sys/sequential v0.6.0/go.mod h1:uyv8EUTrca5PnDsdMGXhZe6CCe8U/UiTWd+lL+7b/Ko=
github.com/moby/sys/sequential v0.7.0 h1:ASQNGNROJSuOO6LL6bPHbKvuZu6NU8P4ldPWk31zj/8=
github.com/moby/sys/sequential v0.7.0/go.mod h1:NfSTAp6V3fw4tmkD62PEcOKeZKquXT8VKCkf7aVR79o=
github.com/moby/sys/signal v0.7.1 h1:PrQxdvxcGijdo6UXXo/lU/TvHUWyPhj7UOpSo8tuvk0=
github.com/moby/sys/signal v0.7.1/go.mod h1:Se1VGehYokAkrSQwL4tDzHvETwUZlnY7S5XtQ50mQp8=
github.com/moby/sys/symlink v0.3.0 h1:GZX89mEZ9u53f97npBy4Rc3vJKj7JBDj/PN2I22GrNU=
github.com/moby/sys/symlink v0.3.0/go.mod h1:3eNdhduHmYPcgsJtZXW1W4XUJdZGBIkttZ8xKqPUJq0=
github.com/moby/sys/user v0.4.0 h1:jhcMKit7SA80hivmFJcbB1vqmw//wU61Zdui2eQXuMs=
github.com/moby/sys/user v0.4.0/go.mod h1:bG+tYYYJgaMtRKgEmuueC0hJEAZWwtIbZTB+85uoHjs=
github.com/moby/sys/user v0.4.1 h1:RgjRlaDKi/Xmyrz4t8lyzXT6v2ooFeO/7xtchmhVWE0=
github.com/moby/sys/user v0.4.1/go.mod h1:E9QsW5WRe1kUAf7kW8hXKwu1uhsZEAdPLYHYSDudF4Y=
github.com/moby/sys/userns v0.1.0 h1:tVLXkFOxVu9A64/yh59slHVv9ahO9UIev4JZusOLG/g=
github.com/moby/sys/userns v0.1.0/go.mod h1:IHUYgu/kao6N8YZlp9Cf444ySSvCmDlmzUcYfDHOl28=
github.com/moby/term v0.5.2 h1:6qk3FJAFDs6i/q3W/pQ97SX192qKfZgGjCQqfCJkgzQ=
@@ -1749,8 +1749,8 @@ github.com/quic-go/qpack v0.6.0 h1:g7W+BMYynC1LbYLSqRt8PBg5Tgwxn214ZZR34VIOjz8=
github.com/quic-go/qpack v0.6.0/go.mod h1:lUpLKChi8njB4ty2bFLX2x4gzDqXwUpaO1DP9qMDZII=
github.com/quic-go/quic-go v0.59.0 h1:OLJkp1Mlm/aS7dpKgTc6cnpynnD2Xg7C1pwL6vy/SAw=
github.com/quic-go/quic-go v0.59.0/go.mod h1:upnsH4Ju1YkqpLXC305eW3yDZ4NfnNbmQRCMWS58IKU=
github.com/rabbitmq/amqp091-go v1.13.0 h1:L8NA1WtF76C6KA3LAoufjfLgbist/If1UQYcsOjtxXA=
github.com/rabbitmq/amqp091-go v1.13.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rabbitmq/amqp091-go v1.14.0 h1:RSaT7aOKt/OrkVUyswPDW29lnRz9psuGmfZFBmLqLek=
github.com/rabbitmq/amqp091-go v1.14.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rclone/Proton-API-Bridge v1.0.4 h1:uGQJRjQC1hVLd5kqLsXc6CWO6oqrVeLoKQYoHapEZDg=
github.com/rclone/Proton-API-Bridge v1.0.4/go.mod h1:VTPBYZotKAeDLlAzxU2O/s14NXk9FxUt9hn1jhH2iY8=
github.com/rclone/go-proton-api v1.0.3 h1:3gBTzR+j0dYiTwtj9yKIdN/aV3W2a8KIPKp0GArojyQ=
@@ -1790,6 +1790,8 @@ github.com/rs/zerolog v1.34.0 h1:k43nTLIwcTVQAncfCw4KZ2VY6ukYoZaBPNOE8txlOeY=
github.com/rs/zerolog v1.34.0/go.mod h1:bJsvje4Z08ROH4Nhs5iH600c3IkWhwp44iRc54W6wYQ=
github.com/ruudk/golang-pdf417 v0.0.0-20181029194003-1af4ab5afa58/go.mod h1:6lfFZQK844Gfx8o5WFuvpxWRwnSoipWe/p622j1v06w=
github.com/ruudk/golang-pdf417 v0.0.0-20201230142125-a7e3863a1245/go.mod h1:pQAZKsJ8yyVxGRWYNEm9oFB8ieLgKFnamEyDmSA0BRk=
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd h1:CmH9+J6ZSsIjUK3dcGsnCnO41eRBOnY12zwkn5qVwgc=
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd/go.mod h1:hPqNNc0+uJM6H+SuU8sEs5K5IQeKccPqeSjfgcKGgPk=
github.com/ryanuber/go-glob v1.0.0 h1:iQh3xXAumdQ+4Ufa5b25cRpC5TYKlno6hsv6Cb3pkBk=
github.com/ryanuber/go-glob v1.0.0/go.mod h1:807d1WSdnB0XRJzKNil9Om6lcp/3a0v4qIHxIXzX/Yc=
github.com/sabhiram/go-gitignore v0.0.0-20210923224102-525f6e181f06 h1:OkMGxebDjyw0ULyrTYWeN0UNCCkmCWfjPnIA2W6oviI=
@@ -1808,8 +1810,8 @@ github.com/seaweedfs/cockroachdb-parser v0.0.0-20260225204133-2f342c5ea564 h1:Tg
github.com/seaweedfs/cockroachdb-parser v0.0.0-20260225204133-2f342c5ea564/go.mod h1:JSKCh6uCHBz91lQYFYHCyTrSVIPge4SUFVn28iwMNB0=
github.com/seaweedfs/go-fuse/v2 v2.9.4 h1:ACyloiuopdhRSjdLLeSWbsVaemMPskORaRF01TY6GyM=
github.com/seaweedfs/go-fuse/v2 v2.9.4/go.mod h1:zABdmWEa6A0bwaBeEOBUeUkGIZlxUhcdv+V1Dcc/U/I=
github.com/seaweedfs/goexif v1.0.3 h1:ve/OjI7dxPW8X9YQsv3JuVMaxEyF9Rvfd04ouL+Bz30=
github.com/seaweedfs/goexif v1.0.3/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/goexif v2.0.0+incompatible h1:x8pckiT12QQhifwhDQpeISgDfsqmQ6VR4LFPQ64JRps=
github.com/seaweedfs/goexif v2.0.0+incompatible/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/raft v1.2.0 h1:Ez4Hw9ifBbTT7wg54DvGHBjw1vRlTb4roH0TKl0Oj9Y=
github.com/seaweedfs/raft v1.2.0/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/secure-systems-lab/go-securesystemslib v0.10.0 h1:l+H5ErcW0PAehBNrBxoGv1jjNpGYdZ9RcheFkB2WI14=
@@ -1820,8 +1822,8 @@ github.com/sergi/go-diff v1.2.0 h1:XU+rvMAioB0UC3q1MFrIQy4Vo5/4VsRDQQXHsEya6xQ=
github.com/sergi/go-diff v1.2.0/go.mod h1:STckp+ISIX8hZLjrqAeVduY0gWCT9IjLuqbuNXdaHfM=
github.com/shibumi/go-pathspec v1.3.0 h1:QUyMZhFo0Md5B8zV8x2tesohbb5kfbpTi9rBnKh5dkI=
github.com/shibumi/go-pathspec v1.3.0/go.mod h1:Xutfslp817l2I1cZvgcfeMQJG5QnU2lh5tVaaMCl3jE=
github.com/shirou/gopsutil/v4 v4.26.6 h1:Mzr/npDtQC/xpeEuQKHZt8Zo9CmPvhTj8nkR8w5TLDs=
github.com/shirou/gopsutil/v4 v4.26.6/go.mod h1:LZ6ewCSkBqUpvSOf+LsTGnRinC6iaNUNMGBtDkJBaLQ=
github.com/shirou/gopsutil/v4 v4.26.7 h1:IXzpHz/dkMRYAhKkOXr1HB6SuzWU3eoyyeWe7g3bNZc=
github.com/shirou/gopsutil/v4 v4.26.7/go.mod h1:5O9FjBiXoTDFatIWjZZosqj4pV0DRtLx598xGbBehzM=
github.com/sigstore/sigstore v1.10.4 h1:ytOmxMgLdcUed3w1SbbZOgcxqwMG61lh1TmZLN+WeZE=
github.com/sigstore/sigstore v1.10.4/go.mod h1:tDiyrdOref3q6qJxm2G+JHghqfmvifB7hw+EReAfnbI=
github.com/sigstore/sigstore-go v1.1.4 h1:wTTsgCHOfqiEzVyBYA6mDczGtBkN7cM8mPpjJj5QvMg=
@@ -1907,8 +1909,8 @@ github.com/tarantool/go-iproto v1.1.0 h1:HULVOIHsiehI+FnHfM7wMDntuzUddO09DKqu2Wn
github.com/tarantool/go-iproto v1.1.0/go.mod h1:LNCtdyZxojUed8SbOiYHoc3v9NvaZTB7p96hUySMlIo=
github.com/tarantool/go-option v1.1.0 h1:ShoOhNsdL41sRpm4hXCRDjV8H0WzPkd4UnKhLKbW//w=
github.com/tarantool/go-option v1.1.0/go.mod h1:hMr9z2JXOWlgdCBpCPSL2nwp8718GKYvNBJ+ZuzJbCo=
github.com/tarantool/go-tarantool/v3 v3.0.0 h1:zsIXS4nvSSXqZWtN/f1Bfuq973j6J85O5U/8RYQCRcE=
github.com/tarantool/go-tarantool/v3 v3.0.0/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/tarantool/go-tarantool/v3 v3.0.1 h1:vaUX4xmVmXh2dIJ/LqlX1MXK3iYqAqV6YiE54Wwl/qg=
github.com/tarantool/go-tarantool/v3 v3.0.1/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/testcontainers/testcontainers-go v0.43.0 h1:oEQx5MW2DGd9z3AeEQfB2lPM0eLs7ztyaGRu75bFo5A=
github.com/testcontainers/testcontainers-go v0.43.0/go.mod h1:+VxkT2NQnKOZPKi6praMuMKYHYyOGXr0XSBSlSMCzFo=
github.com/testcontainers/testcontainers-go/modules/compose v0.42.0 h1:+t1ZN31TD36cwxmeLqGwe7wIdvblBm0Z+vlj4SX8Mv0=
@@ -2030,14 +2032,14 @@ github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e h1:9LPdmD
github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e/go.mod h1:HEUYX/p8966tMUHHT+TsS0hF/Ca/NYwqprC5WXSDMfE=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20221215182650-986f9d10542f/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20230528143953-42c825ace222/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b h1:xeiobG1riqe6dTMZuTcOvwOY0BbFidSk3gRLyVeLfJA=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705 h1:7VKlOrBIQ8L8acJ9wFCt6lWzLDbOhc05VLpL31743tU=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2 h1:e2nGQPGC5OEPBWlMnLPFpdgyqNA8iTOPiZShCrv6944=
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2/go.mod h1:9YzkhlIymWaJGX6KMU3vh5sOf3UKbCXkG/ZdjaI3zNM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.44.0/go.mod h1:oSLwnuilwIpaF5bJJMAofnGgzPJusoI3zWMNb8I+GnM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.47.3/go.mod h1:bWnOIcUHd7+Sl7DN+yhyY1H/I61z53GczvwJgXMgvj0=
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1 h1:WRxyl1UdFD6EmFlujwrHm4sX9ktYb/adZ/SW+WWNOEg=
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1/go.mod h1:b9NEO6mgaiqsnOMkS003uS82XsKh6GL+ZTFfPqXWz+c=
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1 h1:T+fB2ZDHpYIGC7DWjK+rzZHLoexBF0zG/n3Q9DGNskY=
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1/go.mod h1:dJXJ1u00IqO8Vsph8fWmYj8b1E7jyphnvJPqA69emBY=
github.com/ydb-platform/ydb-go-yc v0.12.1 h1:qw3Fa+T81+Kpu5Io2vYHJOwcrYrVjgJlT6t/0dOXJrA=
github.com/ydb-platform/ydb-go-yc v0.12.1/go.mod h1:t/ZA4ECdgPWjAb4jyDe8AzQZB5dhpGbi3iCahFaNwBY=
github.com/ydb-platform/ydb-go-yc-metadata v0.6.1 h1:9E5q8Nsy2RiJMZDNVy0A3KUrIMBPakJ2VgloeWbcI84=
@@ -2657,8 +2659,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.293.0 h1:p9XIWOf63U4OgYx120ZwVU8+vl4XTPmWfgVPnmOAS9w=
google.golang.org/api v0.293.0/go.mod h1:6n5tjEB1gzwniZTepZ0g5u+wM7Bof5GeULCx/zh8ZE0=
google.golang.org/api v0.294.0 h1:8gASjJxdtcIieB3OqbkLcF0FfbXVNqKtU5iozD1ssvA=
google.golang.org/api v0.294.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2792,12 +2794,12 @@ google.golang.org/genproto v0.0.0-20230209215440-0dfe4f8abfcc/go.mod h1:RGgjbofJ
google.golang.org/genproto v0.0.0-20230216225411-c8e22ba71e44/go.mod h1:8B0gmkoRebU8ukX6HP+4wrVQUY1+6PkQ44BSyIlflHA=
google.golang.org/genproto v0.0.0-20230222225845-10f96fb3dbec/go.mod h1:3Dl5ZL0q0isWJt+FVcfpQyirqemEuLAK/iFvg1UP1Hw=
google.golang.org/genproto v0.0.0-20230306155012-7f2fa6fef1f4/go.mod h1:NWraEVixdDnqcqQ30jipen1STv2r/n24Wb7twVTGR4s=
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 h1:YJjbgu+dkp5kUJLfpMyCLfBIWZb/FcJyuLeo1gVBOuo=
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94/go.mod h1:RRHjglSYABVCWpQ7USCpdfhcd9t4PkajvVwyynZizTc=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 h1:admdQBe8jR3VWhBsUrAOaF2Qw6K/+p5pSm1GN8+6Fw4=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800/go.mod h1:FPk7EXUKMtImne7AmknoYjT4QXqKIzzRbeQIXzLk6fQ=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea h1:kVhQEPTpKQahD5+JSBTfBB19wcgQTTjAIn45MBqnyHk=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d h1:C9v1o0/4quuhOAfmRXA2j+we0PqZIp8traLdeogF3Ms=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d/go.mod h1:Wz2wFJntZFmLGo7pLDXZ3wYk5hyc0Mb+SkHhDDXT+lU=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d h1:QwnJwPte4XXAkhPu26LTDIahnsMSUV0kK8HkxbC+Pc4=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d/go.mod h1:WRrQ7/7N19PypuT0fxLOL5Lq0waoiRri4FbtHDEKrGE=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 h1:cYNAzI2sUwhmCcoj9TxvihSrqsxt6uIkj3rDRhSDmW4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
google.golang.org/grpc v1.20.1/go.mod h1:10oTOabMzJvdu6/UiuZezV6QK5dSlG84ov/aaiqXj38=
google.golang.org/grpc v1.21.1/go.mod h1:oYelfM1adQP15Ek0mdvEgi9Df8B9CZIaU1084ijfRaM=
@@ -2838,8 +2840,8 @@ google.golang.org/grpc v1.51.0/go.mod h1:wgNDFcnuBGmxLKI/qn4T+m5BtEBYXJPvibbUPsA
google.golang.org/grpc v1.52.0/go.mod h1:pu6fVzoFb+NBYNAvQL08ic+lvB2IojljRYuun5vorUY=
google.golang.org/grpc v1.53.0/go.mod h1:OnIrk0ipVdj4N5d9IUoFUx72/VlD7+jUsHwZgwSMQpw=
google.golang.org/grpc v1.55.0/go.mod h1:iYEXKGkEBhg1PjZQvoYEVPTDkHo1/bjTnfwTeGONTY8=
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6 h1:HfjjkdGIa8u9sP9EW5WCygy0kQDuTI/Tax4j//t24Fo=
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc v1.85.0-dev h1:HxkDyKIIZPpFnroC56tQv5gNuKTmVvi0t7TzOf5zt7g=
google.golang.org/grpc v1.85.0-dev/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc/cmd/protoc-gen-go-grpc v1.1.0/go.mod h1:6Kw0yEErY5E/yWrBtf03jp27GLLJujG4z/JK95pnjjw=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6 h1:ExN12ndbJ608cboPYflpTny6mXSzPrDLh0iTaVrRrds=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6/go.mod h1:6ytKWczdvnpnO+m+JiG9NjEDzR1FJfsnmJdG7B8QVZ8=
@@ -2861,8 +2863,8 @@ google.golang.org/protobuf v1.27.1/go.mod h1:9q0QmTI4eRPtz6boOQmLYwt+qCgq0jsYwAQ
google.golang.org/protobuf v1.28.0/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.28.1/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.30.0/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
google.golang.org/protobuf v1.36.12 h1:pJOKDDOyeXErUroCihFAd5LQuwXBSpVnKGrj5o/fwxc=
google.golang.org/protobuf v1.36.12/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/alecthomas/kingpin.v2 v2.2.6/go.mod h1:FMv+mEhP44yOT+4EoQTLFTRgOQ1FBLkstjWtayDeSgw=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20180628173108-788fd7840127/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
@@ -2962,8 +2964,8 @@ modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.18.1/go.mod h1:6ho+Gow7oX5V+OiOQ6Tr4xeqbx13UZ6t+Fw9IRUG4d4=
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/sqlite v1.57.0 h1:qNQP6xnx5M0ISNtlnxoOX0+cD5bJ0/gr9aMmndFczzg=
modernc.org/sqlite v1.57.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/strutil v1.1.0/go.mod h1:lstksw84oURvj9y3tn8lGvRxyRC1S2+g5uuIzNfIOBs=
modernc.org/strutil v1.1.1/go.mod h1:DE+MQQ/hjKBZS2zNInV5hhcipt5rLPWkmpbGeW5mmdw=
modernc.org/strutil v1.1.3/go.mod h1:MEHNA7PdEnEwLvspRMtWTNnp2nnyvMfkimT1NKNAGbw=
+84 -2
View File
@@ -9,7 +9,7 @@
# curl -fsSL ... | bash -s -- --version 4.34 --dir /usr/local/bin
#
# Options:
# --component COMP Which binary to install: weed, volume-rust, all (default: weed)
# --component COMP Which binary to install: weed, volume-rust, worker-rust, all (default: weed)
# --version VER Release version tag (default: latest)
# --large-disk Use large disk variant (5-byte offset, 8TB max volume)
# --dir DIR Installation directory (default: /usr/local/bin)
@@ -22,6 +22,7 @@ COMPONENT="weed"
VERSION=""
LARGE_DISK=false
INSTALL_DIR="/usr/local/bin"
WORKER_INSTALLED=false
# Colors (if terminal supports them)
if [ -t 1 ]; then
@@ -128,6 +129,33 @@ rust_asset_name() {
fi
}
# Does a release carry this asset? Answers 0 for yes and 1 for a 404, and stops
# the installer on anything else: a rate limit or a network blip must not read
# as "this release predates the binary" and quietly skip it.
asset_exists() {
local url="$1" code=""
if command -v curl &>/dev/null; then
code="$(curl -sL -o /dev/null -I -w '%{http_code}' "$url" || true)"
elif command -v wget &>/dev/null; then
# --spider's exit status folds 404 in with every other server error, so
# read the status line itself; -S prints one per redirect hop.
code="$(wget -S --spider -q -O /dev/null "$url" 2>&1 | awk '/^ *HTTP\// {c=$2} END {print c}')"
fi
case "$code" in
200) return 0 ;;
404) return 1 ;;
*) error "Could not check ${url} (HTTP ${code:-none}). Retry, or install components one at a time." ;;
esac
}
# Build Rust maintenance worker asset name. No large-disk variant: the worker
# maintains tables through the namespace and never opens a volume file.
worker_asset_name() {
local os="$1" arch="$2"
echo "weed-worker_${os}_${arch}.tar.gz"
}
# Install a single component
install_component() {
local component="$1" os="$2" arch="$3"
@@ -204,10 +232,48 @@ install_component() {
ok "Installed weed-volume to ${INSTALL_DIR}/${dest_name}"
;;
worker-rust)
# Published for linux only: the worker runs beside the cluster it
# maintains, and its dependency tree makes every extra target an
# expensive build.
case "$os" in
linux) ;;
*) error "Rust maintenance worker is not available for ${os}. Supported: linux" ;;
esac
case "$arch" in
amd64|arm64) ;;
*) error "Rust maintenance worker is not available for ${arch}. Supported: amd64, arm64" ;;
esac
asset_name="$(worker_asset_name "$os" "$arch")"
download_url="https://github.com/${REPO}/releases/download/${VERSION}/${asset_name}"
download "$download_url" "${tmpdir}/${asset_name}"
info "Extracting ${asset_name}..."
tar xzf "${tmpdir}/${asset_name}" -C "$tmpdir"
local worker_bin
worker_bin="$(find "$tmpdir" -name 'weed-worker' -type f | head -1)"
if [ -z "$worker_bin" ]; then
error "Could not find weed-worker binary in archive"
fi
chmod +x "$worker_bin"
install_binary "$worker_bin" "weed-worker"
WORKER_INSTALLED=true
ok "Installed weed-worker to ${INSTALL_DIR}/weed-worker"
;;
*)
error "Unknown component: ${component}. Use: weed, volume-rust, all"
error "Unknown component: ${component}. Use: weed, volume-rust, worker-rust, all"
;;
esac
# The trap is per-process, so a later component's would replace this one and
# leave the earlier extraction behind. Clean up here and hand the trap back;
# the error paths above exit, which still fires it.
rm -rf "$tmpdir"
trap - EXIT
}
# Copy binary to install dir, using sudo if needed
@@ -251,6 +317,16 @@ main() {
all)
install_component "weed" "$os" "$arch"
install_component "volume-rust" "$os" "$arch"
# The worker is published for linux amd64/arm64 only, and only by
# releases new enough to carry it; skip either case rather than fail
# an install that has already put two binaries in place.
if [ "$os" != "linux" ] || { [ "$arch" != "amd64" ] && [ "$arch" != "arm64" ]; }; then
warn "Skipping the Rust maintenance worker: no build for ${os}/${arch}"
elif ! asset_exists "https://github.com/${REPO}/releases/download/${VERSION}/$(worker_asset_name "$os" "$arch")"; then
warn "Skipping the Rust maintenance worker: ${VERSION} does not carry one"
else
install_component "worker-rust" "$os" "$arch"
fi
;;
*)
install_component "$COMPONENT" "$os" "$arch"
@@ -265,11 +341,17 @@ main() {
if [ "$COMPONENT" = "volume-rust" ] || [ "$COMPONENT" = "all" ]; then
info " weed-volume: ${INSTALL_DIR}/weed-volume"
fi
if [ "$WORKER_INSTALLED" = true ]; then
info " weed-worker: ${INSTALL_DIR}/weed-worker"
fi
echo ""
info "Quick start:"
info " weed master # Start master server"
info " weed volume -mserver=localhost:9333 # Start Go volume server"
info " weed-volume -mserver localhost:9333 # Start Rust volume server"
if [ "$WORKER_INSTALLED" = true ]; then
info " weed-worker --admin localhost:23646 # Start the Rust maintenance worker"
fi
}
main
@@ -21,6 +21,7 @@ spec:
{{- $nodePorts = .Values.admin.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.admin.service }}
ports:
- name: "http"
port: {{ .Values.admin.port }}
@@ -21,6 +21,7 @@ spec:
{{- $nodePorts = .Values.allInOne.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.allInOne.service }}
internalTrafficPolicy: {{ .Values.allInOne.service.internalTrafficPolicy | default "Cluster" }}
{{- if and (semverCompare ">=1.31-0" .Capabilities.KubeVersion.GitVersion) .Values.allInOne.s3.trafficDistribution }}
trafficDistribution: {{ include "seaweedfs.trafficDistribution" (dict "value" .Values.allInOne.s3.trafficDistribution "Capabilities" .Capabilities) }}
@@ -149,6 +149,8 @@ spec:
{{- if .Values.s3.icebergPort }}
-port.iceberg={{ .Values.s3.icebergPort }} \
{{- end }}
{{- /* rendered even when 0: an if would drop the flag and weed would serve its default */}}
-port.lance={{ .Values.s3.lancePort | default 0 }} \
{{- range .Values.s3.extraArgs }}
{{ . }} \
{{- end }}
@@ -198,6 +200,10 @@ spec:
- containerPort: {{ .Values.s3.icebergPort }}
name: swfs-iceberg
{{- end }}
{{- if .Values.s3.lancePort }}
- containerPort: {{ .Values.s3.lancePort }}
name: swfs-lance
{{- end }}
{{- if .Values.s3.metricsPort }}
- containerPort: {{ .Values.s3.metricsPort }}
name: metrics
@@ -0,0 +1,61 @@
{{- define "seaweedfs.s3.lance.ingress.paths" -}}
paths:
- path: {{ .Values.s3.lanceIngress.path | quote }}
pathType: {{ .Values.s3.lanceIngress.pathType | quote }}
backend:
{{- if semverCompare ">=1.19-0" .Capabilities.KubeVersion.GitVersion }}
service:
name: {{ include "seaweedfs.componentName" (list . "s3") }}
port:
number: {{ .Values.s3.lancePort }}
{{- else }}
serviceName: {{ include "seaweedfs.componentName" (list . "s3") }}
servicePort: {{ .Values.s3.lancePort }}
{{- end }}
{{- end -}}
{{- if and .Values.s3.enabled .Values.s3.lancePort .Values.s3.lanceIngress.enabled }}
{{- $hosts := list }}
{{- if kindIs "slice" .Values.s3.lanceIngress.host }}
{{- $hosts = .Values.s3.lanceIngress.host }}
{{- else if .Values.s3.lanceIngress.host }}
{{- $hosts = list .Values.s3.lanceIngress.host }}
{{- end }}
{{- if semverCompare ">=1.19-0" .Capabilities.KubeVersion.GitVersion }}
apiVersion: networking.k8s.io/v1
{{- else if semverCompare ">=1.14-0" .Capabilities.KubeVersion.GitVersion }}
apiVersion: networking.k8s.io/v1beta1
{{- else }}
apiVersion: extensions/v1beta1
{{- end }}
kind: Ingress
metadata:
name: ingress-{{ include "seaweedfs.fullname" . }}-s3-lance
namespace: {{ .Release.Namespace }}
{{- with .Values.s3.lanceIngress.annotations }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
labels:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3-lance
spec:
{{- if .Values.s3.lanceIngress.className }}
ingressClassName: {{ .Values.s3.lanceIngress.className | quote }}
{{- end }}
tls:
{{ .Values.s3.lanceIngress.tls | default list | toYaml | nindent 6}}
rules:
{{- if $hosts }}
{{- range $host := $hosts }}
- host: {{ $host | quote }}
http:
{{- include "seaweedfs.s3.lance.ingress.paths" $ | nindent 6 }}
{{- end }}
{{- else }}
- http:
{{- include "seaweedfs.s3.lance.ingress.paths" . | nindent 4 }}
{{- end }}
{{- end }}
@@ -21,6 +21,7 @@ spec:
{{- $nodePorts = .Values.s3.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.s3.service }}
internalTrafficPolicy: {{ .Values.s3.internalTrafficPolicy | default "Cluster" }}
{{- $td := .Values.s3.trafficDistribution | default .Values.filer.s3.trafficDistribution }}
{{- if and (semverCompare ">=1.31-0" .Capabilities.KubeVersion.GitVersion) $td }}
@@ -43,6 +44,15 @@ spec:
{{- end }}
protocol: TCP
{{- end }}
{{- if and .Values.s3.enabled .Values.s3.lancePort }}
- name: "swfs-lance"
port: {{ .Values.s3.lancePort }}
targetPort: {{ .Values.s3.lancePort }}
{{- if $nodePorts.lance }}
nodePort: {{ $nodePorts.lance }}
{{- end }}
protocol: TCP
{{- end }}
{{- if and .Values.s3.enabled .Values.s3.httpsPort }}
- name: "swfs-s3-tls"
port: {{ .Values.s3.httpsPort }}
@@ -21,6 +21,7 @@ spec:
{{- $nodePorts = .Values.sftp.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.sftp.service }}
internalTrafficPolicy: {{ .Values.sftp.internalTrafficPolicy | default "Cluster" }}
ports:
- name: "swfs-sftp"
@@ -148,6 +148,15 @@ true
{{- end -}}
{{- end -}}
{{/* Lance namespace URL the worker's Lance container maintains; empty when unreachable */}}
{{- define "seaweedfs.worker.lanceNamespaceUrl" -}}
{{- if .Values.worker.namespaceUrl -}}
{{- .Values.worker.namespaceUrl -}}
{{- else if and .Values.s3.enabled .Values.s3.lancePort -}}
{{- printf "http://%s.%s:%d" (include "seaweedfs.componentName" (list . "s3")) .Release.Namespace (int .Values.s3.lancePort) -}}
{{- end -}}
{{- end -}}
{{/* Return the proper volume image */}}
{{- define "seaweedfs.volume.image" -}}
{{- if .Values.volume.imageOverride -}}
@@ -548,3 +557,23 @@ true
{{- and (eq .value "PreferClose") (semverCompare ">=1.35-0" .Capabilities.KubeVersion.GitVersion) | ternary "PreferSameZone" .value -}}
{{- end -}}
{{- end -}}
{{/*
Render LoadBalancer-specific service fields (loadBalancerClass, loadBalancerIP,
loadBalancerSourceRanges), only when the service type is LoadBalancer.
Usage: {{ include "seaweedfs.service.loadBalancerFields" .Values.s3.service }}
*/}}
{{- define "seaweedfs.service.loadBalancerFields" -}}
{{- if eq (.type | default "ClusterIP") "LoadBalancer" }}
{{- with .loadBalancerClass }}
loadBalancerClass: {{ . }}
{{- end }}
{{- with .loadBalancerIP }}
loadBalancerIP: {{ . }}
{{- end }}
{{- with .loadBalancerSourceRanges }}
loadBalancerSourceRanges:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- end }}
{{- end -}}
@@ -73,6 +73,9 @@
{{- if .Values.s3.icebergPort }}
{{- $ports = append $ports .Values.s3.icebergPort }}
{{- end }}
{{- if .Values.s3.lancePort }}
{{- $ports = append $ports .Values.s3.lancePort }}
{{- end }}
{{- if .Values.s3.metricsPort }}
{{- $ports = append $ports .Values.s3.metricsPort }}
{{- end }}
@@ -97,6 +100,9 @@
{{- if .Values.worker.metricsPort }}
{{- $ports = append $ports .Values.worker.metricsPort }}
{{- end }}
{{- if and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- $ports = append $ports .Values.worker.lanceMetricsPort }}
{{- end }}
{{- $targets = append $targets (dict "component" "worker" "ports" $ports) }}
{{- end }}
@@ -223,6 +223,81 @@ spec:
{{- if .Values.worker.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.worker.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- if include "seaweedfs.worker.lanceNamespaceUrl" . }}
- name: worker-lance
image: {{ template "seaweedfs.worker.image" . }}
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
- name: POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
{{- /* the URL crosses a shell line; metacharacters must arrive as data, not syntax */}}
- name: LANCE_NAMESPACE_URL
value: {{ include "seaweedfs.worker.lanceNamespaceUrl" . | quote }}
command:
- "/bin/sh"
- "-ec"
- |
{{- /* the armv7/386 placeholder is empty; exec of it becomes the shell and exits 0 */}}
if [ ! -s /usr/bin/weed-worker ]; then
echo "the Rust worker is not available on this platform ($(uname -m)); it ships for amd64 and arm64" >&2
exit 1
fi
exec /usr/bin/weed-worker \
--id="$POD_NAME" \
{{- if .Values.worker.adminServer }}
--admin={{ .Values.worker.adminServer }} \
{{- else }}
--admin={{ template "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}:{{ .Values.admin.port }}{{ if .Values.admin.grpcPort }}.{{ .Values.admin.grpcPort }}{{ end }} \
{{- end }}
--namespace="$LANCE_NAMESPACE_URL" \
{{- if .Values.global.seaweedfs.enableSecurity }}
--tls-ca=/usr/local/share/ca-certificates/ca/tls.crt \
--tls-cert=/usr/local/share/ca-certificates/worker/tls.crt \
--tls-key=/usr/local/share/ca-certificates/worker/tls.key \
{{- end }}
{{- if .Values.worker.lanceMetricsPort }}
--metrics-port={{ .Values.worker.lanceMetricsPort }} \
--metrics-ip=0.0.0.0 \
{{- end }}
--max-concurrency={{ .Values.worker.maxExecute }}
{{- if .Values.global.seaweedfs.enableSecurity }}
volumeMounts:
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
- name: worker-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/worker/
{{- end }}
{{- if .Values.worker.lanceMetricsPort }}
ports:
- containerPort: {{ .Values.worker.lanceMetricsPort }}
name: lance-metrics
livenessProbe:
httpGet:
path: /health
port: lance-metrics
initialDelaySeconds: 30
periodSeconds: 60
successThreshold: 1
failureThreshold: 5
timeoutSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: lance-metrics
initialDelaySeconds: 20
periodSeconds: 15
successThreshold: 1
failureThreshold: 3
timeoutSeconds: 10
{{- end }}
{{- if .Values.worker.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.worker.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- end }}
{{- if .Values.worker.sidecars }}
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.worker.sidecars "context" $) | nindent 8 }}
{{- end }}
@@ -12,13 +12,22 @@ metadata:
app.kubernetes.io/component: worker
spec:
clusterIP: None # Headless service
{{- if .Values.worker.metricsPort }}
{{- $lanceMetrics := and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- if or .Values.worker.metricsPort $lanceMetrics }}
ports:
{{- if .Values.worker.metricsPort }}
- name: "metrics"
port: {{ .Values.worker.metricsPort }}
targetPort: {{ .Values.worker.metricsPort }}
protocol: TCP
{{- end }}
{{- if $lanceMetrics }}
- name: "lance-metrics"
port: {{ .Values.worker.lanceMetricsPort }}
targetPort: {{ .Values.worker.lanceMetricsPort }}
protocol: TCP
{{- end }}
{{- end }}
selector:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
@@ -1,6 +1,7 @@
{{- include "seaweedfs.compat" . -}}
{{- if .Values.worker.enabled }}
{{- if .Values.worker.metricsPort }}
{{- $lanceMetrics := and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- if or .Values.worker.metricsPort $lanceMetrics }}
{{- if .Values.global.seaweedfs.monitoring.enabled }}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
@@ -22,9 +23,16 @@ metadata:
{{- end }}
spec:
endpoints:
{{- if .Values.worker.metricsPort }}
- interval: 30s
port: metrics
scrapeTimeout: 5s
{{- end }}
{{- if $lanceMetrics }}
- interval: 30s
port: lance-metrics
scrapeTimeout: 5s
{{- end }}
selector:
matchLabels:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
+38
View File
@@ -1006,6 +1006,9 @@ s3:
# Iceberg catalog REST port (Apache Iceberg REST Catalog API)
# Set to a port number to enable, or 0/null to disable
icebergPort: null
# Lance Namespace port; weed serves 9101 by default, 0 disables it
# (and, unless worker.namespaceUrl points elsewhere, the worker's Lance container)
lancePort: 9101
loggingOverrideLevel: null
# enable user & permission to s3 (need to inject to all services)
enableAuth: false
@@ -1184,11 +1187,16 @@ s3:
# Service settings
service:
type: ClusterIP
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
http: null
https: null
iceberg: null
lance: null
metrics: null
icebergIngress:
@@ -1200,6 +1208,15 @@ s3:
annotations: {}
tls: []
lanceIngress:
enabled: false
className: ""
host: "seaweedfs-lance.cluster.local"
path: "/"
pathType: Prefix
annotations: {}
tls: []
sftp:
enabled: false
imageOverride: null
@@ -1288,6 +1305,10 @@ sftp:
# Service settings
service:
type: ClusterIP
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
sftp: null
@@ -1430,6 +1451,10 @@ admin:
service:
type: ClusterIP
annotations: {}
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
http: null
@@ -1448,6 +1473,15 @@ worker:
metricsPort: 9327
metricsIp: "" # If empty, defaults to 0.0.0.0
# The lance_* jobs run in their own container, /usr/bin/weed-worker.
# amd64/arm64 only; pin mixed clusters with worker.affinity/nodeSelector.
# Lance namespace URL override; empty derives it from s3.lancePort.
namespaceUrl: ""
# Metrics port for the Lance worker container; the Go worker keeps metricsPort
lanceMetricsPort: 9328
# Admin server to connect to
adminServer: ""
@@ -1654,6 +1688,10 @@ allInOne:
annotations: {} # Annotations for the service
type: ClusterIP # Service type (ClusterIP, NodePort, LoadBalancer)
internalTrafficPolicy: Cluster # Internal traffic policy
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
master: null
+34
View File
@@ -27,6 +27,8 @@ service Seaweed {
}
rpc VolumeList (VolumeListRequest) returns (VolumeListResponse) {
}
rpc VolumeListStream (VolumeListRequest) returns (stream VolumeListStreamResponse) {
}
rpc LookupEcVolume (LookupEcVolumeRequest) returns (LookupEcVolumeResponse) {
}
rpc VacuumVolume (VacuumVolumeRequest) returns (VacuumVolumeResponse) {
@@ -406,12 +408,44 @@ message TopologyInfo {
map<string, DiskInfo> diskInfos = 3;
}
message VolumeListRequest {
// Empty and zero take everything. Only the volumes and ec shards listed
// under a disk are selected; the topology and its disk counters are always
// reported in full.
string collection = 1;
repeated uint32 volume_ids = 2;
// The one collection the empty string cannot name. A named collection wins.
bool default_collection_only = 3;
// Empty and zero take everything. Wildcards are supported.
string remote_storage_name = 4;
bool local_volume_only = 5;
// The topology, its disks and their counters alone, without the volumes
// and ec shards. Selecting volumes above contradicts this and is refused.
// A master that predates this field ignores it and answers in full.
bool topology_only = 6;
}
message VolumeListResponse {
TopologyInfo topology_info = 1;
uint64 volume_size_limit_mb = 2;
}
// VolumeListStream answers the same request as VolumeList without either end
// holding every volume in the cluster at once. At 800k volumes the reply is
// 36MB on the wire but 305MB as messages, which the master built in full
// before sending any of it.
message VolumeListStreamResponse {
// Sent once, first, listing no volumes: the topology, its disks and their
// counters. Every message after carries volumes for one of those disks.
VolumeListResponse header = 1;
// Which disk this batch is from. A disk arrives over as many batches as it
// takes, so append rather than assign.
string data_center = 2;
string rack = 3;
string data_node = 4;
string disk_type = 5;
repeated VolumeInformationMessage volume_infos = 6;
repeated VolumeEcShardInformationMessage ec_shard_infos = 7;
}
message LookupEcVolumeRequest {
uint32 volume_id = 1;
}
+3
View File
@@ -539,6 +539,7 @@ message VolumeEcShardsInfoResponse {
uint64 volume_size = 2;
uint64 file_count = 3;
uint64 file_deleted_count = 4;
EcShardConfig ec_shard_config = 10; // the layout this holder serves reads through; a binary predating a field reports it as unset
}
message EcShardInfo {
@@ -612,6 +613,7 @@ message EcShardConfig {
uint32 data_shards = 1; // Number of data shards (e.g., 10)
uint32 parity_shards = 2; // Number of parity shards (e.g., 4)
int64 encode_ts_ns = 3; // encode time (unix nanos); a read served from a shard of a different encode run is rejected
int64 block_size = 4; // uniform block layout: each shard is a single contiguous block of this many bytes; 0 = legacy 1GiB/1MiB two-tier layout
}
// EcBitrotProtection is the entire content of a bitrot checksum sidecar
// (<base>.ecsum for the legacy generation, <base>.ecsum.v<N> for vacuum
@@ -714,6 +716,7 @@ enum VolumeScrubMode {
FULL = 2;
LOCAL = 3;
CHECKSUM = 4; // EC only: verify each local shard's raw bytes against the bitrot checksum sidecar
READS = 5; // like FULL, but EC intervals no shard can serve are reconstructed from parity
}
message ScrubVolumeRequest {
+432 -62
View File
@@ -38,6 +38,7 @@ fn scrub_mode_label(mode: i32) -> &'static str {
2 => "FULL",
3 => "LOCAL",
4 => "CHECKSUM",
5 => "READS",
_ => "UNKNOWN",
}
}
@@ -91,6 +92,57 @@ pub fn load_state_file(
volume_server_pb::VolumeServerState::decode(data.as_slice()).ok()
}
/// One disk location's stake in an EC volume, as seen by the rebuild handler.
struct LocInfo {
dir: String,
idx_dir: String,
shard_count: usize,
has_ecx: bool,
}
/// Picks the location a rebuild should write into — the one holding an `.ecx`
/// and the most shards — and returns every other directory it may have to read
/// from. Shards are only half of what the rebuild needs: a split
/// `-dir`/`-dir.idx` layout keeps `.ecx`/`.ecj`/`.vif` with the INDEX, and on a
/// multi-disk server the chosen disk may hold nothing but shards while this
/// volume's `.vif` or generation-0 `.ecsum` sits on a sibling. Miss those and
/// the layout resolution falls back to 10+4 with the legacy striping and
/// reconstructs through the wrong matrix, so both directories of every other
/// location are listed. The rebuild's own two are passed separately by the
/// caller and dropped here, along with empties and duplicates.
///
/// Returns `None` when no location holds an `.ecx`, i.e. there is nothing to
/// rebuild from.
fn select_rebuild_location(loc_infos: &[LocInfo]) -> Option<(usize, Vec<String>)> {
let mut rebuild_loc_idx: Option<usize> = None;
let mut other_dirs: Vec<String> = Vec::new();
for (i, info) in loc_infos.iter().enumerate() {
let better = info.has_ecx
&& rebuild_loc_idx
.is_none_or(|prev| info.shard_count > loc_infos[prev].shard_count);
if better {
if let Some(prev) = rebuild_loc_idx {
other_dirs.push(loc_infos[prev].dir.clone());
other_dirs.push(loc_infos[prev].idx_dir.clone());
}
rebuild_loc_idx = Some(i);
} else {
other_dirs.push(info.dir.clone());
other_dirs.push(info.idx_dir.clone());
}
}
let rebuild_loc_idx = rebuild_loc_idx?;
let rebuild_dir = &loc_infos[rebuild_loc_idx].dir;
let rebuild_idx_dir = &loc_infos[rebuild_loc_idx].idx_dir;
other_dirs.retain(|d| !d.is_empty() && d != rebuild_dir && d != rebuild_idx_dir);
other_dirs.sort();
other_dirs.dedup();
Some((rebuild_loc_idx, other_dirs))
}
struct WriteThrottler {
bytes_per_second: i64,
last_size_counter: i64,
@@ -1768,11 +1820,14 @@ impl VolumeServer for VolumeGrpcService {
}
Some(store.locations[info.disk_id as usize].directory.clone())
} else {
// The mounted-volume refusal above means no disk
// holds an in-memory claim here.
store
.find_ec_shard_target_location(
&info.collection,
vid,
DATA_SHARDS_COUNT as u32,
&[],
)
.map(|i| store.locations[i].directory.clone())
};
@@ -2382,13 +2437,21 @@ impl VolumeServer for VolumeGrpcService {
)
};
// Check existing .vif for EC shard config (matching Go's MaybeLoadVolumeInfo)
let (data_shards, parity_shards) =
// Check existing .vif for EC shard config (matching Go's MaybeLoadVolumeInfo).
// The block size is recomputed by the encode for the current .dat, so
// only the ratio is carried over from a prior config.
let (data_shards, parity_shards, _) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config(
&dir, &idx_dir, collection, vid,
);
)
.map_err(|e| {
tonic::Status::internal(format!(
"read ec shard config for volume {}: {}",
vid.0, e
))
})?;
if let Err(e) = crate::storage::erasure_coding::ec_encoder::write_ec_files(
let block_size = match crate::storage::erasure_coding::ec_encoder::write_ec_files(
&dir,
&idx_dir,
collection,
@@ -2396,16 +2459,19 @@ impl VolumeServer for VolumeGrpcService {
data_shards as usize,
parity_shards as usize,
) {
// Cleanup partially-created .ecNN and .ecx files on failure (matching Go defer)
let base = crate::storage::volume::volume_file_name(&dir, collection, vid);
let total_shards = data_shards + parity_shards;
for i in 0..total_shards {
let shard_path = format!("{}.ec{:02}", base, i);
let _ = std::fs::remove_file(&shard_path);
Ok(block_size) => block_size,
Err(e) => {
// Cleanup partially-created .ecNN and .ecx files on failure (matching Go defer)
let base = crate::storage::volume::volume_file_name(&dir, collection, vid);
let total_shards = data_shards + parity_shards;
for i in 0..total_shards {
let shard_path = format!("{}.ec{:02}", base, i);
let _ = std::fs::remove_file(&shard_path);
}
let _ = std::fs::remove_file(format!("{}.ecx", base));
return Err(Status::internal(e.to_string()));
}
let _ = std::fs::remove_file(format!("{}.ecx", base));
return Err(Status::internal(e.to_string()));
}
};
// Write .vif file with EC shard metadata
{
@@ -2424,6 +2490,7 @@ impl VolumeServer for VolumeGrpcService {
.duration_since(std::time::UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as i64,
block_size,
}),
..Default::default()
};
@@ -2457,13 +2524,6 @@ impl VolumeServer for VolumeGrpcService {
format!("{}_{}", collection, vid.0)
};
struct LocInfo {
dir: String,
idx_dir: String,
shard_count: usize,
has_ecx: bool,
}
let store = self.state.store.read().unwrap();
let mut loc_infos: Vec<LocInfo> = Vec::new();
@@ -2511,26 +2571,8 @@ impl VolumeServer for VolumeGrpcService {
));
}
// Pick rebuild location: has .ecx and most shards
let mut rebuild_loc_idx: Option<usize> = None;
let mut other_dirs: Vec<String> = Vec::new();
for (i, info) in loc_infos.iter().enumerate() {
if info.has_ecx
&& (rebuild_loc_idx.is_none()
|| info.shard_count > loc_infos[rebuild_loc_idx.unwrap()].shard_count)
{
if let Some(prev) = rebuild_loc_idx {
other_dirs.push(loc_infos[prev].dir.clone());
}
rebuild_loc_idx = Some(i);
} else {
other_dirs.push(info.dir.clone());
}
}
let rebuild_loc_idx = match rebuild_loc_idx {
Some(i) => i,
let (rebuild_loc_idx, other_dirs) = match select_rebuild_location(&loc_infos) {
Some(picked) => picked,
None => {
return Ok(Response::new(
volume_server_pb::VolumeEcShardsRebuildResponse {
@@ -2544,13 +2586,37 @@ impl VolumeServer for VolumeGrpcService {
let rebuild_idx_dir = loc_infos[rebuild_loc_idx].idx_dir.clone();
// Determine data/parity shard config from rebuild dir
let (data_shards, parity_shards) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config(
// The encode-time .dat size resolves the row count the ecx rebuild
// de-stripes with; 0 leaves it to infer from the padded shard extent.
// Both lookups search the sibling disks too: the rebuild writes into one
// location, but a multi-disk server may keep this volume's .vif or its
// generation-0 .ecsum on another, and defaulting to 10+4 with the
// legacy layout would reconstruct through the wrong matrix.
let dat_file_size = crate::storage::erasure_coding::ec_volume::load_vif_info_across_dirs(
&rebuild_dir,
&rebuild_idx_dir,
&other_dirs,
collection,
vid,
)
.ok()
.flatten()
.map(|(v, _)| v.dat_file_size)
.unwrap_or(0);
let (data_shards, parity_shards, block_size) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config_across_dirs(
&rebuild_dir,
&rebuild_idx_dir,
&other_dirs,
collection,
vid,
);
)
.map_err(|e| {
tonic::Status::internal(format!(
"read ec shard config for volume {}: {}",
vid.0, e
))
})?;
let total_shards = data_shards + parity_shards;
// Check which shards are missing (check rebuild dir and all other dirs)
@@ -2589,8 +2655,17 @@ impl VolumeServer for VolumeGrpcService {
// Rebuild missing shards, searching all locations for input shards.
// Pass other_dirs so shards on sibling disks are found even when the
// primary rebuild dir doesn't hold them.
let other_dir_refs: Vec<&str> = other_dirs.iter().map(|s| s.as_str()).collect();
// primary rebuild dir doesn't hold them. This one takes a single flat
// list — the shape Go's RebuildEcFiles uses — so unlike the resolvers
// above it cannot be handed the rebuild's own index directory
// separately, and a split -dir/-dir.idx location keeps its .ecx and
// .vif there. Go's additionalDirs carries that directory for the same
// reason.
let mut rebuild_search_dirs: Vec<String> = other_dirs.clone();
if !rebuild_idx_dir.is_empty() && rebuild_idx_dir != rebuild_dir {
rebuild_search_dirs.push(rebuild_idx_dir.clone());
}
let other_dir_refs: Vec<&str> = rebuild_search_dirs.iter().map(|s| s.as_str()).collect();
crate::storage::erasure_coding::ec_encoder::rebuild_ec_files(
&rebuild_dir,
collection,
@@ -2628,6 +2703,8 @@ impl VolumeServer for VolumeGrpcService {
collection,
vid,
data_shards as usize,
block_size,
dat_file_size,
&ecx_dir_refs,
)
.map_err(|e| Status::internal(format!("RebuildEcxFile: {}", e)))?;
@@ -2652,13 +2729,15 @@ impl VolumeServer for VolumeGrpcService {
// When disk_id > 0: use that specific location.
// When disk_id == 0 (unset): auto-select via
// find_ec_shard_target_location, which prefers a disk that
// already has the EC volume mounted, then a disk that owns the
// .ecx on disk (volume not yet mounted — relevant for
// ec.rebuild, where only the first shard carries .ecx and
// subsequent shards must land on the same disk; see #9212),
// then any HDD, then any disk. Pass the build's default
// data-shard count; the helper takes it as a parameter so
// custom-ratio builds can swap it.
// already owns one of the shards being copied (a retried move
// must overwrite in place, not leave two disks of this server
// claiming the same shard), then a disk that already has the EC
// volume mounted, then a disk that owns the .ecx on disk (volume
// not yet mounted — relevant for ec.rebuild, where only the
// first shard carries .ecx and subsequent shards must land on
// the same disk; see #9212), then any HDD, then any disk. Pass
// the build's default data-shard count; the helper takes it as a
// parameter so custom-ratio builds can swap it.
let (dest_dir, dest_idx_dir) = {
let store = self.state.store.read().unwrap();
let count = store.locations.len();
@@ -2674,10 +2753,27 @@ impl VolumeServer for VolumeGrpcService {
let loc = &store.locations[req.disk_id as usize];
(loc.directory.clone(), loc.idx_directory.clone())
} else {
// A batch whose requested shards are already owned by
// different local disks has no single correct destination:
// writing them all to one disk would duplicate the other
// disks' claims. Refuse so the caller splits the batch per
// shard (or chooses explicitly via disk_id).
let owners = store.ec_shard_owner_disks(vid, &req.shard_ids);
if owners.len() > 1 {
let dirs: Vec<&str> = owners
.iter()
.map(|&i| store.locations[i].directory.as_str())
.collect();
return Err(Status::failed_precondition(format!(
"volume {} shards {:?} are already owned by multiple local disks {:?}: no single destination; copy per shard or pass disk_id",
req.volume_id, req.shard_ids, dirs
)));
}
match store.find_ec_shard_target_location(
&req.collection,
vid,
DATA_SHARDS_COUNT as u32,
&req.shard_ids,
) {
Some(i) => {
let loc = &store.locations[i];
@@ -3014,6 +3110,27 @@ impl VolumeServer for VolumeGrpcService {
Status::internal(format!("mount {}.{}: {}", req.volume_id, shard_id, e))
})?;
}
// A delivery can bring the checksum manifest alongside the shards, but
// the receive path only writes the file. When this server already had
// the volume mounted, the EcVolume in memory keeps whatever protection
// state it resolved at mount — off, for a volume whose sidecar arrives
// now — until a remount. Re-resolve it here, where the shards it
// describes have just been added.
//
// Every per-disk runtime, not just the first: a vid mounts as one
// EcVolume per disk, the delivery lands the .ecsum on one of them, and
// the first-match lookup would leave the siblings reporting no
// protection. Each re-resolves against its own data and index
// directories, so a shared -dir.idx reaches all of them.
// Resolving across every EC metadata directory is what makes that
// reload mean something: startup mirroring gives each shard-bearing
// disk its own .ecx/.ecj/.vif but deliberately not the sidecar, so a
// runtime restricted to its own two directories would find nothing
// however often it reloaded. One delivered copy, reachable from all.
let ec_metadata_dirs = store.ec_metadata_dirs();
for ec_vol in store.find_all_ec_volumes_mut(vid) {
ec_vol.reload_bitrot_sidecar(&ec_metadata_dirs);
}
drop(store);
self.state.volume_state_notify.notify_one();
@@ -3348,6 +3465,8 @@ impl VolumeServer for VolumeGrpcService {
let ecx_dir = ec_vol.ecx_actual_dir().to_string();
let collection = ec_vol.collection.clone();
let vif_dat_file_size = ec_vol.dat_file_size;
let (large_block_size, small_block_size) =
(ec_vol.large_block_size(), ec_vol.small_block_size());
// shard_dirs[i] is guaranteed Some for i in 0..data_shards by
// the check above; collect concrete dirs for the decoder.
let per_shard_dirs: Vec<String> = shard_dirs[..data_shards]
@@ -3381,6 +3500,8 @@ impl VolumeServer for VolumeGrpcService {
vif_dat_file_size,
data_shards,
&per_shard_dirs,
large_block_size as usize,
small_block_size as usize,
)
.map_err(|e| Status::internal(format!("WriteDatFile: {}", e)))?;
@@ -3439,12 +3560,24 @@ impl VolumeServer for VolumeGrpcService {
.walk_ecx_stats()
.map_err(|e| Status::internal(e.to_string()))?;
// The layout this holder serves reads through, as Go reports it: a
// coordinator cannot otherwise tell a holder that understands the
// uniform block layout from one that dropped the unknown .vif field
// and mounted the volume as legacy.
let ec_shard_config = Some(volume_server_pb::EcShardConfig {
data_shards: ec_vol.data_shards,
parity_shards: ec_vol.parity_shards,
encode_ts_ns: 0,
block_size: ec_vol.block_size,
});
Ok(Response::new(
volume_server_pb::VolumeEcShardsInfoResponse {
ec_shard_infos: shard_infos,
volume_size,
file_count,
file_deleted_count,
ec_shard_config,
},
))
}
@@ -3575,7 +3708,12 @@ impl VolumeServer for VolumeGrpcService {
modified_time: dat_modified_secs,
extension: ".dat".to_string(),
});
vol.refresh_remote_write_mode();
vol.refresh_remote_write_mode().map_err(|e| {
Status::internal(format!(
"volume {} failed to refresh write mode: {}",
vid, e
))
})?;
if let Err(e) = vol.save_volume_info() {
return Err(Status::internal(format!(
@@ -3761,10 +3899,40 @@ impl VolumeServer for VolumeGrpcService {
)));
}
if !vol.volume_info.files.is_empty() {
vol.volume_info.files.remove(0);
// Snapshot the remote reference before dropping it: the
// refresh below can fail, and a half-applied transition
// leaves the volume claiming local while the remote backend
// is still attached and the on-disk .vif still says remote
// — a state a retry reads as "already on local disk" and
// refuses to finish.
let removed_remote = if vol.volume_info.files.is_empty() {
None
} else {
Some(vol.volume_info.files.remove(0))
};
// Swaps the read-only sorted map out before the volume is
// published as writable; without it the first write would
// append to the local .dat and then fail to index.
if let Err(e) = vol.refresh_remote_write_mode() {
if let Some(remote) = removed_remote {
vol.volume_info.files.insert(0, remote);
}
// Put the derived flags and the needle map back where
// the restored reference says they belong. Best effort:
// if even this fails the volume stays pinned read-only,
// which is the safe end of the transition.
if let Err(restore_err) = vol.refresh_remote_write_mode() {
tracing::warn!(
volume_id = vid.0,
error = %restore_err,
"tier-down rollback could not restore the remote write mode",
);
}
return Err(Status::internal(format!(
"volume {} failed to refresh write mode: {}",
vid, e
)));
}
vol.refresh_remote_write_mode();
if let Err(e) = vol.save_volume_info() {
return Err(Status::internal(format!(
@@ -4062,7 +4230,7 @@ impl VolumeServer for VolumeGrpcService {
// Validate mode
let mode = req.mode;
match mode {
1 | 2 | 3 => {} // INDEX=1, FULL=2, LOCAL=3
1 | 2 | 3 | 5 => {} // INDEX=1, FULL=2, LOCAL=3, READS=5 (FULL for regular volumes)
_ => {
return Err(Status::invalid_argument(format!(
"unsupported volume scrub mode {}",
@@ -4162,7 +4330,7 @@ impl VolumeServer for VolumeGrpcService {
// Validate mode
let mode = req.mode;
match mode {
1 | 2 | 3 | 4 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4
1 | 2 | 3 | 4 | 5 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4, READS=5
_ => {
return Err(Status::invalid_argument(format!(
"unsupported EC volume scrub mode {}",
@@ -4171,6 +4339,14 @@ impl VolumeServer for VolumeGrpcService {
}
}
// Only the modes that walk needles can be strict about deleted ones.
let force_deleted_needles_check = req.force_deleted_needles_check;
if force_deleted_needles_check && mode != 2 && mode != 5 {
return Err(Status::invalid_argument(
"deleted needle checks are only supported for FULL and READS scrubs",
));
}
// Collect the volume ids under a brief lock, then release it: FULL (mode 2)
// reads remote shards and must not hold the !Send store guard across .await.
let vids: Vec<VolumeId> = {
@@ -4212,8 +4388,8 @@ impl VolumeServer for VolumeGrpcService {
}
}
}
2 => {
// FULL: Go-parity per-needle local+remote walk, PLUS a TEMPORARY
2 | 5 => {
// FULL/READS: Go-parity per-needle local+remote walk, PLUS a TEMPORARY
// local Reed-Solomon parity check. The needle walk only reads
// DATA-shard intervals of LIVE needles, so on its own it can't
// catch silent bitrot in a PARITY shard or an unwalked cold
@@ -4245,8 +4421,13 @@ impl VolumeServer for VolumeGrpcService {
// (1) Per-needle local+remote walk (Go ScrubEcVolume parity).
let (files, mut shard_infos, mut errs) =
crate::server::store_ec::scrub_ec_volume_distributed(&self.state, vid, false)
.await;
crate::server::store_ec::scrub_ec_volume_distributed(
&self.state,
vid,
force_deleted_needles_check,
mode == 5,
)
.await;
total_files += files as u64; // count comes from the needle walk only
// (2) Local parity check, gated on all-shards-local. Blocking RS
@@ -5037,6 +5218,110 @@ mod tests {
use tempfile::TempDir;
use tokio_stream::StreamExt;
fn loc(dir: &str, idx_dir: &str, shard_count: usize, has_ecx: bool) -> LocInfo {
LocInfo {
dir: dir.to_string(),
idx_dir: idx_dir.to_string(),
shard_count,
has_ecx,
}
}
// The rebuild reads its shards from one directory but resolves the volume's
// layout -- ratio and uniform block size -- from the .vif or the
// generation-0 .ecsum, which on a multi-disk server may sit anywhere. Every
// directory that could hold one has to be in the search list, or the
// resolution silently falls back to 10+4 with the legacy striping and
// reconstructs through the wrong matrix.
// The rebuild's own data and index directories are handed to the resolvers
// as their own arguments, so they are deliberately absent from this list --
// unlike Go, whose resolver takes a single directory list and therefore
// carries the rebuild's index directory inside it.
#[test]
fn select_rebuild_location_excludes_the_rebuilds_own_dirs() {
for infos in [
vec![loc("/data1", "/idx1", 3, true)],
vec![loc("/data1", "/data1", 3, true)],
] {
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 0);
assert!(others.is_empty(), "got {:?}", others);
}
}
// The case two reviewers flagged: a sibling holding only shards while its
// index directory holds this volume's .vif.
#[test]
fn select_rebuild_location_searches_a_siblings_index_dir_not_just_its_data_dir() {
let infos = vec![
loc("/data1", "/data1", 5, true),
loc("/data2", "/idx2", 2, false),
];
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 0);
assert_eq!(others, vec!["/data2".to_string(), "/idx2".to_string()]);
}
// Several disks pointed at one index directory is a normal -dir.idx
// deployment; the shared directory is worth searching but only once.
#[test]
fn select_rebuild_location_lists_a_shared_index_dir_once() {
let infos = vec![
loc("/data1", "/data1", 5, true),
loc("/data2", "/shared-idx", 2, false),
loc("/data3", "/shared-idx", 1, false),
];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(
others,
vec![
"/data2".to_string(),
"/data3".to_string(),
"/shared-idx".to_string()
]
);
}
// When the shared index directory is the rebuild's own it drops out, since
// the caller passes it separately.
#[test]
fn select_rebuild_location_omits_a_shared_index_dir_it_rebuilds_into() {
let infos = vec![
loc("/data1", "/shared-idx", 5, true),
loc("/data2", "/shared-idx", 2, false),
];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(others, vec!["/data2".to_string()]);
}
// The winner moves as a fuller location turns up; the one it displaces
// still has to be searched, index directory included.
#[test]
fn select_rebuild_location_keeps_the_displaced_winners_dirs() {
let infos = vec![
loc("/data1", "/idx1", 2, true),
loc("/data2", "/idx2", 9, true),
];
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 1, "the fuller location wins");
assert_eq!(others, vec!["/data1".to_string(), "/idx1".to_string()]);
}
#[test]
fn select_rebuild_location_drops_empty_dirs() {
let infos = vec![loc("/data1", "", 3, true), loc("/data2", "", 1, false)];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(others, vec!["/data2".to_string()]);
}
// Nothing carries an .ecx: there is no index to rebuild the shards against,
// so the caller answers with an empty rebuild rather than guessing.
#[test]
fn select_rebuild_location_is_none_without_an_ecx() {
let infos = vec![loc("/data1", "/idx1", 3, false)];
assert!(select_rebuild_location(&infos).is_none());
}
#[test]
fn test_parse_grpc_address_with_explicit_grpc_port() {
// Format: "ip:port.grpcPort" — used by SeaweedFS for source_data_node
@@ -6053,4 +6338,89 @@ mod tests {
assert!(vif.expire_at_sec >= before + ttl.to_seconds());
assert!(vif.expire_at_sec <= before + ttl.to_seconds() + 5);
}
async fn scrub_ec_volume_1(
service: &VolumeGrpcService,
mode: i32,
) -> volume_server_pb::ScrubEcVolumeResponse {
service
.scrub_ec_volume(Request::new(volume_server_pb::ScrubEcVolumeRequest {
mode,
volume_ids: vec![1],
force_deleted_needles_check: false,
}))
.await
.unwrap()
.into_inner()
}
#[tokio::test]
async fn test_scrub_ec_volume_reads_mode_reconstructs_missing_shard() {
let (service, _tmp) = make_local_service_with_volume("", None);
service
.volume_ec_shards_generate(Request::new(
volume_server_pb::VolumeEcShardsGenerateRequest {
volume_id: 1,
collection: String::new(),
},
))
.await
.unwrap();
service
.volume_ec_shards_mount(Request::new(
volume_server_pb::VolumeEcShardsMountRequest {
volume_id: 1,
collection: String::new(),
shard_ids: (0..14).collect(),
source_disk_type: String::new(),
recover_missing_index: false,
},
))
.await
.unwrap();
// Seed the shard-location cache so the scrub skips the master lookup.
// The explicit-gRPC-port form targets port 1, so remote reads fail fast
// with connection-refused instead of resolving to a live local port.
{
let store = service.state.store.read().unwrap();
let ecv = store.find_ec_volume(VolumeId(1)).unwrap();
let mut locs = ecv.shard_locations.write().unwrap();
for sid in 0u8..14 {
locs.insert(sid, vec!["127.0.0.1:255.1".to_string()]);
}
*ecv.shard_locations_refresh_time.lock().unwrap() =
Some(std::time::Instant::now());
}
// All shards local: FULL is clean.
let resp = scrub_ec_volume_1(&service, 2).await;
assert!(resp.broken_volume_ids.is_empty(), "{:?}", resp.details);
assert_eq!(resp.total_files, 1);
service
.volume_ec_shards_unmount(Request::new(
volume_server_pb::VolumeEcShardsUnmountRequest {
volume_id: 1,
shard_ids: vec![0],
encode_ts_ns: 0,
},
))
.await
.unwrap();
// With a shard unreadable, FULL flags it AND fails the needle...
let resp = scrub_ec_volume_1(&service, 2).await;
assert_eq!(resp.broken_volume_ids, vec![1]);
assert!(resp.broken_shard_infos.iter().any(|s| s.shard_id == 0));
assert!(!resp.details.is_empty());
// ...while READS still reports the shard broken but reconstructs the
// interval from the local survivors, so no needle errors surface.
let resp = scrub_ec_volume_1(&service, 5).await;
assert_eq!(resp.broken_volume_ids, vec![1]);
assert!(resp.broken_shard_infos.iter().any(|s| s.shard_id == 0));
assert!(resp.details.is_empty(), "{:?}", resp.details);
assert_eq!(resp.total_files, 1);
}
}
+50 -4
View File
@@ -1589,12 +1589,15 @@ async fn get_or_head_handler_inner(
}
/// Handle HTTP Range requests. Returns 206 Partial Content or 416 Range Not Satisfiable.
#[derive(Clone, Copy)]
#[derive(Clone, Copy, Debug)]
struct HttpRange {
start: i64,
length: i64,
}
// Returned when the first-byte-pos of every byte-range-spec is at or past the content size.
const RANGE_NO_OVERLAP: &str = "invalid range: failed to overlap";
fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str> {
if s.is_empty() {
return Ok(Vec::new());
@@ -1604,6 +1607,7 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
return Err("invalid range");
}
let mut ranges = Vec::new();
let mut no_overlap = false;
for part in s[PREFIX.len()..].split(',') {
let part = part.trim();
if part.is_empty() {
@@ -1627,9 +1631,13 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
r.length = size - r.start;
} else {
let i = start_str.parse::<i64>().map_err(|_| "invalid range")?;
if i > size || i < 0 {
if i < 0 {
return Err("invalid range");
}
if i >= size {
no_overlap = true;
continue;
}
r.start = i;
if end_str.is_empty() {
r.length = size - r.start;
@@ -1646,6 +1654,9 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
}
ranges.push(r);
}
if no_overlap && ranges.is_empty() {
return Err(RANGE_NO_OVERLAP);
}
Ok(ranges)
}
@@ -1679,7 +1690,15 @@ fn handle_range_request(
let total = data.len() as i64;
let ranges = match parse_range_header(range_str, total) {
Ok(r) => r,
Err(msg) => return range_error_response(headers, msg),
Err(msg) => {
if msg == RANGE_NO_OVERLAP {
headers.insert(
"Content-Range",
format!("bytes */{}", total).parse().unwrap(),
);
}
return range_error_response(headers, msg);
}
};
// Go's ProcessRangeRequest returns nil (empty body) for empty or oversized ranges
@@ -1762,7 +1781,15 @@ fn handle_range_request_from_source(
let total = info.data_size as i64;
let ranges = match parse_range_header(range_str, total) {
Ok(r) => r,
Err(msg) => return range_error_response(headers, msg),
Err(msg) => {
if msg == RANGE_NO_OVERLAP {
headers.insert(
"Content-Range",
format!("bytes */{}", total).parse().unwrap(),
);
}
return range_error_response(headers, msg);
}
};
if ranges.is_empty() {
@@ -3912,6 +3939,25 @@ mod tests {
assert!(parse_url_path("").is_none());
}
#[test]
fn test_parse_range_header_no_overlap() {
assert_eq!(
parse_range_header("bytes=10-", 10).unwrap_err(),
RANGE_NO_OVERLAP
);
assert_eq!(
parse_range_header("bytes=100-", 10).unwrap_err(),
RANGE_NO_OVERLAP
);
// 416 only when every range fails to overlap
let ranges = parse_range_header("bytes=10-,0-1", 10).unwrap();
assert_eq!(ranges.len(), 1);
assert_eq!((ranges[0].start, ranges[0].length), (0, 2));
// an end past the size is clamped, still satisfiable
let ranges = parse_range_header("bytes=5-100", 10).unwrap();
assert_eq!((ranges[0].start, ranges[0].length), (5, 5));
}
#[test]
fn test_extract_jwt_bearer() {
let mut headers = HeaderMap::new();
+56 -26
View File
@@ -398,8 +398,7 @@ async fn do_heartbeat(
// Keep track of what we sent, to generate delta updates
let (initial_hb, initial_volumes) = collect_heartbeat_with_snapshot(config, state);
let mut last_volumes: HashMap<u32, master_pb::VolumeInformationMessage> =
initial_volumes.iter().map(|v| (v.id, v.clone())).collect();
let mut last_volumes: HashMap<u32, VolumeIdentity> = volume_identities(&initial_volumes);
let mut last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -466,8 +465,7 @@ async fn do_heartbeat(
if changed {
let (adjusted_hb, adjusted_volumes) =
collect_heartbeat_with_snapshot(config, state);
last_volumes =
adjusted_volumes.iter().map(|v| (v.id, v.clone())).collect();
last_volumes = volume_identities(&adjusted_volumes);
last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -501,7 +499,7 @@ async fn do_heartbeat(
s.maybe_adjust_volume_max();
}
let (current_hb, current_volumes) = collect_heartbeat_with_snapshot(config, state);
last_volumes = current_volumes.iter().map(|v| (v.id, v.clone())).collect();
last_volumes = volume_identities(&current_volumes);
last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -530,7 +528,7 @@ async fn do_heartbeat(
return Ok(None);
}
let held_volumes = collect_volume_snapshot(config, state);
let current_volumes: HashMap<u32, _> = held_volumes.iter().map(|v| (v.id, v.clone())).collect();
let current_volumes = volume_identities(&held_volumes);
let current_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -541,29 +539,13 @@ async fn do_heartbeat(
for (id, vol) in &current_volumes {
if !last_volumes.contains_key(id) {
new_vols.push(master_pb::VolumeShortInformationMessage {
id: *id,
collection: vol.collection.clone(),
version: vol.version,
replica_placement: vol.replica_placement,
ttl: vol.ttl,
disk_type: vol.disk_type.clone(),
disk_id: vol.disk_id,
});
new_vols.push(vol.to_short_message(*id));
}
}
for (id, vol) in &last_volumes {
if !current_volumes.contains_key(id) {
del_vols.push(master_pb::VolumeShortInformationMessage {
id: *id,
collection: vol.collection.clone(),
version: vol.version,
replica_placement: vol.replica_placement,
ttl: vol.ttl,
disk_type: vol.disk_type.clone(),
disk_id: vol.disk_id,
});
del_vols.push(vol.to_short_message(*id));
}
}
@@ -744,6 +726,54 @@ fn parse_bool_property(value: Option<&String>) -> bool {
.unwrap_or(true)
}
/// What a mount or unmount delta has to name, which is far less than the
/// information message the heartbeat carries. A server holding millions of
/// volumes cannot keep a whole message for each just to notice one leave; the
/// Go report state keeps the same fields for the same reason.
#[derive(Clone)]
struct VolumeIdentity {
collection: String,
disk_type: String,
version: u32,
replica_placement: u32,
ttl: u32,
disk_id: u32,
}
impl VolumeIdentity {
fn of(v: &master_pb::VolumeInformationMessage) -> Self {
Self {
collection: v.collection.clone(),
disk_type: v.disk_type.clone(),
version: v.version,
replica_placement: v.replica_placement,
ttl: v.ttl,
disk_id: v.disk_id,
}
}
fn to_short_message(&self, id: u32) -> master_pb::VolumeShortInformationMessage {
master_pb::VolumeShortInformationMessage {
id,
collection: self.collection.clone(),
version: self.version,
replica_placement: self.replica_placement,
ttl: self.ttl,
disk_type: self.disk_type.clone(),
disk_id: self.disk_id,
}
}
}
fn volume_identities(
volumes: &[master_pb::VolumeInformationMessage],
) -> HashMap<u32, VolumeIdentity> {
volumes
.iter()
.map(|v| (v.id, VolumeIdentity::of(v)))
.collect()
}
/// Collect volume information into a Heartbeat message.
fn collect_heartbeat_with_snapshot(
config: &HeartbeatConfig,
@@ -1578,7 +1608,7 @@ mod tests {
let (_, volume) = store.find_volume_mut(VolumeId(17)).unwrap();
volume.set_read_only().unwrap();
volume.volume_info.files.push(Default::default());
volume.refresh_remote_write_mode();
volume.refresh_remote_write_mode().unwrap();
}
let heartbeat = build_heartbeat(&test_config(), &mut store);
@@ -1944,7 +1974,7 @@ mod tests {
key: "volumes/71.dat".to_string(),
..Default::default()
});
volume.refresh_remote_write_mode();
volume.refresh_remote_write_mode().unwrap();
let heartbeat = build_heartbeat(&test_config(), &mut store);
+331 -122
View File
@@ -33,7 +33,9 @@ use std::sync::Arc;
use std::time::{Duration, Instant};
use futures::future::join_all;
use futures::stream::{self, StreamExt};
use reed_solomon_erasure::galois_8::ReedSolomon;
use tokio::sync::Semaphore;
use tonic::Request;
use crate::pb::master_pb::{self, seaweed_client::SeaweedClient, LookupEcVolumeRequest};
@@ -49,6 +51,19 @@ use crate::storage::store_ec_reconcile::EcVolumeMissingIndex;
use crate::storage::types::*;
use crate::storage::volume::volume_file_name;
/// Bounds the fan-out of a single needle read. Mirrors Go's
/// `ecIntervalReadConcurrency`.
const INTERVAL_READ_CONCURRENCY: usize = 8;
/// Bounds the bytes EC recovery holds in flight across every concurrent read.
/// Recovery is the one read path that multiplies the served bytes — it keeps an
/// interval-sized buffer per shard alive until Reed-Solomon runs — and a peer
/// that is slow to fail holds each of them for the whole gRPC timeout, so a
/// burst of reads during a network blip walked the server into an OOM. Mirrors
/// Go's `ecRecoverBudget`.
const EC_RECOVER_BUDGET: usize = 256 << 20;
static EC_RECOVER_SEM: Semaphore = Semaphore::const_new(EC_RECOVER_BUDGET);
/// One interval's data after Phase A.
enum IntervalResult {
/// Already read from a locally-mounted shard.
@@ -90,7 +105,7 @@ pub async fn read_ec_shard_needle_distributed(
// intervals, and read any locally-mounted shard intervals. We must
// not `.await` while holding this guard (std::sync::RwLockReadGuard
// is !Send).
let snapshot = match snapshot_under_lock(state, vid, needle_id)? {
let mut snapshot = match snapshot_under_lock(state, vid, needle_id)? {
Some(s) => s,
None => return Ok(None),
};
@@ -106,7 +121,9 @@ pub async fn read_ec_shard_needle_distributed(
let mut shard_locations = snapshot.cached_locations.clone();
if any_remote
&& needs_refresh(
&& claim_shard_locations_refresh(
state,
vid,
&shard_locations,
snapshot.cache_refreshed_at,
snapshot.data_shards as usize,
@@ -117,17 +134,21 @@ pub async fn read_ec_shard_needle_distributed(
Ok(fresh) => {
// A complete reply merges into the cache; an incomplete one
// (< data_shards) is left unwritten — keep the prior cache.
if let Some(merged) =
write_back_shard_locations(state, vid, fresh, snapshot.data_shards as usize)
match write_back_shard_locations(state, vid, fresh, snapshot.data_shards as usize)
{
shard_locations = merged;
Some(merged) => shard_locations = merged,
// An incomplete reply leaves the cache unwritten and its refresh
// time unadvanced, so the mark this refresh consumed goes back.
None => mark_shard_locations_stale(state, vid),
}
}
Err(e) => {
// Lookup failed — proceed with cached values. If cache
// is empty, the remote fetch below will fail and we
// surface a NotFound (matching Go's behavior when no
// locations are known).
// locations are known). The mark goes back: nothing
// answered for it, and the map stays disproved.
mark_shard_locations_stale(state, vid);
tracing::warn!(
"ec lookup failed for volume {}: {} — using cached locations ({} entries)",
vid.0,
@@ -138,39 +159,56 @@ pub async fn read_ec_shard_needle_distributed(
}
}
// Phase C — fetch missing intervals, reconstructing when the
// direct peer read fails.
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(snapshot.intervals.len());
for res in snapshot.intervals {
match res {
IntervalResult::Local(buf) => assembled.push(buf),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
} => {
let (buf, is_deleted) = fetch_one_interval(
state,
vid,
needle_id,
// Phase C — fetch missing intervals, reconstructing when the direct peer
// read fails. Blocks that follow each other in the .dat live on different
// shards, so a needle spanning several of them costs one round trip per
// block when fetched in sequence; `buffered` keeps the order while letting
// INTERVAL_READ_CONCURRENCY of them fly at once.
let data_shards = snapshot.data_shards as usize;
let parity_shards = snapshot.parity_shards as usize;
let encode_ts_ns = snapshot.encode_ts_ns;
let intervals = std::mem::take(&mut snapshot.intervals);
let fetched: Vec<io::Result<(Vec<u8>, bool)>> = stream::iter(intervals.into_iter().map(|res| {
let shard_locations = &shard_locations;
async move {
match res {
IntervalResult::Local(buf) => Ok((buf, false)),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
&shard_locations,
snapshot.data_shards as usize,
snapshot.parity_shards as usize,
snapshot.encode_ts_ns,
)
.await?;
// A peer reports the needle deleted (a cross-server window where the
// local index still shows it live): treat as not-found rather than
// serving zeros, mirroring Go's ErrorDeleted.
if is_deleted {
return Ok(None);
} => {
fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
encode_ts_ns,
)
.await
}
assembled.push(buf);
}
}
}))
.buffered(INTERVAL_READ_CONCURRENCY)
.collect()
.await;
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(fetched.len());
for res in fetched {
let (buf, is_deleted) = res?;
// A peer reports the needle deleted (a cross-server window where the
// local index still shows it live): treat as not-found rather than
// serving zeros, mirroring Go's ErrorDeleted.
if is_deleted {
return Ok(None);
}
assembled.push(buf);
}
// Phase D — assemble and parse the Needle. Mirrors the tail of
@@ -208,7 +246,9 @@ pub async fn read_ec_shard_needle_distributed(
/// without decoding (so genuine shard faults are reported rather than healed).
/// Mirrors Go's `Store.ScrubEcVolume`. Returns (rows walked, broken shards,
/// errors). `force_deleted_needles_check` disables the benign delete-state
/// size-mismatch suppression.
/// size-mismatch suppression. `recover_unreadable` (READS mode) rebuilds an
/// unreadable interval from the surviving shards: the same shards are reported
/// broken, but only needles parity can no longer recover become errors.
///
/// Shard locations are refreshed once up front. Each needle is then processed via
/// `scrub_snapshot_under_lock` + lock-drop + no-reconstruct `read_remote_ec_shard_interval`,
@@ -217,10 +257,19 @@ pub async fn scrub_ec_volume_distributed(
state: &Arc<VolumeServerState>,
vid: VolumeId,
force_deleted_needles_check: bool,
recover_unreadable: bool,
) -> (i64, Vec<crate::pb::volume_server_pb::EcShardInfo>, Vec<String>) {
// Phase A — under the Store read lock, run the index scrub and grab the
// paths/scalars + shard-location staleness; release the lock before any await.
let (ecx_path, collection, seed_errs, cached_locations, cache_refreshed_at, data_shards, total_shards) = {
let (
ecx_path,
collection,
seed_errs,
cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
) = {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
Some(v) => v,
@@ -255,10 +304,18 @@ pub async fn scrub_ec_volume_distributed(
// cachedLookupEcShardLocations). A partial reply (< data_shards locations, a
// master mid-recovery) or a failed lookup is a hard, retryable error — never
// overwrite a good cache with a partial map or storm a down master per needle.
if needs_refresh(&cached_locations, cache_refreshed_at, data_shards, total_shards) {
if claim_shard_locations_refresh(
state,
vid,
&cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
) {
match cached_lookup_ec_shard_locations(state, vid).await {
Ok(fresh) => {
if write_back_shard_locations(state, vid, fresh, data_shards).is_none() {
mark_shard_locations_stale(state, vid);
return (
0,
Vec::new(),
@@ -270,11 +327,12 @@ pub async fn scrub_ec_volume_distributed(
}
}
Err(e) => {
mark_shard_locations_stale(state, vid);
return (
0,
Vec::new(),
vec![format!("failed to locate shard via master grpc: {}", e)],
)
);
}
}
}
@@ -340,8 +398,9 @@ pub async fn scrub_ec_volume_distributed(
}
};
// Read each interval local-then-remote WITHOUT reconstructing: we verify
// the shards are valid, we do not heal them. Locations refreshed above.
// Read each interval local-then-remote. Neither read decodes: the point is to
// find shards that are themselves broken, not to heal around them. READS then
// rebuilds what it could not read. Locations refreshed above.
let n_intervals = snapshot.intervals.len();
let mut data: Vec<u8> = Vec::with_capacity(snapshot.actual_size);
for (i, res) in snapshot.intervals.iter().enumerate() {
@@ -371,11 +430,9 @@ pub async fn scrub_ec_volume_distributed(
// -> the delete-state suppression (mirrors Go's pre-zeroed buffer).
Ok((_, true)) => data.resize(data.len() + *ssize, 0),
Ok((buf, false)) => data.extend_from_slice(&buf),
Err(_) => {
errs.push(format!(
"failed to read EC shard {} for needle {} on volume {} (interval {}/{})",
shard_id, id.0, vid.0, i + 1, n_intervals
));
Err(read_err) => {
// The shard is broken whether or not the needle survives it,
// so report it either way.
broken_shards.insert(
*shard_id,
crate::pb::volume_server_pb::EcShardInfo {
@@ -386,7 +443,41 @@ pub async fn scrub_ec_volume_distributed(
..Default::default()
},
);
break;
if !recover_unreadable {
errs.push(format!(
"failed to read EC shard {} for needle {} on volume {} (interval {}/{}): {}",
shard_id, id.0, vid.0, i + 1, n_intervals, read_err
));
break;
}
match recover_one_remote_ec_shard_interval(
state,
vid,
id,
*shard_id,
*shard_offset,
*ssize,
&locations,
data_shards,
total_shards - data_shards,
snapshot.encode_ts_ns,
)
.await
{
// Same as the direct read above: a holder reporting the
// needle deleted is authoritative and answers with no
// bytes, so zero-fill and let the delete-state
// suppression have it.
Ok((_, true)) => data.resize(data.len() + *ssize, 0),
Ok((buf, false)) => data.extend_from_slice(&buf),
Err(e) => {
errs.push(format!(
"failed to recover EC shard {} for needle {} on volume {} (interval {}/{}): {}",
shard_id, id.0, vid.0, i + 1, n_intervals, e
));
break;
}
}
}
}
}
@@ -508,7 +599,7 @@ fn read_local_intervals(
) -> Vec<IntervalResult> {
let mut interval_results = Vec::with_capacity(intervals.len());
for interval in intervals {
let (shard_id, shard_offset) = interval.to_shard_id_and_offset(ecv.data_shards);
let (shard_id, shard_offset) = ecv.interval_to_shard_id_and_offset(interval);
let buf_size = interval.size as usize;
let local = ecv.shards.get(shard_id as usize).and_then(|s| s.as_ref());
match local {
@@ -571,6 +662,7 @@ fn build_snapshot(
fn needs_refresh(
locations: &HashMap<ShardId, Vec<String>>,
refreshed_at: Option<Instant>,
stale: bool,
data_shards: usize,
total_shards: usize,
) -> bool {
@@ -580,16 +672,53 @@ fn needs_refresh(
None => return true,
};
let shard_count = locations.len();
if shard_count < data_shards && age < Duration::from_secs(11) {
return false;
// A complete map is trusted longest. One short of data_shards, or one a
// failed read has just disproved, is re-checked promptly: until it is, reads
// keep aiming at a location the shard has left.
let ttl = if stale || shard_count < data_shards {
Duration::from_secs(11)
} else if shard_count == total_shards {
Duration::from_secs(37 * 60)
} else {
Duration::from_secs(7 * 60)
};
age >= ttl
}
/// Mark the cached shard map for a prompt re-check after a read failed against
/// one of its locations. Go drops the entry outright in `forgetShardId`, which
/// costs it the direct read until the map is re-learned; here the entry stays
/// (a dead peer just fails fast on the next attempt) and only the freshness
/// window is cut, so a shard that has moved is picked up in seconds either way.
fn mark_shard_locations_stale(state: &Arc<VolumeServerState>, vid: VolumeId) {
let store = state.store.read().unwrap();
if let Some(ecv) = store.find_ec_volume(vid) {
*ecv.shard_locations_stale.lock().unwrap() = true;
}
if shard_count == total_shards && age < Duration::from_secs(37 * 60) {
return false;
}
/// Decide whether the cached map is due a master lookup and, when it is, consume
/// its stale mark in the same critical section. A mark raised from here on
/// belongs to the next refresh: the read that raised it has disproved the map
/// this lookup is about to install.
fn claim_shard_locations_refresh(
state: &Arc<VolumeServerState>,
vid: VolumeId,
locations: &HashMap<ShardId, Vec<String>>,
refreshed_at: Option<Instant>,
data_shards: usize,
total_shards: usize,
) -> bool {
let store = state.store.read().unwrap();
let Some(ecv) = store.find_ec_volume(vid) else {
return needs_refresh(locations, refreshed_at, false, data_shards, total_shards);
};
let mut stale = ecv.shard_locations_stale.lock().unwrap();
let refresh = needs_refresh(locations, refreshed_at, *stale, data_shards, total_shards);
if refresh {
*stale = false;
}
if shard_count >= data_shards && age < Duration::from_secs(7 * 60) {
return false;
}
true
refresh
}
async fn cached_lookup_ec_shard_locations(
@@ -725,6 +854,9 @@ async fn fetch_one_interval(
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
}
}
@@ -732,7 +864,7 @@ async fn fetch_one_interval(
// Reconstruct: fan-out reads to every other shard at the same
// (shard_offset, size). Mirrors `recoverOneRemoteEcShardInterval`.
let buf = recover_one_remote_ec_shard_interval(
recover_one_remote_ec_shard_interval(
state,
vid,
needle_id,
@@ -744,8 +876,7 @@ async fn fetch_one_interval(
parity_shards,
expected_encode_ts_ns,
)
.await?;
Ok((buf, false))
.await
}
async fn read_remote_ec_shard_interval(
@@ -905,7 +1036,7 @@ async fn recover_one_remote_ec_shard_interval(
data_shards: usize,
parity_shards: usize,
expected_encode_ts_ns: i64,
) -> io::Result<Vec<u8>> {
) -> io::Result<(Vec<u8>, bool)> {
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards).map_err(|e| {
io::Error::new(
@@ -914,89 +1045,136 @@ async fn recover_one_remote_ec_shard_interval(
)
})?;
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap. An interval whose
// fan-out outgrows the whole budget takes all of it and so runs alone,
// rather than waiting on permits that can never be granted.
let _permit = EC_RECOVER_SEM
.acquire_many((size * data_shards).min(EC_RECOVER_BUDGET) as u32)
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
),
)
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
// Phase 0: seed bufs from LOCALLY mounted shards. If this node
// already holds enough sibling shards, reconstruction completes
// without any peer fan-out — and even with a cold/incomplete
// shard_locations cache or a failed master lookup, local
// survivors still contribute. Mirrors Go's
// recoverOneRemoteEcShardInterval behaviour, which is implicitly
// local-aware because the Store fan-out targets ALL known
// locations (including the caller's own server address); the
// Rust port had been remote-only, so reconstructing with a cold
// cache failed even when enough siblings were on disk.
// survivors still contribute.
let mut available = 0usize;
{
let store = state.store.read().unwrap();
if let Some(ecv) = store.find_ec_volume(vid) {
for sid in 0..total_shards {
if sid as ShardId == shard_id_to_recover {
continue;
}
if let Some(Some(shard)) = ecv.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
}
for sid in 0..total_shards {
if available >= data_shards {
break;
}
if sid as ShardId == shard_id_to_recover {
continue;
}
// Resolve the shard together with the EcVolume on the disk that owns
// it: a reconciled volume has its shards split across data dirs. A
// shard from a different encode run must not be fed to Reed-Solomon;
// lenient only when the caller carries no identity (pre-upgrade).
// Mirrors Go's `readLocalEcShardInterval`.
let owner = match store.find_ec_volume_with_shard(vid, sid as u32) {
Some(ecv) if expected_encode_ts_ns == 0 || ecv.encode_ts_ns == expected_encode_ts_ns => ecv,
_ => continue,
};
if let Some(Some(shard)) = owner.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
available += 1;
}
}
}
}
// Phase 1: remote fan-out — one task per known shard location
// we DON'T already have locally and DON'T need to recover.
let mut tasks = Vec::new();
for (sid, locs) in shard_locations {
if *sid == shard_id_to_recover || locs.is_empty() {
continue;
}
if bufs[*sid as usize].is_some() {
continue;
}
let sid = *sid;
let locs = locs.clone();
let state = state.clone();
tasks.push(async move {
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
sid,
shard_offset,
size,
expected_encode_ts_ns,
)
.await;
(sid, res)
});
}
let results = join_all(tasks).await;
// Phase 1: remote fan-out over the shard locations we DON'T already have
// locally and DON'T need to recover. Reconstruction consumes data_shards
// shards, so reading every remaining one holds a third more buffers than
// that and asks a third more of peers that may already be struggling: fetch
// what is still missing, and widen only if some of those reads fail.
let mut candidates: Vec<(ShardId, Vec<String>)> = shard_locations
.iter()
.filter(|(sid, locs)| {
**sid != shard_id_to_recover
&& (**sid as usize) < total_shards
&& !locs.is_empty()
&& bufs[**sid as usize].is_none()
})
.map(|(sid, locs)| (*sid, locs.clone()))
.collect();
for (sid, res) in results {
match res {
// Exclude a deleted shard from reconstruction (Go gates on a full
// read): feeding the empty/zero buffer into Reed-Solomon would
// corrupt the recovered shard.
Ok((buf, is_deleted)) => {
if !is_deleted && (sid as usize) < total_shards {
bufs[sid as usize] = Some(buf);
}
}
Err(e) => {
tracing::debug!(
"recover: read {}.{} for needle {} failed: {}",
vid.0,
sid,
let mut any_deleted = false;
while available < data_shards && !candidates.is_empty() {
let rest = candidates.split_off((data_shards - available).min(candidates.len()));
let wave = std::mem::replace(&mut candidates, rest);
let results = join_all(wave.into_iter().map(|(sid, locs)| {
let state = state.clone();
async move {
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
e
);
sid,
shard_offset,
size,
expected_encode_ts_ns,
)
.await;
(sid, res)
}
}))
.await;
for (sid, res) in results {
match res {
// Exclude a deleted shard from reconstruction (Go gates on a full
// read): feeding the empty/zero buffer into Reed-Solomon would
// corrupt the recovered shard.
Ok((buf, is_deleted)) => {
if is_deleted {
any_deleted = true;
continue;
}
bufs[sid as usize] = Some(buf);
available += 1;
}
Err(e) => {
tracing::debug!(
"recover: read {}.{} for needle {} failed: {}",
vid.0,
sid,
needle_id,
e
);
mark_shard_locations_stale(state, vid);
}
}
}
if any_deleted {
// every shard of a deleted needle answers deleted, so another wave cannot help
break;
}
}
let available = bufs.iter().filter(|b| b.is_some()).count();
if available < data_shards {
// A holder reporting the needle deleted is authoritative -- deletes are
// never invented and never undone -- so answer that rather than the
// failure to gather shards of a needle that is gone.
if any_deleted {
return Ok((Vec::new(), true));
}
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
@@ -1017,7 +1195,7 @@ async fn recover_one_remote_ec_shard_interval(
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok(buf),
Some(buf) => Ok((buf, any_deleted)),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
@@ -1255,3 +1433,34 @@ async fn drain_copy_stream(
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
fn locations(count: usize) -> HashMap<ShardId, Vec<String>> {
(0..count)
.map(|sid| (sid as ShardId, vec!["127.0.0.1:8080".to_string()]))
.collect()
}
#[test]
fn needs_refresh_re_checks_a_map_a_failed_read_disproved() {
let just_now = Some(Instant::now());
let aged = Some(Instant::now() - Duration::from_secs(12));
// A complete map is trusted for a long time, and one shard short still
// outlasts a 12-second gap.
assert!(!needs_refresh(&locations(14), aged, false, 10, 14));
assert!(!needs_refresh(&locations(13), aged, false, 10, 14));
// Disproved by a read, the same maps are re-checked within seconds.
assert!(needs_refresh(&locations(14), aged, true, 10, 14));
assert!(needs_refresh(&locations(13), aged, true, 10, 14));
// But the mark buys one prompt re-check, not a lookup per read.
assert!(!needs_refresh(&locations(14), just_now, true, 10, 14));
// A map short of the data shards is re-checked promptly regardless.
assert!(needs_refresh(&locations(9), aged, false, 10, 14));
// An unrefreshed cache always looks up.
assert!(needs_refresh(&locations(0), None, false, 10, 14));
}
}
+126 -14
View File
@@ -843,7 +843,8 @@ impl DiskLocation {
// .ecx open error, .ecj create error, malformed .vif) would
// have to panic via unwrap(). Build the EcVolume up front and
// propagate the error to the caller.
if !self.ec_volumes.contains_key(&vid) {
let created = !self.ec_volumes.contains_key(&vid);
if created {
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid)
.map_err(VolumeError::Io)?;
self.ec_volumes.insert(vid, ec_vol);
@@ -866,9 +867,27 @@ impl DiskLocation {
}
for &shard_id in shard_ids {
// A mount retry re-listing a shard this volume already holds:
// keep the existing registration (mirrors Go's AddEcVolumeShard
// added=false) — re-adding would replace a serving fd and bump
// the ec_shards gauge without growing the mounted count.
if ec_vol.has_shard(shard_id as u8) {
continue;
}
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id as u8);
shard.disk_type = ec_vol.disk_type.clone();
ec_vol.add_shard(shard).map_err(VolumeError::Io)?;
if let Err(e) = ec_vol.add_shard(shard) {
// The shard was dropped (its descriptors closed) inside the
// failed add. If this call just created the EcVolume and it
// holds nothing, remove it too — a zero-shard registration
// would advertise a mount that serves no data while pinning
// its descriptors.
let now_empty = ec_vol.shard_count() == 0;
if created && now_empty {
self.ec_volumes.remove(&vid);
}
return Err(VolumeError::Io(e));
}
crate::metrics::VOLUME_GAUGE
.with_label_values(&[collection, "ec_shards"])
.inc();
@@ -921,27 +940,32 @@ impl DiskLocation {
// double-count for those filenames.
let mut seen: HashSet<String> = HashSet::new();
let mut entries: Vec<String> = Vec::new();
for ent in fs::read_dir(&self.directory)? {
let ent = ent?;
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
if seen.insert(name.clone()) {
entries.push(name);
}
}
if self.idx_directory != self.directory {
for ent in fs::read_dir(&self.idx_directory)? {
// Keep only the shard and index files this scan acts on: a disk of
// regular volumes has millions of .dat/.idx/.vif names that would
// otherwise each cost a String here and a slot in the sort below.
let mut collect = |dir: &str| -> io::Result<()> {
for ent in fs::read_dir(dir)? {
let ent = ent?;
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let ext = &name[dot..];
if parse_ec_shard_extension(ext).is_none() && ext != ".ecx" {
continue;
}
if seen.insert(name.clone()) {
entries.push(name);
}
}
Ok(())
};
collect(&self.directory)?;
if self.idx_directory != self.directory {
collect(&self.idx_directory)?;
}
entries.sort();
@@ -1774,6 +1798,94 @@ mod tests {
}
}
/// A refused shard (a 0-byte file beside an index with entries) on the
/// FIRST mount of a volume must not leave the just-created zero-shard
/// EcVolume registered — it would advertise a mount serving no data,
/// pin the .ecx/.ecj descriptors, and make placement's mounted tier
/// prefer this disk. A volume that already holds shards keeps them
/// (the mount RPC's pre-existing first-error-aborts contract).
#[test]
fn test_mount_ec_shards_refused_shard_removes_created_empty_volume() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
// An index with one 16-byte entry and a 0-byte shard: the mount is
// refused, and the EcVolume created for it must be unregistered.
std::fs::write(format!("{}/pics_9.ecx", dir), [0u8; 16]).unwrap();
std::fs::write(format!("{}/pics_9.ec00", dir), b"").unwrap();
let err = loc
.mount_ec_shards(VolumeId(9), "pics", &[0], "")
.expect_err("a 0-byte shard beside an index with entries must refuse the mount");
assert!(
err.to_string().contains("empty (0 bytes)"),
"want the empty-shard refusal, got: {}",
err
);
assert!(
loc.find_ec_volume(VolumeId(9)).is_none(),
"a refused first mount must not leave a zero-shard EcVolume registered",
);
// With a valid shard mounted, a later refused shard keeps the
// existing registration intact.
std::fs::write(format!("{}/pics_9.ec01", dir), b"good bytes").unwrap();
loc.mount_ec_shards(VolumeId(9), "pics", &[1], "").unwrap();
loc.mount_ec_shards(VolumeId(9), "pics", &[0], "")
.expect_err("the 0-byte shard stays refused");
assert_eq!(
loc.find_ec_volume(VolumeId(9)).map(|v| v.shard_count()),
Some(1),
"an existing volume keeps its valid shards when a later shard is refused",
);
}
/// A mount retry re-listing an already mounted shard must keep the
/// existing registration and not bump the ec_shards gauge — the Rust
/// twin of Go's AddEcVolumeShard added=false handling.
#[test]
fn test_mount_ec_shards_duplicate_keeps_registration_and_gauge() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
// A collection name unique to this test: the gauge is process-global
// and sibling tests running in parallel touch other labels.
std::fs::write(format!("{}/dupmount_11.ec00", dir), b"shard bytes").unwrap();
let gauge = crate::metrics::VOLUME_GAUGE.with_label_values(&["dupmount", "ec_shards"]);
let before = gauge.get();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "").unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.expect("a duplicate mount must succeed as a no-op");
assert_eq!(
loc.find_ec_volume(VolumeId(11)).map(|v| v.shard_count()),
Some(1),
);
assert_eq!(
gauge.get(),
before + 1.0,
"the duplicate mount must not bump the ec_shards gauge",
);
}
#[test]
fn test_disk_location_persists_directory_uuid_and_tags() {
let tmp = TempDir::new().unwrap();
@@ -465,6 +465,28 @@ pub fn resolve_status(
}
}
/// Whether a generation-matching sidecar agrees with the geometry the volume is
/// mounted with. Both files record the layout the generation was encoded with,
/// so a disagreement means one of them is wrong and reads through the other
/// would land at the wrong shard offsets — the caller fails the mount rather
/// than merely dropping protection. A sidecar that records no EC config has
/// nothing to contradict.
pub fn geometry_matches(
prot: &EcBitrotProtection,
data_shards: usize,
parity_shards: usize,
block_size: i64,
) -> bool {
match &prot.ec_shard_config {
None => true,
Some(cfg) => {
cfg.data_shards as usize == data_shards
&& cfg.parity_shards as usize == parity_shards
&& cfg.block_size == block_size
}
}
}
/// Returns the [`EcShardChecksums`] entry for a shard id, or `None`.
pub fn shard_checksums(prot: &EcBitrotProtection, shard_id: u32) -> Option<&EcShardChecksums> {
prot.shards.iter().find(|s| s.shard_id == shard_id)
@@ -539,11 +561,12 @@ fn read_full_at(f: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
/// Builds the `EcShardConfig` proto for the given layout. The bitrot sidecar
/// carries its own top-level encode_uuid, so the nested config leaves it empty.
pub fn ec_shard_config(data_shards: u32, parity_shards: u32) -> EcShardConfig {
pub fn ec_shard_config(data_shards: u32, parity_shards: u32, block_size: i64) -> EcShardConfig {
EcShardConfig {
data_shards,
parity_shards,
encode_ts_ns: 0,
block_size,
}
}
@@ -567,6 +590,7 @@ mod tests {
data_shards: 10,
parity_shards: 4,
encode_ts_ns: 0,
block_size: 0,
}),
shards: vec![
EcShardChecksums {
@@ -712,7 +736,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4)),
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
shards: vec![EcShardChecksums {
shard_id: 0,
covered_size: covered,
@@ -750,7 +774,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4)),
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
shards: vec![EcShardChecksums {
shard_id: 0,
covered_size: 5,
@@ -794,7 +818,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4)),
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
shards,
encode_uuid: vec![0u8; 16],
}
@@ -80,6 +80,8 @@ pub fn write_dat_file_from_shards(
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let dirs: Vec<String> = (0..data_shards).map(|_| dir.to_string()).collect();
write_dat_file_from_shards_with_dirs(
@@ -90,6 +92,8 @@ pub fn write_dat_file_from_shards(
encoded_dat_file_size,
data_shards,
&dirs,
large_block_size,
small_block_size,
)
}
@@ -113,7 +117,10 @@ pub fn write_dat_file_from_shards(
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Pass zero when the .vif does
/// not record the encode-time size to infer the layout from the shard
/// size.
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
@@ -122,6 +129,8 @@ pub fn write_dat_file_from_shards_with_dirs(
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
write_dat_file(
dat_dir,
@@ -131,8 +140,8 @@ pub fn write_dat_file_from_shards_with_dirs(
encoded_dat_file_size,
data_shards,
shard_dirs,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
large_block_size,
small_block_size,
)
}
@@ -412,7 +421,9 @@ mod tests {
// Encode to EC
let data_shards = 10;
let parity_shards = 4;
ec_encoder::write_ec_files(dir, dir, "", VolumeId(1), data_shards, parity_shards).unwrap();
let block_size =
ec_encoder::write_ec_files(dir, dir, "", VolumeId(1), data_shards, parity_shards)
.unwrap();
// Delete original .dat and .idx
std::fs::remove_file(format!("{}/1.dat", dir)).unwrap();
@@ -426,6 +437,8 @@ mod tests {
original_dat_size as i64,
original_dat_size as i64,
data_shards,
block_size as usize,
block_size as usize,
)
.unwrap();
write_idx_file_from_ec_index(dir, "", VolumeId(1)).unwrap();
@@ -472,7 +485,16 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
// No shard files exist, so de-striping must fail and publish nothing:
// neither the final .dat nor a partial .dat.tmp may remain.
let res = write_dat_file_from_shards(dir, "", VolumeId(7), 100, 100, 10);
let res = write_dat_file_from_shards(
dir,
"",
VolumeId(7),
100,
100,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
);
assert!(res.is_err());
assert!(!std::path::Path::new(&format!("{}/7.dat", dir)).exists());
assert!(!std::path::Path::new(&format!("{}/7.dat.tmp", dir)).exists());
@@ -525,6 +547,7 @@ mod tests {
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
)
@@ -616,6 +639,7 @@ mod tests {
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
)
@@ -25,6 +25,10 @@ use crate::storage::volume::volume_file_name;
///
/// Creates .ec00-.ec13 files in the same directory.
/// Also creates a sorted .ecx index from the .idx file.
///
/// Always encodes with the uniform block layout, sized for this .dat, and
/// returns the block size so the caller can persist it to .vif. Mirrors Go's
/// WriteEcFiles.
pub fn write_ec_files(
dir: &str,
idx_dir: &str,
@@ -32,7 +36,7 @@ pub fn write_ec_files(
volume_id: VolumeId,
data_shards: usize,
parity_shards: usize,
) -> io::Result<()> {
) -> io::Result<i64> {
let base = volume_file_name(dir, collection, volume_id);
let dat_path = format!("{}.dat", base);
let idx_base = volume_file_name(idx_dir, collection, volume_id);
@@ -66,7 +70,7 @@ pub fn write_ec_files(
.map(|_| ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64))
.collect();
// Encode in large blocks, then small blocks
let block_size = uniform_block_size(dat_size, data_shards);
encode_dat_file(
&dat_file,
dat_size,
@@ -75,8 +79,9 @@ pub fn write_ec_files(
&mut builders,
data_shards,
parity_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
ENCODE_BUFFER_SIZE,
block_size as usize,
block_size as usize,
)?;
// Close all shards
@@ -103,6 +108,7 @@ pub fn write_ec_files(
ec_shard_config: Some(ec_bitrot::ec_shard_config(
data_shards as u32,
parity_shards as u32,
block_size,
)),
shards: shard_checksums,
encode_uuid: ec_bitrot::new_encode_uuid(),
@@ -120,7 +126,19 @@ pub fn write_ec_files(
);
}
Ok(())
Ok(block_size)
}
/// uniform_block_size returns the per-shard block size of the uniform layout
/// for a .dat of the given size: ceil(dat_file_size/data_shards) rounded up to
/// a whole small block. For every input this equals the legacy layout's padded
/// shard size, so only the byte placement differs between the two layouts,
/// never the shard length. Mirrors Go's UniformBlockSize.
pub fn uniform_block_size(dat_file_size: i64, data_shards: usize) -> i64 {
let small = ERASURE_CODING_SMALL_BLOCK_SIZE as i64;
let per_shard = (dat_file_size + data_shards as i64 - 1) / data_shards as i64;
let blocks = ((per_shard + small - 1) / small).max(1);
blocks * small
}
/// Rebuild missing EC shard files from existing shards using Reed-Solomon reconstruct.
@@ -370,7 +388,7 @@ pub fn verify_ec_shards(
}
/// Write sorted .ecx index from .idx file.
fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::Result<()> {
pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::Result<()> {
if !std::path::Path::new(idx_path).exists() {
return Err(io::Error::new(
io::ErrorKind::NotFound,
@@ -421,6 +439,8 @@ pub fn rebuild_ecx_file(
collection: &str,
volume_id: VolumeId,
data_shards: usize,
block_size: i64,
dat_file_size: i64,
additional_dirs: &[&str],
) -> io::Result<()> {
use crate::storage::needle::needle::get_actual_size;
@@ -463,10 +483,42 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
// multiple (mirrors the ecdFileSize-1 fallback in the read path).
let (large_block, small_block) = if block_size > 0 {
(block_size, block_size)
} else {
(
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
)
};
// The row count the de-stripe walks with. The encode-time .dat size is the
// authority — the same value the read path divides by data_shards — and the
// padded extent is only a fallback: under the legacy layout a shard that is
// an exact large-block multiple reads as one row too many, which
// re-interprets its last large row as small blocks and scrambles the
// recovered offsets. Subtracting one keeps that fallback on the safe side of
// the boundary, exactly as the read path's own fallback does.
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
};
// Read version from superblock (first byte of logical data)
let mut sb_buf = [0u8; SUPER_BLOCK_SIZE];
read_from_data_shards(&shards, &mut sb_buf, 0, data_shards)?;
read_from_data_shards(
&shards,
&mut sb_buf,
0,
data_shards,
locate_shard_size,
large_block,
small_block,
)?;
let version = Version(sb_buf[0]);
// Walk needles starting after superblock
@@ -475,10 +527,30 @@ pub fn rebuild_ecx_file(
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
while offset + header_size as i64 <= total_data_size {
// Read needle header (cookie + needle_id + size = 16 bytes)
// Read needle header (cookie + needle_id + size = 16 bytes).
// A read failure is NOT the end of the data — every offset in
// range maps into the shards, so an error means a truncated or
// unreadable shard. Publishing the entries collected so far as
// a successful .ecx would hand out a silently incomplete
// recovery index; propagate instead. (The scan still ends
// normally on the zero-cookie tail below.)
let mut header_buf = [0u8; NEEDLE_HEADER_SIZE];
if read_from_data_shards(&shards, &mut header_buf, offset as u64, data_shards).is_err() {
break;
if let Err(e) = read_from_data_shards(
&shards,
&mut header_buf,
offset as u64,
data_shards,
locate_shard_size,
large_block,
small_block,
) {
for s in &mut shards {
s.close();
}
return Err(io::Error::new(
e.kind(),
format!("scan needle header at offset {}: {}", offset, e),
));
}
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
@@ -532,58 +604,83 @@ pub fn rebuild_ecx_file(
Ok(())
}
/// Read bytes from EC data shards at a logical offset in the .dat file.
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
logical_offset: u64,
data_shards: usize,
locate_shard_size: i64,
large_block_size: i64,
small_block_size: i64,
) -> io::Result<()> {
let small_block = ERASURE_CODING_SMALL_BLOCK_SIZE as u64;
let data_shards_u64 = data_shards as u64;
let mut bytes_read = 0u64;
let mut remaining = buf.len() as u64;
let mut current_offset = logical_offset;
while remaining > 0 {
// Determine which shard and at what shard-offset this logical offset maps to.
// The data is interleaved: large blocks first, then small blocks.
// For simplicity, use the small block size for all calculations since
// large blocks are multiples of small blocks.
let row_size = small_block * data_shards_u64;
let row_index = current_offset / row_size;
let row_offset = current_offset % row_size;
let shard_index = (row_offset / small_block) as usize;
let shard_offset = row_index * small_block + (row_offset % small_block);
if shard_index >= data_shards {
let intervals = crate::storage::erasure_coding::ec_locate::locate_data(
logical_offset as i64,
Size(buf.len() as i32),
locate_shard_size,
data_shards as u32,
large_block_size,
small_block_size,
);
let mut bytes_read = 0usize;
for interval in &intervals {
let (shard_id, shard_offset) =
interval.to_shard_id_and_offset(data_shards as u32, large_block_size, small_block_size);
if shard_id as usize >= data_shards {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
"shard index out of range",
));
}
// How many bytes can we read from this position in this shard block
let bytes_left_in_block = small_block - (row_offset % small_block);
let to_read = remaining.min(bytes_left_in_block) as usize;
let dest = &mut buf[bytes_read as usize..bytes_read as usize + to_read];
shards[shard_index].read_at(dest, shard_offset)?;
bytes_read += to_read as u64;
remaining -= to_read as u64;
current_offset += to_read as u64;
let to_read = interval.size as usize;
let dest = &mut buf[bytes_read..bytes_read + to_read];
// Exact-read semantics: read_at may legally return fewer bytes
// than requested, and treating a short read as complete leaves
// the tail of `dest` as whatever the buffer held before. Loop
// until filled; zero bytes inside the mapped range means the
// shard is truncated — an error, not an end.
let mut filled = 0usize;
while filled < to_read {
let n = shards[shard_id as usize]
.read_at(&mut dest[filled..], shard_offset as u64 + filled as u64)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"short read from data shard {}: {} of {} bytes at offset {}",
shard_id, filled, to_read, shard_offset
),
));
}
filled += n;
}
bytes_read += to_read;
}
if bytes_read != buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"short read from data shards",
));
}
Ok(())
}
/// Buffer size for one encode sub-batch per shard, mirroring Go's 256KB
/// bufferSize in WriteEcFiles. A block is processed in block_size/buffer_size
/// sub-batches, so memory stays at total_shards * 256KB no matter how large
/// the uniform block is.
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// Encode the .dat file data into shard files.
///
/// Uses a two-phase approach matching Go's ec_encoder.go:
/// 1. Process as many large blocks (1GB) as possible
/// 2. Process remaining data with small blocks (1MB)
/// 1. Process as many large blocks as possible
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
@@ -593,44 +690,50 @@ pub(crate) fn encode_dat_file(
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
buffer_size: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let total_shards = data_shards + parity_shards;
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
.map(|_| vec![0u8; buffer_size])
.collect();
let mut remaining = dat_size;
let mut offset: u64 = 0;
// Phase 1: Process large blocks (1GB each) while enough data remains
// Phase 1: process whole large-block rows while enough data remains
let large_row_size = large_block_size * data_shards;
while remaining >= large_row_size as i64 {
encode_one_batch(
encode_data(
dat_file,
offset,
large_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
parity_shards,
)?;
offset += large_row_size as u64;
remaining -= large_row_size as i64;
}
// Phase 2: Process remaining data with small blocks (1MB each)
// Phase 2: process remaining data with small blocks
let small_row_size = small_block_size * data_shards;
while remaining > 0 {
let to_process = remaining.min(small_row_size as i64);
encode_one_batch(
encode_data(
dat_file,
offset,
small_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
parity_shards,
)?;
offset += to_process as u64;
remaining -= to_process;
@@ -639,61 +742,71 @@ pub(crate) fn encode_dat_file(
Ok(())
}
/// Encode one batch (row) of data.
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
encode_one_batch(
dat_file,
row_offset + (b * buffer_size) as u64,
block_size,
rs,
buffers,
shards,
builders,
data_shards,
)?;
}
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
) -> io::Result<()> {
let total_shards = data_shards + parity_shards;
// Each batch allocates block_size * total_shards bytes.
// With large blocks (1 GiB) this is 14 GiB -- guard against OOM.
let total_alloc = block_size.checked_mul(total_shards).ok_or_else(|| {
io::Error::new(
io::ErrorKind::InvalidInput,
"block_size * shard count overflows usize",
)
})?;
// Large-block encoding uses 1 GiB * 14 shards = 14 GiB; allow up to 16 GiB.
const MAX_BATCH_ALLOC: usize = 16 * 1024 * 1024 * 1024; // 16 GiB safety limit
if total_alloc > MAX_BATCH_ALLOC {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"batch allocation too large ({} bytes, limit {} bytes); block_size={} shards={}",
total_alloc, MAX_BATCH_ALLOC, block_size, total_shards,
),
));
}
// Allocate buffers for all shards
let mut buffers: Vec<Vec<u8>> = (0..total_shards).map(|_| vec![0u8; block_size]).collect();
// Read data shards from .dat file
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
let read_offset = offset + (i * block_size) as u64;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buffers[i], read_offset)?;
}
#[cfg(not(unix))]
{
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(read_offset))?;
f.read(&mut buffers[i])?;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
}
}
// Encode parity shards
rs.encode(&mut buffers).map_err(|e| {
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
@@ -710,6 +823,29 @@ fn encode_one_batch(
Ok(())
}
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let mut n = 0;
while n < buf.len() {
#[cfg(unix)]
let r = {
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buf[n..], offset + n as u64)?
};
#[cfg(not(unix))]
let r = {
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(offset + n as u64))?;
f.read(&mut buf[n..])?
};
if r == 0 {
break;
}
n += r;
}
Ok(n)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -1020,7 +1156,7 @@ mod tests {
// Without additional_dirs, rebuild must fail: shards 1, 3, 6 are not
// in primary and the full logical .dat content can't be reconstructed.
let res = rebuild_ecx_file(&primary, "", VolumeId(1), 10, &[]);
let res = rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[]);
assert!(
res.is_err(),
"ecx rebuild without additional_dirs must fail when data shards are on another disk"
@@ -1031,7 +1167,7 @@ mod tests {
);
// With additional_dirs pointing at the secondary, the rebuild must succeed.
rebuild_ecx_file(&primary, "", VolumeId(1), 10, &[secondary.as_str()]).unwrap();
rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[secondary.as_str()]).unwrap();
assert!(
std::path::Path::new(&ecx_path).exists(),
@@ -1043,6 +1179,118 @@ mod tests {
);
}
// A uniform-layout volume (block size > 1MiB) must have its .ecx rebuilt
// through the recorded geometry; the legacy 1MiB mapping would scan
// garbage past the first block boundary.
#[test]
fn test_rebuild_ecx_file_uniform_layout() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for i in 1u64..=12 {
let data: Vec<u8> = (0..2 << 20)
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
.collect();
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
let block_size = write_ec_files(&dir, &dir, "", VolumeId(2), 10, 4).unwrap();
assert!(
block_size > ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
"fixture must diverge from the legacy layout"
);
let ecx_path = format!("{}/2.ecx", dir);
let canonical = std::fs::read(&ecx_path).unwrap();
std::fs::remove_file(&ecx_path).unwrap();
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
let rebuilt = std::fs::read(&ecx_path).unwrap();
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
}
// A truncated data shard must FAIL the .ecx rebuild, not publish the
// entries scanned so far as a successful (silently incomplete) index.
#[test]
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(3),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for i in 1u64..=12 {
let data: Vec<u8> = (0..2 << 20)
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
.collect();
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
let block_size = write_ec_files(&dir, &dir, "", VolumeId(3), 10, 4).unwrap();
let ecx_path = format!("{}/3.ecx", dir);
std::fs::remove_file(&ecx_path).unwrap();
// Truncate shard 0 to just the superblock: the scan's very first
// needle-header read (offset SUPER_BLOCK_SIZE, shard 0 under the
// uniform layout) lands in the missing region. The pre-fix code
// broke the scan there and published an EMPTY .ecx as success.
let shard_path = format!("{}/3.ec00", dir);
let f = std::fs::OpenOptions::new()
.write(true)
.open(&shard_path)
.unwrap();
f.set_len(crate::storage::super_block::SUPER_BLOCK_SIZE as u64)
.unwrap();
drop(f);
let res = rebuild_ecx_file(&dir, "", VolumeId(3), 10, block_size, 0, &[]);
assert!(res.is_err(), "rebuild over a truncated shard must fail");
assert!(
!std::path::Path::new(&ecx_path).exists(),
"a failed rebuild must not leave a partial .ecx behind"
);
}
#[test]
fn test_reed_solomon_basic() {
let data_shards = 10;
@@ -18,21 +18,26 @@ pub struct Interval {
}
impl Interval {
pub fn to_shard_id_and_offset(&self, data_shards: u32) -> (ShardId, i64) {
pub fn to_shard_id_and_offset(
&self,
data_shards: u32,
large_block_size: i64,
small_block_size: i64,
) -> (ShardId, i64) {
let data_shards_usize = data_shards as usize;
let shard_id = (self.block_index % data_shards_usize) as ShardId;
let row_index = self.block_index / data_shards_usize;
let block_size = if self.is_large_block {
ERASURE_CODING_LARGE_BLOCK_SIZE as i64
large_block_size
} else {
ERASURE_CODING_SMALL_BLOCK_SIZE as i64
small_block_size
};
let mut offset = row_index as i64 * block_size + self.inner_block_offset;
if !self.is_large_block {
// Small blocks come after large blocks in the shard file
offset += self.large_block_rows_count as i64 * ERASURE_CODING_LARGE_BLOCK_SIZE as i64;
offset += self.large_block_rows_count as i64 * large_block_size;
}
(shard_id, offset)
@@ -42,7 +47,14 @@ impl Interval {
/// Locate the EC shard intervals needed to read data at the given offset and size.
///
/// `shard_size` is the size of a single shard file.
pub fn locate_data(offset: i64, size: Size, shard_size: i64, data_shards: u32) -> Vec<Interval> {
pub fn locate_data(
offset: i64,
size: Size,
shard_size: i64,
data_shards: u32,
large_block_size: i64,
small_block_size: i64,
) -> Vec<Interval> {
let mut intervals = Vec::new();
let data_size = size.0 as i64;
@@ -50,17 +62,14 @@ pub fn locate_data(offset: i64, size: Size, shard_size: i64, data_shards: u32) -
return intervals;
}
let large_block_size = ERASURE_CODING_LARGE_BLOCK_SIZE as i64;
let small_block_size = ERASURE_CODING_SMALL_BLOCK_SIZE as i64;
let large_row_size = large_block_size * data_shards as i64;
let small_row_size = small_block_size * data_shards as i64;
// Number of large block rows
let n_large_block_rows = if shard_size > 0 {
((shard_size - 1) / large_block_size) as usize
} else {
0
};
// Number of large block rows. Mirrors Go's shardDatSize/largeBlockLength:
// the caller's ecd-size fallback already subtracts 1 to disambiguate the
// exact-multiple case, so no further -1 here — a shard size that IS an
// exact multiple (dat_file_size path) means real full large rows.
let n_large_block_rows = (shard_size / large_block_size) as usize;
let large_section_size = n_large_block_rows as i64 * large_row_size;
let mut remaining_offset = offset;
@@ -150,7 +159,11 @@ mod tests {
is_large_block: true,
large_block_rows_count: 1,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert_eq!(shard_id, 0);
assert_eq!(offset, 100);
@@ -162,7 +175,11 @@ mod tests {
is_large_block: true,
large_block_rows_count: 1,
};
let (shard_id, _offset) = interval.to_shard_id_and_offset(data_shards);
let (shard_id, _offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert_eq!(shard_id, 5);
// Block index 12 (data_shards=10) → row_index 1, shard_id 2
@@ -173,7 +190,11 @@ mod tests {
is_large_block: true,
large_block_rows_count: 5,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert_eq!(shard_id, 2); // 12 % 10 = 2
assert_eq!(offset, large_block_size + 200); // row 1 offset + inner_block_offset
@@ -185,7 +206,11 @@ mod tests {
is_large_block: true,
large_block_rows_count: 2,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert_eq!(shard_id, 0);
assert_eq!(offset, ERASURE_CODING_LARGE_BLOCK_SIZE as i64); // row 1 offset
}
@@ -193,7 +218,14 @@ mod tests {
#[test]
fn test_locate_data_small_file() {
// Small file: 100 bytes at offset 50, shard size = 1MB
let intervals = locate_data(50, Size(100), 1024 * 1024, 10);
let intervals = locate_data(
50,
Size(100),
1024 * 1024,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert!(!intervals.is_empty());
// Should be a single small block interval (no large block rows for 1MB shard)
@@ -203,7 +235,14 @@ mod tests {
#[test]
fn test_locate_data_empty() {
let intervals = locate_data(0, Size(0), 1024 * 1024, 10);
let intervals = locate_data(
0,
Size(0),
1024 * 1024,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
assert!(intervals.is_empty());
}
@@ -216,7 +255,11 @@ mod tests {
is_large_block: false,
large_block_rows_count: 2,
};
let (_shard_id, offset) = interval.to_shard_id_and_offset(10);
let (_shard_id, offset) = interval.to_shard_id_and_offset(
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
// Should be after 2 large block rows
assert_eq!(offset, 2 * ERASURE_CODING_LARGE_BLOCK_SIZE as i64);
}
File diff suppressed because it is too large Load Diff
+74 -41
View File
@@ -14,7 +14,10 @@ use std::path::Path;
use std::sync::atomic::{AtomicI64, AtomicU64, Ordering};
mod compact_map;
pub mod file_pool;
pub mod sorted_file;
use compact_map::CompactMap;
use sorted_file::SortedFileNeedleMap;
use redb::{Database, Durability, ReadableDatabase, ReadableTable, TableDefinition};
@@ -750,10 +753,12 @@ impl RedbNeedleMap {
Ok(())
}
/// Look up a needle.
pub fn get(&self, key: NeedleId) -> Option<NeedleValue> {
/// Look up a needle. A redb failure is an ERROR, not an absent needle:
/// answering "not found" would turn a database problem into a read miss
/// and let a delete report success without recording a tombstone.
pub fn get(&self, key: NeedleId) -> io::Result<Option<NeedleValue>> {
let key_u64: u64 = key.into();
self.get_internal(key_u64).ok().flatten()
self.get_internal(key_u64)
}
/// Internal get that returns io::Result for error propagation.
@@ -939,33 +944,31 @@ impl RedbNeedleMap {
}
/// Collect all entries as a Vec for iteration (used by volume.rs iter patterns).
pub fn collect_entries(&self) -> Vec<(NeedleId, NeedleValue)> {
pub fn collect_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
let mut result = Vec::new();
let txn: redb::ReadTransaction = match self.db.begin_read() {
Ok(t) => t,
Err(_) => return result,
};
let table = match txn.open_table(NEEDLE_TABLE) {
Ok(t) => t,
Err(_) => return result,
};
let iter = match table.iter() {
Ok(i) => i,
Err(_) => return result,
};
let txn: redb::ReadTransaction = self
.db
.begin_read()
.map_err(|e| io::Error::other(format!("redb begin_read: {e}")))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {e}")))?;
let iter = table
.iter()
.map_err(|e| io::Error::other(format!("redb iter: {e}")))?;
for entry in iter {
if let Ok((key_guard, val_guard)) = entry {
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if bytes.len() == PACKED_NEEDLE_VALUE_SIZE {
let mut arr = [0u8; PACKED_NEEDLE_VALUE_SIZE];
arr.copy_from_slice(bytes);
let nv = unpack_needle_value(&arr);
result.push((NeedleId(key_u64), nv));
}
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb entry: {e}")))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if bytes.len() == PACKED_NEEDLE_VALUE_SIZE {
let mut arr = [0u8; PACKED_NEEDLE_VALUE_SIZE];
arr.copy_from_slice(bytes);
let nv = unpack_needle_value(&arr);
result.push((NeedleId(key_u64), nv));
}
}
result
Ok(result)
}
}
@@ -977,6 +980,10 @@ impl RedbNeedleMap {
pub enum NeedleMap {
InMemory(CompactNeedleMap),
Redb(RedbNeedleMap),
/// Read-only volumes — including every cloud-tiered one — search the sorted
/// `.sdx` on disk instead of holding an index in RAM. Mirrors Go's
/// `SortedFileNeedleMap`.
SortedFile(SortedFileNeedleMap),
}
impl NeedleMap {
@@ -985,14 +992,18 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.put(key, offset, size),
NeedleMap::Redb(nm) => nm.put(key, offset, size),
NeedleMap::SortedFile(nm) => nm.put(key, offset, size),
}
}
/// Look up a needle.
pub fn get(&self, key: NeedleId) -> Option<NeedleValue> {
/// Look up a needle. Disk- and database-backed maps report their own
/// failures rather than folding them into "not found" — see the notes on
/// `RedbNeedleMap::get` and `SortedFileNeedleMap::get`.
pub fn get(&self, key: NeedleId) -> io::Result<Option<NeedleValue>> {
match self {
NeedleMap::InMemory(nm) => nm.get(key),
NeedleMap::InMemory(nm) => Ok(nm.get(key)),
NeedleMap::Redb(nm) => nm.get(key),
NeedleMap::SortedFile(nm) => nm.get(key),
}
}
@@ -1001,6 +1012,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.delete(key, offset),
NeedleMap::Redb(nm) => nm.delete(key, offset),
NeedleMap::SortedFile(nm) => nm.delete(key, offset),
}
}
@@ -1009,6 +1021,9 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.set_idx_file(file, offset),
NeedleMap::Redb(nm) => nm.set_idx_file(file, offset),
// The sorted map borrows its .idx per append, so there is no
// long-lived writer to install.
NeedleMap::SortedFile(_) => {}
}
}
@@ -1017,6 +1032,8 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.has_idx_writer(),
NeedleMap::Redb(nm) => nm.has_idx_writer(),
// Appends open the .idx on demand, so one is always available.
NeedleMap::SortedFile(_) => true,
}
}
@@ -1025,6 +1042,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.content_size(),
NeedleMap::Redb(nm) => nm.content_size(),
NeedleMap::SortedFile(nm) => nm.content_size(),
}
}
@@ -1033,6 +1051,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.deleted_size(),
NeedleMap::Redb(nm) => nm.deleted_size(),
NeedleMap::SortedFile(nm) => nm.deleted_size(),
}
}
@@ -1041,6 +1060,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.file_count(),
NeedleMap::Redb(nm) => nm.file_count(),
NeedleMap::SortedFile(nm) => nm.file_count(),
}
}
@@ -1049,6 +1069,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.deleted_count(),
NeedleMap::Redb(nm) => nm.deleted_count(),
NeedleMap::SortedFile(nm) => nm.deleted_count(),
}
}
@@ -1057,6 +1078,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.max_file_key(),
NeedleMap::Redb(nm) => nm.max_file_key(),
NeedleMap::SortedFile(nm) => nm.max_file_key(),
}
}
@@ -1067,6 +1089,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.max_needle_end(),
NeedleMap::Redb(nm) => nm.max_needle_end(),
NeedleMap::SortedFile(nm) => nm.max_needle_end(),
}
}
@@ -1075,6 +1098,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.index_file_size(),
NeedleMap::Redb(nm) => nm.index_file_size(),
NeedleMap::SortedFile(nm) => nm.index_file_size(),
}
}
@@ -1083,6 +1107,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.sync(),
NeedleMap::Redb(nm) => nm.sync(),
NeedleMap::SortedFile(nm) => nm.sync(),
}
}
@@ -1091,6 +1116,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.close(),
NeedleMap::Redb(nm) => nm.close(),
NeedleMap::SortedFile(nm) => nm.close(),
}
}
@@ -1099,6 +1125,7 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.save_to_idx(path),
NeedleMap::Redb(nm) => nm.save_to_idx(path),
NeedleMap::SortedFile(nm) => nm.save_to_idx(path),
}
}
@@ -1110,22 +1137,28 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.ascending_visit(f),
NeedleMap::Redb(nm) => nm.ascending_visit(f),
NeedleMap::SortedFile(nm) => nm.ascending_visit(f),
}
}
/// Iterate all entries. Returns a Vec of (NeedleId, NeedleValue) pairs.
/// For InMemory this collects via ascending visit; for Redb it reads from disk.
pub fn iter_entries(&self) -> Vec<(NeedleId, NeedleValue)> {
/// For InMemory this collects via ascending visit; the disk-backed maps read
/// it back off disk, so a truncated .sdx or a redb read fault surfaces here
/// as an error. Compaction treats the result as the complete live set, so a
/// partial scan must never be mistaken for an empty tail.
pub fn iter_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
match self {
NeedleMap::InMemory(nm) => {
let mut entries = Vec::new();
// The visitor never fails, so neither can this.
let _ = nm.ascending_visit(|id, nv| {
entries.push((id, *nv));
Ok(())
});
entries
Ok(entries)
}
NeedleMap::Redb(nm) => nm.collect_entries(),
NeedleMap::SortedFile(nm) => nm.iter_entries(),
}
}
}
@@ -1280,13 +1313,13 @@ mod tests {
nm.put(NeedleId(2), Offset::from_actual_offset(128), Size(200))
.unwrap();
let v1 = nm.get(NeedleId(1)).unwrap();
let v1 = nm.get(NeedleId(1)).unwrap().unwrap();
assert_eq!(v1.size, Size(100));
let v2 = nm.get(NeedleId(2)).unwrap();
let v2 = nm.get(NeedleId(2)).unwrap().unwrap();
assert_eq!(v2.size, Size(200));
assert!(nm.get(NeedleId(99)).is_none());
assert!(nm.get(NeedleId(99)).unwrap().is_none());
}
#[test]
@@ -1311,7 +1344,7 @@ mod tests {
assert_eq!(nm.deleted_size(), 100);
// Deleted entry should have negated size
let nv = nm.get(NeedleId(1)).unwrap();
let nv = nm.get(NeedleId(1)).unwrap().unwrap();
assert_eq!(nv.size, Size(-100));
}
@@ -1384,9 +1417,9 @@ mod tests {
let mut cursor = Cursor::new(idx_data);
let nm = RedbNeedleMap::load_from_idx(db_path.to_str().unwrap(), &mut cursor, Version::current()).unwrap();
assert!(nm.get(NeedleId(1)).is_some());
assert!(nm.get(NeedleId(2)).is_none()); // deleted and removed
assert!(nm.get(NeedleId(3)).is_some());
assert!(nm.get(NeedleId(1)).unwrap().is_some());
assert!(nm.get(NeedleId(2)).unwrap().is_none()); // deleted and removed
assert!(nm.get(NeedleId(3)).unwrap().is_some());
assert_eq!(nm.file_count(), 2);
}
@@ -1503,7 +1536,7 @@ mod tests {
let mut nm = NeedleMap::InMemory(CompactNeedleMap::new());
nm.put(NeedleId(1), Offset::from_actual_offset(0), Size(100))
.unwrap();
assert_eq!(nm.get(NeedleId(1)).unwrap().size, Size(100));
assert_eq!(nm.get(NeedleId(1)).unwrap().unwrap().size, Size(100));
assert_eq!(nm.file_count(), 1);
}
@@ -1514,7 +1547,7 @@ mod tests {
let mut nm = NeedleMap::Redb(RedbNeedleMap::new(db_path.to_str().unwrap()).unwrap());
nm.put(NeedleId(1), Offset::from_actual_offset(0), Size(100))
.unwrap();
assert_eq!(nm.get(NeedleId(1)).unwrap().size, Size(100));
assert_eq!(nm.get(NeedleId(1)).unwrap().unwrap().size, Size(100));
assert_eq!(nm.file_count(), 1);
}
}
@@ -0,0 +1,247 @@
//! Bounded pool of open index-file descriptors.
//!
//! Read-only volumes — cloud-tiered ones above all — outnumber writable ones by
//! orders of magnitude on a large server, and a volume that pins its `.idx` and
//! `.sdx` for the life of the process costs two descriptors whether or not
//! anybody reads it. At ~600K volumes per server that alone exhausts any fd
//! limit. Neither file is needed except while a lookup is in flight, so
//! [`SortedFileNeedleMap`](super::sorted_file::SortedFileNeedleMap) borrows them
//! from this pool: an idle volume holds nothing, a busy one keeps its handles
//! hot rather than paying an `open()` per needle.
//!
//! Mirrors Go's `weed/storage/needle_map_file_pool.go`. Handles are handed out
//! as `Arc<File>`, so an eviction cannot close a descriptor a reader still
//! holds — the file closes when the last borrower drops its `Arc`.
use std::collections::{BTreeMap, HashMap};
use std::fs::{File, OpenOptions};
use std::io;
use std::sync::{Arc, Mutex, OnceLock};
/// Descriptors the pool keeps open. Matches Go's `maxPooledIndexFiles`.
pub const MAX_POOLED_INDEX_FILES: usize = 1024;
struct Entry {
file: Arc<File>,
tick: u64,
}
#[derive(Default)]
struct Inner {
entries: HashMap<String, Entry>,
/// Recency order, oldest tick first, so eviction is a `pop_first`.
order: BTreeMap<u64, String>,
next_tick: u64,
}
pub struct IndexFilePool {
capacity: usize,
inner: Mutex<Inner>,
}
/// Writable and read-only handles for the same path are pooled separately so a
/// read never depends on the file being openable for write — a volume served
/// off a read-only mount still answers lookups.
fn pool_key(path: &str, writable: bool) -> String {
if writable {
format!("{path}\0rw")
} else {
path.to_string()
}
}
impl IndexFilePool {
pub fn new(capacity: usize) -> Self {
IndexFilePool {
capacity: capacity.max(1),
inner: Mutex::new(Inner::default()),
}
}
/// Hand out an open handle for `path`, reusing the pooled one when there is
/// one. The descriptor lives as long as the returned `Arc`.
pub fn borrow(&self, path: &str, writable: bool) -> io::Result<Arc<File>> {
let key = pool_key(path, writable);
if let Some(file) = self.touch(&key) {
return Ok(file);
}
// Opened outside the lock: a cold open blocks on disk, and holding a
// process-wide mutex across it would serialize every volume's lookups.
let file = Arc::new(OpenOptions::new().read(true).write(writable).open(path)?);
Ok(self.insert(key, file))
}
/// Forget the pooled handles for `path`, so a later rename or delete of that
/// path cannot be served from a descriptor on the old inode.
pub fn discard(&self, path: &str) {
let mut inner = self.inner.lock().unwrap();
for key in [pool_key(path, false), pool_key(path, true)] {
if let Some(entry) = inner.entries.remove(&key) {
inner.order.remove(&entry.tick);
}
}
}
/// Descriptors currently pooled. Test-only visibility into the bound.
#[cfg(test)]
pub fn pooled_count(&self) -> usize {
self.inner.lock().unwrap().entries.len()
}
fn touch(&self, key: &str) -> Option<Arc<File>> {
let mut inner = self.inner.lock().unwrap();
let tick = inner.next_tick;
let entry = inner.entries.get_mut(key)?;
let file = entry.file.clone();
let old_tick = std::mem::replace(&mut entry.tick, tick);
inner.order.remove(&old_tick);
inner.order.insert(tick, key.to_string());
inner.next_tick += 1;
Some(file)
}
fn insert(&self, key: String, file: Arc<File>) -> Arc<File> {
let mut inner = self.inner.lock().unwrap();
if let Some(entry) = inner.entries.get(&key) {
// Another borrower opened the same path first; keep one descriptor.
return entry.file.clone();
}
let tick = inner.next_tick;
inner.next_tick += 1;
inner.order.insert(tick, key.clone());
inner.entries.insert(
key,
Entry {
file: file.clone(),
tick,
},
);
while inner.entries.len() > self.capacity {
let Some((_, oldest)) = inner.order.pop_first() else {
break;
};
inner.entries.remove(&oldest);
}
file
}
}
/// Process-wide pool shared by every read-only volume on this server.
pub fn pooled_index_files() -> &'static IndexFilePool {
static POOL: OnceLock<IndexFilePool> = OnceLock::new();
POOL.get_or_init(|| IndexFilePool::new(MAX_POOLED_INDEX_FILES))
}
/// Descriptors this process holds on `.idx`/`.sdx` files under `dir`, read from
/// `/proc/self/fd` where it exists and from `lsof` otherwise. `None` when
/// neither is available, so a caller can skip rather than assert vacuously.
#[cfg(test)]
pub(crate) fn open_index_fds(dir: &std::path::Path) -> Option<usize> {
let prefix = std::fs::canonicalize(dir).unwrap_or_else(|_| dir.to_path_buf());
let is_index = |target: &std::path::Path| {
target.starts_with(&prefix)
&& matches!(
target.extension().and_then(|e| e.to_str()),
Some("idx") | Some("sdx")
)
};
if let Ok(entries) = std::fs::read_dir("/proc/self/fd") {
return Some(
entries
.filter_map(|e| std::fs::read_link(e.ok()?.path()).ok())
.filter(|target| is_index(target))
.count(),
);
}
let out = std::process::Command::new("lsof")
.args(["-p", &std::process::id().to_string(), "-F", "n"])
.output()
.ok()?;
Some(
String::from_utf8_lossy(&out.stdout)
.lines()
.filter_map(|line| line.strip_prefix('n'))
.filter(|line| is_index(std::path::Path::new(line)))
.count(),
)
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::Write;
fn write_file(dir: &std::path::Path, name: &str, contents: &[u8]) -> String {
let path = dir.join(name);
let mut f = File::create(&path).unwrap();
f.write_all(contents).unwrap();
path.to_str().unwrap().to_string()
}
#[test]
fn evicted_handle_stays_usable_for_its_borrower() {
let dir = tempfile::tempdir().unwrap();
let first = write_file(dir.path(), "first", b"first");
let second = write_file(dir.path(), "second", b"second");
let pool = IndexFilePool::new(1);
let borrowed = pool.borrow(&first, false).unwrap();
// Pushes the single slot over, evicting the entry still in use.
let _other = pool.borrow(&second, false).unwrap();
assert_eq!(pool.pooled_count(), 1);
let mut buf = [0u8; 5];
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
borrowed.read_exact_at(&mut buf, 0).unwrap();
}
assert_eq!(&buf, b"first");
}
#[test]
fn borrow_reuses_the_pooled_handle() {
let dir = tempfile::tempdir().unwrap();
let path = write_file(dir.path(), "idx", b"x");
let pool = IndexFilePool::new(4);
let a = pool.borrow(&path, false).unwrap();
let b = pool.borrow(&path, false).unwrap();
assert!(Arc::ptr_eq(&a, &b));
assert_eq!(pool.pooled_count(), 1);
// A writable handle is pooled separately from the read-only one.
let w = pool.borrow(&path, true).unwrap();
assert!(!Arc::ptr_eq(&a, &w));
assert_eq!(pool.pooled_count(), 2);
}
#[test]
fn discard_drops_both_handles() {
let dir = tempfile::tempdir().unwrap();
let path = write_file(dir.path(), "idx", b"x");
let pool = IndexFilePool::new(4);
let _r = pool.borrow(&path, false).unwrap();
let _w = pool.borrow(&path, true).unwrap();
assert_eq!(pool.pooled_count(), 2);
pool.discard(&path);
assert_eq!(pool.pooled_count(), 0);
}
#[test]
fn pool_stays_within_capacity() {
let dir = tempfile::tempdir().unwrap();
let pool = IndexFilePool::new(3);
for i in 0..10 {
let path = write_file(dir.path(), &format!("f{i}"), b"x");
let _ = pool.borrow(&path, false).unwrap();
}
assert_eq!(pool.pooled_count(), 3);
}
}
File diff suppressed because it is too large Load Diff
+231 -14
View File
@@ -259,11 +259,25 @@ impl Store {
/// Returns the index of the disk that should receive a new EC
/// shard / index file for `(collection, vid)`. Selection order:
///
/// 0. a disk that already owns one of `shard_ids` (in-memory claim),
/// 1. a disk that already has the EC volume mounted (in-memory state),
/// 2. a disk that owns the `.ecx` file on disk (volume not yet mounted),
/// 3. any HDD with free space,
/// 4. any disk with free space.
///
/// Step 0 keeps the per-server invariant that a shard id is owned by
/// at most one disk. Steps 1-4 only know the volume, and a multi-disk
/// server legitimately mounts the same vid on several disks, so the
/// free-count tie-break alone can send a re-copy of a shard the server
/// already holds (a retried `ec.balance` / `ec.rebuild` move) to a
/// sibling disk. Both disks then claim the same (vid, shard) and
/// report it to the master from two disk ids, and which claimant
/// serves reads or survives a later unmount/delete of the shard id
/// becomes an accident of location order. Overwriting in place is what
/// the caller meant, so an owning disk wins ahead of the space filters
/// too — a re-copy needs no new shard slot, and a genuinely full disk
/// fails the write rather than silently splitting the claim.
///
/// Step 2 is the missing primitive that pinned subsequent shards to
/// the first-shard disk during `ec.rebuild`: rebuild only sets
/// `CopyEcxFile=true` on the first shard, then relies on auto-select
@@ -286,20 +300,26 @@ impl Store {
collection: &str,
vid: VolumeId,
data_shard_count: u32,
shard_ids: &[u32],
) -> Option<usize> {
const TIER_ANY_DISK: u8 = 1;
const TIER_HDD: u8 = 2;
const TIER_ECX_ON_DISK: u8 = 3;
const TIER_MOUNTED: u8 = 4;
const TIER_OWNS_SHARD: u8 = 5;
let mut best: Option<(usize, u8, i64)> = None;
// (index, tier, owned shard count, free shard slots)
let mut best: Option<(usize, u8, usize, i64)> = None;
for (i, loc) in self.locations.iter().enumerate() {
if loc.is_disk_space_low.load(Ordering::Relaxed) {
continue;
}
let owned = owned_ec_shard_count(loc, vid, shard_ids);
let free = ec_free_shard_count(loc, data_shard_count);
if free <= 0 {
continue;
if owned == 0 {
if loc.is_disk_space_low.load(Ordering::Relaxed) {
continue;
}
if free <= 0 {
continue;
}
}
let mut tier = TIER_ANY_DISK;
if loc.disk_type == DiskType::HardDrive {
@@ -311,15 +331,41 @@ impl Store {
if loc.has_ec_volume(vid) {
tier = TIER_MOUNTED;
}
if owned > 0 {
tier = TIER_OWNS_SHARD;
}
let better = match best {
None => true,
Some((_, b_tier, b_free)) => tier > b_tier || (tier == b_tier && free > b_free),
// owned only separates disks inside TIER_OWNS_SHARD; it is 0
// everywhere else, so this falls through to the free-count
// tie-break for the other tiers.
Some((_, b_tier, b_owned, b_free)) => {
tier > b_tier
|| (tier == b_tier
&& (owned > b_owned || (owned == b_owned && free > b_free)))
}
};
if better {
best = Some((i, tier, free));
best = Some((i, tier, owned, free));
}
}
best.map(|(i, _, _)| i)
best.map(|(i, _, _, _)| i)
}
/// Returns the distinct disk indexes that already own one of `shard_ids`
/// for `vid`, in location order. More than one owner means the batch has
/// no single correct destination — whichever disk receives it would
/// duplicate a sibling disk's claim — so batch callers must split by
/// owner (`volume_ec_shards_copy` refuses such a batch instead of
/// guessing). Mirrors `Store.EcShardOwnerDisks` in
/// `weed/storage/store_ec.go`.
pub fn ec_shard_owner_disks(&self, vid: VolumeId, shard_ids: &[u32]) -> Vec<usize> {
self.locations
.iter()
.enumerate()
.filter(|(_, loc)| owned_ec_shard_count(loc, vid, shard_ids) > 0)
.map(|(i, _)| i)
.collect()
}
/// Create a new volume, placing it on the location with the most free space.
@@ -804,6 +850,37 @@ impl Store {
}
/// Find an EC volume across all locations (mutable).
/// Every per-disk `EcVolume` this store maps for `vid`. A vid can mount on
/// N disks as N distinct runtimes, and the first-match `find_ec_volume_mut`
/// hides the siblings — so anything that has to reach the whole volume,
/// rather than any one runtime of it, iterates this instead.
/// Every directory on this server that could hold an EC volume's metadata —
/// each disk's data and index directory. Startup mirroring gives each
/// shard-bearing disk its own .ecx/.ecj/.vif, but the checksum sidecar is
/// not mirrored and a repair delivers exactly one copy, so a runtime looking
/// only at its own two directories cannot see it. Handing this list to the
/// sidecar resolution keeps one authoritative copy reachable from every
/// runtime rather than duplicating a file that is rewritten as shards are
/// repaired and generations published.
pub fn ec_metadata_dirs(&self) -> Vec<String> {
let mut dirs: Vec<String> = Vec::with_capacity(self.locations.len() * 2);
for loc in &self.locations {
for dir in [&loc.directory, &loc.idx_directory] {
if !dir.is_empty() && !dirs.iter().any(|d| d == dir) {
dirs.push(dir.clone());
}
}
}
dirs
}
pub fn find_all_ec_volumes_mut(&mut self, vid: VolumeId) -> Vec<&mut EcVolume> {
self.locations
.iter_mut()
.filter_map(|loc| loc.find_ec_volume_mut(vid))
.collect()
}
pub fn find_ec_volume_mut(&mut self, vid: VolumeId) -> Option<&mut EcVolume> {
for loc in &mut self.locations {
if let Some(ecv) = loc.find_ec_volume_mut(vid) {
@@ -1302,6 +1379,20 @@ fn ec_free_shard_count(loc: &DiskLocation, data_shard_count: u32) -> i64 {
free
}
/// Reports how many of `shard_ids` this disk already claims for `vid`, per
/// the in-memory registration the read path and heartbeats use.
///
/// Mirrors `ownedEcShardCount` in `weed/storage/store_ec.go`.
fn owned_ec_shard_count(loc: &DiskLocation, vid: VolumeId, shard_ids: &[u32]) -> usize {
let Some(ecv) = loc.find_ec_volume(vid) else {
return 0;
};
shard_ids
.iter()
.filter(|&&shard_id| ecv.has_shard(shard_id as u8))
.count()
}
// ============================================================================
// Tests
// ============================================================================
@@ -1840,7 +1931,7 @@ mod tests {
let base = volume_file_name(&store.locations[2].idx_directory, collection, vid);
std::fs::write(format!("{}.ecx", base), vec![0u8; 20]).unwrap();
let got = store.find_ec_shard_target_location(collection, vid, 10);
let got = store.find_ec_shard_target_location(collection, vid, 10, &[]);
assert_eq!(
got,
Some(2),
@@ -1982,7 +2073,7 @@ mod tests {
let base = volume_file_name(&store.locations[2].idx_directory, collection, vid);
std::fs::write(format!("{}.ecx", base), vec![0u8; 20]).unwrap();
let got = store.find_ec_shard_target_location(collection, vid, 10);
let got = store.find_ec_shard_target_location(collection, vid, 10, &[]);
assert_eq!(got, Some(1), "expected the mounted disk to win; got {:?}", got);
}
@@ -1991,7 +2082,7 @@ mod tests {
#[test]
fn test_find_ec_shard_target_location_falls_through_to_hdd_when_nothing_matches() {
let (store, _tmp) = make_ec_target_test_store(2);
let got = store.find_ec_shard_target_location("grafana-loki", VolumeId(3333), 10);
let got = store.find_ec_shard_target_location("grafana-loki", VolumeId(3333), 10, &[]);
assert!(got.is_some(), "expected an HDD fallback");
assert_eq!(store.locations[got.unwrap()].disk_type, DiskType::HardDrive);
}
@@ -2007,7 +2098,7 @@ mod tests {
.max_volume_count
.store(0, Ordering::Relaxed);
let got = store.find_ec_shard_target_location("grafana-loki", VolumeId(4444), 10);
let got = store.find_ec_shard_target_location("grafana-loki", VolumeId(4444), 10, &[]);
assert_eq!(
got,
Some(0),
@@ -2045,7 +2136,7 @@ mod tests {
.mount_ec_shards(vid, collection, &[0], "")
.unwrap();
let got = store.find_ec_shard_target_location(collection, vid, 10);
let got = store.find_ec_shard_target_location(collection, vid, 10, &[]);
assert_eq!(
got,
Some(1),
@@ -2053,4 +2144,130 @@ mod tests {
got,
);
}
/// Per-server invariant: a shard id is owned by at most one disk.
///
/// A multi-disk server legitimately mounts one vid on several disks, each
/// holding a disjoint subset of the shards, so the mounted tier ties and
/// the free-count tie-break decides — and it points at whichever disk
/// happens to be emptier, not at the disk that already has this shard. A
/// re-copy of a shard the server already holds (a retried `ec.balance` /
/// `ec.rebuild` move) then lands a second copy on the sibling disk, and
/// both disks register the same shard id.
#[test]
fn test_find_ec_shard_target_location_pins_to_the_disk_owning_the_shard() {
let (mut store, _tmp) = make_ec_target_test_store(2);
let collection = "grafana-loki";
let vid = VolumeId(8888);
// Disk 0 owns shards 0 and 1, disk 1 owns shard 2 — so disk 1 is the
// emptier of the two and wins the free-count tie-break.
let base0 = volume_file_name(&store.locations[0].directory, collection, vid);
std::fs::write(format!("{}.ec00", base0), b"x").unwrap();
std::fs::write(format!("{}.ec01", base0), b"x").unwrap();
store.locations[0]
.mount_ec_shards(vid, collection, &[0, 1], "")
.unwrap();
let base1 = volume_file_name(&store.locations[1].directory, collection, vid);
std::fs::write(format!("{}.ec02", base1), b"x").unwrap();
store.locations[1]
.mount_ec_shards(vid, collection, &[2], "")
.unwrap();
assert_eq!(
store.find_ec_shard_target_location(collection, vid, 10, &[0]),
Some(0),
"a copy of shard 0 left its owning disk",
);
assert_eq!(
store.find_ec_shard_target_location(collection, vid, 10, &[2]),
Some(1),
"a copy of shard 2 left its owning disk",
);
// A shard no disk owns yet is placed by the unchanged waterfall: both
// disks have it mounted, so the emptier one wins.
assert_eq!(
store.find_ec_shard_target_location(collection, vid, 10, &[7]),
Some(1),
"placement of an unclaimed shard changed",
);
}
/// The second half of the invariant: a disk with no free shard slots
/// still wins for a shard it already owns. Re-copying that shard
/// overwrites bytes the disk is already accounted for, while routing to a
/// sibling splits the claim across two disks.
#[test]
fn test_find_ec_shard_target_location_owning_disk_wins_when_full() {
let (mut store, _tmp) = make_ec_target_test_store(2);
store.locations[0]
.max_volume_count
.store(1, Ordering::Relaxed);
let collection = "grafana-loki";
let vid = VolumeId(9999);
let base = volume_file_name(&store.locations[0].directory, collection, vid);
std::fs::write(format!("{}.ec00", base), b"x").unwrap();
store.locations[0]
.mount_ec_shards(vid, collection, &[0], "")
.unwrap();
// Fill disk 0 past its shard-slot budget so ec_free_shard_count is 0.
let filler = VolumeId(10000);
let filler_base = volume_file_name(&store.locations[0].directory, collection, filler);
let filler_shards: Vec<u32> = (0..10).collect();
for shard_id in &filler_shards {
std::fs::write(format!("{}.ec{:02}", filler_base, shard_id), b"x").unwrap();
}
store.locations[0]
.mount_ec_shards(filler, collection, &filler_shards, "")
.unwrap();
assert_eq!(
store.find_ec_shard_target_location(collection, vid, 10, &[0]),
Some(0),
"a copy of shard 0 left its owning disk when the disk was full",
);
// A shard it does not own still respects the space filter.
assert_eq!(
store.find_ec_shard_target_location(collection, vid, 10, &[7]),
Some(1),
"an unclaimed shard should go to the disk with free slots",
);
}
/// Mixed-owner batch contract: a batch whose requested shards are
/// already owned by different disks reports every owner, so
/// `volume_ec_shards_copy` can refuse it rather than rank the owners
/// into one destination and duplicate the loser's claim.
#[test]
fn test_ec_shard_owner_disks() {
let (mut store, _tmp) = make_ec_target_test_store(3);
let collection = "grafana-loki";
let vid = VolumeId(11111);
let base0 = volume_file_name(&store.locations[0].directory, collection, vid);
std::fs::write(format!("{}.ec00", base0), b"x").unwrap();
std::fs::write(format!("{}.ec01", base0), b"x").unwrap();
store.locations[0]
.mount_ec_shards(vid, collection, &[0, 1], "")
.unwrap();
let base1 = volume_file_name(&store.locations[1].directory, collection, vid);
std::fs::write(format!("{}.ec02", base1), b"x").unwrap();
store.locations[1]
.mount_ec_shards(vid, collection, &[2], "")
.unwrap();
assert_eq!(
store.ec_shard_owner_disks(vid, &[0, 2]),
vec![0, 1],
"a batch owned by two disks must report both owners",
);
// A batch on one disk, with or without unowned extras, has one owner.
assert_eq!(store.ec_shard_owner_disks(vid, &[0, 1, 7]), vec![0]);
// A wholly unowned batch reports none — fresh placement stays allowed.
assert!(store.ec_shard_owner_disks(vid, &[7, 8]).is_empty());
}
}
File diff suppressed because it is too large Load Diff
+65
View File
@@ -4666,6 +4666,70 @@ dependencies = [
"prost 0.14.4",
]
[[package]]
name = "protoc-bin-vendored"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d1c381df33c98266b5f08186583660090a4ffa0889e76c7e9a5e175f645a67fa"
dependencies = [
"protoc-bin-vendored-linux-aarch_64",
"protoc-bin-vendored-linux-ppcle_64",
"protoc-bin-vendored-linux-s390_64",
"protoc-bin-vendored-linux-x86_32",
"protoc-bin-vendored-linux-x86_64",
"protoc-bin-vendored-macos-aarch_64",
"protoc-bin-vendored-macos-x86_64",
"protoc-bin-vendored-win32",
]
[[package]]
name = "protoc-bin-vendored-linux-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c350df4d49b5b9e3ca79f7e646fde2377b199e13cfa87320308397e1f37e1a4c"
[[package]]
name = "protoc-bin-vendored-linux-ppcle_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a55a63e6c7244f19b5c6393f025017eb5d793fd5467823a099740a7a4222440c"
[[package]]
name = "protoc-bin-vendored-linux-s390_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1dba5565db4288e935d5330a07c264a4ee8e4a5b4a4e6f4e83fad824cc32f3b0"
[[package]]
name = "protoc-bin-vendored-linux-x86_32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8854774b24ee28b7868cd71dccaae8e02a2365e67a4a87a6cd11ee6cdbdf9cf5"
[[package]]
name = "protoc-bin-vendored-linux-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b38b07546580df720fa464ce124c4b03630a6fb83e05c336fea2a241df7e5d78"
[[package]]
name = "protoc-bin-vendored-macos-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "89278a9926ce312e51f1d999fee8825d324d603213344a9a706daa009f1d8092"
[[package]]
name = "protoc-bin-vendored-macos-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "81745feda7ccfb9471d7a4de888f0652e806d5795b61480605d4943176299756"
[[package]]
name = "protoc-bin-vendored-win32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "quick-xml"
version = "0.39.4"
@@ -5379,6 +5443,7 @@ dependencies = [
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"protoc-bin-vendored",
"tokio",
"tokio-stream",
"tonic",
+27
View File
@@ -12,6 +12,17 @@ one.
`core` knows nothing about any job. A second worker is a new crate beside
`lance` that depends on it, not a fork of the protocol.
## Building
`core` compiles `plugin.proto` with the protoc that protoc-bin-vendored ships,
the way seaweed-volume does, so it needs no system install.
The lance crates compile protos of their own, in their own build-script
processes, which nothing our build script sets can reach. They need a protoc of
their own: either one on PATH — `brew install protobuf`, `apt install
protobuf-compiler` — or `PROTOC` naming one. CI points it at the vendored
binary for the runner's platform, resolved from the version in `Cargo.lock`.
## Running
cargo run -p weed-lance-worker -- --admin 127.0.0.1:23646
@@ -20,6 +31,22 @@ The admin's *HTTP* address is what an operator has; the gRPC port is derived
from it the way the Go side does. Dialling the HTTP port fails as "frame with
invalid size", which reads like a protocol bug rather than a wrong port.
The binary is `weed-worker`, not `weed-lance-worker`: it is the Rust side of
`weed worker`, and lance is the first family of jobs it carries rather than the
only one it ever will.
Released builds do not need a toolchain. The worker ships inside the SeaweedFS
image, beside the Rust volume server, under the verb that mirrors
`volume-rust`:
docker run chrislusf/seaweedfs worker-rust --admin admin:23646
and as `weed-worker_linux_{amd64,arm64}.tar.gz` on each GitHub release. Both are
linux amd64/arm64 only — lance, arrow and datafusion make every extra target an
expensive build, and the worker runs beside the cluster it maintains. On an
architecture without a build the image carries an empty placeholder and the
entrypoint says so rather than failing as "not found".
## Metrics
cargo run -p weed-lance-worker -- --admin 127.0.0.1:23646 --metrics-port 9328
+4
View File
@@ -21,3 +21,7 @@ tracing.workspace = true
[build-dependencies]
tonic-build.workspace = true
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry. The same crate seaweed-volume uses.
protoc-bin-vendored = "3"
+8
View File
@@ -1,4 +1,12 @@
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Use the protoc that ships with protoc-bin-vendored rather than a system
// one, so the build needs no package manager and always sees the same
// version. An explicit PROTOC still wins, for packagers supplying their own
// and for the lance crates, whose own build scripts read the same variable.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
// Compiled straight out of the Go tree, the way seaweed-volume already reads
// filer.proto, so the contract cannot drift from a vendored copy.
tonic_build::configure()
+4 -1
View File
@@ -7,8 +7,11 @@ description = "SeaweedFS maintenance worker for Lance tables"
[lib]
name = "weed_lance_worker"
# The binary is not named for lance: it is the Rust side of `weed worker`, and
# the job families it registers will outgrow this crate. When a second one
# arrives the bin target moves to a crate of its own under the same name.
[[bin]]
name = "weed-lance-worker"
name = "weed-worker"
path = "src/main.rs"
[dependencies]
+2 -2
View File
@@ -13,8 +13,8 @@ use weed_lance_worker::metrics::LanceMetrics;
/// language and an operator should not have to learn a second set of names.
#[derive(Parser, Debug)]
#[command(
name = "weed-lance-worker",
about = "SeaweedFS maintenance worker for Lance tables"
name = "weed-worker",
about = "SeaweedFS maintenance worker"
)]
struct Args {
/// Admin server gRPC address.
+1 -1
View File
@@ -139,7 +139,7 @@ The telemetry server exposes these Prometheus metrics:
### Cluster Metrics
- `seaweedfs_telemetry_total_clusters`: Total unique clusters (30 days)
- `seaweedfs_telemetry_active_clusters`: Active clusters (7 days)
- `seaweedfs_telemetry_confirmed_clusters`: Active clusters seen on 2+ distinct days — one-shot reports don't count, and the version/OS distributions in `/api/stats` are computed over these
- `seaweedfs_telemetry_confirmed_clusters`: Active clusters seen on 7+ distinct days — one-shot reports don't count, and the version/OS distributions in `/api/stats` are computed over these
### Per-Cluster Metrics
- `seaweedfs_telemetry_volume_servers{cluster_id}`: Volume servers per cluster
+44 -2
View File
@@ -2,6 +2,7 @@ package api
import (
"bytes"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
@@ -11,6 +12,9 @@ import (
protobuf "google.golang.org/protobuf/proto"
)
// promauto registers on the global registry: one storage per test binary.
var testHandler = NewHandler(storage.NewPrometheusStorage())
func validReport() *proto.TelemetryData {
return &proto.TelemetryData{
TopologyId: "38422678-6a0d-4482-aa33-65b90010ac47",
@@ -43,8 +47,7 @@ func marshalReport(t *testing.T, data *proto.TelemetryData) []byte {
}
func TestCollectTelemetryValidation(t *testing.T) {
// promauto registers on the global registry: one storage per test binary.
h := NewHandler(storage.NewPrometheusStorage())
h := testHandler
t.Run("valid report accepted", func(t *testing.T) {
if w := postCollect(t, h, marshalReport(t, validReport()), "application/x-protobuf"); w.Code != http.StatusOK {
@@ -96,3 +99,42 @@ func TestCollectTelemetryValidation(t *testing.T) {
}
})
}
// The dashboard looks confirmation windows up by the select's string value,
// so the stats JSON must key confirmed_by_days by decimal strings and carry
// unmet thresholds as zeros.
func TestStatsSerializedThresholds(t *testing.T) {
data := validReport()
data.TopologyId = "49533789-7b1e-4593-bb44-76ca1121bd58"
if w := postCollect(t, testHandler, marshalReport(t, data), "application/x-protobuf"); w.Code != http.StatusOK {
t.Fatalf("collect: got %d: %s", w.Code, w.Body.String())
}
req := httptest.NewRequest(http.MethodGet, "/api/stats", nil)
w := httptest.NewRecorder()
testHandler.GetStats(w, req)
if w.Code != http.StatusOK {
t.Fatalf("stats: got %d", w.Code)
}
var stats struct {
ConfirmedByDays map[string]int `json:"confirmed_by_days"`
}
if err := json.Unmarshal(w.Body.Bytes(), &stats); err != nil {
t.Fatalf("decode: %v", err)
}
if len(stats.ConfirmedByDays) != 5 {
t.Fatalf("confirmed_by_days = %v, want the 5 thresholds", stats.ConfirmedByDays)
}
for _, key := range []string{"1", "3", "7", "14", "30"} {
if _, ok := stats.ConfirmedByDays[key]; !ok {
t.Errorf("confirmed_by_days missing %q: %v", key, stats.ConfirmedByDays)
}
}
if stats.ConfirmedByDays["1"] < 1 {
t.Errorf("fresh cluster missing from the 1-day count: %v", stats.ConfirmedByDays)
}
if stats.ConfirmedByDays["30"] != 0 {
t.Errorf("unmet threshold not zero: %v", stats.ConfirmedByDays)
}
}
+30 -2
View File
@@ -57,6 +57,15 @@ func (h *Handler) ServeIndex(w http.ResponseWriter, r *http.Request) {
color: #666;
margin-top: 5px;
}
.stat-label select {
border: none;
background: none;
color: inherit;
font: inherit;
padding: 0;
cursor: pointer;
text-decoration: underline dotted;
}
.chart-container {
background: white;
padding: 20px;
@@ -129,7 +138,14 @@ func (h *Handler) ServeIndex(w http.ResponseWriter, r *http.Request) {
</div>
<div class="stat-card">
<div class="stat-value" id="confirmedInstances">-</div>
<div class="stat-label">Confirmed Clusters (2+ days)</div>
<div class="stat-label">Confirmed Clusters
(<select id="confirmDays" onchange="updateConfirmed()">
<option value="1">1+</option>
<option value="3">3+</option>
<option value="7" selected>7+</option>
<option value="14">14+</option>
<option value="30">30+</option>
</select> days)</div>
</div>
<div class="stat-card">
<div class="stat-value" id="totalVersions">-</div>
@@ -229,14 +245,26 @@ func (h *Handler) ServeIndex(w http.ResponseWriter, r *http.Request) {
}
}
let latestStats = {};
function updateStats(stats) {
latestStats = stats;
document.getElementById('totalInstances').textContent = stats.total_instances || 0;
document.getElementById('activeInstances').textContent = stats.active_instances || 0;
document.getElementById('confirmedInstances').textContent = stats.confirmed_instances || 0;
updateConfirmed();
document.getElementById('totalVersions').textContent = Object.keys(stats.versions || {}).length;
document.getElementById('totalOS').textContent = Object.keys(stats.os_distribution || {}).length;
}
// Servers from before confirmed_by_days fall back to the fixed
// 7-day count.
function updateConfirmed() {
const days = document.getElementById('confirmDays').value;
const byDays = latestStats.confirmed_by_days || {};
const count = byDays[days] !== undefined ? byDays[days] : latestStats.confirmed_instances;
document.getElementById('confirmedInstances').textContent = count || 0;
}
function updateCharts(stats) {
createPieChart('versionChart', 'Version Distribution', stats.versions || {});
createPieChart('osChart', 'Operating System Distribution', stats.os_distribution || {});
+1 -1
View File
@@ -1,6 +1,6 @@
module github.com/seaweedfs/seaweedfs/telemetry/server
go 1.25.8
go 1.26
require (
github.com/prometheus/client_golang v1.24.1
+21 -6
View File
@@ -44,13 +44,18 @@ func TestConfirmedClusters(t *testing.T) {
t.Fatalf("fallback distribution missing active cluster: %v", v)
}
// Give cluster A a sample from yesterday: now seen on 2 distinct days.
// Give cluster A samples from the six previous days: now seen on 7
// distinct days.
s.mu.Lock()
id := "aaaaaaaa-0000-0000-0000-000000000001"
s.histories[id] = append([]HistorySample{{
Ts: time.Now().AddDate(0, 0, -1).Unix(),
TotalDiskBytes: 50,
}}, s.histories[id]...)
var older []HistorySample
for offset := -6; offset < 0; offset++ {
older = append(older, HistorySample{
Ts: time.Now().AddDate(0, 0, offset).Unix(),
TotalDiskBytes: 50,
})
}
s.histories[id] = append(older, s.histories[id]...)
s.mu.Unlock()
// A one-shot cluster B arrives (like an injected report): it counts as
@@ -60,7 +65,7 @@ func TestConfirmedClusters(t *testing.T) {
}
stats = statsOf(t, s)
if stats["active_instances"] != 2 || stats["confirmed_instances"] != 1 {
t.Fatalf("day two: active=%v confirmed=%v, want 2/1", stats["active_instances"], stats["confirmed_instances"])
t.Fatalf("day seven: active=%v confirmed=%v, want 2/1", stats["active_instances"], stats["confirmed_instances"])
}
v := stats["versions"].(map[string]int)
if v["4.40"] != 1 {
@@ -69,4 +74,14 @@ func TestConfirmedClusters(t *testing.T) {
if _, ok := v["9.99"]; ok {
t.Errorf("one-shot cluster polluted the distribution: %v", v)
}
// Cluster A meets the 1/3/7-day thresholds, B only the 1-day one, and
// unmet thresholds are present as zero so the dashboard can show them.
byDays := stats["confirmed_by_days"].(map[int]int)
want := map[int]int{1: 2, 3: 1, 7: 1, 14: 0, 30: 0}
for threshold, expected := range want {
if got, ok := byDays[threshold]; !ok || got != expected {
t.Errorf("confirmed_by_days[%d] = %v (present=%v), want %d", threshold, got, ok, expected)
}
}
}
+8 -4
View File
@@ -8,7 +8,11 @@ import (
// confirmDays is how many distinct UTC days a cluster must have reported
// on before it counts as confirmed in the aggregated stats.
const confirmDays = 2
const confirmDays = 7
// confirmThresholds are the confirmation windows the dashboard lets the
// viewer pick between; confirmDays is the one everything else is built on.
var confirmThresholds = []int{1, 3, 7, 14, 30}
// activeDays is how recently a cluster must have reported to count as active.
const activeDays = 7
@@ -43,9 +47,9 @@ func (s *PrometheusStorage) appendHistory(data *proto.TelemetryData, receivedAt
}
// seriesHistories picks the clusters the fleet-wide series are built from: the
// confirmed ones. A cluster that only ever reported on one day is usually a CI
// or test cluster that lived for a minute, and those arrive faster than they
// age out, so counting them makes every fleet total climb forever. Falls back to
// confirmed ones. A cluster that reported for less than a week is usually a CI
// or test cluster, and those arrive faster than they age out, so counting them
// makes every fleet total climb forever. Falls back to
// all clusters while none is confirmed yet, so a fresh server still draws its
// charts. Callers must hold s.mu.
func (s *PrometheusStorage) seriesHistories() map[string][]HistorySample {
+18 -11
View File
@@ -15,7 +15,8 @@ func TestGetMetricsSumsEachDay(t *testing.T) {
seedSamples(s, "daily", HistorySample{TotalDiskBytes: 300, VolumeServerCount: 3},
-9, -8, -7, -6, -5, -4, -3, -2, -1, 0)
// Stopped reporting past the active window: counts on its own days only.
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9}, -9, -8)
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9},
-14, -13, -12, -11, -10, -9, -8)
metrics, err := s.GetMetrics(10)
if err != nil {
@@ -40,7 +41,8 @@ func TestGetMetricsSumsEachDay(t *testing.T) {
func TestGetMetricsCarriesSkippedDaysForward(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
seedSamples(s, "gappy", HistorySample{TotalDiskBytes: 500, VolumeServerCount: 5}, -3, -1)
seedSamples(s, "gappy", HistorySample{TotalDiskBytes: 500, VolumeServerCount: 5},
-8, -7, -6, -5, -4, -3, -1)
metrics, err := s.GetMetrics(4)
if err != nil {
@@ -56,21 +58,23 @@ func TestGetMetricsCarriesSkippedDaysForward(t *testing.T) {
// out of nothing on its first day of data.
func TestGetMetricsWindowStartsAtOldestSample(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
seedSamples(s, "recent", HistorySample{TotalDiskBytes: 100, VolumeServerCount: 1}, -2, -1, 0)
seedSamples(s, "recent", HistorySample{TotalDiskBytes: 100, VolumeServerCount: 1},
-6, -5, -4, -3, -2, -1, 0)
metrics, err := s.GetMetrics(30)
if err != nil {
t.Fatal(err)
}
if got := metrics["dates"].([]string); len(got) != 3 {
t.Errorf("dates = %v, want the 3 days with history, not 30", got)
if got := metrics["dates"].([]string); len(got) != 7 {
t.Errorf("dates = %v, want the 7 days with history, not 30", got)
}
if got := metrics["disk_usage"].([]uint64); !equal(got, []uint64{100, 100, 100}) {
if got := metrics["disk_usage"].([]uint64); !equal(got, []uint64{100, 100, 100, 100, 100, 100, 100}) {
t.Errorf("disk_usage = %v, want no leading zero days", got)
}
// History reaching past the requested window still clips to the window.
seedSamples(s, "old", HistorySample{TotalDiskBytes: 50, VolumeServerCount: 1}, -40, -39)
seedSamples(s, "old", HistorySample{TotalDiskBytes: 50, VolumeServerCount: 1},
-45, -44, -43, -42, -41, -40, -39)
metrics, err = s.GetMetrics(10)
if err != nil {
t.Fatal(err)
@@ -86,8 +90,10 @@ func TestGetMetricsAgreesWithClusterSizes(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
seedSamples(s, "daily", HistorySample{TotalDiskBytes: 300, VolumeServerCount: 3},
-9, -8, -7, -6, -5, -4, -3, -2, -1, 0)
seedSamples(s, "lagging", HistorySample{TotalDiskBytes: 200, VolumeServerCount: 2}, -3, -2)
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9}, -9, -8)
seedSamples(s, "lagging", HistorySample{TotalDiskBytes: 200, VolumeServerCount: 2},
-8, -7, -6, -5, -4, -3, -2)
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9},
-14, -13, -12, -11, -10, -9, -8)
metrics, err := s.GetMetrics(10)
if err != nil {
@@ -111,7 +117,8 @@ func TestGetMetricsAgreesWithClusterSizes(t *testing.T) {
func TestGetMetricsExcludesUnconfirmedClusters(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
seedSamples(s, "real", HistorySample{TotalDiskBytes: 300, VolumeServerCount: 3}, -3, -2, -1, 0)
seedSamples(s, "real", HistorySample{TotalDiskBytes: 300, VolumeServerCount: 3},
-6, -5, -4, -3, -2, -1, 0)
for _, id := range []string{"ci-1", "ci-2", "ci-3"} {
seedSamples(s, id, HistorySample{TotalDiskBytes: 5, VolumeServerCount: 14}, -1)
}
@@ -128,7 +135,7 @@ func TestGetMetricsExcludesUnconfirmedClusters(t *testing.T) {
}
}
// Until any cluster has two days of history the charts fall back to every
// Until any cluster has a week of history the charts fall back to every
// cluster, so a fresh server doesn't serve empty series.
func TestGetMetricsFallsBackWhenNoneConfirmed(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
+17 -6
View File
@@ -53,7 +53,7 @@ func newPrometheusStorage(reg prometheus.Registerer) *PrometheusStorage {
}),
confirmedClusters: promauto.NewGauge(prometheus.GaugeOpts{
Name: "seaweedfs_telemetry_confirmed_clusters",
Help: "Active clusters seen on at least 2 distinct days (last 7 days)",
Help: "Active clusters seen on at least 7 distinct days (last 7 days)",
}),
volumeServerCount: promauto.NewGaugeVec(prometheus.GaugeOpts{
Name: "seaweedfs_telemetry_volume_servers",
@@ -207,6 +207,10 @@ func (s *PrometheusStorage) updateStats() {
totalInstances := 0
activeInstances := 0
confirmedInstances := 0
confirmedByDays := make(map[int]int, len(confirmThresholds))
for _, threshold := range confirmThresholds {
confirmedByDays[threshold] = 0
}
versionsAll := make(map[string]int)
osAll := make(map[string]int)
versionsConfirmed := make(map[string]int)
@@ -220,10 +224,16 @@ func (s *PrometheusStorage) updateStats() {
activeInstances++
versionsAll[instance.TelemetryData.Version]++
osAll[instance.TelemetryData.Os]++
// A cluster is confirmed once seen on >=2 distinct UTC days
// (histories hold one sample per day), so one-shot reports
// can't skew the distributions below.
if len(s.histories[instance.TelemetryData.TopologyId]) >= confirmDays {
// A cluster is confirmed once seen on confirmDays distinct UTC
// days (histories hold one sample per day), so short-lived
// clusters can't skew the distributions below.
daysSeen := len(s.histories[instance.TelemetryData.TopologyId])
for _, threshold := range confirmThresholds {
if daysSeen >= threshold {
confirmedByDays[threshold]++
}
}
if daysSeen >= confirmDays {
confirmedInstances++
versionsConfirmed[instance.TelemetryData.Version]++
osConfirmed[instance.TelemetryData.Os]++
@@ -231,7 +241,7 @@ func (s *PrometheusStorage) updateStats() {
}
}
// Before any cluster has two days of history (fresh server with no
// Before any cluster has a week of history (fresh server with no
// prior state), fall back to all active clusters so the dashboard
// distributions aren't empty.
versions, osDistribution := versionsConfirmed, osConfirmed
@@ -249,6 +259,7 @@ func (s *PrometheusStorage) updateStats() {
"total_instances": totalInstances,
"active_instances": activeInstances,
"confirmed_instances": confirmedInstances,
"confirmed_by_days": confirmedByDays,
"versions": versions,
"os_distribution": osDistribution,
}
+10 -8
View File
@@ -25,11 +25,13 @@ func TestClusterSizeSeries(t *testing.T) {
// Reported every day of the window.
seedSamples(s, "daily", HistorySample{TotalDiskBytes: 300, VolumeServerCount: 3},
-9, -8, -7, -6, -5, -4, -3, -2, -1, 0)
// Reported two days ago and not since: still active, so its size is held
// to the right edge instead of dropping out of the stack.
seedSamples(s, "lagging", HistorySample{TotalDiskBytes: 200, VolumeServerCount: 2}, -3, -2)
// Stopped reporting two days ago: still active, so its size is held to
// the right edge instead of dropping out of the stack.
seedSamples(s, "lagging", HistorySample{TotalDiskBytes: 200, VolumeServerCount: 2},
-8, -7, -6, -5, -4, -3, -2)
// Stopped reporting past the active window: its own days only.
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9}, -9, -8)
seedSamples(s, "gone", HistorySample{TotalDiskBytes: 900, VolumeServerCount: 9},
-14, -13, -12, -11, -10, -9, -8)
// One day of history only: unconfirmed, so it stays out of the stack.
seedSamples(s, "oneshot", HistorySample{TotalDiskBytes: 400, VolumeServerCount: 4}, -1)
@@ -48,7 +50,7 @@ func TestClusterSizeSeries(t *testing.T) {
if got := byId["daily"].Disk; !equal(got, []uint64{300, 300, 300, 300, 300, 300, 300, 300, 300, 300}) {
t.Errorf("daily = %v, want 300 every day", got)
}
if got := byId["lagging"].Disk; !equal(got, []uint64{0, 0, 0, 0, 0, 0, 200, 200, 200, 200}) {
if got := byId["lagging"].Disk; !equal(got, []uint64{0, 200, 200, 200, 200, 200, 200, 200, 200, 200}) {
t.Errorf("lagging = %v, want its size carried to the right edge", got)
}
if got := byId["gone"].Disk; !equal(got, []uint64{900, 900, 0, 0, 0, 0, 0, 0, 0, 0}) {
@@ -62,7 +64,7 @@ func TestClusterSizeSeries(t *testing.T) {
if got := byId["daily"].Servers; !equal(got, []uint64{3, 3, 3, 3, 3, 3, 3, 3, 3, 3}) {
t.Errorf("daily servers = %v, want 3 every day", got)
}
if got := byId["lagging"].Servers; !equal(got, []uint64{0, 0, 0, 0, 0, 0, 2, 2, 2, 2}) {
if got := byId["lagging"].Servers; !equal(got, []uint64{0, 2, 2, 2, 2, 2, 2, 2, 2, 2}) {
t.Errorf("lagging servers = %v, want carried to the right edge", got)
}
if got := byId["gone"].Servers; !equal(got, []uint64{9, 9, 0, 0, 0, 0, 0, 0, 0, 0}) {
@@ -89,10 +91,10 @@ func TestClusterSizeSeries(t *testing.T) {
if series.Other == nil || series.Other.Count != 2 {
t.Fatalf("other = %+v, want 2 clusters", series.Other)
}
if !equal(series.Other.Disk, []uint64{900, 900, 0, 0, 0, 0, 200, 200, 200, 200}) {
if !equal(series.Other.Disk, []uint64{900, 1100, 200, 200, 200, 200, 200, 200, 200, 200}) {
t.Errorf("other = %v, want lagging+gone summed per day", series.Other.Disk)
}
if !equal(series.Other.Servers, []uint64{9, 9, 0, 0, 0, 0, 2, 2, 2, 2}) {
if !equal(series.Other.Servers, []uint64{9, 11, 2, 2, 2, 2, 2, 2, 2, 2}) {
t.Errorf("other servers = %v, want lagging+gone summed per day", series.Other.Servers)
}
if series.ClusterCount != 3 || series.TotalDisk != 500 || series.TotalServers != 5 {
+15 -15
View File
@@ -10,31 +10,31 @@ func TestVersionSeries(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
// Upgraded mid-window: its band leaves the old version for the new one.
seedSamples(s, "upgraded", HistorySample{Version: "4.39"}, -4, -3)
seedSamples(s, "upgraded", HistorySample{Version: "4.39"}, -6, -5, -4, -3)
seedSamples(s, "upgraded", HistorySample{Version: "4.40"}, -2, -1, 0)
// Reported every day on the same version.
seedSamples(s, "steady", HistorySample{Version: "4.40"}, -4, -3, -2, -1, 0)
seedSamples(s, "steady", HistorySample{Version: "4.40"}, -6, -5, -4, -3, -2, -1, 0)
// One day of history only: unconfirmed, so it stays out of the stack.
seedSamples(s, "oneshot", HistorySample{Version: "4.40"}, -1)
// Confirmed, but its samples predate versions being recorded.
seedSamples(s, "versionless", HistorySample{}, -9, -8)
seedSamples(s, "versionless", HistorySample{}, -13, -12, -11, -10, -9, -8, -7)
series := s.GetVersionSeries(10, 0)
// The versionless days are dropped before the axis is built, so the chart
// spans the days a version is known for instead of climbing out of blanks.
if len(series.Dates) != 5 {
t.Fatalf("dates = %v, want the 5 days with versions", series.Dates)
if len(series.Dates) != 7 {
t.Fatalf("dates = %v, want the 7 days with versions", series.Dates)
}
if len(series.Versions) != 2 ||
series.Versions[0].Version != "4.39" || series.Versions[1].Version != "4.40" {
t.Fatalf("versions = %+v, want 4.39 then 4.40", series.Versions)
}
if got := series.Versions[0].Clusters; !equal(got, []uint64{1, 1, 0, 0, 0}) {
t.Errorf("4.39 = %v, want the upgraded cluster's first two days", got)
if got := series.Versions[0].Clusters; !equal(got, []uint64{1, 1, 1, 1, 0, 0, 0}) {
t.Errorf("4.39 = %v, want the upgraded cluster's first four days", got)
}
if got := series.Versions[1].Clusters; !equal(got, []uint64{1, 1, 2, 2, 2}) {
t.Errorf("4.40 = %v, want steady plus upgraded from day 3", got)
if got := series.Versions[1].Clusters; !equal(got, []uint64{1, 1, 1, 1, 2, 2, 2}) {
t.Errorf("4.40 = %v, want steady plus upgraded from day 5", got)
}
if series.TotalClusters != 2 {
t.Errorf("total_clusters = %d, want 2", series.TotalClusters)
@@ -48,12 +48,12 @@ func TestVersionSeriesHoldsForwardAndLimits(t *testing.T) {
s := newPrometheusStorage(prometheus.NewRegistry())
// Stopped reporting past the active window: its own days only.
seedSamples(s, "gone", HistorySample{Version: "3.97"}, -9, -8)
seedSamples(s, "gone", HistorySample{Version: "3.97"}, -14, -13, -12, -11, -10, -9, -8)
// Reported every day of the window.
seedSamples(s, "daily", HistorySample{Version: "4.40"}, -9, -8, -7, -6, -5, -4, -3, -2, -1, 0)
// Reported three days ago and not since: still active, so it holds its
// version to the right edge instead of dropping out of the stack.
seedSamples(s, "lagging", HistorySample{Version: "4.30"}, -3, -2)
// Stopped reporting two days ago: still active, so it holds its version
// to the right edge instead of dropping out of the stack.
seedSamples(s, "lagging", HistorySample{Version: "4.30"}, -8, -7, -6, -5, -4, -3, -2)
series := s.GetVersionSeries(10, 0)
if len(series.Dates) != 10 {
@@ -66,7 +66,7 @@ func TestVersionSeriesHoldsForwardAndLimits(t *testing.T) {
if got := byVersion["3.97"]; !equal(got, []uint64{1, 1, 0, 0, 0, 0, 0, 0, 0, 0}) {
t.Errorf("3.97 = %v, want nothing after its last report", got)
}
if got := byVersion["4.30"]; !equal(got, []uint64{0, 0, 0, 0, 0, 0, 1, 1, 1, 1}) {
if got := byVersion["4.30"]; !equal(got, []uint64{0, 1, 1, 1, 1, 1, 1, 1, 1, 1}) {
t.Errorf("4.30 = %v, want carried to the right edge", got)
}
if got := byVersion["4.40"]; !equal(got, []uint64{1, 1, 1, 1, 1, 1, 1, 1, 1, 1}) {
@@ -85,7 +85,7 @@ func TestVersionSeriesHoldsForwardAndLimits(t *testing.T) {
if series.Other == nil || series.Other.Count != 2 {
t.Fatalf("other = %+v, want 2 versions", series.Other)
}
if !equal(series.Other.Clusters, []uint64{1, 1, 0, 0, 0, 0, 1, 1, 1, 1}) {
if !equal(series.Other.Clusters, []uint64{1, 2, 1, 1, 1, 1, 1, 1, 1, 1}) {
t.Errorf("other = %v, want 3.97+4.30 summed per day", series.Other.Clusters)
}
if series.TotalClusters != 2 {
+2 -2
View File
@@ -933,7 +933,7 @@ func copyFileContents(src, dst string) error {
// chaosDataNodes lists the data nodes from a fresh master topology snapshot.
func chaosDataNodes(commandEnv *shell.CommandEnv) []*master_pb.DataNodeInfo {
var resp *master_pb.VolumeListResponse
err := commandEnv.MasterClient.WithClient(false, func(client master_pb.SeaweedClient) error {
err := commandEnv.MasterClient.WithClient(context.Background(), false, func(client master_pb.SeaweedClient) error {
var e error
resp, e = client.VolumeList(context.Background(), &master_pb.VolumeListRequest{})
return e
@@ -955,7 +955,7 @@ func chaosDataNodes(commandEnv *shell.CommandEnv) []*master_pb.DataNodeInfo {
func masterEcGenerations(commandEnv *shell.CommandEnv, volumeId uint32) map[int64]bool {
generations := map[int64]bool{}
var resp *master_pb.VolumeListResponse
err := commandEnv.MasterClient.WithClient(false, func(client master_pb.SeaweedClient) error {
err := commandEnv.MasterClient.WithClient(context.Background(), false, func(client master_pb.SeaweedClient) error {
var e error
resp, e = client.VolumeList(context.Background(), &master_pb.VolumeListRequest{})
return e
@@ -200,7 +200,7 @@ func disksWithShards(testDir string, volumeId uint32) int {
func nodeVolumeDiskCounts(t *testing.T, commandEnv *shell.CommandEnv) map[string]int {
t.Helper()
var resp *master_pb.VolumeListResponse
err := commandEnv.MasterClient.WithClient(false, func(client master_pb.SeaweedClient) error {
err := commandEnv.MasterClient.WithClient(context.Background(), false, func(client master_pb.SeaweedClient) error {
var e error
resp, e = client.VolumeList(context.Background(), &master_pb.VolumeListRequest{})
return e
@@ -370,7 +370,7 @@ func removeTwoShardFiles(t *testing.T, testDir string, volumeId uint32) []int {
func masterEcShardIds(commandEnv *shell.CommandEnv, volumeId uint32) map[int]bool {
ids := map[int]bool{}
var resp *master_pb.VolumeListResponse
err := commandEnv.MasterClient.WithClient(false, func(client master_pb.SeaweedClient) error {
err := commandEnv.MasterClient.WithClient(context.Background(), false, func(client master_pb.SeaweedClient) error {
var e error
resp, e = client.VolumeList(context.Background(), &master_pb.VolumeListRequest{})
return e
+45 -1
View File
@@ -8,6 +8,7 @@ import (
"os/exec"
"path/filepath"
"strconv"
"strings"
"sync"
"syscall"
"testing"
@@ -90,6 +91,8 @@ func startDLMTestCluster(t testing.TB) *dlmTestCluster {
require.NoError(t, c.startVolume(configDir))
require.NoError(t, c.waitForTCP(fmt.Sprintf("127.0.0.1:%d", c.volumePort), 30*time.Second),
"volume not ready\n%s", c.tailLog("volume"))
require.NoError(t, c.waitForVolumeRegistered(30*time.Second),
"volume server registration\n%s", c.tailLog("master"))
// Start 2 filers
for i := 0; i < 2; i++ {
@@ -261,7 +264,10 @@ func (c *dlmTestCluster) tailLog(name string) string {
}
func (c *dlmTestCluster) copyLogsForCI() {
ciLogDir := "/tmp/seaweedfs-fuse-dlm-logs"
// One subdirectory per test: a flat layout lets every teardown overwrite
// the previous test's logs, so the CI artifact only ever shows the last
// cluster, never the failing one.
ciLogDir := filepath.Join("/tmp/seaweedfs-fuse-dlm-logs", strings.ReplaceAll(c.t.Name(), "/", "_"))
os.MkdirAll(ciLogDir, 0755)
logsDir := filepath.Join(c.baseDir, "logs")
entries, err := os.ReadDir(logsDir)
@@ -348,6 +354,44 @@ func (c *dlmTestCluster) waitForFilerCount(expected int, timeout time.Duration)
return fmt.Errorf("timed out waiting for %d filers in group %q", expected, filerGroup)
}
// waitForVolumeRegistered waits until the volume server shows up in the master
// topology with its slots reported. An open volume port only means the process
// is listening: until the master has elected itself and accepted a heartbeat,
// an assign fails and the first write on a fresh mount surfaces it as ENOSPC.
func (c *dlmTestCluster) waitForVolumeRegistered(timeout time.Duration) error {
addr := fmt.Sprintf("127.0.0.1:%d", c.masterGrpcPort)
conn, err := grpc.NewClient(addr, grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
return err
}
defer conn.Close()
client := master_pb.NewSeaweedClient(conn)
deadline := time.Now().Add(timeout)
for time.Now().Before(deadline) {
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
resp, err := client.VolumeList(ctx, &master_pb.VolumeListRequest{})
cancel()
if err == nil && resp.TopologyInfo != nil {
var slots int64
for _, dc := range resp.TopologyInfo.DataCenterInfos {
for _, rack := range dc.RackInfos {
for _, dn := range rack.DataNodeInfos {
for _, disk := range dn.DiskInfos {
slots += disk.MaxVolumeCount
}
}
}
}
if slots > 0 {
return nil
}
}
time.Sleep(200 * time.Millisecond)
}
return fmt.Errorf("volume server not registered with the master within %v", timeout)
}
// waitForLockRingConverged verifies that both filers have a consistent view of
// the lock ring by acquiring the same lock through each filer and checking
// mutual exclusion. Adapted from test/s3/distributed_lock/.
+49 -1
View File
@@ -3,12 +3,14 @@
package fuse_p2p
import (
"context"
"fmt"
"net"
"os"
"os/exec"
"path/filepath"
"strconv"
"strings"
"sync"
"syscall"
"testing"
@@ -16,7 +18,10 @@ import (
"github.com/seaweedfs/seaweedfs/test/testutil"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/stretchr/testify/require"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
)
// p2pTestCluster manages a minimal SeaweedFS cluster exercising the peer
@@ -100,6 +105,8 @@ func startP2PTestCluster(t testing.TB, numMounts int) *p2pTestCluster {
require.NoError(t, c.startVolume(configDir))
require.NoError(t, c.waitForTCP(c.volumeCmd, "volume",
fmt.Sprintf("127.0.0.1:%d", c.volumePort), 30*time.Second))
require.NoError(t, c.waitForVolumeRegistered(30*time.Second),
"volume server registration\n%s", c.tailLog("master"))
require.NoError(t, c.startFiler(configDir))
require.NoError(t, c.waitForTCP(c.filerCmd, "filer",
@@ -272,7 +279,10 @@ func (c *p2pTestCluster) tailLogFull(name string) string {
}
func (c *p2pTestCluster) copyLogsForCI() {
ciLogDir := "/tmp/seaweedfs-fuse-p2p-logs"
// One subdirectory per test: a flat layout lets every teardown overwrite
// the previous test's logs, so the CI artifact only ever shows the last
// cluster, never the failing one.
ciLogDir := filepath.Join("/tmp/seaweedfs-fuse-p2p-logs", strings.ReplaceAll(c.t.Name(), "/", "_"))
os.MkdirAll(ciLogDir, 0755)
logsDir := filepath.Join(c.baseDir, "logs")
entries, err := os.ReadDir(logsDir)
@@ -288,6 +298,44 @@ func (c *p2pTestCluster) copyLogsForCI() {
}
}
// waitForVolumeRegistered waits until the volume server shows up in the master
// topology with its slots reported. An open volume port only means the process
// is listening: until the master has elected itself and accepted a heartbeat,
// an assign fails and the first write on a fresh mount surfaces it as ENOSPC.
func (c *p2pTestCluster) waitForVolumeRegistered(timeout time.Duration) error {
addr := fmt.Sprintf("127.0.0.1:%d", c.masterGrpcPort)
conn, err := grpc.NewClient(addr, grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
return err
}
defer conn.Close()
client := master_pb.NewSeaweedClient(conn)
deadline := time.Now().Add(timeout)
for time.Now().Before(deadline) {
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
resp, err := client.VolumeList(ctx, &master_pb.VolumeListRequest{})
cancel()
if err == nil && resp.TopologyInfo != nil {
var slots int64
for _, dc := range resp.TopologyInfo.DataCenterInfos {
for _, rack := range dc.RackInfos {
for _, dn := range rack.DataNodeInfos {
for _, disk := range dn.DiskInfos {
slots += disk.MaxVolumeCount
}
}
}
}
if slots > 0 {
return nil
}
}
time.Sleep(200 * time.Millisecond)
}
return fmt.Errorf("volume server not registered with the master within %v", timeout)
}
// waitForTCP polls addr until it accepts a connection, OR the supplied
// subprocess exits — whichever comes first. Short-circuiting on child
// exit turns a 30 s-spin-on-dead-process into an immediate failure with
+65
View File
@@ -10,6 +10,7 @@ import (
"os"
"os/exec"
"path/filepath"
"regexp"
"runtime"
"strconv"
"strings"
@@ -36,6 +37,9 @@ type masterNode struct {
// peersStr overrides the cluster-wide peer list for this node, so a test
// can start a master that only knows about a subset of the cluster.
peersStr string
// raftBootstrap starts the master with -raftBootstrap, the way the helm
// chart renders master.raftBootstrap on every master, every restart.
raftBootstrap bool
}
// MasterCluster manages a 3-node master raft cluster for integration tests.
@@ -153,6 +157,13 @@ func (mc *MasterCluster) SetNodePeers(i int, peers string) {
mc.nodes[i].peersStr = peers
}
// SetRaftBootstrap makes node i start with -raftBootstrap.
func (mc *MasterCluster) SetRaftBootstrap(i int) {
mc.mu.Lock()
defer mc.mu.Unlock()
mc.nodes[i].raftBootstrap = true
}
// StartNode starts the master process at the given index (0–2).
func (mc *MasterCluster) StartNode(i int) {
mc.t.Helper()
@@ -187,6 +198,9 @@ func (mc *MasterCluster) StartNode(i int) {
if mc.raftHashicorp {
args = append(args, "-raftHashicorp")
}
if n.raftBootstrap {
args = append(args, "-raftBootstrap")
}
n.cmd = exec.Command(mc.weedBinary, args...)
n.cmd.Dir = mc.baseDir
@@ -386,6 +400,57 @@ func (mc *MasterCluster) WaitForNodeReady(i int, timeout time.Duration) error {
return fmt.Errorf("node %d not ready within %v", i, timeout)
}
// LogContains reports whether node i's log holds the given text.
func (mc *MasterCluster) LogContains(i int, text string) bool {
b, err := os.ReadFile(mc.nodes[i].logFile)
if err != nil {
return false
}
return strings.Contains(string(b), text)
}
// topologyIdLine matches every line a master logs when it learns a TopologyId,
// whichever raft implementation applied it.
var topologyIdLine = regexp.MustCompile(`TopologyId[^:]*: ([0-9a-f-]{36})`)
// NodeTopologyIds returns the TopologyIds node i has logged. /dir/status is
// proxied to the leader, so a master's own view of the cluster identity is only
// visible in its log, and that is where a fork shows up.
func (mc *MasterCluster) NodeTopologyIds(i int) []string {
b, err := os.ReadFile(mc.nodes[i].logFile)
if err != nil {
return nil
}
var ids []string
for _, m := range topologyIdLine.FindAllStringSubmatch(string(b), -1) {
ids = append(ids, m[1])
}
return ids
}
// WaitForNodeTopologyIds waits until every master has logged a TopologyId and
// returns what each one saw.
func (mc *MasterCluster) WaitForNodeTopologyIds(timeout time.Duration) ([3][]string, error) {
var ids [3][]string
deadline := time.Now().Add(timeout)
for {
missing := -1
for i := range 3 {
ids[i] = mc.NodeTopologyIds(i)
if len(ids[i]) == 0 {
missing = i
}
}
if missing < 0 {
return ids, nil
}
if !time.Now().Before(deadline) {
return ids, fmt.Errorf("master %d logged no TopologyId within %v", missing, timeout)
}
time.Sleep(waitTick)
}
}
// DumpLogs prints the tail of all master logs.
func (mc *MasterCluster) DumpLogs() {
for i := range 3 {
+59
View File
@@ -151,3 +151,62 @@ func peerCountExcludingSelf(peers []string, self string) int {
}
return count
}
// TestRaftBootstrapKeepsExistingCluster covers a master restarting under
// -raftBootstrap, the way the helm chart renders it on every master on every
// roll. Bootstrapping is genesis: seeding a second cluster over committed raft
// state mints a rival TopologyId, and the split-brain guard then Fatals every
// master that still holds the first one.
func TestRaftBootstrapKeepsExistingCluster(t *testing.T) {
for _, impl := range raftImplementations {
t.Run(impl.name, func(t *testing.T) {
mc := NewMasterCluster(t, impl.raftHashicorp)
for i := range 3 {
mc.SetRaftBootstrap(i)
mc.StartNode(i)
}
before, err := mc.WaitForTopologyId(waitTimeout)
if err != nil {
mc.DumpLogs()
t.Fatalf("cluster did not mint a TopologyId: %v", err)
}
for i := range 3 {
mc.StopNode(i)
}
for i := range 3 {
mc.StartNode(i)
}
after, err := mc.WaitForTopologyId(waitTimeout)
if err != nil {
mc.DumpLogs()
t.Fatalf("cluster did not come back after a restart: %v", err)
}
if after != before {
mc.DumpLogs()
t.Fatalf("-raftBootstrap re-seeded the cluster: TopologyId %s became %s", before, after)
}
// The leader answers for the whole cluster, so a follower that
// forked is only visible in its own log.
seen, err := mc.WaitForNodeTopologyIds(waitTimeout)
if err != nil {
mc.DumpLogs()
t.Fatal(err)
}
for i, ids := range seen {
for _, id := range ids {
if id != before {
mc.DumpLogs()
t.Fatalf("master %d saw TopologyId %s, want %s", i, id, before)
}
}
if mc.LogContains(i, "Split-brain detected") {
mc.DumpLogs()
t.Fatalf("master %d hit the split-brain guard", i)
}
}
})
}
}
+19 -1
View File
@@ -15,6 +15,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/storage/volume_info"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
)
@@ -339,7 +340,14 @@ func (v *VolumeServer) VolumeEcShardsInfo(ctx context.Context, req *volume_serve
}
}
resp := &volume_server_pb.VolumeEcShardsInfoResponse{}
// Answer with the layout out of the .vif that was actually delivered here,
// the way a real holder answers from the context it mounted the shards
// with. A coordinator uses this to tell a server that understands the
// shard block layout from one that never knew the field, so a fake that
// always reported "unset" would look like a pre-upgrade server.
resp := &volume_server_pb.VolumeEcShardsInfoResponse{
EcShardConfig: v.ecShardConfigFromVif(req.VolumeId),
}
prefix := fmt.Sprintf("%d.ec", req.VolumeId)
entries, _ := os.ReadDir(v.baseDir)
for _, entry := range entries {
@@ -372,6 +380,16 @@ func (v *VolumeServer) VolumeEcShardsInfo(ctx context.Context, req *volume_serve
return resp, nil
}
// ecShardConfigFromVif reads the EC layout out of the .vif this server was
// given, which distribution ships to every holder alongside its shards.
func (v *VolumeServer) ecShardConfigFromVif(volumeID uint32) *volume_server_pb.EcShardConfig {
vi, _, found, err := volume_info.MaybeLoadVolumeInfo(v.filePath(volumeID, ".vif"))
if err != nil || !found {
return nil
}
return vi.GetEcShardConfig()
}
func (v *VolumeServer) VolumeDelete(ctx context.Context, req *volume_server_pb.VolumeDeleteRequest) (*volume_server_pb.VolumeDeleteResponse, error) {
v.mu.Lock()
v.deleteRequests = append(v.deleteRequests, req)
+7 -5
View File
@@ -156,8 +156,9 @@ func TestRenameObjectSourceIfMatch(t *testing.T) {
assert.True(t, objectExists(t, client, bucketName, "target.txt"))
}
// TestRenameObjectOntoDirectory: a key that already holds other objects is a
// directory, and an object must not be allowed to replace one.
// TestRenameObjectOntoDirectory: S3 keys are flat, so a key that other keys are
// nested under is still a key of its own. The rename writes it without disturbing
// them - it does not replace the directory, it stores the object on it.
func TestRenameObjectOntoDirectory(t *testing.T) {
client := getS3Client(t)
bucketName := getNewBucketName()
@@ -172,9 +173,10 @@ func TestRenameObjectOntoDirectory(t *testing.T) {
Key: aws.String("target"),
RenameSource: aws.String(createRenameSource("source.txt")),
})
requireRenameStatus(t, err, 409)
assert.True(t, objectExists(t, client, bucketName, "source.txt"))
assert.True(t, objectExists(t, client, bucketName, "target/child.txt"))
require.NoError(t, err)
assert.False(t, objectExists(t, client, bucketName, "source.txt"))
assert.Equal(t, "content", getObjectBody(t, getObject(t, client, bucketName, "target")))
assert.Equal(t, "child", getObjectBody(t, getObject(t, client, bucketName, "target/child.txt")))
}
// TestRenameObjectDirectorySource: a directory can be named without a trailing
+546
View File
@@ -0,0 +1,546 @@
package example
import (
"bytes"
"io"
"net/http"
"sort"
"strings"
"testing"
"time"
"github.com/aws/aws-sdk-go/aws"
"github.com/aws/aws-sdk-go/aws/awserr"
v1credentials "github.com/aws/aws-sdk-go/aws/credentials"
v1signer "github.com/aws/aws-sdk-go/aws/signer/v4"
"github.com/aws/aws-sdk-go/service/s3"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// TestS3PrefixObjectKeys covers keys that are a strict prefix of other keys: S3's
// namespace is flat, so "collision/foo" and "collision/foo/bar" are independent
// objects that coexist in either write order.
func TestS3PrefixObjectKeys(t *testing.T) {
if testing.Short() {
t.Skip("Skipping integration test in short mode")
}
cluster, err := startMiniCluster(t)
require.NoError(t, err)
defer cluster.Stop()
put := func(t *testing.T, bucket, key string, body []byte) {
t.Helper()
_, err := cluster.s3Client.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
Body: bytes.NewReader(body),
})
require.NoError(t, err, "put %s", key)
}
// read checks both paths a client reaches an object by, since a directory entry
// carrying an object is served by neither the directory nor the plain object path
// alone.
read := func(t *testing.T, bucket, key string, want []byte) {
t.Helper()
head, err := cluster.s3Client.HeadObject(&s3.HeadObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
})
require.NoError(t, err, "head %s", key)
assert.Equal(t, int64(len(want)), aws.Int64Value(head.ContentLength), "head %s", key)
get, err := cluster.s3Client.GetObject(&s3.GetObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
})
require.NoError(t, err, "get %s", key)
defer get.Body.Close()
got, err := io.ReadAll(get.Body)
require.NoError(t, err, "read %s", key)
assert.Equal(t, want, got, "get %s", key)
}
// gone checks the key answers as absent rather than lingering on a directory
// entry that outlived the object.
gone := func(t *testing.T, bucket, key string) {
t.Helper()
_, err := cluster.s3Client.GetObject(&s3.GetObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
})
var missing awserr.RequestFailure
require.ErrorAs(t, err, &missing, "get %s", key)
assert.Equal(t, http.StatusNotFound, missing.StatusCode(), "get %s", key)
_, err = cluster.s3Client.HeadObject(&s3.HeadObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
})
require.ErrorAs(t, err, &missing, "head %s", key)
assert.Equal(t, http.StatusNotFound, missing.StatusCode(), "head %s", key)
}
listKeys := func(t *testing.T, bucket string) []string {
t.Helper()
resp, err := cluster.s3Client.ListObjectsV2(&s3.ListObjectsV2Input{Bucket: aws.String(bucket)})
require.NoError(t, err)
keys := collectKeys(resp.Contents)
sort.Strings(keys)
return keys
}
body := []byte("prefix object")
// Distinct bodies, so a read that resolves to the wrong entry cannot pass.
nested := []byte("nested under the prefix object")
// The reported order: the nested key is written first, so the prefix key has to
// land on a path the filer already holds a directory at.
t.Run("ChildFirst", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-child-first-")
put(t, bucket, "collision/foo/bar", nested)
put(t, bucket, "collision/foo", body)
assert.Equal(t, []string{"collision/foo", "collision/foo/bar"}, listKeys(t, bucket))
read(t, bucket, "collision/foo", body)
read(t, bucket, "collision/foo/bar", nested)
})
// The opposite order used to keep the prefix key's data but hide the key.
t.Run("PrefixFirst", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-first-")
put(t, bucket, "collision/foo", body)
put(t, bucket, "collision/foo/bar", nested)
assert.Equal(t, []string{"collision/foo", "collision/foo/bar"}, listKeys(t, bucket))
read(t, bucket, "collision/foo", body)
read(t, bucket, "collision/foo/bar", nested)
})
// An empty object leaves no chunks, content or mime behind, so it is the case a
// promoted directory carries no other trace of.
t.Run("EmptyObject", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-empty-")
put(t, bucket, "a/foo/bar", nil)
put(t, bucket, "a/foo", nil)
put(t, bucket, "b/foo", nil)
put(t, bucket, "b/foo/bar", nil)
assert.Equal(t, []string{"a/foo", "a/foo/bar", "b/foo", "b/foo/bar"}, listKeys(t, bucket))
read(t, bucket, "a/foo", []byte{})
read(t, bucket, "b/foo", []byte{})
})
// The key has no trailing slash, and the keys nested under it still roll up into
// their own CommonPrefix.
t.Run("Delimiter", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-delimiter-")
put(t, bucket, "collision/foo/bar", nested)
put(t, bucket, "collision/foo", body)
put(t, bucket, "collision/other", body)
resp, err := cluster.s3Client.ListObjectsV2(&s3.ListObjectsV2Input{
Bucket: aws.String(bucket),
Prefix: aws.String("collision/"),
Delimiter: aws.String("/"),
})
require.NoError(t, err)
keys := collectKeys(resp.Contents)
sort.Strings(keys)
assert.Equal(t, []string{"collision/foo", "collision/other"}, keys)
assert.Equal(t, []string{"collision/foo/"}, collectPrefixes(resp.CommonPrefixes))
// Listing the prefix itself names only what is under it.
resp, err = cluster.s3Client.ListObjectsV2(&s3.ListObjectsV2Input{
Bucket: aws.String(bucket),
Prefix: aws.String("collision/foo/"),
Delimiter: aws.String("/"),
})
require.NoError(t, err)
assert.Equal(t, []string{"collision/foo/bar"}, collectKeys(resp.Contents))
assert.Empty(t, collectPrefixes(resp.CommonPrefixes))
})
// The key and the CommonPrefix its nested keys fold into come off one filer
// entry, so a page boundary must not drop either of them.
t.Run("Paged", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-paged-")
for _, key := range []string{"foo", "foo/bar", "foobar", "other", "zed", "zed/a"} {
put(t, bucket, key, body)
}
for _, maxKeys := range []int64{1, 2, 3, 4, 5} {
var keys, prefixes []string
var token *string
for page := 0; page < 12; page++ {
resp, err := cluster.s3Client.ListObjectsV2(&s3.ListObjectsV2Input{
Bucket: aws.String(bucket),
Delimiter: aws.String("/"),
MaxKeys: aws.Int64(maxKeys),
ContinuationToken: token,
})
require.NoError(t, err)
// One entry over the budget is the documented same-entry exception;
// anything more means the unsigned budget wrapped.
assert.LessOrEqual(t, int64(len(resp.Contents)+len(resp.CommonPrefixes)), maxKeys+1,
"maxKeys=%d page %d", maxKeys, page)
keys = append(keys, collectKeys(resp.Contents)...)
prefixes = append(prefixes, collectPrefixes(resp.CommonPrefixes)...)
if !aws.BoolValue(resp.IsTruncated) {
token = nil
break
}
token = resp.NextContinuationToken
require.NotNil(t, token, "a truncated page must name where to resume")
}
require.Nil(t, token, "maxKeys=%d did not finish", maxKeys)
assert.Equal(t, []string{"foo", "foobar", "other", "zed"}, keys, "maxKeys=%d", maxKeys)
assert.Equal(t, []string{"foo/", "zed/"}, prefixes, "maxKeys=%d", maxKeys)
}
})
// Versioning reaches a prefix object from two directions: a suspended bucket
// writes the null version at the key's own path, and a bucket versioned later
// finds one already sitting there. Both leave a key that is a directory with
// version history beside it.
t.Run("Versioned", func(t *testing.T) {
setVersioning := func(t *testing.T, bucket, status string) {
t.Helper()
_, err := cluster.s3Client.PutBucketVersioning(&s3.PutBucketVersioningInput{
Bucket: aws.String(bucket),
VersioningConfiguration: &s3.VersioningConfiguration{Status: aws.String(status)},
})
require.NoError(t, err)
}
// Both write orders, in a bucket that is versioned and in one where versioning
// was suspended - the suspended one is the case that writes at the key's path.
for _, state := range []string{"Enabled", "Suspended"} {
bucket := createTestBucket(t, cluster, "test-prefix-"+strings.ToLower(state)+"-")
setVersioning(t, bucket, "Enabled")
if state == "Suspended" {
setVersioning(t, bucket, "Suspended")
}
put(t, bucket, "child/foo/bar", nested)
put(t, bucket, "child/foo", body)
put(t, bucket, "prefix/foo", body)
put(t, bucket, "prefix/foo/bar", nested)
assert.Equal(t, []string{"child/foo", "child/foo/bar", "prefix/foo", "prefix/foo/bar"},
listKeys(t, bucket), state)
for _, key := range []string{"child/foo", "prefix/foo"} {
read(t, bucket, key, body)
}
for _, key := range []string{"child/foo/bar", "prefix/foo/bar"} {
read(t, bucket, key, nested)
}
}
// A prefix object written before versioning is the key's null version. Removing
// that version by id must not take the keys nested under it with it.
bucket := createTestBucket(t, cluster, "test-prefix-nullversion-")
put(t, bucket, "collision/foo/bar", nested)
put(t, bucket, "collision/foo", body)
setVersioning(t, bucket, "Enabled")
newer := []byte("written after versioning was enabled")
versioned, err := cluster.s3Client.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("collision/foo"),
Body: bytes.NewReader(newer),
})
require.NoError(t, err)
for _, v := range []struct {
id string
want []byte
}{{"null", body}, {aws.StringValue(versioned.VersionId), newer}} {
got, err := cluster.s3Client.GetObject(&s3.GetObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("collision/foo"),
VersionId: aws.String(v.id),
})
require.NoError(t, err, "get version %s", v.id)
body, err := io.ReadAll(got.Body)
require.NoError(t, err)
got.Body.Close()
assert.Equal(t, v.want, body, "get version %s", v.id)
}
_, err = cluster.s3Client.DeleteObject(&s3.DeleteObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("collision/foo"),
VersionId: aws.String("null"),
})
require.NoError(t, err, "the null version sits on a directory other keys live in")
read(t, bucket, "collision/foo", newer)
read(t, bucket, "collision/foo/bar", nested)
remaining, err := cluster.s3Client.ListObjectVersions(&s3.ListObjectVersionsInput{
Bucket: aws.String(bucket),
Prefix: aws.String("collision/foo"),
})
require.NoError(t, err)
for _, v := range remaining.Versions {
if aws.StringValue(v.Key) != "collision/foo" {
// collision/foo/bar predates versioning too, and keeps its null version.
continue
}
assert.NotEqual(t, "null", aws.StringValue(v.VersionId), "the null version was deleted")
}
})
// The two listings walk the tree differently, and a prefix object is the entry
// they disagree about: it is a directory the version listing descends through and
// a key at the same time. They have to name the same keys and the same prefixes.
t.Run("VersionListingMatchesObjectListing", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-parity-")
for _, key := range []string{"foo", "foo/bar", "other", "a/foo", "a/foo/bar", "a/z"} {
put(t, bucket, key, body)
}
for _, q := range []struct{ prefix, delimiter string }{
{"", ""},
{"foo/", ""},
{"a/", ""},
{"a/foo/", ""},
{"", "/"},
{"a/", "/"},
} {
name := "prefix=" + q.prefix + " delimiter=" + q.delimiter
objects, err := cluster.s3Client.ListObjectsV2(&s3.ListObjectsV2Input{
Bucket: aws.String(bucket),
Prefix: aws.String(q.prefix),
Delimiter: aws.String(q.delimiter),
})
require.NoError(t, err, name)
versions, err := cluster.s3Client.ListObjectVersions(&s3.ListObjectVersionsInput{
Bucket: aws.String(bucket),
Prefix: aws.String(q.prefix),
Delimiter: aws.String(q.delimiter),
})
require.NoError(t, err, name)
versionKeys := make([]string, 0, len(versions.Versions))
for _, v := range versions.Versions {
versionKeys = append(versionKeys, aws.StringValue(v.Key))
}
versionPrefixes := make([]string, 0, len(versions.CommonPrefixes))
for _, p := range versions.CommonPrefixes {
versionPrefixes = append(versionPrefixes, aws.StringValue(p.Prefix))
}
sort.Strings(versionKeys)
sort.Strings(versionPrefixes)
objectKeys := collectKeys(objects.Contents)
objectPrefixes := collectPrefixes(objects.CommonPrefixes)
sort.Strings(objectKeys)
sort.Strings(objectPrefixes)
assert.Equal(t, objectKeys, versionKeys, "keys, %s", name)
assert.Equal(t, objectPrefixes, versionPrefixes, "prefixes, %s", name)
}
})
// A prefix object written before versioning is the key's null version, so the
// version written after it has to take the latest flag off it.
t.Run("VersionedAfterPrefixObject", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-versioned-")
put(t, bucket, "collision/foo/bar", nested)
put(t, bucket, "collision/foo", body)
_, err := cluster.s3Client.PutBucketVersioning(&s3.PutBucketVersioningInput{
Bucket: aws.String(bucket),
VersioningConfiguration: &s3.VersioningConfiguration{Status: aws.String("Enabled")},
})
require.NoError(t, err)
newer := []byte("written after versioning was enabled")
put(t, bucket, "collision/foo", newer)
resp, err := cluster.s3Client.ListObjectVersions(&s3.ListObjectVersionsInput{Bucket: aws.String(bucket)})
require.NoError(t, err)
latest := map[string]int{}
var nullSize int64 = -1
for _, v := range resp.Versions {
if aws.BoolValue(v.IsLatest) {
latest[aws.StringValue(v.Key)]++
}
if aws.StringValue(v.Key) == "collision/foo" && aws.StringValue(v.VersionId) == "null" {
nullSize = aws.Int64Value(v.Size)
assert.False(t, aws.BoolValue(v.IsLatest), "the newer version is the latest one")
}
}
assert.Equal(t, 1, latest["collision/foo"], "exactly one version of a key is the latest")
assert.Equal(t, int64(len(body)), nullSize, "the null version keeps the prefix object's size")
read(t, bucket, "collision/foo", newer)
read(t, bucket, "collision/foo/bar", nested)
})
// A key that other keys are nested under is a copy source and a copy destination
// like any other. The keys nested under either end are not part of the copy.
t.Run("Copy", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-copy-")
copyObject := func(t *testing.T, src, dst string) {
t.Helper()
_, err := cluster.s3Client.CopyObject(&s3.CopyObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(dst),
CopySource: aws.String(bucket + "/" + src),
})
require.NoError(t, err, "copy %s to %s", src, dst)
}
put(t, bucket, "collision/foo", body)
put(t, bucket, "collision/foo/bar", nested)
// Out of a prefix object, into a key of its own.
copyObject(t, "collision/foo", "plain")
read(t, bucket, "plain", body)
read(t, bucket, "collision/foo", body)
read(t, bucket, "collision/foo/bar", nested)
// Into a key that other keys are nested under.
put(t, bucket, "target/child", nested)
copyObject(t, "plain", "target")
read(t, bucket, "target", body)
read(t, bucket, "target/child", nested)
// And between two of them.
put(t, bucket, "other", []byte("copied between prefix keys"))
copyObject(t, "other", "collision/foo")
read(t, bucket, "collision/foo", []byte("copied between prefix keys"))
read(t, bucket, "collision/foo/bar", nested)
assert.Equal(t, []string{"collision/foo", "collision/foo/bar", "other", "plain", "target", "target/child"},
listKeys(t, bucket))
})
// Rename moves the object off the key without moving the keys nested under it,
// which is not what the filer's atomic rename of a directory would do.
t.Run("Rename", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-rename-")
renameObject := func(t *testing.T, src, dst string) {
t.Helper()
req, _ := http.NewRequest(http.MethodPut, cluster.s3Endpoint+"/"+bucket+"/"+dst+"?renameObject=", nil)
req.Header.Set("x-amz-rename-source", "/"+bucket+"/"+src)
signer := v1signer.NewSigner(v1credentials.NewStaticCredentials(testAccessKey, testSecretKey, ""))
_, err := signer.Sign(req, nil, "s3", testRegion, time.Now())
require.NoError(t, err)
resp, err := (&http.Client{Timeout: 20 * time.Second}).Do(req)
require.NoError(t, err, "rename %s to %s", src, dst)
defer resp.Body.Close()
io.Copy(io.Discard, resp.Body)
require.Equal(t, http.StatusOK, resp.StatusCode, "rename %s to %s", src, dst)
}
put(t, bucket, "collision/foo", body)
put(t, bucket, "collision/foo/bar", nested)
// Off a prefix object: the key goes, the keys under it stay.
renameObject(t, "collision/foo", "moved")
read(t, bucket, "moved", body)
read(t, bucket, "collision/foo/bar", nested)
gone(t, bucket, "collision/foo")
// Onto a key other keys are nested under.
put(t, bucket, "target/child", nested)
renameObject(t, "moved", "target")
read(t, bucket, "target", body)
read(t, bucket, "target/child", nested)
gone(t, bucket, "moved")
assert.Equal(t, []string{"collision/foo/bar", "target", "target/child"}, listKeys(t, bucket))
})
// A directory SeaweedFS keeps its own state in is not a prefix a key can be
// stored on: the object would replace that state with its own.
t.Run("ReservedDirectory", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-reserved-")
_, err := cluster.s3Client.PutBucketVersioning(&s3.PutBucketVersioningInput{
Bucket: aws.String(bucket),
VersioningConfiguration: &s3.VersioningConfiguration{Status: aws.String("Enabled")},
})
require.NoError(t, err)
put(t, bucket, "foo", body)
_, err = cluster.s3Client.PutBucketVersioning(&s3.PutBucketVersioningInput{
Bucket: aws.String(bucket),
VersioningConfiguration: &s3.VersioningConfiguration{Status: aws.String("Suspended")},
})
require.NoError(t, err)
_, err = cluster.s3Client.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("foo.versions"),
Body: bytes.NewReader(body),
})
// Not just any error: a permanently impossible write must not come back as a
// 500 the SDK retries.
var refused awserr.RequestFailure
require.ErrorAs(t, err, &refused, "the version history of foo is not a prefix of foo.versions")
assert.Equal(t, http.StatusConflict, refused.StatusCode())
assert.Equal(t, "ExistingObjectIsDirectory", refused.Code())
versions, err := cluster.s3Client.ListObjectVersions(&s3.ListObjectVersionsInput{Bucket: aws.String(bucket)})
require.NoError(t, err)
require.Len(t, versions.Versions, 1)
assert.Equal(t, "foo", aws.StringValue(versions.Versions[0].Key))
read(t, bucket, "foo", body)
// The multipart staging folder is the other one, and an in-flight upload has
// to survive the attempt.
staging := createTestBucket(t, cluster, "test-prefix-uploads-")
created, err := cluster.s3Client.CreateMultipartUpload(&s3.CreateMultipartUploadInput{
Bucket: aws.String(staging),
Key: aws.String("mp.bin"),
})
require.NoError(t, err)
_, err = cluster.s3Client.PutObject(&s3.PutObjectInput{
Bucket: aws.String(staging),
Key: aws.String(".uploads"),
Body: bytes.NewReader(body),
})
require.ErrorAs(t, err, &refused, "the multipart staging folder is not a prefix of .uploads")
assert.Equal(t, http.StatusConflict, refused.StatusCode())
assert.Equal(t, "ExistingObjectIsDirectory", refused.Code())
uploads, err := cluster.s3Client.ListMultipartUploads(&s3.ListMultipartUploadsInput{Bucket: aws.String(staging)})
require.NoError(t, err)
require.Len(t, uploads.Uploads, 1)
assert.Equal(t, "mp.bin", aws.StringValue(uploads.Uploads[0].Key))
_, err = cluster.s3Client.AbortMultipartUpload(&s3.AbortMultipartUploadInput{
Bucket: aws.String(staging),
Key: aws.String("mp.bin"),
UploadId: created.UploadId,
})
require.NoError(t, err)
})
// Either key can be deleted without touching the other.
t.Run("Delete", func(t *testing.T) {
bucket := createTestBucket(t, cluster, "test-prefix-delete-")
put(t, bucket, "collision/foo/bar", nested)
put(t, bucket, "collision/foo", body)
_, err := cluster.s3Client.DeleteObject(&s3.DeleteObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("collision/foo"),
})
require.NoError(t, err)
assert.Equal(t, []string{"collision/foo/bar"}, listKeys(t, bucket))
read(t, bucket, "collision/foo/bar", nested)
gone(t, bucket, "collision/foo")
put(t, bucket, "collision/foo", body)
_, err = cluster.s3Client.DeleteObject(&s3.DeleteObjectInput{
Bucket: aws.String(bucket),
Key: aws.String("collision/foo/bar"),
})
require.NoError(t, err)
assert.Equal(t, []string{"collision/foo"}, listKeys(t, bucket))
read(t, bucket, "collision/foo", body)
gone(t, bucket, "collision/foo/bar")
})
}
+5 -5
View File
@@ -26,16 +26,16 @@ nginx (:9000)
v
SeaweedFS S3 (:8333, -s3.externalUrl=http://localhost:9000)
| externalHost = "localhost:9000" (parsed at startup)
| extractHostHeader() returns "localhost:9000"
| extractHostHeaderCandidates() tries "localhost:9000" first
| Matches what AWS CLI signed with
v
Signature verification succeeds
```
**Note:** When `-s3.externalUrl` is configured, direct access to the backend
port (8333) will fail signature verification because the client signs with a
different Host header than what `externalUrl` specifies. This is expected —
all S3 traffic should go through the proxy.
**Note:** `-s3.externalUrl` is tried first, not exclusively. A client that
dials the backend port (8333) directly still verifies against the host it
actually signed, so a mixed topology of proxied and in-cluster clients works
with the flag set.
## Prerequisites
@@ -0,0 +1,231 @@
package retention
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/aws/aws-sdk-go-v2/aws"
"github.com/aws/aws-sdk-go-v2/service/s3"
"github.com/aws/aws-sdk-go-v2/service/s3/types"
"github.com/aws/smithy-go"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func requireAccessDenied(t *testing.T, err error, msg string) {
t.Helper()
require.Error(t, err, msg)
var apiErr smithy.APIError
require.True(t, errors.As(err, &apiErr), "expected an API error, got %T", err)
assert.Equal(t, "AccessDenied", apiErr.ErrorCode(), msg)
}
// A key ending in "/" is stored as the filer directory rather than as an object
// beside it, and is deleted the unversioned way. Object Lock still covers it: the
// gateway lists it as an object and serves retention set on it.
func TestObjectLockDirectoryMarker(t *testing.T) {
client := getS3Client(t)
bucketName := getNewBucketName()
createBucketWithObjectLock(t, client, bucketName)
defer deleteBucket(t, client, bucketName)
retainUntil := time.Now().Add(24 * time.Hour)
t.Run("retention headers are honored, not dropped", func(t *testing.T) {
key := "records/evidence/"
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("marker"),
ObjectLockMode: types.ObjectLockModeCompliance,
ObjectLockRetainUntilDate: aws.Time(retainUntil),
})
require.NoError(t, err)
_, err = client.DeleteObject(context.TODO(), &s3.DeleteObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
requireAccessDenied(t, err, "a retained marker must not be deletable")
_, err = client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
assert.NoError(t, err, "the marker must still be there after the refused delete")
})
t.Run("retention set through PutObjectRetention is honored", func(t *testing.T) {
key := "records/ledger/"
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("marker"),
})
require.NoError(t, err)
_, err = client.PutObjectRetention(context.TODO(), &s3.PutObjectRetentionInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Retention: &types.ObjectLockRetention{
Mode: types.ObjectLockRetentionModeCompliance,
RetainUntilDate: aws.Time(retainUntil),
},
})
require.NoError(t, err)
_, err = client.DeleteObject(context.TODO(), &s3.DeleteObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
requireAccessDenied(t, err, "retention the gateway serves back must also block the delete")
_, err = client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
assert.NoError(t, err, "the marker must still be there after the refused delete")
})
t.Run("multi-object delete is refused too", func(t *testing.T) {
key := "records/batch/"
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("marker"),
ObjectLockMode: types.ObjectLockModeCompliance,
ObjectLockRetainUntilDate: aws.Time(retainUntil),
})
require.NoError(t, err)
resp, err := client.DeleteObjects(context.TODO(), &s3.DeleteObjectsInput{
Bucket: aws.String(bucketName),
Delete: &types.Delete{Objects: []types.ObjectIdentifier{{Key: aws.String(key)}}},
})
require.NoError(t, err)
assert.Empty(t, resp.Deleted, "a retained marker must not be reported deleted")
require.Len(t, resp.Errors, 1)
assert.Equal(t, key, aws.ToString(resp.Errors[0].Key))
assert.Equal(t, "AccessDenied", aws.ToString(resp.Errors[0].Code))
_, err = client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
assert.NoError(t, err, "the marker must survive the batch delete")
})
t.Run("invalid lock headers are rejected, as on a regular key", func(t *testing.T) {
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("records/bad-mode/"),
Body: strings.NewReader("marker"),
ObjectLockMode: "INVALID_MODE",
ObjectLockRetainUntilDate: aws.Time(retainUntil),
})
require.Error(t, err)
_, err = client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("records/no-date/"),
Body: strings.NewReader("marker"),
ObjectLockMode: types.ObjectLockModeGovernance,
})
require.Error(t, err)
})
// Past 1KiB a trailing-slash key is a real versioned object, and the delete
// takes the whole history at once, so an older retained version has to block
// it even when the version on top carries no retention of its own.
t.Run("a retained version under an unretained one still blocks", func(t *testing.T) {
key := "records/history/"
body := strings.Repeat("x", 2048)
first, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader(body),
ObjectLockMode: types.ObjectLockModeCompliance,
ObjectLockRetainUntilDate: aws.Time(retainUntil),
})
require.NoError(t, err)
require.NotNil(t, first.VersionId)
_, err = client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader(body),
})
require.NoError(t, err)
_, err = client.DeleteObject(context.TODO(), &s3.DeleteObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
requireAccessDenied(t, err, "the retained version underneath must block the delete")
_, err = client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
VersionId: first.VersionId,
})
assert.NoError(t, err, "the retained version must survive")
})
// mkdir replaces the marker entry outright, taking its lock metadata with it,
// so a plain PUT over a retained marker has to be refused - including once the
// key has grown a version history that the latest-version lookup would find
// instead of the marker.
t.Run("a plain PUT cannot replace a retained marker", func(t *testing.T) {
key := "records/overwrite/"
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("marker"),
ObjectLockMode: types.ObjectLockModeCompliance,
ObjectLockRetainUntilDate: aws.Time(retainUntil),
})
require.NoError(t, err)
_, err = client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("replacement"),
})
requireAccessDenied(t, err, "a retained marker must not be replaceable")
_, err = client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader(strings.Repeat("x", 2048)),
})
require.NoError(t, err, "a versioned write of the same key adds a version")
_, err = client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("replacement"),
})
requireAccessDenied(t, err, "the history must not hide the marker's own lock")
})
t.Run("an unretained marker still deletes", func(t *testing.T) {
key := "records/plain/"
_, err := client.PutObject(context.TODO(), &s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
Body: strings.NewReader("marker"),
})
require.NoError(t, err)
_, err = client.DeleteObject(context.TODO(), &s3.DeleteObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String(key),
})
require.NoError(t, err)
})
}
@@ -0,0 +1,130 @@
package s3api
import (
"context"
"errors"
"testing"
"time"
"github.com/aws/aws-sdk-go-v2/aws"
"github.com/aws/aws-sdk-go-v2/service/s3"
"github.com/aws/smithy-go"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func requireAPIErrorCode(t *testing.T, err error, expected string) {
t.Helper()
require.Error(t, err)
var apiErr smithy.APIError
require.True(t, errors.As(err, &apiErr), "expected a smithy.APIError, got %T: %v", err, err)
assert.Equal(t, expected, apiErr.ErrorCode())
}
// TestConditionalReadsOfMissingObject verifies that a missing key stays a missing key
// under If-Match and If-Unmodified-Since instead of surfacing as 412.
// reproduces issue #10984
func TestConditionalReadsOfMissingObject(t *testing.T) {
client := getS3Client(t)
bucketName := getNewBucketName()
createBucket(t, client, bucketName)
defer deleteBucket(t, client, bucketName)
existing := putObject(t, client, bucketName, "etag-source", "content")
require.NotNil(t, existing.ETag)
future := aws.Time(time.Now().Add(24 * time.Hour))
missing := aws.String("conditional-missing")
t.Run("HeadObject If-Match", func(t *testing.T) {
_, err := client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName), Key: missing, IfMatch: existing.ETag,
})
requireAPIErrorCode(t, err, "NotFound")
})
t.Run("HeadObject If-Unmodified-Since", func(t *testing.T) {
_, err := client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName), Key: missing, IfUnmodifiedSince: future,
})
requireAPIErrorCode(t, err, "NotFound")
})
t.Run("GetObject If-Match", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: missing, IfMatch: existing.ETag,
})
requireAPIErrorCode(t, err, "NoSuchKey")
})
t.Run("GetObject If-Unmodified-Since", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: missing, IfUnmodifiedSince: future,
})
requireAPIErrorCode(t, err, "NoSuchKey")
})
t.Run("GetObject stale If-Match on a live object stays 412", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: aws.String("etag-source"),
IfMatch: aws.String(`"0000000000000000000000000000dead"`),
})
requireAPIErrorCode(t, err, "PreconditionFailed")
})
}
// TestConditionalReadsOfNamedVersion verifies that a conditional GET or HEAD of an
// explicit versionId is evaluated against that version rather than the latest one,
// including when the latest version is a delete marker.
func TestConditionalReadsOfNamedVersion(t *testing.T) {
client := getS3Client(t)
bucketName := getNewBucketName()
createBucket(t, client, bucketName)
defer deleteBucket(t, client, bucketName)
enableVersioning(t, client, bucketName)
key := "conditional-read-version"
v1 := putObject(t, client, bucketName, key, "content-v1")
require.NotNil(t, v1.ETag)
require.NotNil(t, v1.VersionId)
v2 := putObject(t, client, bucketName, key, "content-v2")
require.NotNil(t, v2.ETag)
require.NotEqual(t, *v1.ETag, *v2.ETag)
t.Run("If-Match matches the named version, not the latest", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: aws.String(key),
VersionId: v1.VersionId, IfMatch: v1.ETag,
})
require.NoError(t, err)
})
t.Run("If-Match against the latest ETag fails on the named version", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: aws.String(key),
VersionId: v1.VersionId, IfMatch: v2.ETag,
})
requireAPIErrorCode(t, err, "PreconditionFailed")
})
_, err := client.DeleteObject(context.TODO(), &s3.DeleteObjectInput{
Bucket: aws.String(bucketName), Key: aws.String(key),
})
require.NoError(t, err)
t.Run("named version survives a delete marker on the latest", func(t *testing.T) {
_, err := client.HeadObject(context.TODO(), &s3.HeadObjectInput{
Bucket: aws.String(bucketName), Key: aws.String(key),
VersionId: v1.VersionId, IfMatch: v1.ETag,
})
require.NoError(t, err)
})
t.Run("delete marker latest is a missing object", func(t *testing.T) {
_, err := client.GetObject(context.TODO(), &s3.GetObjectInput{
Bucket: aws.String(bucketName), Key: aws.String(key), IfMatch: v1.ETag,
})
requireAPIErrorCode(t, err, "NoSuchKey")
})
}
+107
View File
@@ -1,6 +1,7 @@
package tus
import (
"bufio"
"bytes"
"context"
"encoding/base64"
@@ -996,3 +997,109 @@ func TestTusAbortedPatchKeepsStoredChunks(t *testing.T) {
require.NoError(t, err)
assert.Equal(t, testData, body, "the resumed upload should read back whole")
}
// TestTusConcurrentPatchRefused checks that a PATCH sent while another PATCH
// on the same session is still consuming its body is refused with 423 Locked.
// Before the per-session claim, both were accepted at the same offset, recorded
// the range twice, and completion failed on the duplicate while HEAD reported
// the upload fully received - the file was never created and every stored byte
// became garbage when the session expired.
func TestTusConcurrentPatchRefused(t *testing.T) {
if testing.Short() {
t.Skip("Skipping integration test in short mode")
}
ctx, cancel := context.WithTimeout(context.Background(), 180*time.Second)
defer cancel()
cluster, err := startTestCluster(t, ctx)
require.NoError(t, err)
defer func() {
cluster.Stop()
os.RemoveAll(cluster.dataDir)
}()
const subChunkSize = 4 * 1024 * 1024
testData := make([]byte, 3*subChunkSize)
for i := range testData {
testData[i] = byte(i % 251)
}
targetPath := "/raced/video.bin"
client := &http.Client{}
createReq, err := http.NewRequest(http.MethodPost, cluster.TusURL()+targetPath, nil)
require.NoError(t, err)
createReq.Header.Set("Tus-Resumable", TusVersion)
createReq.Header.Set("Upload-Length", strconv.Itoa(len(testData)))
createResp, err := client.Do(createReq)
require.NoError(t, err)
createResp.Body.Close()
require.Equal(t, http.StatusCreated, createResp.StatusCode)
uploadLocation := createResp.Header.Get("Location")
// PATCH A promises one sub-chunk and stalls mid-body, holding the session.
conn, err := net.Dial("tcp", "127.0.0.1:"+testFilerPort)
require.NoError(t, err)
defer conn.Close()
// bound the raw reads below so a filer that never answers fails here
require.NoError(t, conn.SetDeadline(time.Now().Add(60*time.Second)))
_, err = fmt.Fprintf(conn, "PATCH %s HTTP/1.1\r\nHost: 127.0.0.1:%s\r\nTus-Resumable: %s\r\nContent-Type: application/offset+octet-stream\r\nUpload-Offset: 0\r\nContent-Length: %d\r\n\r\n",
uploadLocation, testFilerPort, TusVersion, subChunkSize)
require.NoError(t, err)
_, err = conn.Write(testData[:1024*1024])
require.NoError(t, err)
time.Sleep(2 * time.Second)
// PATCH B is the client's retry of the same range while A is in flight.
retryReq, err := http.NewRequest(http.MethodPatch, cluster.FullURL(uploadLocation), bytes.NewReader(testData[:subChunkSize]))
require.NoError(t, err)
retryReq.Header.Set("Tus-Resumable", TusVersion)
retryReq.Header.Set("Upload-Offset", "0")
retryReq.Header.Set("Content-Type", "application/offset+octet-stream")
retryResp, err := client.Do(retryReq)
require.NoError(t, err)
retryResp.Body.Close()
require.Equal(t, http.StatusLocked, retryResp.StatusCode, "a concurrent PATCH must be refused, not recorded twice")
// Finish PATCH A and read its response.
_, err = conn.Write(testData[1024*1024 : subChunkSize])
require.NoError(t, err)
respReader := bufio.NewReader(conn)
respA, err := http.ReadResponse(respReader, nil)
require.NoError(t, err)
respA.Body.Close()
require.Equal(t, http.StatusNoContent, respA.StatusCode)
headReq, err := http.NewRequest(http.MethodHead, cluster.FullURL(uploadLocation), nil)
require.NoError(t, err)
headReq.Header.Set("Tus-Resumable", TusVersion)
headResp, err := client.Do(headReq)
require.NoError(t, err)
headResp.Body.Close()
require.Equal(t, http.StatusOK, headResp.StatusCode)
currentOffset, err := strconv.Atoi(headResp.Header.Get("Upload-Offset"))
require.NoError(t, err)
require.Equal(t, subChunkSize, currentOffset, "only PATCH A's sub-chunk should be recorded")
patchReq, err := http.NewRequest(http.MethodPatch, cluster.FullURL(uploadLocation), bytes.NewReader(testData[currentOffset:]))
require.NoError(t, err)
patchReq.Header.Set("Tus-Resumable", TusVersion)
patchReq.Header.Set("Upload-Offset", strconv.Itoa(currentOffset))
patchReq.Header.Set("Content-Type", "application/offset+octet-stream")
patchResp, err := client.Do(patchReq)
require.NoError(t, err)
patchResp.Body.Close()
require.Equal(t, http.StatusNoContent, patchResp.StatusCode)
getResp, err := client.Get(cluster.FilerURL() + targetPath)
require.NoError(t, err)
defer getResp.Body.Close()
require.Equal(t, http.StatusOK, getResp.StatusCode)
body, err := io.ReadAll(getResp.Body)
require.NoError(t, err)
assert.Equal(t, testData, body, "the raced upload should complete with intact content")
}
@@ -3,6 +3,7 @@ package volume_server_grpc_test
import (
"context"
"net/http"
"strings"
"testing"
"time"
@@ -414,3 +415,87 @@ func TestScrubEcVolumeIndexCorruptEcx(t *testing.T) {
t.Fatalf("expected broken volume after ECX corruption")
}
}
// scrubEcUntilLocated runs an EC scrub, retrying while the server is still waiting on
// the master for the shard locations a distributed scrub needs.
func scrubEcUntilLocated(t *testing.T, ctx context.Context, grpcClient volume_server_pb.VolumeServerClient, volumeID uint32, mode volume_server_pb.VolumeScrubMode) *volume_server_pb.ScrubEcVolumeResponse {
t.Helper()
deadline := time.Now().Add(60 * time.Second)
for {
resp, err := grpcClient.ScrubEcVolume(ctx, &volume_server_pb.ScrubEcVolumeRequest{
VolumeIds: []uint32{volumeID},
Mode: mode,
})
if err != nil {
t.Fatalf("ScrubEcVolume %s failed: %v", mode, err)
}
waiting := false
for _, d := range resp.GetDetails() {
if strings.Contains(d, "failed to locate shard via master grpc") {
waiting = true
break
}
}
if !waiting {
return resp
}
if time.Now().After(deadline) {
t.Fatalf("master never reported EC shard locations for volume %d: %v", volumeID, resp.GetDetails())
}
time.Sleep(time.Second)
}
}
// With a shard gone, FULL cannot read the intervals that lived on it, while READS
// rebuilds them from parity. Either way the missing shard has to be reported: a
// volume that scrubs clean is a volume nobody repairs.
func TestScrubEcVolumeReadsRecoversMissingShard(t *testing.T) {
if testing.Short() {
t.Skip("skipping integration test in short mode")
}
clusterHarness := framework.StartVolumeCluster(t, matrix.P1())
conn, grpcClient := framework.DialVolumeServer(t, clusterHarness.VolumeGRPCAddress())
defer conn.Close()
const volumeID = uint32(216)
const missingShard = uint32(0)
httpClient := framework.NewHTTPClient()
ecSetup(t, grpcClient, httpClient, clusterHarness.VolumeAdminURL(), volumeID)
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute)
defer cancel()
if _, err := grpcClient.VolumeEcShardsUnmount(ctx, &volume_server_pb.VolumeEcShardsUnmountRequest{
VolumeId: volumeID,
ShardIds: []uint32{missingShard},
}); err != nil {
t.Fatalf("VolumeEcShardsUnmount shard %d failed: %v", missingShard, err)
}
if _, err := grpcClient.VolumeEcShardsDelete(ctx, &volume_server_pb.VolumeEcShardsDeleteRequest{
VolumeId: volumeID,
ShardIds: []uint32{missingShard},
}); err != nil {
t.Fatalf("VolumeEcShardsDelete shard %d failed: %v", missingShard, err)
}
fullResp := scrubEcUntilLocated(t, ctx, grpcClient, volumeID, volume_server_pb.VolumeScrubMode_FULL)
assertOnlyBrokenShard(t, "FULL", fullResp, missingShard)
if len(fullResp.GetDetails()) == 0 {
t.Fatalf("FULL should report the needles it could not read")
}
readsResp := scrubEcUntilLocated(t, ctx, grpcClient, volumeID, volume_server_pb.VolumeScrubMode_READS)
assertOnlyBrokenShard(t, "READS", readsResp, missingShard)
if len(readsResp.GetDetails()) != 0 {
t.Fatalf("READS should rebuild every needle from parity, got: %v", readsResp.GetDetails())
}
}
func assertOnlyBrokenShard(t *testing.T, mode string, resp *volume_server_pb.ScrubEcVolumeResponse, shardID uint32) {
t.Helper()
infos := resp.GetBrokenShardInfos()
if len(infos) != 1 || infos[0].GetShardId() != shardID {
t.Fatalf("%s reported broken shards %v, want only shard %d (details: %v)", mode, infos, shardID, resp.GetDetails())
}
}
+22 -2
View File
@@ -180,11 +180,31 @@ func TestRenameOverExisting(t *testing.T) {
// they indict different layers; a second look says whether it persists.
time.Sleep(200 * time.Millisecond)
fi2, err2 := os.Stat(src)
t.Fatalf("stat of the renamed-away source did not return not-exist: stat=%v err=%v; 200ms later stat=%v err=%v",
describeFileInfo(fi), err, describeFileInfo(fi2), err2)
// A listing reads no per-path cache and the mount's own forgets
// within a second, so a name that survives both is back on the filer.
listed := dirNames(t, dir)
time.Sleep(2 * time.Second)
_, errLater := os.Stat(src)
t.Fatalf("stat of the renamed-away source did not return not-exist: stat=%v err=%v; 200ms later stat=%v err=%v; past the path cache err=%v; %s lists %v",
describeFileInfo(fi), err, describeFileInfo(fi2), err2, errLater, dir, listed)
}
}
// dirNames lists a directory for a failure message, reporting the error in
// place of the names rather than failing a test that is already failing.
func dirNames(t *testing.T, dir string) []string {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
return []string{"readdir: " + err.Error()}
}
names := make([]string, 0, len(entries))
for _, entry := range entries {
names = append(names, entry.Name())
}
return names
}
func TestRenameAcrossDirectories(t *testing.T) {
dir := testRoot(t)
from := filepath.Join(dir, "from")
+147 -37
View File
@@ -2,10 +2,13 @@ package dash
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net"
"net/http"
"net/url"
"sort"
"strings"
"sync"
@@ -34,6 +37,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/s3api"
"github.com/seaweedfs/seaweedfs/weed/s3api/lifecycle_xml"
"github.com/seaweedfs/seaweedfs/weed/s3api/policy_engine"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3_constants"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3lifecycle"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3lifecycle/scheduler"
@@ -157,11 +161,15 @@ type AdminServer struct {
s3TablesManager *s3tables.Manager
icebergPort int
lancePort int
// s3PublicEndpoint is the client-facing S3 address from admin.toml
// (s3.public_endpoint); discovered S3 servers are the fallback.
s3PublicEndpoint string
}
// Type definitions moved to types.go
func NewAdminServer(masters string, filerGroup string, templateFS http.FileSystem, dataDir string, icebergPort, lancePort int) *AdminServer {
func NewAdminServer(masters string, filerGroup string, templateFS http.FileSystem, dataDir string, icebergPort, lancePort int, s3PublicEndpoint string) *AdminServer {
grpcDialOption := security.LoadClientTLS(util.GetViper(), "grpc.admin")
// Create master client with multiple master support
@@ -198,6 +206,7 @@ func NewAdminServer(masters string, filerGroup string, templateFS http.FileSyste
s3TablesManager: newS3TablesManager(),
icebergPort: icebergPort,
lancePort: lancePort,
s3PublicEndpoint: normalizeS3PublicEndpoint(s3PublicEndpoint),
pluginLock: lockManager,
adminPresenceLock: presenceLock,
bgCancel: bgCancel,
@@ -259,33 +268,25 @@ func NewAdminServer(masters string, filerGroup string, templateFS http.FileSyste
maintenanceConfig = maintenance.DefaultMaintenanceConfig()
}
// Apply new defaults to handle schema changes (like enabling by default)
schema := maintenance.GetMaintenanceConfigSchema()
if err := schema.ApplyDefaultsToProtobuf(maintenanceConfig); err != nil {
glog.Warningf("Failed to apply schema defaults to loaded config: %v", err)
}
// Force enable maintenance system for new default behavior
// This handles the case where old configs had Enabled=false as default
if !maintenanceConfig.Enabled {
glog.V(1).Infof("Enabling maintenance system (new default behavior)")
maintenanceConfig.Enabled = true
}
glog.V(1).Infof("Maintenance system initialized with persistent configuration (enabled: %v)", maintenanceConfig.Enabled)
glog.V(1).Infof("Maintenance system initialized with persistent configuration (enabled: %v)", maintenanceConfig.GetEnabled())
} else {
maintenanceConfig = maintenance.DefaultMaintenanceConfig()
glog.V(1).Infof("No data directory configured, maintenance system will run in memory-only mode (enabled: %v)", maintenanceConfig.Enabled)
glog.V(1).Infof("No data directory configured, maintenance system will run in memory-only mode (enabled: %v)", maintenanceConfig.GetEnabled())
}
// Load saved task configurations from persistence. This has to run before the maintenance
// manager is created: creating it applies the maintenance policy to the registered
// detectors and schedulers, while this call replaces each task's whole config object, so
// running it afterwards would discard what the policy just applied. Both read the same
// persisted task config files, so the policy ends up as the last writer and stays
// authoritative for the task types it covers.
server.loadTaskConfigurationsFromPersistence()
// Always initialize maintenance manager
server.InitMaintenanceManager(maintenanceConfig)
// Load saved task configurations from persistence
server.loadTaskConfigurationsFromPersistence()
// Start maintenance manager if enabled
if maintenanceConfig.Enabled {
if maintenanceConfig.GetEnabled() {
go func() {
// Give master client a bit of time to connect before starting scans
time.Sleep(2 * time.Second)
@@ -293,6 +294,8 @@ func NewAdminServer(masters string, filerGroup string, templateFS http.FileSyste
glog.Errorf("Failed to start maintenance manager: %v", err)
}
}()
} else {
glog.V(0).Infof("Maintenance system is disabled by configuration, not starting the maintenance manager")
}
pluginOpts := adminplugin.Options{
@@ -432,7 +435,72 @@ func (s *AdminServer) publishMaintenanceMetrics(ctx context.Context) {
}
}
// workerFleetTotals aggregates connected workers and their task slots across
// BOTH worker registries the admin server keeps: the legacy maintenance-worker
// registry (workers that register over the worker gRPC stream) and the plugin
// worker registry (workers started as `weed worker`). Reading only the legacy
// one reported zero workers on clusters that run the admin and the workers as
// separate components, where no legacy worker ever registers.
//
// A worker can appear in both registries: `weed mini` starts both runtimes from
// one working directory, so they share the persisted worker ID. Merging by ID
// keeps such a worker counted once, and its slots are taken from the legacy
// registry, which is where they were accounted for before.
func (s *AdminServer) workerFleetTotals() (workers, usedSlots, maxSlots int) {
var legacySlots map[string]maintenance.WorkerSlots
if s.maintenanceManager != nil {
legacySlots = s.maintenanceManager.GetWorkerSlots()
}
return mergeWorkerFleetTotals(legacySlots, s.GetPluginWorkers())
}
// mergeWorkerFleetTotals unions the legacy and plugin worker registries by
// worker ID. Plugin workers report their slots in the heartbeat, so one that
// has connected but not yet sent a heartbeat adds to the worker count with zero
// slots until its first heartbeat lands.
func mergeWorkerFleetTotals(legacySlots map[string]maintenance.WorkerSlots, pluginWorkers []*adminplugin.WorkerSession) (workers, usedSlots, maxSlots int) {
for _, slots := range legacySlots {
workers++
usedSlots += slots.Used
maxSlots += slots.Max
}
for _, session := range pluginWorkers {
if session == nil {
continue
}
if _, counted := legacySlots[session.WorkerID]; counted {
continue
}
workers++
if heartbeat := session.Heartbeat; heartbeat != nil {
used := int(heartbeat.DetectionSlotsUsed) + int(heartbeat.ExecutionSlotsUsed)
max := int(heartbeat.DetectionSlotsTotal) + int(heartbeat.ExecutionSlotsTotal)
// A worker's self-reported slots are untrusted input; a stale or
// misbehaving one should not be able to drive the aggregate gauge
// negative, matching the same defensiveness as registry.go's own
// slot arithmetic.
if used < 0 {
used = 0
}
if max < 0 {
max = 0
}
usedSlots += used
maxSlots += max
}
}
return
}
func (s *AdminServer) collectMaintenanceMetrics() {
// Published before the maintenanceManager guard below: plugin workers are
// tracked independently of the maintenance manager.
workers, usedSlots, maxSlots := s.workerFleetTotals()
stats_collect.AdminWorkersConnected.Set(float64(workers))
stats_collect.AdminWorkerSlots.WithLabelValues("used").Set(float64(usedSlots))
stats_collect.AdminWorkerSlots.WithLabelValues("max").Set(float64(maxSlots))
if s.maintenanceManager == nil {
return
}
@@ -456,11 +524,6 @@ func (s *AdminServer) collectMaintenanceMetrics() {
} else {
stats_collect.AdminMaintenanceNextScanTimestampSeconds.Set(0)
}
workers, usedSlots, maxSlots := s.maintenanceManager.GetWorkerSlotTotals()
stats_collect.AdminWorkersConnected.Set(float64(workers))
stats_collect.AdminWorkerSlots.WithLabelValues("used").Set(float64(usedSlots))
stats_collect.AdminWorkerSlots.WithLabelValues("max").Set(float64(maxSlots))
}
// loadTaskConfigurationsFromPersistence loads saved task configurations from protobuf files
@@ -823,6 +886,7 @@ func (s *AdminServer) GetS3Buckets() ([]S3Bucket, error) {
Owner: owner,
LifecycleRuleCount: lifecycleRuleCount,
LifecycleEnabledCount: lifecycleEnabledCount,
PolicyStatementCount: extractPolicyStatementCountFromEntry(resp.Entry),
}
buckets = append(buckets, bucket)
}
@@ -928,6 +992,7 @@ func (s *AdminServer) GetBucketDetails(bucketName string) (*BucketDetails, error
details.Bucket.ObjectLockDuration = objectLockDuration
details.Bucket.Owner = owner
details.Bucket.LifecycleRuleCount, details.Bucket.LifecycleEnabledCount = extractLifecycleCountsFromEntry(bucketResp.Entry)
details.Bucket.PolicyStatementCount = extractPolicyStatementCountFromEntry(bucketResp.Entry)
return nil
})
@@ -1168,14 +1233,8 @@ func (s *AdminServer) DeleteS3Bucket(bucketName string) error {
// Then delete bucket directory recursively from filer
// Use same parameters as s3.bucket.delete shell command and S3 API
return s.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
_, err := client.DeleteEntry(ctx, &filer_pb.DeleteEntryRequest{
Directory: filerConfig.BucketsPath,
Name: bucketName,
IsDeleteData: false, // Collection already deleted, just remove metadata
IsRecursive: true,
IgnoreRecursiveError: true, // Same as S3 API and shell command
})
if err != nil {
// The collection is already gone, so this only has to drop the metadata.
if err := filer_pb.DoRemove(ctx, client, filerConfig.BucketsPath, bucketName, false, true, true, false, nil); err != nil {
return fmt.Errorf("failed to delete bucket: %w", err)
}
@@ -1488,6 +1547,32 @@ func (s *AdminServer) GetClusterS3Servers() (*ClusterS3ServersData, error) {
}, nil
}
// GetS3Endpoint returns the configured address clients reach the S3 gateway
// at, or "". S3 servers register only their gRPC address with the master, so
// the client-facing address cannot be discovered and must be configured.
func (s *AdminServer) GetS3Endpoint() string {
return s.s3PublicEndpoint
}
// normalizeS3PublicEndpoint trims a trailing slash and drops, with a warning,
// a value that is not a plain absolute http or https URL, so the file browser
// hides its URL actions instead of copying broken links. Checking the string
// for "?" and "#" rather than the parsed query and fragment also catches
// delimiters with nothing after them, which url.Parse stores as empty.
func normalizeS3PublicEndpoint(endpoint string) string {
endpoint = strings.TrimRight(endpoint, "/")
if endpoint == "" {
return ""
}
u, err := url.Parse(endpoint)
if err != nil || (u.Scheme != "http" && u.Scheme != "https") || u.Host == "" || u.User != nil || strings.ContainsAny(endpoint, "?#") {
// the value is not echoed: it may hold credentials in userinfo or a query
glog.Warningf("ignoring s3.public_endpoint: expecting an http:// or https:// URL with a host and no credentials, query, or fragment")
return ""
}
return endpoint
}
// GetAllFilers method moved to client_management.go
// GetVolumeDetails method moved to volume_management.go
@@ -1514,6 +1599,7 @@ func (as *AdminServer) GetConfigInfo(w http.ResponseWriter, r *http.Request) {
configInfo["master_address"] = string(currentMaster)
configInfo["cache_expiration"] = as.cacheExpiration.String()
configInfo["filer_cache_expiration"] = as.filerCacheExpiration.String()
configInfo["s3_public_endpoint"] = as.s3PublicEndpoint
// Add maintenance system info
if as.maintenanceManager != nil {
@@ -1531,13 +1617,13 @@ func (as *AdminServer) GetConfigInfo(w http.ResponseWriter, r *http.Request) {
}
// StartWorkerGrpcServer starts the worker gRPC server
func (s *AdminServer) StartWorkerGrpcServer(grpcPort int) error {
func (s *AdminServer) StartWorkerGrpcServer(grpcPort int, listener net.Listener) error {
if s.workerGrpcServer != nil {
return fmt.Errorf("worker gRPC server is already running")
}
s.workerGrpcServer = NewWorkerGrpcServer(s)
return s.workerGrpcServer.StartWithTLS(grpcPort)
return s.workerGrpcServer.StartWithTLS(grpcPort, listener)
}
// StopWorkerGrpcServer stops the worker gRPC server
@@ -1779,7 +1865,16 @@ func (s *AdminServer) ListPluginSchedulerStates() ([]adminplugin.SchedulerJobTyp
// InitMaintenanceManager initializes the maintenance manager
func (s *AdminServer) InitMaintenanceManager(config *maintenance.MaintenanceConfig) {
s.maintenanceManager = maintenance.NewMaintenanceManager(s, config)
// Hand the real config store to the manager so that, if it has to build the maintenance policy
// itself, it reads the persisted task configs instead of compiled-in defaults. Only pass it when
// a data directory is actually configured: an unconfigured store has nothing to read, and a typed
// nil pointer would satisfy the loaders' type assertion and then panic on use.
var configPersistence interface{}
if s.configPersistence != nil && s.configPersistence.IsConfigured() {
configPersistence = s.configPersistence
}
s.maintenanceManager = maintenance.NewMaintenanceManager(s, config, configPersistence)
// Set up task persistence if config persistence is available
if s.configPersistence != nil {
@@ -1794,7 +1889,7 @@ func (s *AdminServer) InitMaintenanceManager(config *maintenance.MaintenanceConf
}
}
glog.V(1).Infof("Maintenance manager initialized (enabled: %v)", config.Enabled)
glog.V(1).Infof("Maintenance manager initialized (enabled: %v)", config.GetEnabled())
}
// GetMaintenanceManager returns the maintenance manager
@@ -2032,6 +2127,21 @@ func extractLifecycleCountsFromEntry(entry *filer_pb.Entry) (ruleCount, enabledC
return
}
// extractPolicyStatementCountFromEntry returns the number of statements in
// the bucket's policy, or 0 if it has none or the stored JSON can't be
// parsed. Forgiving on parse failure, same as extractLifecycleCountsFromEntry.
func extractPolicyStatementCountFromEntry(entry *filer_pb.Entry) int {
policyJSON := entry.Extended[s3api.BUCKET_POLICY_METADATA_KEY]
if len(policyJSON) == 0 {
return 0
}
var doc policy_engine.PolicyDocument
if err := json.Unmarshal(policyJSON, &doc); err != nil {
return 0
}
return len(doc.Statement)
}
// GetConfigPersistence returns the config persistence manager
func (as *AdminServer) GetConfigPersistence() *ConfigPersistence {
return as.configPersistence
+101
View File
@@ -13,6 +13,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/filer"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/s3api"
"github.com/seaweedfs/seaweedfs/weed/s3api/policy_engine"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3_constants"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3lifecycle"
)
@@ -257,6 +258,106 @@ func validateBucketLifecycleRules(rules []BucketLifecycleRule) error {
return nil
}
// ShowBucketPolicy returns the policy document for a specific bucket, or
// {"bucket": ..., "policy": null} if the bucket has none.
func (s *AdminServer) ShowBucketPolicy(w http.ResponseWriter, r *http.Request) {
bucketName := mux.Vars(r)["bucket"]
if bucketName == "" {
writeJSONError(w, http.StatusBadRequest, "Bucket name is required")
return
}
policy, raw, err := s.GetBucketPolicy(bucketName)
if err != nil {
writeJSONError(w, bucketPolicyErrorStatus(err), "Failed to get bucket policy: "+err.Error())
return
}
resp := map[string]interface{}{
"bucket": bucketName,
"policy": policy,
}
if policy == nil && len(raw) > 0 {
// Stored bytes the decoder rejects: hand them to the JSON tab so
// the operator can fix or delete the document.
resp["policy_text"] = string(raw)
}
writeJSON(w, http.StatusOK, resp)
}
// UpdateBucketPolicy replaces the bucket policy for a bucket.
func (s *AdminServer) UpdateBucketPolicy(w http.ResponseWriter, r *http.Request) {
if !requireSessionCSRFToken(w, r) {
return
}
bucketName := mux.Vars(r)["bucket"]
if bucketName == "" {
writeJSONError(w, http.StatusBadRequest, "Bucket name is required")
return
}
var req struct {
Policy *policy_engine.PolicyDocument `json:"policy"`
}
if err := decodeJSONBody(newJSONMaxReader(w, r), &req); err != nil {
writeJSONError(w, http.StatusBadRequest, "Invalid request: "+err.Error())
return
}
if req.Policy == nil {
writeJSONError(w, http.StatusBadRequest, "policy is required; use DELETE to clear a bucket policy")
return
}
if err := s.SetBucketPolicy(bucketName, req.Policy); err != nil {
writeJSONError(w, bucketPolicyErrorStatus(err), "Failed to update bucket policy: "+err.Error())
return
}
writeJSON(w, http.StatusOK, map[string]interface{}{
"message": "Bucket policy updated successfully",
"bucket": bucketName,
})
}
// RemoveBucketPolicy clears the bucket policy for a bucket. Named
// "Remove", not "Delete", because (*AdminServer).DeleteBucketPolicy is the
// data-layer method this handler calls.
func (s *AdminServer) RemoveBucketPolicy(w http.ResponseWriter, r *http.Request) {
if !requireSessionCSRFToken(w, r) {
return
}
bucketName := mux.Vars(r)["bucket"]
if bucketName == "" {
writeJSONError(w, http.StatusBadRequest, "Bucket name is required")
return
}
if err := s.DeleteBucketPolicy(bucketName); err != nil {
writeJSONError(w, bucketPolicyErrorStatus(err), "Failed to delete bucket policy: "+err.Error())
return
}
writeJSON(w, http.StatusOK, map[string]interface{}{
"message": "Bucket policy deleted successfully",
"bucket": bucketName,
})
}
// bucketPolicyErrorStatus keeps a request for a bucket that does not exist,
// or an invalid policy document, out of the 5xx bucket where a client
// would retry it. Mirrors bucketLifecycleErrorStatus.
func bucketPolicyErrorStatus(err error) int {
if errors.Is(err, ErrBucketNotFound) {
return http.StatusNotFound
}
if errors.Is(err, ErrInvalidBucketPolicy) {
return http.StatusBadRequest
}
return http.StatusInternalServerError
}
// CreateBucket creates a new S3 bucket
func (s *AdminServer) CreateBucket(w http.ResponseWriter, r *http.Request) {
var req CreateBucketRequest
+161
View File
@@ -0,0 +1,161 @@
package dash
import (
"context"
"encoding/json"
"errors"
"fmt"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/s3api"
"github.com/seaweedfs/seaweedfs/weed/s3api/policy_engine"
"github.com/seaweedfs/seaweedfs/weed/s3api/s3_constants"
)
// ErrInvalidBucketPolicy wraps a validation failure from SetBucketPolicy so
// callers (the HTTP handler) can map it to 400 instead of 500 without
// resorting to matching on the error string.
var ErrInvalidBucketPolicy = errors.New("invalid bucket policy")
// GetBucketPolicy returns the policy document stored on a bucket's filer
// entry, or (nil, nil, nil) if the bucket has no policy — that is not an
// error, it just means the caller (e.g. the admin UI) should show an empty
// editor instead of special-casing a 404. Stored bytes the current decoder
// rejects come back as (nil, raw, nil): a 500 here would leave the UI
// unable to show, fix, or even delete the one policy an operator most
// needs to remove — and DeleteBucketPolicy never reads the document.
func (s *AdminServer) GetBucketPolicy(bucketName string) (*policy_engine.PolicyDocument, []byte, error) {
filerConfig, err := s.getFilerConfig()
if err != nil {
return nil, nil, fmt.Errorf("get filer configuration: %w", err)
}
var doc *policy_engine.PolicyDocument
var raw []byte
err = s.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
resp, err := filer_pb.LookupEntry(context.Background(), client, &filer_pb.LookupDirectoryEntryRequest{
Directory: filerConfig.BucketsPath,
Name: bucketName,
})
if err != nil {
if errors.Is(err, filer_pb.ErrNotFound) {
return fmt.Errorf("%w: %s", ErrBucketNotFound, bucketName)
}
return fmt.Errorf("look up bucket %s: %w", bucketName, err)
}
policyJSON := resp.Entry.Extended[s3api.BUCKET_POLICY_METADATA_KEY]
if len(policyJSON) == 0 {
return nil
}
raw = policyJSON
var parsed policy_engine.PolicyDocument
if err := json.Unmarshal(policyJSON, &parsed); err == nil {
doc = &parsed
}
return nil
})
if err != nil {
return nil, nil, err
}
return doc, raw, nil
}
// SetBucketPolicy validates and stores a bucket policy, applying the exact
// same validation the S3 gateway's PutBucketPolicy enforces
// (policy_engine.ValidatePolicy + policy_engine.ValidateBucketPolicy), so
// the admin UI and the S3 API never disagree about what's a valid policy.
//
// Propagation to every S3 gateway is automatic: writing the
// s3-bucket-policy extended attribute drives the filer metadata log, which
// each gateway's onBucketMetadataChange subscription watches to rebuild its
// bucket policy cache and to maintain the advanced-IAM
// "bucket-policy:<bucket>" mirror (mirrorBucketPolicyToIAM in
// weed/s3api/s3api_bucket_policy_handlers.go). No separate notify step is
// needed here.
func (s *AdminServer) SetBucketPolicy(bucketName string, doc *policy_engine.PolicyDocument) error {
if err := policy_engine.ValidatePolicy(doc); err != nil {
return fmt.Errorf("%w: %w", ErrInvalidBucketPolicy, err)
}
if err := policy_engine.ValidateBucketPolicy(doc, bucketName); err != nil {
return fmt.Errorf("%w: %w", ErrInvalidBucketPolicy, err)
}
policyJSON, err := json.Marshal(doc)
if err != nil {
return fmt.Errorf("marshal policy document: %w", err)
}
if len(policyJSON) > policy_engine.MaxBucketPolicySize {
return fmt.Errorf("%w: bucket policy is %d bytes, which exceeds the %d byte limit", ErrInvalidBucketPolicy, len(policyJSON), policy_engine.MaxBucketPolicySize)
}
return s.writeBucketPolicy(bucketName, policyJSON)
}
// DeleteBucketPolicy clears the bucket policy stored on a bucket's filer
// entry. Deleting a policy that doesn't exist is a success, matching
// DeleteBucketLifecycle's idempotent behavior. This cannot go through
// SetBucketPolicy: validation there rejects a nil document.
func (s *AdminServer) DeleteBucketPolicy(bucketName string) error {
return s.writeBucketPolicy(bucketName, nil)
}
// writeBucketPolicy patches the policy key on the bucket's filer entry; a
// nil policyJSON clears it.
func (s *AdminServer) writeBucketPolicy(bucketName string, policyJSON []byte) error {
filerConfig, err := s.getFilerConfig()
if err != nil {
return fmt.Errorf("get filer configuration: %w", err)
}
return s.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
// PATCH_EXTENDED is a no-op on a missing entry, so the existence
// check has to happen here rather than fall out of the write.
if _, err := filer_pb.LookupEntry(context.Background(), client, &filer_pb.LookupDirectoryEntryRequest{
Directory: filerConfig.BucketsPath,
Name: bucketName,
}); err != nil {
if errors.Is(err, filer_pb.ErrNotFound) {
return fmt.Errorf("%w: %s", ErrBucketNotFound, bucketName)
}
return fmt.Errorf("look up bucket %s: %w", bucketName, err)
}
bucketPath := filerConfig.BucketsPath + "/" + bucketName
resp, err := client.ObjectTransaction(context.Background(), &filer_pb.ObjectTransactionRequest{
LockKey: bucketPath,
RouteKey: s3_constants.ObjectWriteRouteKeyPrefix + bucketPath,
Mutations: []*filer_pb.ObjectMutation{bucketPolicyMutation(filerConfig.BucketsPath, bucketName, policyJSON)},
})
if err != nil {
return fmt.Errorf("write bucket policy: %w", err)
}
if resp.Error != "" {
return fmt.Errorf("write bucket policy: %s", resp.Error)
}
return nil
})
}
// bucketPolicyMutation patches only the policy key rather than writing the
// whole entry back: the filer re-reads and merges under the bucket path
// lock, so a concurrent owner/quota/versioning/lifecycle change is
// preserved instead of being reverted by a stale snapshot. Same pattern as
// bucketLifecycleMutation. A nil/empty policyJSON clears the key.
func bucketPolicyMutation(bucketsPath, bucketName string, policyJSON []byte) *filer_pb.ObjectMutation {
mutation := &filer_pb.ObjectMutation{
Type: filer_pb.ObjectMutation_PATCH_EXTENDED,
Directory: bucketsPath,
Name: bucketName,
}
if len(policyJSON) > 0 {
mutation.SetExtended = map[string][]byte{
s3api.BUCKET_POLICY_METADATA_KEY: policyJSON,
}
return mutation
}
mutation.DeleteExtended = []string{s3api.BUCKET_POLICY_METADATA_KEY}
return mutation
}
+168
View File
@@ -0,0 +1,168 @@
package dash
import (
"encoding/json"
"errors"
"fmt"
"net/http"
"strings"
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/s3api"
"github.com/seaweedfs/seaweedfs/weed/s3api/policy_engine"
)
func validBucketPolicyJSON(bucket string) []byte {
return []byte(fmt.Sprintf(`{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Principal":"*","Action":"s3:GetObject","Resource":"arn:aws:s3:::%s/*"}]}`, bucket))
}
func validBucketPolicyDoc(bucket string) *policy_engine.PolicyDocument {
var doc policy_engine.PolicyDocument
if err := json.Unmarshal(validBucketPolicyJSON(bucket), &doc); err != nil {
panic(err)
}
return &doc
}
func TestBucketPolicyMutation_SetsPolicy(t *testing.T) {
policyJSON := validBucketPolicyJSON("mybucket")
m := bucketPolicyMutation("/buckets", "mybucket", policyJSON)
if m.Type != filer_pb.ObjectMutation_PATCH_EXTENDED {
t.Fatalf("expected a PATCH_EXTENDED mutation, got %v", m.Type)
}
if m.Directory != "/buckets" || m.Name != "mybucket" {
t.Fatalf("expected the mutation to target /buckets/mybucket, got %s/%s", m.Directory, m.Name)
}
if got := m.SetExtended[s3api.BUCKET_POLICY_METADATA_KEY]; string(got) != string(policyJSON) {
t.Fatalf("expected the policy key to carry the marshaled document, got %q", got)
}
if len(m.DeleteExtended) != 0 {
t.Fatalf("expected no key deletions when saving a policy, got %v", m.DeleteExtended)
}
}
func TestBucketPolicyMutation_ClearsKey(t *testing.T) {
m := bucketPolicyMutation("/buckets", "mybucket", nil)
if len(m.SetExtended) != 0 {
t.Fatalf("expected no key writes when clearing, got %v", m.SetExtended)
}
if len(m.DeleteExtended) != 1 || m.DeleteExtended[0] != s3api.BUCKET_POLICY_METADATA_KEY {
t.Fatalf("expected only the policy key to be cleared, got %v", m.DeleteExtended)
}
}
// A whole-entry write would have carried the rest of the bucket entry with
// it; the patch must name only the key it owns, so a concurrent owner,
// quota, or lifecycle change survives.
func TestBucketPolicyMutation_TouchesOnlyPolicyKey(t *testing.T) {
for _, m := range []*filer_pb.ObjectMutation{
bucketPolicyMutation("/buckets", "mybucket", validBucketPolicyJSON("mybucket")),
bucketPolicyMutation("/buckets", "mybucket", nil),
} {
if m.Entry != nil {
t.Fatal("expected the mutation to carry no entry snapshot")
}
if m.SetContent {
t.Fatal("expected the mutation to leave entry content alone")
}
for k := range m.SetExtended {
if k != s3api.BUCKET_POLICY_METADATA_KEY {
t.Fatalf("unexpected key written: %s", k)
}
}
for _, k := range m.DeleteExtended {
if k != s3api.BUCKET_POLICY_METADATA_KEY {
t.Fatalf("unexpected key deleted: %s", k)
}
}
}
}
func TestExtractPolicyStatementCountFromEntry(t *testing.T) {
tests := []struct {
name string
entry *filer_pb.Entry
want int
}{
{"no extended attrs", &filer_pb.Entry{}, 0},
{"absent key", &filer_pb.Entry{Extended: map[string][]byte{"other": []byte("x")}}, 0},
{"one statement", &filer_pb.Entry{Extended: map[string][]byte{
s3api.BUCKET_POLICY_METADATA_KEY: validBucketPolicyJSON("b"),
}}, 1},
{"three statements", &filer_pb.Entry{Extended: map[string][]byte{
s3api.BUCKET_POLICY_METADATA_KEY: []byte(`{"Version":"2012-10-17","Statement":[
{"Effect":"Allow","Principal":"*","Action":"s3:GetObject","Resource":"arn:aws:s3:::b/*"},
{"Effect":"Allow","Principal":"*","Action":"s3:PutObject","Resource":"arn:aws:s3:::b/*"},
{"Effect":"Deny","Principal":"*","Action":"s3:DeleteObject","Resource":"arn:aws:s3:::b/*"}
]}`),
}}, 3},
{"garbage bytes", &filer_pb.Entry{Extended: map[string][]byte{
s3api.BUCKET_POLICY_METADATA_KEY: []byte("not json"),
}}, 0},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := extractPolicyStatementCountFromEntry(tt.entry); got != tt.want {
t.Errorf("extractPolicyStatementCountFromEntry() = %d, want %d", got, tt.want)
}
})
}
}
func TestBucketPolicyErrorStatus(t *testing.T) {
if got := bucketPolicyErrorStatus(fmt.Errorf("%w: mybucket", ErrBucketNotFound)); got != http.StatusNotFound {
t.Fatalf("expected a missing bucket to map to 404, got %d", got)
}
if got := bucketPolicyErrorStatus(fmt.Errorf("%w: bad statement", ErrInvalidBucketPolicy)); got != http.StatusBadRequest {
t.Fatalf("expected an invalid policy to map to 400, got %d", got)
}
if got := bucketPolicyErrorStatus(errors.New("filer unreachable")); got != http.StatusInternalServerError {
t.Fatalf("expected an unrelated failure to stay 500, got %d", got)
}
}
func TestSetBucketPolicy_RejectsOversized(t *testing.T) {
// A resource list long enough to blow the cap: this must fail before
// any filer call, which is what makes it testable without one.
doc := validBucketPolicyDoc("mybucket")
doc.Statement[0].Sid = strings.Repeat("x", policy_engine.MaxBucketPolicySize+1)
err := (&AdminServer{}).SetBucketPolicy("mybucket", doc)
if err == nil {
t.Fatal("expected an oversized bucket policy to be rejected")
}
if !errors.Is(err, ErrInvalidBucketPolicy) {
t.Fatalf("expected an ErrInvalidBucketPolicy, got: %v", err)
}
}
func TestSetBucketPolicy_RejectsForeignResource(t *testing.T) {
// Proves the shared policy_engine.ValidateBucketPolicy validator is
// actually wired in: this is exactly the check the S3 gateway applies.
doc := validBucketPolicyDoc("mybucket")
doc.Statement[0].Resource = policy_engine.NewStringOrStringSlicePtr("arn:aws:s3:::other-bucket/*")
err := (&AdminServer{}).SetBucketPolicy("mybucket", doc)
if err == nil {
t.Fatal("expected a policy referencing a different bucket to be rejected")
}
if !errors.Is(err, ErrInvalidBucketPolicy) {
t.Fatalf("expected an ErrInvalidBucketPolicy, got: %v", err)
}
}
func TestSetBucketPolicy_RejectsMissingPrincipal(t *testing.T) {
doc := validBucketPolicyDoc("mybucket")
doc.Statement[0].Principal = nil
err := (&AdminServer{}).SetBucketPolicy("mybucket", doc)
if err == nil {
t.Fatal("expected a policy with no Principal to be rejected")
}
if !errors.Is(err, ErrInvalidBucketPolicy) {
t.Fatalf("expected an ErrInvalidBucketPolicy, got: %v", err)
}
}
+1 -1
View File
@@ -20,7 +20,7 @@ import (
// WithMasterClient executes a function with a master client connection
func (s *AdminServer) WithMasterClient(f func(client master_pb.SeaweedClient) error) error {
return s.masterClient.WithClient(false, f)
return s.masterClient.WithClient(context.Background(), false, f)
}
// WithFilerClient executes a function with a filer client connection
+110 -152
View File
@@ -14,6 +14,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/pb/worker_pb"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/ec_balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/vacuum"
"google.golang.org/protobuf/encoding/protojson"
@@ -29,6 +30,7 @@ const (
VacuumTaskConfigFile = "task_vacuum.pb"
ECTaskConfigFile = "task_erasure_coding.pb"
BalanceTaskConfigFile = "task_balance.pb"
EcBalanceTaskConfigFile = "task_ec_balance.pb"
ReplicationTaskConfigFile = "task_replication.pb"
// JSON reference files
@@ -36,6 +38,7 @@ const (
VacuumTaskConfigJSONFile = "task_vacuum.json"
ECTaskConfigJSONFile = "task_erasure_coding.json"
BalanceTaskConfigJSONFile = "task_balance.json"
EcBalanceTaskConfigJSONFile = "task_ec_balance.json"
ReplicationTaskConfigJSONFile = "task_replication.json"
// Task persistence subdirectories and settings
@@ -53,6 +56,7 @@ type (
VacuumTaskConfig = worker_pb.VacuumTaskConfig
ErasureCodingTaskConfig = worker_pb.ErasureCodingTaskConfig
BalanceTaskConfig = worker_pb.BalanceTaskConfig
EcBalanceTaskConfig = worker_pb.EcBalanceTaskConfig
ReplicationTaskConfig = worker_pb.ReplicationTaskConfig
)
@@ -155,8 +159,19 @@ func (cp *ConfigPersistence) LoadMaintenanceConfig() (*MaintenanceConfig, error)
if configData, err := os.ReadFile(configPath); err == nil {
var config MaintenanceConfig
if err := proto.Unmarshal(configData, &config); err == nil {
// Fill in fields added to the schema after this file was written. The
// enabled flag tracks presence, so an explicitly persisted false survives
// this, while a file from before presence tracking (where an operator's
// explicit false and the field's absence look the same on the wire) keeps
// the enabled default rather than silently switching maintenance off.
if err := maintenance.GetMaintenanceConfigSchema().ApplyDefaultsToProtobuf(&config); err != nil {
glog.Warningf("Failed to apply schema defaults to loaded maintenance config: %v", err)
}
if config.Enabled == nil {
config.Enabled = proto.Bool(true)
}
// Always populate policy from separate task configuration files
config.Policy = buildPolicyFromTaskConfigs()
config.Policy = cp.buildPolicyFromTaskConfigs()
return &config, nil
}
}
@@ -268,6 +283,28 @@ func (cp *ConfigPersistence) RestoreConfig(filename, backupName string) error {
return nil
}
// Default task policies. These derive from each task's own NewDefaultConfig() so that a
// task type has exactly one definition of its defaults. They used to be hand-written copies
// here, and had drifted from the values the tasks themselves and the admin UI schema use:
// vacuum scanned every 24h instead of 2h, balance every 6h instead of 30m with a 0.1 instead
// of 0.2 imbalance threshold, and erasure coding every 168h instead of 1h with a 0.90 instead
// of 0.95 fullness ratio and a 1024MB instead of 30MB minimum volume size.
func defaultVacuumTaskPolicy() *worker_pb.TaskPolicy {
return vacuum.NewDefaultConfig().ToTaskPolicy()
}
func defaultErasureCodingTaskPolicy() *worker_pb.TaskPolicy {
return erasure_coding.NewDefaultConfig().ToTaskPolicy()
}
func defaultBalanceTaskPolicy() *worker_pb.TaskPolicy {
return balance.NewDefaultConfig().ToTaskPolicy()
}
func defaultEcBalanceTaskPolicy() *worker_pb.TaskPolicy {
return ec_balance.NewDefaultConfig().ToTaskPolicy()
}
// SaveVacuumTaskConfig saves vacuum task configuration to protobuf file
func (cp *ConfigPersistence) SaveVacuumTaskConfig(config *VacuumTaskConfig) error {
return cp.saveTaskConfig(VacuumTaskConfigFile, config)
@@ -288,28 +325,14 @@ func (cp *ConfigPersistence) LoadVacuumTaskConfig() (*VacuumTaskConfig, error) {
}
// Return default config if no valid config found
return &VacuumTaskConfig{
GarbageThreshold: 0.3,
MinVolumeAgeHours: 24,
}, nil
return defaultVacuumTaskPolicy().GetVacuumConfig(), nil
}
// LoadVacuumTaskPolicy loads complete vacuum task policy from protobuf file
func (cp *ConfigPersistence) LoadVacuumTaskPolicy() (*worker_pb.TaskPolicy, error) {
if cp.dataDir == "" {
// Return default policy if no data directory
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 2,
RepeatIntervalSeconds: 24 * 3600, // 24 hours in seconds
CheckIntervalSeconds: 6 * 3600, // 6 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_VacuumConfig{
VacuumConfig: &worker_pb.VacuumTaskConfig{
GarbageThreshold: 0.3,
MinVolumeAgeHours: 24,
},
},
}, nil
return defaultVacuumTaskPolicy(), nil
}
confDir := filepath.Join(cp.dataDir, ConfigSubdir)
@@ -318,18 +341,7 @@ func (cp *ConfigPersistence) LoadVacuumTaskPolicy() (*worker_pb.TaskPolicy, erro
// Check if file exists
if _, err := os.Stat(configPath); os.IsNotExist(err) {
// Return default policy if file doesn't exist
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 2,
RepeatIntervalSeconds: 24 * 3600, // 24 hours in seconds
CheckIntervalSeconds: 6 * 3600, // 6 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_VacuumConfig{
VacuumConfig: &worker_pb.VacuumTaskConfig{
GarbageThreshold: 0.3,
MinVolumeAgeHours: 24,
},
},
}, nil
return defaultVacuumTaskPolicy(), nil
}
// Read file
@@ -371,32 +383,14 @@ func (cp *ConfigPersistence) LoadErasureCodingTaskConfig() (*ErasureCodingTaskCo
}
// Return default config if no valid config found
return &ErasureCodingTaskConfig{
FullnessRatio: 0.9,
QuietForSeconds: 3600,
MinVolumeSizeMb: 1024,
CollectionFilter: "",
}, nil
return defaultErasureCodingTaskPolicy().GetErasureCodingConfig(), nil
}
// LoadErasureCodingTaskPolicy loads complete EC task policy from protobuf file
func (cp *ConfigPersistence) LoadErasureCodingTaskPolicy() (*worker_pb.TaskPolicy, error) {
if cp.dataDir == "" {
// Return default policy if no data directory
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 1,
RepeatIntervalSeconds: 168 * 3600, // 1 week in seconds
CheckIntervalSeconds: 24 * 3600, // 24 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_ErasureCodingConfig{
ErasureCodingConfig: &worker_pb.ErasureCodingTaskConfig{
FullnessRatio: 0.9,
QuietForSeconds: 3600,
MinVolumeSizeMb: 1024,
CollectionFilter: "",
},
},
}, nil
return defaultErasureCodingTaskPolicy(), nil
}
confDir := filepath.Join(cp.dataDir, ConfigSubdir)
@@ -405,20 +399,7 @@ func (cp *ConfigPersistence) LoadErasureCodingTaskPolicy() (*worker_pb.TaskPolic
// Check if file exists
if _, err := os.Stat(configPath); os.IsNotExist(err) {
// Return default policy if file doesn't exist
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 1,
RepeatIntervalSeconds: 168 * 3600, // 1 week in seconds
CheckIntervalSeconds: 24 * 3600, // 24 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_ErasureCodingConfig{
ErasureCodingConfig: &worker_pb.ErasureCodingTaskConfig{
FullnessRatio: 0.9,
QuietForSeconds: 3600,
MinVolumeSizeMb: 1024,
CollectionFilter: "",
},
},
}, nil
return defaultErasureCodingTaskPolicy(), nil
}
// Read file
@@ -460,28 +441,14 @@ func (cp *ConfigPersistence) LoadBalanceTaskConfig() (*BalanceTaskConfig, error)
}
// Return default config if no valid config found
return &BalanceTaskConfig{
ImbalanceThreshold: 0.1,
MinServerCount: 2,
}, nil
return defaultBalanceTaskPolicy().GetBalanceConfig(), nil
}
// LoadBalanceTaskPolicy loads complete balance task policy from protobuf file
func (cp *ConfigPersistence) LoadBalanceTaskPolicy() (*worker_pb.TaskPolicy, error) {
if cp.dataDir == "" {
// Return default policy if no data directory
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 1,
RepeatIntervalSeconds: 6 * 3600, // 6 hours in seconds
CheckIntervalSeconds: 12 * 3600, // 12 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_BalanceConfig{
BalanceConfig: &worker_pb.BalanceTaskConfig{
ImbalanceThreshold: 0.1,
MinServerCount: 2,
},
},
}, nil
return defaultBalanceTaskPolicy(), nil
}
confDir := filepath.Join(cp.dataDir, ConfigSubdir)
@@ -490,18 +457,7 @@ func (cp *ConfigPersistence) LoadBalanceTaskPolicy() (*worker_pb.TaskPolicy, err
// Check if file exists
if _, err := os.Stat(configPath); os.IsNotExist(err) {
// Return default policy if file doesn't exist
return &worker_pb.TaskPolicy{
Enabled: true,
MaxConcurrent: 1,
RepeatIntervalSeconds: 6 * 3600, // 6 hours in seconds
CheckIntervalSeconds: 12 * 3600, // 12 hours in seconds
TaskConfig: &worker_pb.TaskPolicy_BalanceConfig{
BalanceConfig: &worker_pb.BalanceTaskConfig{
ImbalanceThreshold: 0.1,
MinServerCount: 2,
},
},
}, nil
return defaultBalanceTaskPolicy(), nil
}
// Read file
@@ -523,6 +479,60 @@ func (cp *ConfigPersistence) LoadBalanceTaskPolicy() (*worker_pb.TaskPolicy, err
return nil, fmt.Errorf("failed to unmarshal balance task configuration")
}
// SaveEcBalanceTaskPolicy saves complete EC balance task policy to protobuf file
func (cp *ConfigPersistence) SaveEcBalanceTaskPolicy(policy *worker_pb.TaskPolicy) error {
return cp.saveTaskConfig(EcBalanceTaskConfigFile, policy)
}
// LoadEcBalanceTaskConfig loads EC balance task configuration from protobuf file
func (cp *ConfigPersistence) LoadEcBalanceTaskConfig() (*EcBalanceTaskConfig, error) {
if taskPolicy, err := cp.LoadEcBalanceTaskPolicy(); err == nil && taskPolicy != nil {
if ecBalanceConfig := taskPolicy.GetEcBalanceConfig(); ecBalanceConfig != nil {
return ecBalanceConfig, nil
}
}
// Return default config if no valid config found
return defaultEcBalanceTaskPolicy().GetEcBalanceConfig(), nil
}
// LoadEcBalanceTaskPolicy loads complete EC balance task policy from protobuf file.
// ec_balance is registered like the other maintenance tasks and ec_balance.LoadConfigFromPersistence
// asserts on this accessor, so without it the task could never be configured at all.
func (cp *ConfigPersistence) LoadEcBalanceTaskPolicy() (*worker_pb.TaskPolicy, error) {
if cp.dataDir == "" {
// Return default policy if no data directory
return defaultEcBalanceTaskPolicy(), nil
}
confDir := filepath.Join(cp.dataDir, ConfigSubdir)
configPath := filepath.Join(confDir, EcBalanceTaskConfigFile)
// Check if file exists
if _, err := os.Stat(configPath); os.IsNotExist(err) {
// Return default policy if file doesn't exist
return defaultEcBalanceTaskPolicy(), nil
}
// Read file
configData, err := os.ReadFile(configPath)
if err != nil {
return nil, fmt.Errorf("failed to read EC balance task config file: %w", err)
}
// Try to unmarshal as TaskPolicy
var policy worker_pb.TaskPolicy
if err := proto.Unmarshal(configData, &policy); err == nil {
// Validate that it's actually a TaskPolicy with EC balance config
if policy.GetEcBalanceConfig() != nil {
glog.V(1).Infof("Loaded EC balance task policy from %s", configPath)
return &policy, nil
}
}
return nil, fmt.Errorf("failed to unmarshal EC balance task configuration")
}
// SaveReplicationTaskConfig saves replication task configuration to protobuf file
func (cp *ConfigPersistence) SaveReplicationTaskConfig(config *ReplicationTaskConfig) error {
return cp.saveTaskConfig(ReplicationTaskConfigFile, config)
@@ -632,6 +642,8 @@ func (cp *ConfigPersistence) SaveTaskPolicy(taskType string, policy *worker_pb.T
return cp.SaveErasureCodingTaskPolicy(policy)
case "balance":
return cp.SaveBalanceTaskPolicy(policy)
case "ec_balance":
return cp.SaveEcBalanceTaskPolicy(policy)
case "replication":
return cp.SaveReplicationTaskPolicy(policy)
}
@@ -687,67 +699,13 @@ func (cp *ConfigPersistence) GetConfigInfo() map[string]interface{} {
return info
}
// buildPolicyFromTaskConfigs loads task configurations from separate files and builds a MaintenancePolicy
func buildPolicyFromTaskConfigs() *worker_pb.MaintenancePolicy {
policy := &worker_pb.MaintenancePolicy{
GlobalMaxConcurrent: 4,
DefaultRepeatIntervalSeconds: 6 * 3600, // 6 hours in seconds
DefaultCheckIntervalSeconds: 12 * 3600, // 12 hours in seconds
TaskPolicies: make(map[string]*worker_pb.TaskPolicy),
}
// Load vacuum task configuration
if vacuumConfig := vacuum.LoadConfigFromPersistence(nil); vacuumConfig != nil {
policy.TaskPolicies["vacuum"] = &worker_pb.TaskPolicy{
Enabled: vacuumConfig.Enabled,
MaxConcurrent: int32(vacuumConfig.MaxConcurrent),
RepeatIntervalSeconds: int32(vacuumConfig.ScanIntervalSeconds),
CheckIntervalSeconds: int32(vacuumConfig.ScanIntervalSeconds),
TaskConfig: &worker_pb.TaskPolicy_VacuumConfig{
VacuumConfig: &worker_pb.VacuumTaskConfig{
GarbageThreshold: float64(vacuumConfig.GarbageThreshold),
MinVolumeAgeHours: int32((vacuumConfig.MinVolumeAgeSeconds + 3599) / 3600), // round up so sub-hour values don't become 0
},
},
}
}
// Load erasure coding task configuration
if ecConfig := erasure_coding.LoadConfigFromPersistence(nil); ecConfig != nil {
policy.TaskPolicies["erasure_coding"] = &worker_pb.TaskPolicy{
Enabled: ecConfig.Enabled,
MaxConcurrent: int32(ecConfig.MaxConcurrent),
RepeatIntervalSeconds: int32(ecConfig.ScanIntervalSeconds),
CheckIntervalSeconds: int32(ecConfig.ScanIntervalSeconds),
TaskConfig: &worker_pb.TaskPolicy_ErasureCodingConfig{
ErasureCodingConfig: &worker_pb.ErasureCodingTaskConfig{
FullnessRatio: float64(ecConfig.FullnessRatio),
QuietForSeconds: int32(ecConfig.QuietForSeconds),
MinVolumeSizeMb: int32(ecConfig.MinSizeMB),
CollectionFilter: ecConfig.CollectionFilter,
},
},
}
}
// Load balance task configuration
if balanceConfig := balance.LoadConfigFromPersistence(nil); balanceConfig != nil {
policy.TaskPolicies["balance"] = &worker_pb.TaskPolicy{
Enabled: balanceConfig.Enabled,
MaxConcurrent: int32(balanceConfig.MaxConcurrent),
RepeatIntervalSeconds: int32(balanceConfig.ScanIntervalSeconds),
CheckIntervalSeconds: int32(balanceConfig.ScanIntervalSeconds),
TaskConfig: &worker_pb.TaskPolicy_BalanceConfig{
BalanceConfig: &worker_pb.BalanceTaskConfig{
ImbalanceThreshold: float64(balanceConfig.ImbalanceThreshold),
MinServerCount: int32(balanceConfig.MinServerCount),
},
},
}
}
glog.V(1).Infof("Built maintenance policy from separate task configs - %d task policies loaded", len(policy.TaskPolicies))
return policy
// buildPolicyFromTaskConfigs builds the maintenance policy from the persisted task configs.
//
// The body lives in weed/admin/maintenance because the maintenance manager needs the same
// policy when it has to build one itself, and this package already imports that one. Keeping
// a second copy here is what let the two drift apart in the first place.
func (cp *ConfigPersistence) buildPolicyFromTaskConfigs() *worker_pb.MaintenancePolicy {
return maintenance.BuildPolicyFromTaskConfigs(cp)
}
// SaveTaskDetail saves detailed task information to disk
+22 -1
View File
@@ -12,6 +12,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/base"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/vacuum"
"google.golang.org/protobuf/proto"
"google.golang.org/protobuf/types/known/timestamppb"
)
@@ -30,6 +31,20 @@ type TomlConfig interface {
// directory loss and override admin UI edits on restart. Absent keys keep
// their persisted values.
func (cp *ConfigPersistence) ApplyMaintenanceConfigFromToml(v TomlConfig) error {
var maintenanceConf *MaintenanceConfig
maintenanceChanged := false
if k := "maintenance.enabled"; v.IsSet(k) {
conf, err := cp.LoadMaintenanceConfig()
if err != nil {
return fmt.Errorf("load maintenance config: %w", err)
}
conf.Enabled = proto.Bool(v.GetBool(k))
// the policy lives in the per-task config files; don't snapshot it here
conf.Policy = nil
maintenanceConf = conf
maintenanceChanged = true
}
vacuumConf := vacuum.LoadConfigFromPersistence(cp)
vacuumChanged := applyBaseConfigFromToml(v, "maintenance.vacuum.", &vacuumConf.BaseConfig)
if k := "maintenance.vacuum.garbage_threshold"; v.IsSet(k) {
@@ -84,13 +99,19 @@ func (cp *ConfigPersistence) ApplyMaintenanceConfigFromToml(v TomlConfig) error
ecChanged = true
}
if !vacuumChanged && !balanceChanged && !ecChanged {
if !maintenanceChanged && !vacuumChanged && !balanceChanged && !ecChanged {
return nil
}
if !cp.IsConfigured() {
return fmt.Errorf("admin.toml maintenance settings require -dataDir to persist")
}
if maintenanceChanged {
if err := cp.SaveMaintenanceConfig(maintenanceConf); err != nil {
return fmt.Errorf("save maintenance config: %w", err)
}
glog.V(0).Infof("Applied [maintenance] settings from admin.toml (enabled: %v)", maintenanceConf.GetEnabled())
}
if vacuumChanged {
if err := cp.SaveVacuumTaskPolicy(vacuumConf.ToTaskPolicy()); err != nil {
return fmt.Errorf("save vacuum task config: %w", err)
+30
View File
@@ -90,6 +90,36 @@ preferred_tags = "Fast, ssd"
}
}
func TestApplyMaintenanceConfigFromTomlEnabledToggle(t *testing.T) {
dir := t.TempDir()
cp := NewConfigPersistence(dir)
if err := cp.ApplyMaintenanceConfigFromToml(tomlConfig(t, "[maintenance]\nenabled = false\n")); err != nil {
t.Fatalf("apply: %v", err)
}
conf, err := cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if conf.GetEnabled() {
t.Errorf("maintenance still enabled after [maintenance] enabled = false")
}
if conf.ScanIntervalSeconds != 30*60 {
t.Errorf("scan interval = %d, want default 1800 kept alongside the toggle", conf.ScanIntervalSeconds)
}
if err := cp.ApplyMaintenanceConfigFromToml(tomlConfig(t, "[maintenance]\nenabled = true\n")); err != nil {
t.Fatalf("apply: %v", err)
}
conf, err = cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if !conf.GetEnabled() {
t.Errorf("maintenance still disabled after [maintenance] enabled = true")
}
}
func TestApplyMaintenanceConfigFromTomlNoKeys(t *testing.T) {
dir := t.TempDir()
cp := NewConfigPersistence(dir)
+4 -1
View File
@@ -73,9 +73,12 @@ func (s *AdminServer) recordDashboardSample() {
}
}
sample.tasks = float64(active)
sample.workers = float64(stats.ActiveWorkers)
}
}
// Counted across both worker registries, same as the Prometheus gauge, so
// the card isn't stuck at 0 on clusters that only run plugin workers.
workers, _, _ := s.workerFleetTotals()
sample.workers = float64(workers)
s.dashSamplesMu.Lock()
s.dashSamples = append(s.dashSamples, sample)
+1 -1
View File
@@ -561,7 +561,7 @@ func (s *AdminServer) GetEcVolumeDetails(volumeID uint32, sortBy string, sortOrd
// Get detailed EC shard information for the specific volume via gRPC
err := s.WithMasterClient(func(client master_pb.SeaweedClient) error {
resp, err := client.VolumeList(context.Background(), &master_pb.VolumeListRequest{VolumeId: volumeID})
resp, err := client.VolumeList(context.Background(), &master_pb.VolumeListRequest{VolumeIds: []uint32{volumeID}})
if err != nil {
return err
}
+66 -3
View File
@@ -2,6 +2,7 @@ package dash
import (
"context"
"net/url"
"path"
"strings"
"time"
@@ -46,6 +47,7 @@ type FileBrowserData struct {
BucketName string `json:"bucket_name"`
IsTableBucketPath bool `json:"is_table_bucket_path"`
TableBucketName string `json:"table_bucket_name"`
S3Endpoint string `json:"s3_endpoint,omitempty"`
// Pagination fields
PageSize int `json:"page_size"`
HasNextPage bool `json:"has_next_page"`
@@ -53,8 +55,11 @@ type FileBrowserData struct {
CurrentLastFileName string `json:"current_last_file_name"` // Cursor from current request (for page size changes)
}
// GetFileBrowser retrieves file browser data for a given path with cursor-based pagination
func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize int) (*FileBrowserData, error) {
// GetFileBrowser retrieves file browser data for a given path with cursor-based
// pagination. A non-empty prefix limits the listing to entries whose name starts
// with it, so callers that only want part of a large directory don't have to page
// through all of it.
func (s *AdminServer) GetFileBrowser(dir string, prefix string, lastFileName string, pageSize int) (*FileBrowserData, error) {
if dir == "" {
dir = "/"
}
@@ -76,7 +81,7 @@ func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize i
// Fetch entries starting from the cursor (lastFileName)
stream, err := client.ListEntries(context.Background(), &filer_pb.ListEntriesRequest{
Directory: dir,
Prefix: "",
Prefix: prefix,
Limit: uint32(fetchLimit),
StartFromFileName: lastFileName,
InclusiveStartFrom: false, // Don't include the cursor file itself
@@ -185,6 +190,7 @@ func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize i
bucketName := ""
isTableBucketPath := false
tableBucketName := ""
isRegularBucket := false
if strings.HasPrefix(dir, "/buckets/") {
isBucketPath = true
pathParts := strings.Split(strings.Trim(dir, "/"), "/")
@@ -202,6 +208,8 @@ func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize i
if s3tables.IsTableBucketEntry(resp.Entry) {
isTableBucketPath = true
tableBucketName = bucketName
} else {
isRegularBucket = true
}
return nil
}); err != nil {
@@ -210,6 +218,11 @@ func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize i
}
}
s3Endpoint := ""
if isRegularBucket {
s3Endpoint = s.GetS3Endpoint()
}
return &FileBrowserData{
CurrentPath: dir,
ParentPath: parentPath,
@@ -221,6 +234,7 @@ func (s *AdminServer) GetFileBrowser(dir string, lastFileName string, pageSize i
BucketName: bucketName,
IsTableBucketPath: isTableBucketPath,
TableBucketName: tableBucketName,
S3Endpoint: s3Endpoint,
// Pagination metadata
PageSize: pageSize,
HasNextPage: hasNextPage,
@@ -269,3 +283,52 @@ func (s *AdminServer) generateBreadcrumbs(dir string) []BreadcrumbItem {
return breadcrumbs
}
// S3ObjectURL builds the path-style S3 URL for a filer path under /buckets/,
// percent-encoding each path segment. Returns "" for other paths.
func S3ObjectURL(endpoint, fullPath string) string {
rel, ok := strings.CutPrefix(fullPath, "/buckets/")
if !ok || rel == "" || endpoint == "" {
return ""
}
segments := strings.Split(rel, "/")
for i, segment := range segments {
segments[i] = url.PathEscape(segment)
}
return strings.TrimRight(endpoint, "/") + "/" + strings.Join(segments, "/")
}
// GetS3ObjectURL returns the S3 URL an object under a regular bucket is served
// at, or "" when the path is not a bucket object, the bucket is an S3 Tables
// bucket, or no S3 endpoint is known.
func (s *AdminServer) GetS3ObjectURL(fullPath string) string {
rel, ok := strings.CutPrefix(fullPath, "/buckets/")
if !ok {
return ""
}
bucketName, key, found := strings.Cut(rel, "/")
if !found || key == "" {
return ""
}
endpoint := s.GetS3Endpoint()
if endpoint == "" {
return ""
}
if err := s.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
resp, err := filer_pb.LookupEntry(context.Background(), client, &filer_pb.LookupDirectoryEntryRequest{
Directory: "/buckets",
Name: bucketName,
})
if err != nil {
return err
}
if s3tables.IsTableBucketEntry(resp.Entry) {
endpoint = ""
}
return nil
}); err != nil {
glog.V(1).Infof("object url bucket lookup failed for %s: %v", bucketName, err)
return ""
}
return S3ObjectURL(endpoint, fullPath)
}
+92
View File
@@ -93,6 +93,98 @@ func TestGenerateBreadcrumbs(t *testing.T) {
}
}
// TestS3ObjectURL verifies path-style S3 URL construction with per-segment
// percent-encoding
func TestS3ObjectURL(t *testing.T) {
tests := []struct {
name string
endpoint string
fullPath string
expected string
}{
{
name: "simple object",
endpoint: "https://s3.example.com",
fullPath: "/buckets/public-images/homarr/example.png",
expected: "https://s3.example.com/public-images/homarr/example.png",
},
{
name: "space hash and question mark in key",
endpoint: "https://s3.example.com",
fullPath: "/buckets/b/a b#c?d.txt",
expected: "https://s3.example.com/b/a%20b%23c%3Fd.txt",
},
{
name: "non-ascii key",
endpoint: "https://s3.example.com",
fullPath: "/buckets/b/图片.png",
expected: "https://s3.example.com/b/%E5%9B%BE%E7%89%87.png",
},
{
name: "percent in key",
endpoint: "https://s3.example.com",
fullPath: "/buckets/b/100%.txt",
expected: "https://s3.example.com/b/100%25.txt",
},
{
name: "trailing slash on endpoint",
endpoint: "https://s3.example.com/",
fullPath: "/buckets/b/k",
expected: "https://s3.example.com/b/k",
},
{
name: "not a bucket path",
endpoint: "https://s3.example.com",
fullPath: "/topics/t/k",
expected: "",
},
{
name: "empty endpoint",
endpoint: "",
fullPath: "/buckets/b/k",
expected: "",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := S3ObjectURL(tt.endpoint, tt.fullPath); got != tt.expected {
t.Errorf("S3ObjectURL(%q, %q) = %q, expected %q", tt.endpoint, tt.fullPath, got, tt.expected)
}
})
}
}
// TestNormalizeS3PublicEndpoint verifies unusable configured endpoints are
// dropped rather than producing broken links
func TestNormalizeS3PublicEndpoint(t *testing.T) {
tests := []struct {
endpoint string
expected string
}{
{"https://s3.example.com", "https://s3.example.com"},
{"http://10.0.0.1:8333/", "http://10.0.0.1:8333"},
{"http://[::1]:8333", "http://[::1]:8333"},
{"https://proxy.example.com/s3", "https://proxy.example.com/s3"},
{"", ""},
{"/", ""},
{"s3.example.com", ""},
{"ftp://s3.example.com", ""},
{"https://", ""},
{"https://s3.example.com?x=1", ""},
{"https://s3.example.com/?", ""},
{"https://s3.example.com#frag", ""},
{"https://s3.example.com/#", ""},
{"http://user:pass@s3.example.com", ""},
}
for _, tt := range tests {
if got := normalizeS3PublicEndpoint(tt.endpoint); got != tt.expected {
t.Errorf("normalizeS3PublicEndpoint(%q) = %q, expected %q", tt.endpoint, got, tt.expected)
}
}
}
// TestPathHandlingWithForwardSlashes verifies that the production code
// correctly handles paths with forward slashes (not OS-specific backslashes)
func TestPathHandlingWithForwardSlashes(t *testing.T) {
@@ -0,0 +1,73 @@
package dash
import (
"testing"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/vacuum"
)
// TestLoadMaintenanceConfigHonoursPersistedTaskConfigs guards against a regression: buildPolicyFromTaskConfigs used to call
// LoadConfigFromPersistence(nil), which can never satisfy the loaders' type assertion, so every
// task silently fell back to its compiled-in defaults (Enabled: true) and a task disabled in the
// admin UI kept being scheduled.
func TestLoadMaintenanceConfigHonoursPersistedTaskConfigs(t *testing.T) {
dir := t.TempDir()
cp := NewConfigPersistence(dir)
// A maintenance.pb must exist, otherwise LoadMaintenanceConfig returns early with defaults.
if err := cp.SaveMaintenanceConfig(DefaultMaintenanceConfig()); err != nil {
t.Fatalf("save maintenance config: %v", err)
}
// Disable balance and vacuum the way the admin UI does, and change a value that is not a bool
// so a fallback to defaults cannot pass by coincidence.
disabledBalance := balance.NewDefaultConfig()
disabledBalance.Enabled = false
disabledBalance.MinServerCount = 7
if err := cp.SaveBalanceTaskPolicy(disabledBalance.ToTaskPolicy()); err != nil {
t.Fatalf("save balance policy: %v", err)
}
disabledVacuum := vacuum.NewDefaultConfig()
disabledVacuum.Enabled = false
if err := cp.SaveVacuumTaskPolicy(disabledVacuum.ToTaskPolicy()); err != nil {
t.Fatalf("save vacuum policy: %v", err)
}
config, err := cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if config.Policy == nil {
t.Fatal("policy is nil, want it populated from the persisted task configs")
}
balancePolicy := config.Policy.TaskPolicies["balance"]
if balancePolicy == nil {
t.Fatal("no balance task policy in the built maintenance policy")
}
if balancePolicy.Enabled {
t.Error("balance enabled = true, want false from the persisted config")
}
if got := balancePolicy.GetBalanceConfig().GetMinServerCount(); got != 7 {
t.Errorf("balance min server count = %d, want persisted 7", got)
}
vacuumPolicy := config.Policy.TaskPolicies["vacuum"]
if vacuumPolicy == nil {
t.Fatal("no vacuum task policy in the built maintenance policy")
}
if vacuumPolicy.Enabled {
t.Error("vacuum enabled = true, want false from the persisted config")
}
// erasure_coding was never saved, so it keeps the loader's default of enabled.
ecPolicy := config.Policy.TaskPolicies["erasure_coding"]
if ecPolicy == nil {
t.Fatal("no erasure_coding task policy in the built maintenance policy")
}
if !ecPolicy.Enabled {
t.Error("erasure_coding enabled = false, want the default true for a config never saved")
}
}
+187
View File
@@ -0,0 +1,187 @@
package dash
import (
"testing"
"github.com/seaweedfs/seaweedfs/weed/admin/maintenance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/vacuum"
"github.com/seaweedfs/seaweedfs/weed/worker/types"
"google.golang.org/protobuf/proto"
)
// The task definitions these tests configure are process-global, so put them back the way a
// fresh process would have them. Passing no config store makes every task fall back to its
// own NewDefaultConfig, which is exactly the state package init left them in.
func restoreGlobalTaskState(t *testing.T) {
t.Helper()
t.Cleanup(func() {
tasks.GetGlobalConfigUpdateRegistry().UpdateAllConfigs(nil)
})
}
// TestDisabledTaskIsNotScannedAfterStartup walks the admin server's startup sequence over a
// data directory that has a disabled balance task saved in it, and checks the end state that
// actually matters: the balance detector reports disabled, so ScanWithTaskDetectors skips it.
//
// This is the whole reported bug in one test. The reporter disabled balance, and the
// scanner kept detecting balance tasks, cancelling them and re-detecting them. Two separate
// defects had to line up for the disabled flag to survive to here: the policy had to be built
// from the persisted configs rather than from a nil store, and the policy had to reach
// detector.IsEnabled() rather than dying in a failed type assertion.
func TestDisabledTaskIsNotScannedAfterStartup(t *testing.T) {
restoreGlobalTaskState(t)
dir := t.TempDir()
cp := NewConfigPersistence(dir)
// What the admin writes when a user turns balance off, and leaves vacuum on.
disabledBalance := balance.NewDefaultConfig()
disabledBalance.Enabled = false
if err := cp.SaveBalanceTaskPolicy(disabledBalance.ToTaskPolicy()); err != nil {
t.Fatalf("save balance policy: %v", err)
}
enabledVacuum := vacuum.NewDefaultConfig()
enabledVacuum.Enabled = true
if err := cp.SaveVacuumTaskPolicy(enabledVacuum.ToTaskPolicy()); err != nil {
t.Fatalf("save vacuum policy: %v", err)
}
// The admin server's startup sequence, in order:
// loadTaskConfigurationsFromPersistence, then InitMaintenanceManager.
tasks.GetGlobalConfigUpdateRegistry().UpdateAllConfigs(cp)
maintenanceConfig, err := cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
manager := maintenance.NewMaintenanceManager(nil, maintenanceConfig, cp)
if manager == nil {
t.Fatal("NewMaintenanceManager returned nil")
}
registry := tasks.GetGlobalTypesRegistry()
balanceDetector := registry.GetDetector(types.TaskTypeBalance)
if balanceDetector == nil {
t.Fatal("no balance detector registered")
}
if balanceDetector.IsEnabled() {
t.Error("balance detector reports enabled after startup over a data directory where " +
"balance is saved as disabled; the scanner will keep detecting and cancelling balance tasks")
}
vacuumDetector := registry.GetDetector(types.TaskTypeVacuum)
if vacuumDetector == nil {
t.Fatal("no vacuum detector registered")
}
if !vacuumDetector.IsEnabled() {
t.Error("vacuum detector reports disabled although vacuum is saved as enabled; " +
"the fix must not switch off tasks the user left on")
}
// Tasks the user never touched keep their compiled-in default of enabled rather than
// being switched off by a policy entry built from a config that was never saved.
for _, taskType := range []types.TaskType{types.TaskTypeErasureCoding, types.TaskTypeECBalance} {
detector := registry.GetDetector(taskType)
if detector == nil {
t.Fatalf("no %s detector registered", taskType)
}
if !detector.IsEnabled() {
t.Errorf("%s detector reports disabled although its config was never saved", taskType)
}
}
}
// TestLoadMaintenanceConfigKeepsPersistedEnabledFalse covers the top of issue-shaped startup:
// an operator persisted enabled=false, and the load path used to lose it twice over — once to
// ApplyDefaultsToProtobuf treating the bool zero value as unset, and once to an explicit
// force-enable "migration" block. The persisted flag must come back as saved, while fields the
// old file never carried still pick up their schema defaults.
func TestLoadMaintenanceConfigKeepsPersistedEnabledFalse(t *testing.T) {
dir := t.TempDir()
cp := NewConfigPersistence(dir)
// A file that only knows about the enabled flag: everything else zero.
if err := cp.SaveMaintenanceConfig(&MaintenanceConfig{Enabled: proto.Bool(false)}); err != nil {
t.Fatalf("save maintenance config: %v", err)
}
loaded, err := cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if loaded.GetEnabled() {
t.Error("persisted enabled=false came back true; the maintenance system cannot be disabled")
}
if loaded.ScanIntervalSeconds != 30*60 {
t.Errorf("scan interval = %d, want schema default 1800 filled in", loaded.ScanIntervalSeconds)
}
// And with no file at all, the default is enabled.
fresh, err := NewConfigPersistence(t.TempDir()).LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if !fresh.GetEnabled() {
t.Error("maintenance not enabled by default when nothing is persisted")
}
}
// TestLoadMaintenanceConfigTreatsLegacyAbsentEnabledAsOn: a maintenance.pb written before
// enabled tracked presence carries no enabled field on the wire whether the old default left
// it false or an operator unchecked it — the two are indistinguishable. Such files must keep
// the enabled default rather than silently switching maintenance off on upgrade; only a file
// that explicitly persists the toggle may disable it.
func TestLoadMaintenanceConfigTreatsLegacyAbsentEnabledAsOn(t *testing.T) {
dir := t.TempDir()
cp := NewConfigPersistence(dir)
// What an old writer produced for enabled=false plus a tuned scan interval: the bool is
// simply absent from the wire.
if err := cp.SaveMaintenanceConfig(&MaintenanceConfig{ScanIntervalSeconds: 15 * 60}); err != nil {
t.Fatalf("save maintenance config: %v", err)
}
loaded, err := cp.LoadMaintenanceConfig()
if err != nil {
t.Fatalf("load maintenance config: %v", err)
}
if !loaded.GetEnabled() {
t.Error("legacy config without an enabled field loads as disabled; " +
"upgrading would silently switch the maintenance system off")
}
if loaded.ScanIntervalSeconds != 15*60 {
t.Errorf("scan interval = %d, want persisted 900 kept", loaded.ScanIntervalSeconds)
}
}
// TestPolicyMirrorsWhatTheDetectorsReport checks that the maintenance policy the queue and
// the scanner run on agrees with the detectors. A disagreement means one of the two paths
// into the task configs has gone stale again.
func TestPolicyMirrorsWhatTheDetectorsReport(t *testing.T) {
restoreGlobalTaskState(t)
dir := t.TempDir()
cp := NewConfigPersistence(dir)
disabledBalance := balance.NewDefaultConfig()
disabledBalance.Enabled = false
if err := cp.SaveBalanceTaskPolicy(disabledBalance.ToTaskPolicy()); err != nil {
t.Fatalf("save balance policy: %v", err)
}
tasks.GetGlobalConfigUpdateRegistry().UpdateAllConfigs(cp)
policy := cp.buildPolicyFromTaskConfigs()
for taskType, detector := range tasks.GetGlobalTypesRegistry().GetAllDetectors() {
policyEnabled := maintenance.IsTaskEnabled(policy, maintenance.MaintenanceTaskType(taskType))
if policyEnabled != detector.IsEnabled() {
t.Errorf("%s: policy says enabled=%v but the detector says enabled=%v",
taskType, policyEnabled, detector.IsEnabled())
}
}
}
+57
View File
@@ -0,0 +1,57 @@
package dash
import (
"context"
weediam "github.com/seaweedfs/seaweedfs/weed/iam"
"github.com/seaweedfs/seaweedfs/weed/iam/integration"
"github.com/seaweedfs/seaweedfs/weed/glog"
)
// principalRoleArn builds the ARN SeaweedFS assigns a role by default (when
// its RoleDefinition.RoleArn isn't explicitly set) - see
// weed/iam/integration/iam_manager.go's CreateRole. ListRoles only returns
// role names, so this reconstructs the well-known default rather than
// fetching every role's stored definition just to populate a suggestion list.
func principalRoleArn(roleName string) string {
return "arn:aws:iam::role/" + roleName
}
// GetPrincipalSuggestions returns candidate ARNs for the policy editor's
// Principal/NotPrincipal autocomplete: one per S3 user, plus one per IAM
// role. Service accounts are deliberately not listed separately - a service
// account is just an additional credential for its parent user, so its ARN
// is identical to the one already suggested for that user.
//
// Role listing is best-effort: if the filer or role store is unavailable,
// the error is logged and suggestions fall back to users only, since an
// incomplete autocomplete list is far less disruptive than blocking policy
// editing over a suggestions-only feature.
func (s *AdminServer) GetPrincipalSuggestions(ctx context.Context) ([]string, error) {
var suggestions []string
users, err := s.GetObjectStoreUsers(ctx)
if err != nil {
return nil, err
}
for _, u := range users {
suggestions = append(suggestions, weediam.UserArn(u.Username))
}
roleStore, err := integration.NewFilerRoleStore(nil, func() string { return s.GetFilerAddress() })
if err != nil {
glog.Warningf("GetPrincipalSuggestions: failed to create role store: %v", err)
return suggestions, nil
}
roleNames, err := roleStore.ListRoles(ctx, s.GetFilerAddress())
if err != nil {
glog.Warningf("GetPrincipalSuggestions: failed to list roles: %v", err)
return suggestions, nil
}
for _, roleName := range roleNames {
suggestions = append(suggestions, principalRoleArn(roleName))
}
return suggestions, nil
}
@@ -0,0 +1,11 @@
package dash
import "testing"
func TestPrincipalRoleArn(t *testing.T) {
got := principalRoleArn("S3ReadOnlyRole")
want := "arn:aws:iam::role/S3ReadOnlyRole"
if got != want {
t.Fatalf("principalRoleArn() = %q, want %q", got, want)
}
}
+43 -9
View File
@@ -37,6 +37,12 @@ type S3TablesBucketSummary struct {
// Format is empty for a bucket created before formats were declared. Such a
// bucket takes tables of either format, which is what it always did.
Format string `json:"format,omitempty"`
// PolicyStatementCount is the number of statements in the table bucket's
// resource policy, or 0 if it has none. Unrelated to the S3 bucket
// policy mechanism (policy_engine.PolicyDocument / s3-bucket-policy):
// S3 Tables stores its own s3tables.PolicyDocument under the
// s3tables.policy extended attribute.
PolicyStatementCount int `json:"policy_statement_count"`
}
type S3TablesNamespacesData struct {
@@ -144,11 +150,12 @@ func (s *AdminServer) GetS3TablesBucketsData(ctx context.Context) (S3TablesBucke
continue
}
buckets = append(buckets, S3TablesBucketSummary{
ARN: arn,
Name: entry.Entry.Name,
OwnerAccountID: metadata.OwnerAccountID,
CreatedAt: metadata.CreatedAt,
Format: metadata.Format,
ARN: arn,
Name: entry.Entry.Name,
OwnerAccountID: metadata.OwnerAccountID,
CreatedAt: metadata.CreatedAt,
Format: metadata.Format,
PolicyStatementCount: extractS3TablesPolicyStatementCountFromEntry(entry.Entry),
})
}
return nil
@@ -165,6 +172,23 @@ func (s *AdminServer) GetS3TablesBucketsData(ctx context.Context) (S3TablesBucke
}, nil
}
// extractS3TablesPolicyStatementCountFromEntry returns the number of
// statements in the table bucket's resource policy, or 0 if it has none or
// the stored JSON can't be parsed. Forgiving on parse failure, matching
// extractPolicyStatementCountFromEntry (the S3 bucket policy equivalent in
// admin_server.go, which is a different, unrelated policy mechanism).
func extractS3TablesPolicyStatementCountFromEntry(entry *filer_pb.Entry) int {
policyJSON := entry.Extended[s3tables.ExtendedKeyPolicy]
if len(policyJSON) == 0 {
return 0
}
var doc s3tables.PolicyDocument
if err := json.Unmarshal(policyJSON, &doc); err != nil {
return 0
}
return len(doc.Statement)
}
// observedRowCounts collects what workers last reported for these tables. For a
// format admin cannot read, this is the only row count that exists.
func (s *AdminServer) observedRowCounts(bucketArn string, namespaceParts []string, tables []s3tables.TableSummary) map[string]string {
@@ -1019,8 +1043,11 @@ func (s *AdminServer) GetS3TablesBucketPolicy(w http.ResponseWriter, r *http.Req
getReq := &s3tables.GetTableBucketPolicyRequest{TableBucketARN: bucketArn}
var resp s3tables.GetTableBucketPolicyResponse
if err := s.executeS3TablesOperation(r.Context(), "GetTableBucketPolicy", getReq, &resp); err != nil {
writeS3TablesError(w, err)
return
// No policy is a normal state for the UI (empty editor), not an error.
if !isS3TablesNoSuchPolicy(err) {
writeS3TablesError(w, err)
return
}
}
writeJSON(w, http.StatusOK, map[string]interface{}{"policy": resp.ResourcePolicy})
}
@@ -1087,8 +1114,10 @@ func (s *AdminServer) GetS3TablesTablePolicy(w http.ResponseWriter, r *http.Requ
getReq := &s3tables.GetTablePolicyRequest{TableBucketARN: bucketArn, Namespace: namespaceParts, Name: name}
var resp s3tables.GetTablePolicyResponse
if err := s.executeS3TablesOperation(r.Context(), "GetTablePolicy", getReq, &resp); err != nil {
writeS3TablesError(w, err)
return
if !isS3TablesNoSuchPolicy(err) {
writeS3TablesError(w, err)
return
}
}
writeJSON(w, http.StatusOK, map[string]interface{}{"policy": resp.ResourcePolicy})
}
@@ -1201,6 +1230,11 @@ func writeS3TablesError(w http.ResponseWriter, err error) {
writeJSONError(w, s3TablesErrorStatus(err), parseS3TablesErrorMessage(err))
}
func isS3TablesNoSuchPolicy(err error) bool {
var s3Err *s3tables.S3TablesError
return errors.As(err, &s3Err) && s3Err.Type == s3tables.ErrCodeNoSuchPolicy
}
func s3TablesErrorStatus(err error) int {
var s3Err *s3tables.S3TablesError
if errors.As(err, &s3Err) {
@@ -0,0 +1,184 @@
package dash
import (
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb/worker_pb"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/ec_balance"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/worker/tasks/vacuum"
"google.golang.org/protobuf/proto"
)
// TestLoadTaskPolicyDefaultsMatchTaskDefaults pins the persistence layer's "nothing saved
// yet" defaults to each task's own NewDefaultConfig(). They used to be a second, hand-written
// copy and had drifted: with a data directory but no config file on disk, vacuum ran on a 24h
// scan interval instead of 2h, balance on 6h with a 0.1 imbalance threshold instead of 30m
// with 0.2, and erasure coding on 168h with a 0.90 fullness ratio and a 1024MB minimum volume
// size instead of 1h with 0.95 and 30MB - none of which is what the admin UI shows as the
// default for those fields.
func TestLoadTaskPolicyDefaultsMatchTaskDefaults(t *testing.T) {
cases := []struct {
name string
want *worker_pb.TaskPolicy
load func(cp *ConfigPersistence) (*worker_pb.TaskPolicy, error)
}{
{
name: "vacuum",
want: vacuum.NewDefaultConfig().ToTaskPolicy(),
load: func(cp *ConfigPersistence) (*worker_pb.TaskPolicy, error) { return cp.LoadVacuumTaskPolicy() },
},
{
name: "erasure_coding",
want: erasure_coding.NewDefaultConfig().ToTaskPolicy(),
load: func(cp *ConfigPersistence) (*worker_pb.TaskPolicy, error) {
return cp.LoadErasureCodingTaskPolicy()
},
},
{
name: "balance",
want: balance.NewDefaultConfig().ToTaskPolicy(),
load: func(cp *ConfigPersistence) (*worker_pb.TaskPolicy, error) { return cp.LoadBalanceTaskPolicy() },
},
{
name: "ec_balance",
want: ec_balance.NewDefaultConfig().ToTaskPolicy(),
load: func(cp *ConfigPersistence) (*worker_pb.TaskPolicy, error) { return cp.LoadEcBalanceTaskPolicy() },
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
// Both no-file branches have to agree with the task's own defaults: no data
// directory at all, and a data directory that has never been written to.
for _, cp := range []*ConfigPersistence{NewConfigPersistence(""), NewConfigPersistence(t.TempDir())} {
got, err := tc.load(cp)
if err != nil {
t.Fatalf("load %s policy: %v", tc.name, err)
}
if !proto.Equal(got, tc.want) {
t.Errorf("%s default policy (dataDir=%q) =\n %v\nwant NewDefaultConfig().ToTaskPolicy() =\n %v",
tc.name, cp.GetDataDir(), got, tc.want)
}
}
})
}
}
// TestLoadTaskConfigDefaultsMatchTaskDefaults covers the narrower Load*TaskConfig accessors,
// which carried a third copy of the same defaults.
func TestLoadTaskConfigDefaultsMatchTaskDefaults(t *testing.T) {
cp := NewConfigPersistence(t.TempDir())
vacuumConfig, err := cp.LoadVacuumTaskConfig()
if err != nil {
t.Fatalf("load vacuum config: %v", err)
}
if want := vacuum.NewDefaultConfig().ToTaskPolicy().GetVacuumConfig(); !proto.Equal(vacuumConfig, want) {
t.Errorf("vacuum default config = %v, want %v", vacuumConfig, want)
}
ecConfig, err := cp.LoadErasureCodingTaskConfig()
if err != nil {
t.Fatalf("load erasure coding config: %v", err)
}
if want := erasure_coding.NewDefaultConfig().ToTaskPolicy().GetErasureCodingConfig(); !proto.Equal(ecConfig, want) {
t.Errorf("erasure coding default config = %v, want %v", ecConfig, want)
}
balanceConfig, err := cp.LoadBalanceTaskConfig()
if err != nil {
t.Fatalf("load balance config: %v", err)
}
if want := balance.NewDefaultConfig().ToTaskPolicy().GetBalanceConfig(); !proto.Equal(balanceConfig, want) {
t.Errorf("balance default config = %v, want %v", balanceConfig, want)
}
}
// TestEcBalanceTaskPolicyRoundTrip checks the accessor ec_balance.LoadConfigFromPersistence
// asserts on. Before it existed, ec_balance was the one registered maintenance task whose
// configuration could not be persisted at all.
func TestEcBalanceTaskPolicyRoundTrip(t *testing.T) {
cp := NewConfigPersistence(t.TempDir())
saved := ec_balance.NewDefaultConfig()
saved.Enabled = false
saved.MinServerCount = 9
saved.ImbalanceThreshold = 0.42
saved.CollectionFilter = "pictures"
if err := cp.SaveEcBalanceTaskPolicy(saved.ToTaskPolicy()); err != nil {
t.Fatalf("save ec_balance policy: %v", err)
}
loaded := ec_balance.LoadConfigFromPersistence(cp)
if loaded == nil {
t.Fatal("ec_balance.LoadConfigFromPersistence returned nil")
}
if loaded.Enabled {
t.Error("ec_balance enabled = true, want the persisted false")
}
if loaded.MinServerCount != 9 {
t.Errorf("ec_balance min server count = %d, want the persisted 9", loaded.MinServerCount)
}
if loaded.ImbalanceThreshold != 0.42 {
t.Errorf("ec_balance imbalance threshold = %v, want the persisted 0.42", loaded.ImbalanceThreshold)
}
if loaded.CollectionFilter != "pictures" {
t.Errorf("ec_balance collection filter = %q, want the persisted %q", loaded.CollectionFilter, "pictures")
}
// The generic dispatcher the maintenance manager uses has to know the type too, and
// has to write the same file the dedicated loader reads.
dispatched := NewConfigPersistence(t.TempDir())
if err := dispatched.SaveTaskPolicy("ec_balance", saved.ToTaskPolicy()); err != nil {
t.Fatalf("SaveTaskPolicy(ec_balance): %v", err)
}
roundTripped := ec_balance.LoadConfigFromPersistence(dispatched)
if roundTripped == nil || roundTripped.MinServerCount != 9 {
t.Errorf("SaveTaskPolicy(ec_balance) did not round-trip through the ec_balance store: %+v", roundTripped)
}
}
// TestBuildPolicyKeepsTaskSpecificFields guards the fields the hand-written policy builder
// used to drop on the floor: the erasure coding preferred tags and replica placement, and
// the balance IO rate limit. Building each entry from the task's own ToTaskPolicy() keeps
// them, so a value set in admin.toml survives into the maintenance policy.
func TestBuildPolicyKeepsTaskSpecificFields(t *testing.T) {
cp := NewConfigPersistence(t.TempDir())
ecConfig := erasure_coding.NewDefaultConfig()
ecConfig.PreferredTags = []string{"ssd", "archive"}
ecConfig.ReplicaPlacement = "020"
if err := cp.SaveErasureCodingTaskPolicy(ecConfig.ToTaskPolicy()); err != nil {
t.Fatalf("save erasure coding policy: %v", err)
}
balanceConfig := balance.NewDefaultConfig()
balanceConfig.IoBytePerSecond = 5 << 20
if err := cp.SaveBalanceTaskPolicy(balanceConfig.ToTaskPolicy()); err != nil {
t.Fatalf("save balance policy: %v", err)
}
policy := cp.buildPolicyFromTaskConfigs()
ecPolicy := policy.TaskPolicies["erasure_coding"].GetErasureCodingConfig()
if ecPolicy == nil {
t.Fatal("no erasure coding config in the built policy")
}
if got := ecPolicy.GetReplicaPlacement(); got != "020" {
t.Errorf("erasure coding replica placement = %q, want the persisted %q", got, "020")
}
if got := ecPolicy.GetPreferredTags(); len(got) != 2 || got[0] != "ssd" || got[1] != "archive" {
t.Errorf("erasure coding preferred tags = %v, want the persisted [ssd archive]", got)
}
balancePolicy := policy.TaskPolicies["balance"].GetBalanceConfig()
if balancePolicy == nil {
t.Fatal("no balance config in the built policy")
}
if got := balancePolicy.GetIoBytePerSecond(); got != 5<<20 {
t.Errorf("balance IO limit = %d, want the persisted %d", got, 5<<20)
}
}
+2 -10
View File
@@ -277,11 +277,7 @@ func (p *TopicRetentionPurger) deleteDirectoryRecursively(client filer_pb.Seawee
}
} else {
// Delete file
_, err = client.DeleteEntry(context.Background(), &filer_pb.DeleteEntryRequest{
Directory: dirPath,
Name: resp.Entry.Name,
})
if err != nil {
if err := filer_pb.DoRemove(context.Background(), client, dirPath, resp.Entry.Name, false, false, false, false, nil); err != nil {
return fmt.Errorf("failed to delete file %s: %v", entryPath, err)
}
}
@@ -291,11 +287,7 @@ func (p *TopicRetentionPurger) deleteDirectoryRecursively(client filer_pb.Seawee
parentDir := path.Dir(dirPath)
dirName := path.Base(dirPath)
_, err = client.DeleteEntry(context.Background(), &filer_pb.DeleteEntryRequest{
Directory: parentDir,
Name: dirName,
})
if err != nil {
if err := filer_pb.DoRemove(context.Background(), client, parentDir, dirName, false, false, false, false, nil); err != nil {
return fmt.Errorf("failed to delete directory %s: %v", dirPath, err)
}
+6
View File
@@ -98,6 +98,12 @@ type S3Bucket struct {
LifecycleRuleCount int `json:"lifecycle_rule_count"`
LifecycleEnabledCount int `json:"lifecycle_enabled_count"`
// PolicyStatementCount is the number of statements in the bucket policy,
// or 0 if the bucket has none. A policy document can't have zero
// statements (see policy_engine.ValidatePolicy), so >0 is a faithful
// "has a policy" flag.
PolicyStatementCount int `json:"policy_statement_count"`
}
type S3Object struct {

Some files were not shown because too many files have changed in this diff Show More