Compare commits

...
Author SHA1 Message Date
Chris Lu a6e07181b2 admin: scrape the metrics ports the cluster advertises
Replaces scraping each node's service port with the dedicated Prometheus
listener each node now advertises. Nodes started without -metricsPort
advertise 0 and are skipped, so nothing is ever fetched from a
client-facing port.

Endpoints are deduplicated by address because a combined "weed server"
advertises one listener for all of its components; series are attributed
by metric name, which already identifies the component.
2026-09-14 23:59:38 -07:00
Chris Lu 51831d6850 admin: derive interval rates and latency quantiles from scrapes
Raw counters and histograms are process-lifetime cumulative, so
charting them directly is meaningless. Counters now become per-second
rates and histograms become p50/p95/p99, both computed from the delta
against the previous scrape, with counter resets skipped. Quantiles use
linear interpolation within the matching bucket, as Prometheus does.
2026-09-14 23:59:38 -07:00
Chris Lu d235dd280b admin: gather the admin's own registry into the metrics store
The admin's maintenance and worker metrics live in the local
stats.Gather registry, so record them directly under the admin/local
source instead of scraping over HTTP. Adds metricsStore.match for
prefix/metric lookups.
2026-09-14 23:59:38 -07:00
Chris Lu 7b9332953d admin: add server-side SVG chart renderer
Extends the existing sparklineSVG approach into a full chart helper
with axes, gridlines, multi-series lines/areas, legends, threshold
lines, and unit formatters (bytes, bps, ms, pct). No JS chart library;
safe to inline in templ pages.
2026-09-14 23:59:38 -07:00
Chris Lu 85147522a9 admin: scrape per-server /metrics into the store
Adds a 15s scrape loop that fetches /metrics from every discovered
master, volume, filer, and S3 server (addresses come from the
existing topology + ListClusterNodes helpers) and records each
series into the in-memory store. Parses Prometheus text exposition
via prometheus/common/expfmt.
2026-09-14 23:59:38 -07:00
Chris Lu 207bc0b75a admin: add in-memory metrics series store
Bounded ring buffer (240 samples) keyed by source/metric[/labels],
reusing the dashSample ring pattern. No persistence; powers the
upcoming monitoring charts.
2026-09-14 23:59:38 -07:00
Chris Lu 2f26d5779b mini: wire the shared metrics port, and tolerate an unset one
weed mini builds MasterOptions directly and never set metricsHttpPort, so
reading it in toMasterOption dereferenced nil and crashed startup. Point
mini's master, volume and filer at its single -metricsPort listener, the
same way weed server does, and treat an unset port as disabled so a
partially initialised MasterOptions cannot panic again.

Reproduced with 'weed mini -dir=... -s3.port=...', which is what the S3
filer-group and delete-regression suites start.
2026-09-14 23:59:32 -07:00
Chris Lu 14bc6e5e4f servers: advertise the configured metricsPort to the master
Filer, S3 and broker report it on the KeepConnected registration, volume
servers in their heartbeat, and masters return their own in
GetMasterConfiguration. MasterClient gains SetMetricsPort so the eight
callers that have no metrics listener are untouched, and the volume
server takes it as a constructor argument because its heartbeat goroutine
starts there.

In combined "weed server" one metrics listener serves the whole shared
registry, so every component advertises the same port.
2026-09-14 23:32:32 -07:00
Chris Lu dfde24f3ee pb: add metrics_port so nodes can advertise their metrics listener
Each server already has a -metricsPort Prometheus listener, but nothing
advertises it, so a central scraper cannot find it. Adds metrics_port to
KeepConnectedRequest and ListClusterNodes (filer, S3, broker), Heartbeat
and DataNodeInfo (volume servers), and GetMasterConfigurationResponse
(masters).

Note the existing metrics_address fields are unrelated: they carry the
Prometheus push gateway that servers push to, not a scrape target.
2026-09-14 23:23:25 -07:00
github-actions[bot] e735c12869 docs: regenerate star history chart 2026-09-15 00:48:17 +00:00
e4ca0d09e7 s3: preserve versions for POST policy uploads (#11316)
* s3: preserve versions for POST policy uploads

Route POST policy uploads through the existing version-aware write helpers
and validate promoted Object Lock headers before writing.

Return the generated version ID when versioning is enabled, return
x-amz-version-id: null when versioning is suspended, and omit the header
when versioning has never been enabled.

* s3: reuse versioning helpers in POST policy handler

Route the POST policy handler through the existing getVersioningState
and isObjectLockEnabled helpers instead of open-coding the object-lock
forces-versioning-enabled rule, matching the PUT path.

Drop the x-amz-version-id: null response header for suspended
versioning; the PUT handler omits it and the S3 PutObject sample
response for suspended buckets does not include it. Trim the moved
fileSize comment.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:36:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
01433e801d build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0 (#11306)
* build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0

Bumps [github.com/redis/go-redis/v9](https://github.com/redis/go-redis) from 9.21.0 to 9.22.0.
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/redis/go-redis/compare/v9.21.0...v9.22.0)

---
updated-dependencies:
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.22.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* test(redis_conf): track go-redis 9.22.0 default read timeout of 5s

go-redis 9.22.0 raised the default ReadTimeout from 3s to 5s (part of the
cross-SDK configuration alignment). Update TestUnsetKeepsGoRedisDefaults to
expect the new default so the bump in #11306 stops failing CI.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:16:35 -07:00
1f61097d4d helm: grant List to the generated read-only S3 identity (#11318)
* helm: grant List to the generated read-only S3 identity

The chart's anvReadOnly identity only carried the Read action, so its
credentials could GetObject and HeadObject but every ListObjects request
was denied: List is a separate action and the identity check is an exact
match. Add List so the read-only credentials can list buckets and objects.
Writes stay denied. Update the README example to match.

Bump the chart to 4.47.1. The chart label is part of the s3 and
all-in-one pod templates, so the upgrade rolls the gateways and they
reload the identity config, which is only read at startup.

Fixes #11317

* helm: roll standalone S3 and all-in-one on s3 config changes

Mirror the filer checksum/s3config pod annotation in the standalone S3
and all-in-one deployments so a changed generated S3 secret triggers a
rollout during a normal helm upgrade without relying on a chart version
bump.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:11:22 -07:00
Chris LuandGitHub 0e82b4e351 s3: populate Initiated timestamp in ListMultipartUploads (#11313)
* s3: populate Initiated timestamp in ListMultipartUploads

ListMultipartUploads returned each upload with only Key and UploadId,
omitting the Initiated timestamp. Clients such as GeeseFS rely on this
field to expire stale uploads and crash on its absence. Set Initiated
from the upload directory entry creation time so repeated listings
preserve the original initiation time.

* test/s3: verify Initiated timestamp in ListMultipartUploads

Add an integration test that initiates a multipart upload, lists it,
and asserts the Initiated field is populated and preserved across
repeated listings rather than reflecting the listing time.
2026-09-14 16:03:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0bd048b76f build(deps): bump google.golang.org/api from 0.296.0 to 0.297.0 (#11307)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.296.0 to 0.297.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.296.0...v0.297.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.297.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 15:09:14 -07:00
Chris LuandGitHub c997e54096 admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback (#11314)
* admin: extract isFlagExplicitlySet helper from applyViperFallback

No behavior change; the inline flag-visit check becomes a reusable
helper so the upcoming bind-address default can share it.

* admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback

PR #11185 made the admin HTTP server default to 127.0.0.1 to stop
exposing the unauthenticated admin API on the network by accident.
That also locked out operators who already authenticated with
-adminPassword: their UI became unreachable from the network after
upgrade unless they added -ip=0.0.0.0 (see #11303).

An authenticated deployment is safe to expose, so auto-upgrade the -ip
default to 0.0.0.0 when -adminPassword or [https.admin] mTLS is
configured. The loopback default stays for the unauthenticated case, so
the unauthenticated API is never exposed on the network. An explicit
-ip is always honored.

The worker gRPC control plane has no password auth (only mTLS), so it
must not follow the HTTP upgrade. Give it a separate bind address that
stays on loopback unless -ip is explicit, so adminPassword no longer
re-exposes the unauthenticated worker stream.

* admin: hint loopback-only bind in startup banner

When the admin server binds to loopback (the default for the
unauthenticated case), print a one-line hint that it is not reachable
from other hosts and how to expose it. This helps operators who, after
the #11185 loopback default, can no longer reach the UI from another
machine quickly see the cause and the fix without reading the docs.

* admin: keep worker gRPC on loopback, decouple from https.admin mTLS

The worker gRPC auto-upgrade to 0.0.0.0 was gated on hasMTLS, which
reads the https.admin (HTTP) mTLS config. The worker gRPC mTLS comes
from grpc.admin + grpc.ca, a separate config, so:

- https.admin mTLS without grpc.admin mTLS widened the worker gRPC to
  0.0.0.0 unauthenticated (re-exposing the control plane), and
- grpc.admin mTLS without https.admin mTLS left the worker gRPC on
  loopback, blocking authenticated remote workers.

Drop the worker gRPC auto-upgrade entirely. The worker gRPC keeps the
raw -ip value (loopback by default), matching the pre-existing
behavior; an operator who wants remote workers sets -ip explicitly.
Only the HTTP admin listener auto-upgrades to 0.0.0.0 when
authenticated.

Addresses review feedback on #11314 from Devin and Greptile.
2026-09-14 14:04:35 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2aa6af033d build(deps): bump github.com/go-sql-driver/mysql from 1.10.0 to 1.10.1 (#11308)
Bumps [github.com/go-sql-driver/mysql](https://github.com/go-sql-driver/mysql) from 1.10.0 to 1.10.1.
- [Release notes](https://github.com/go-sql-driver/mysql/releases)
- [Changelog](https://github.com/go-sql-driver/mysql/blob/master/CHANGELOG.md)
- [Commits](https://github.com/go-sql-driver/mysql/compare/v1.10.0...v1.10.1)

---
updated-dependencies:
- dependency-name: github.com/go-sql-driver/mysql
  dependency-version: 1.10.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 14:04:10 -07:00
Chris LuandGitHub 02749c1192 s3api: configurable trusted-proxy allowlist for aws:SourceIp (#11302) (#11315)
* s3api: add TrustedProxies allowlist helper for aws:SourceIp extraction

Introduces a policy_engine.TrustedProxies type that parses a
comma-separated list of bare IPs and CIDRs (mirroring Guard.UpdateWhiteList)
and extracts the client IP for aws:SourceIp condition evaluation.

When the direct TCP peer is in the allowlist, X-Forwarded-For is walked
right-to-left skipping trusted hops (then X-Real-Ip); otherwise the direct
peer address is returned. This is the building block for restoring
configurable forwarded-header trust removed in b88156f (#11231), as
proposed in #11302.

* s3api: honor trusted-proxy allowlist in bucket/IAM policy engine

Make ExtractConditionValuesFromRequest a method on *PolicyEngine so it
can use the engine TrustedProxies when resolving aws:SourceIp. With no
allowlist configured the behavior is unchanged from b88156f: the direct
TCP peer is used and forwarded headers are ignored. When an allowlist is
configured via SetTrustedProxies, requests from a trusted peer honor
X-Forwarded-For (right-to-left) then X-Real-Ip.

Update the two call sites (auth_credentials.go, s3api_bucket_policy_engine.go)
and the engine tests to the method form, and add a regression test for the
trusted-proxy path.

* s3api: honor trusted-proxy allowlist in IAM role/session policies

Make extractRequestContext and extractSourceIP methods on
*S3IAMIntegration so they can use the integration TrustedProxies when
resolving aws:SourceIp. With no allowlist configured the behavior is
unchanged from b88156f: the direct TCP peer is used and forwarded
headers are ignored. When an allowlist is configured via
SetTrustedProxies, requests from a trusted peer honor X-Forwarded-For
(right-to-left) then X-Real-Ip.

Update the call site in isActionExplicitlyDeniedByIAM to type-assert
the integration and use the method, and add a regression test for the
trusted-proxy path.

* s3api: load [s3.trusted_proxies] from security.toml and wire to engines

Read s3.trusted_proxies.white_list (comma-separated IPs/CIDRs) from
security.toml and propagate the allowlist to the bucket policy engine,
the IAM policy engine (persisted across rebuilds via
IdentityAccessManagement.SetTrustedProxies), and the IAM integration.
Reloaded on SIGHUP alongside the JWT signing keys. Document the new
section in the scaffold security.toml.

Closes #11302.

* s3api: harden TrustedProxies parsing and X-Forwarded-For traversal

Canonicalize bare IP entries (via net.ParseIP + String) so non-canonical
IPv6 allowlist entries such as 2001:0db8::1 match peers rendered as
2001:db8::1, and log+skip unparseable bare entries instead of storing
them inertly.

When walking X-Forwarded-For right-to-left, stop at the first malformed
(non-empty, unparseable) entry instead of skipping it, and only fall
back to the leftmost valid IP when the chain was well-formed. This
prevents a malformed hop from masking a forged IP to its left.

Addresses review feedback on #11315.

* s3api: make TrustedProxies reload race-free via atomic.Pointer

Store the trusted-proxy allowlist behind sync/atomic.Pointer in
PolicyEngine and S3IAMIntegration so SIGHUP reloads (which swap the
allowlist) cannot race with concurrent request handlers reading it.
This mirrors the existing Guard guardState pattern. The
IdentityAccessManagement copy is already protected by iam.m.

Addresses review feedback on #11315.
2026-09-14 13:54:26 -07:00
ac03d3fd78 shell: warn when fs.mergeVolumes source holds only orphan needles (#11310)
* shell: warn when fs.mergeVolumes source holds only orphan needles

fs.mergeVolumes traverses filer entries, so a source volume whose
needles are all orphans — filer entries lost to a crashed write or a
wiped filer store — produces only the plan header and exits 0: no move,
no skip, no error. Operators read that as a successful merge while the
real cleanup (volume.fsck) never runs, and dat>idx volumes keep coming
back read-only after restarts.

Count the source-volume needles seen during traversal and, when a plan
source was never seen but its index still reports needles, print a
warning pointing at volume.fsck. Dry-run warns too.

* shell: make needle counting concurrency-safe and count manifest sub-chunks

TraverseBfs runs its callbacks from five workers, so the plain
needlesSeen map raced between source-heavy merges (fatal concurrent
map writes). All increments now funnel through a mutex-guarded
recordSeen closure.

Manifest sub-chunks that live on planned source volumes are now
recorded too — rewriteManifestChunk visits them (including dry-run
and capacity-skipped ones) but previously never marked their source,
which produced false 'orphan needles' warnings for sources whose
chunks were all reached through manifests.

* shell: extract sourceNeedleCounter so the concurrency test covers the production path

The orphan-warning recording was a closure local to Do, so
TestWarnUnreferencedSources_ConcurrentRecording could only exercise a
test-local copy of it — a regression in the production mutex would pass
the test. Lift the map and mutex into a sourceNeedleCounter type with
record/count methods and use it from Do and the test, so the -race test
now drives the actual recording path. Trim the verbose comments added
with the warning while here.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:48 -07:00
adaf3534fa rust: clippy-clean both crates and adopt the std APIs the 1.91 MSRV allows (#11312)
* rust: apply clippy --fix to both crates

The mechanical part of a clippy sweep: `cargo clippy --all-targets --fix`
on seaweed-volume and the seaweed-worker workspace, hand-reviewed. Both
manifests declare their MSRV (1.91.1 and 1.94.1), so every suggestion
clippy applied is within it: the collapsible_if sites become let chains
(1.88, edition 2024), `% n == 0` becomes is_multiple_of (1.87),
chunks_exact with a constant becomes as_chunks (1.88), repeat().take()
becomes repeat_n (1.82), and io::Error::new(Other, ..) becomes
io::Error::other (1.74). The rest is redundant clones, borrows, casts,
closures and field names.

Nothing here changes behaviour. The three let_and_return sites in
needle_map.rs and store_ec.rs deserve a note: the `let result = ..;
result` shape was a deliberate edition-2021 workaround to drop a redb
guard before the table it borrows. Edition 2024 drops tail-expression
temporaries before locals, which is why clippy now flags it, and the
two comments that described the workaround say so instead.

Manual edits on top of the tool output: the blocks clippy rewrote are
re-indented the way rustfmt lays them out (only those blocks — the
crate is not rustfmt-clean and a whole-crate fmt would bury this diff),
the blank lines let_and_return left behind are removed, and the CRC
legacy_value test compares against a literal worked out from the
original shift formula rather than restating rotate_right.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: clear the clippy warnings --fix cannot apply, and say why the rest stay

Hand fixes for the lints clippy only reports. Behaviour is unchanged
throughout; each rewrite is the one clippy names.

- needless_range_loop (7): index loops over shard vectors become
  iterator loops. Where the old code indexed `v[..n]` the new loop
  iterates `v[..n]` so an undersized vector still panics the same way.
- field_reassign_with_default (6): struct literals with `..Default`.
- redundant_pattern_matching (3): `if let Err(_) = guard.check()` becomes
  `.is_err()`, which also releases the read guard at the end of the
  condition instead of at the end of the block.
- manual_strip (2), manual_checked_ops, format_in_format_args,
  redundant_locals, wrong_self_convention (to_vif takes self by value,
  so it is into_vif; CompactEntry is Copy, so to_needle_value takes self).
- type_complexity (2): `OrphanShardLoad` and `RawNeedleEntry` name two
  tuples that were spelled out inline.
- new_without_default: CompactNeedleMap gets a Default that calls new().
- suspicious_open_options: a test helper spells out `.truncate(false)`,
  which is what `.create(true).write(true)` already did.

What stays, and the attribute that says so:

- too_many_arguments (10): `#[expect]` on each function. Folding 8–15
  parameters into a struct is a design change, not a lint fix.
- await_holding_lock / readonly_write_lock: one test holds the store
  write guard across a sleep on purpose, as a barrier that parks the
  copy task at the mount block. `#[expect(.., reason = ..)]` records it.
- module_inception: needle/needle.rs mirrors the Go package layout.

Two lints become crate-wide policy in `[lints.clippy]`, with the reason
next to each: result_large_err, because every RPC path returns
tonic::Status (176 bytes) and boxing it would change every handler
signature; and needless_update, because `..Default::default()` on a
protobuf message literal is what lets a proto gain a field without
touching every constructor (all 11 sites are pb messages). The worker
workspace gets the same table and its members opt in with
`lints.workspace = true`; its generated plugin.rs also allows
large_enum_variant on prost's oneof enums.

Both crates are now clean under `cargo clippy --all-targets -- -D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: use the std APIs the 1.91 MSRV already pays for

The crate declares rust-version 1.91.1, so a few things the code still
worked around are plain std now. All of them come from the 1.85–1.91
release notes; nothing here needs a newer toolchain than the manifest
already requires.

- std::sync::LazyLock (1.80) replaces the lazy_static! block in
  metrics.rs, and the lazy_static dependency goes. Every use site reads
  the same through Deref, so no caller changes.
- Duration::from_mins / from_hours (1.91) replace `from_secs(v * 60)`
  and `from_secs(v * 3600)` in the option parser and the shard-location
  refresh TTLs. One difference for the parser: an absurd count that
  overflows u64 seconds now panics in release builds too, where the
  multiplication used to wrap.
- Result::flatten (1.89) replaces `.and_then(|r| r)` on the replication
  join handle.
- OsStr::display (1.87) replaces `to_string_lossy()` where the name was
  only being formatted; the output is byte-identical.
- `#[allow]` becomes `#[expect]` (1.81) on the suppressions that are
  meant to be permanent, so a suppression that stops being needed
  becomes a warning rather than lingering. Doing that found four that
  already had: dead_code on ChunkManifest, base_name and last_io_error,
  and too_many_arguments on read_from_data_shards, which is down to
  seven parameters. Those attributes are deleted. The three allows that
  depend on cfg (a unix-only mutation, a linux-only field set, a
  profiling-only parameter) stay as allow, because expect would be
  unfulfilled on the other platforms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out clippy step to both Rust workflows

Both crates are warning-free under `cargo clippy --all-targets
-D warnings` now. Whether that becomes a gate is a policy call, so the
step is present but commented out; uncommenting it is the whole change.
The comment points at the `[lints.clippy]` table where crate-wide
exceptions are recorded, so the gate does not become a reason to
sprinkle allows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: guard parse_duration against overflow panics

Duration::from_mins/from_hours panic when the count overflows u64
seconds. Use checked_mul so an oversized CLI value falls back to the
parser default instead of crashing volume startup.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:29 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
49f20489e4 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.20.1 to 1.20.4 (#11305)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.20.1 to 1.20.4.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.20.1...service/mq/v1.20.4)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.20.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:13:14 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bdec508da9 build(deps): bump golang.org/x/image from 0.45.0 to 0.46.0 (#11304)
Bumps [golang.org/x/image](https://github.com/golang/image) from 0.45.0 to 0.46.0.
- [Commits](https://github.com/golang/image/compare/v0.45.0...v0.46.0)

---
updated-dependencies:
- dependency-name: golang.org/x/image
  dependency-version: 0.46.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:45:16 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fd33c07843 build(deps): bump github/codeql-action from 4.37.9 to 4.38.0 (#11311)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.9 to 4.38.0.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.9...v4.38.0)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:41:36 -07:00
Chris LuandGitHub cf38c01978 admin: bind worker gRPC listener to -ip instead of wildcard (#11300)
* admin: bind worker gRPC listener to -ip instead of wildcard

The worker/plugin gRPC control plane called net.Listen("tcp", ":port")
directly, so it wildcard-bound every interface and ignored the -ip setting.
A cluster bound to loopback still exposed the unauthenticated
WorkerService/PluginControlService streams on 0.0.0.0. Bind through
util.JoinHostPort(bindIp, port) so the listener honors -ip like the
master, filer, and volume gRPC listeners.

* admin: warn when worker gRPC is exposed off loopback without mTLS

The worker gRPC stream has no password auth, so grpc.admin mTLS is the
only effective control once the listener leaves loopback. An operator who
sets -adminPassword and binds -ip=0.0.0.0 authenticates the HTTP API but
still exposes the unauthenticated worker control plane. Log a startup
warning naming the port and the mTLS knobs so the exposure is not silent.

* admin: address review on worker gRPC bind fix

- mini: reserve the admin gRPC port with util.JoinHostPort so an IPv6
  bindIp (e.g. ::1) does not form an invalid unbracketed address and
  lose the reservation.
- worker gRPC: track whether grpc.admin mTLS credentials actually loaded
  rather than only whether they were configured, and gate the
  non-loopback exposure warning on that. A cert/key that fails to load
  now still warns instead of silently suppressing.
2026-09-13 21:48:14 -07:00
Chris LuandGitHub f4bad510c9 test/ec: pin rack in seedAndSpread volume.grow to stop silent no-ops (#11299)
seedAndSpread() calls `volume.grow -dataNode X` without pinning the rack.
The master's grow picks the rack by weighted-random when -rack is unset,
and only one of the three racks holds the requested data node, so an
unpinned grow lands on the wrong rack two times out of three. The
VolumeGrow RPC swallows the "No matching data node" failure for
non-cache collections, so those grows count as success without creating
a volume. The per-server cap (maxGrowsPerServer=4) is then exhausted by
silent no-ops before the volumes ever spread, and seedAndSpread times
out with "volumes never spread across >=2 disks on all 3 nodes".

Pin -dataCenter dc1 and -rack rack{i} alongside -dataNode so every grow
reaches the target node. This removes the timing-sensitive assumption
that made TestECVacuumDuplicateShardClaimAcrossDisks flaky.
2026-09-13 21:05:55 -07:00
Chris LuandGitHub 15d9f6c6fe rust-volume: fix Windows build of find_needle_from_ecx (#11298)
* rust-volume: fix Windows build of find_needle_from_ecx

The .ecx binary-search fallback path used on non-Unix targets
(Seek + Read, both &mut self receivers) requires the ecx_file
binding to be mutable. On Unix the read_exact_at path takes &self,
so the mut would be unused there — gate that warning with
#[cfg_attr(unix, allow(unused_mut))].

Without this the build-rust-volume-windows CI job fails with
E0596 at ec_volume.rs:1033, breaking the weed-volume_windows_amd64
release asset.

* rust-volume: use positional seek_read for .ecx lookups on Windows

The previous fix (making ecx_file mut) compiled but left the Windows
fallback using Seek + Read on the shared .ecx file cursor. Concurrent
find_needle_from_ecx calls could interleave seek/read and read the
wrong index entry, corrupting the binary search (raised by Devin and
Greptile review on the PR).

Switch the Windows path to std::os::windows::fs::FileExt::seek_read,
which is positional (offset passed via OVERLAPPED, cursor untouched)
and takes &self — so the binding no longer needs mut, and concurrent
callers on the cached handle can't interfere. Mirrors the existing
read_exact_at helper in storage::volume.

Add a compile_error fallback for non-unix/non-windows targets to
match the convention in storage::volume.
2026-09-13 20:08:01 -07:00
ea179963c0 filer: clean up manifest resolve error propagation and add webdav tes… (#11297)
filer: clean up manifest resolve error propagation and add webdav test (#78)

Drop GitHub issue references from comments and trim verbose comments.
Replace the viewFromChunksOrErr helper with the existing
NonOverlappingVisibleIntervals + ViewFromVisibleIntervals at the stream
call sites, and add a WebDavFile.Read regression test for the manifest
resolution failure path.

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-13 18:34:09 -07:00
github-actions[bot] c507336000 4.47 2026-09-14 01:31:55 +00:00
github-actions[bot] 3c492b5ab1 docs: regenerate star history chart 2026-09-14 00:52:08 +00:00
Chris LuandGitHub 38c14d3c13 filer: apply SSRF guard to the lazy-remote fetch/list/delete paths (#11294)
* filer: add guarded remote-storage client builder hook for lazy fetch

The lazy-remote fetch path (maybeLazyFetchFromRemote) resolved its
remote-storage client through the unguarded shared cache, bypassing the
SSRF chokepoint (BuildGuardedRemoteStorageClient) that the CVE-2026-73080
remediation wired into the volume, filer stream and s3 stream dial paths.

Add a RemoteStorageClientBuilder hook on Filer plus conf-only lookups on
FilerRemoteStorage, and route the lazy fetch through the builder when set
(endpoint deny-list + DNS-rebinding-safe dialer), falling back to the
shared cache otherwise. The filer server wires the builder in a follow-up.

* filer: route lazy directory listing through the guarded remote client

maybeLazyListFromRemote shared the unguarded client resolution of the
fetch path, so a caller-supplied remote endpoint was dialed without the
SSRF deny-list or rebinding-safe dialer. Resolve the conf and build the
client through buildRemoteStorageClient so the same guard covers listing.

* filer: route lazy remote delete through the guarded remote client

maybeDeleteFromRemote issued outbound DELETE/RemoveDirectory requests
through the unguarded client, giving a write-side SSRF to a caller-chosen
endpoint. Resolve the conf and build the client through
buildRemoteStorageClient so the endpoint deny-list and rebinding-safe
dialer apply to the delete path as well.

* filer server: wire the guarded remote client builder into the filer

Set Filer.BuildGuardedRemoteClient to BuildGuardedRemoteStorageClient and
forward AllowUntrustedRemoteEndpoints so the lazy-remote fetch, list and
delete paths apply the same SSRF endpoint checks as the volume and
streaming read paths.

* filer: test lazy fetch honors the guarded remote client builder

Add a regression test that sets BuildGuardedRemoteClient to a rejecting
builder and asserts maybeLazyFetchFromRemote returns no entry without
reaching the remote, covering the SSRF guard wired in the prior commits.

* filer: skip remote client for local-only lazy deletes

maybeDeleteFromRemote resolved and validated the mount's remote client
before checking entry.Remote, so a local-only file (no Remote entry) under
a mount whose endpoint the guard rejects failed to delete: the guard
error aborted the metadata deletion, leaving a file that needs no remote
operation undeletable. Move the local-only check ahead of client
construction so only remote-backed files and directories pay the guard.

* filer: build the guarded remote client inside the lazy singleflight

The lazy fetch and list paths built the guarded client before their
singleflight blocks, so concurrent requests for the same key each
allocated a fresh SDK client and HTTP transport even though only one
remote operation ran. Move client construction inside the singleflight
so the deduplicated operation builds it once, matching the per-request
guard semantics of the sibling streaming paths without the duplicate
transport churn.

* filer: test guarded rejection for the lazy list and delete paths

Add regression tests that set BuildGuardedRemoteClient to a rejecting
builder and assert the lazy list does not reach the remote, a
remote-backed file delete is blocked, and a local-only file under a
rejected mount still deletes (covering the local-only fix).

* filer: decouple lazy guarded-client build from the first caller's context

Building the guarded client inside the singleflight made concurrent
fetches share the first caller's context. If that caller canceled while
endpoint DNS validation was running, the builder returned an error and
published a not-found result to other callers whose contexts were still
valid. Build with context.WithoutCancel so the guard's DNS validation
is not tied to any single caller's cancellation, matching the list
path's existing decoupling for the remote operation itself.

* filer: reject remote-storage confs that dial blocked endpoints at load

The filer's lazy-fetch / lazy-list / remote-delete paths resolve remote
storage clients by name from FilerRemoteStorage.storageNameToConf and
dial them via remote_storage.GetRemoteStorage, which bypasses the SSRF
deny-list the volume server (BuildGuardedRemoteStorageClient) and the
filer's own direct-read path apply. A RemoteConf planted under
/etc/remote with a loopback / private / IMDS S3 endpoint is reloaded into
storageNameToConf on the next metadata-change event and then dialed on
the next cache miss — server-side request forgery from the filer.

Apply the volume server's SSRF deny-list at conf load time, the single
chokepoint that populates storageNameToConf:

- Add RemoteStorageConfValidator, injected into FilerRemoteStorage by
  the filer server (the filer package cannot import the server package).
  A conf that fails validation is dropped from storageNameToConf, so the
  name-based client resolution on the lazy paths returns "not found"
  instead of dialing the blocked endpoint.
- Add ValidateRemoteConfForLoad in weed_server, which mirrors
  BuildGuardedRemoteStorageClient's gcs credential + endpoint checks
  (validateRemoteEndpoint via guardedRemoteClient) without building a
  client. allowUntrusted skips the check, mirroring the volume server
  opt-out (-filer.allowUntrustedRemoteEndpoints).
- The filer server injects the validator at construction.

A conf whose type dials a fixed provider host (no caller-supplied
endpoint) passes; only caller-influenced endpoints are denied.

* filer: skip DNS resolution in the load-time SSRF validator

ValidateRemoteConfForLoad resolved hostnames during /etc/remote reload,
so a transient DNS failure (2s timeout) dropped the conf from the fresh
map that replaces the live map, disabling a working mount until the next
metadata event. The build-time guard (BuildGuardedRemoteStorageClient)
already re-resolves and re-validates the endpoint at dial time with the
rebinding-safe dialer, so DNS at load is redundant for security.

Split the static checks (scheme, IMDS hostnames, IP-literal blocked
addresses, gcs credentials) into validateRemoteEndpointForLoad, which
does no DNS. Hostname endpoints pass at load and are caught at dial if
they resolve to a blocked address. This preserves fail-fast for
statically-blocked confs (loopback IPs, IMDS hostnames) without letting
transient DNS failures disable mounts.

* filer: accept empty S3 endpoints in the guarded remote client builder

guardedRemoteClient returned ok=true with an empty endpoint for a
standard AWS S3 config (no custom S3Endpoint), so
BuildGuardedRemoteStorageClient and ValidateRemoteConfForLoad rejected
it with "remote endpoint is empty" — breaking standard AWS S3 mounts on
the lazy paths and the sibling streaming read paths that already use the
guarded builder.

An empty endpoint is not caller-supplied: the AWS SDK derives the
regional endpoint from the region, so there is nothing for the SSRF
guard to validate. Return ok=false for empty S3-compatible endpoints so
the builder falls through to the shared unguarded cache, matching the
historical behavior for standard AWS S3.
2026-09-13 14:43:55 -07:00
Chris LuandGitHub 92c379e5b4 filer: accept gcs credentials file paths in the guarded remote client builder (#11296)
* filer: accept gcs credentials file paths in the guarded remote client builder

checkGcsCredentials rejected all filesystem paths, so a gcs mount
configured with remote.configure -gcs.appCredentialsFile (which stores
a path in GcsGoogleApplicationCredentials) was rejected by
BuildGuardedRemoteStorageClient with "gcs credentials must be inline
JSON". This broke existing gcs mounts on the volume, filer, and s3
remote-mount read paths that use the guarded builder.

Read and validate the file content instead of rejecting the path,
mirroring what the gcs client itself does in MakeWithHTTPClient. A path
that does not exist or does not contain valid gcs credentials is still
rejected before any client is built. guardedRemoteClient now reads the
file to extract the token exchange URL for the SSRF deny-list, so the
rebinding-safe dialer still guards the token endpoint.

* filer: resolve gcs credential paths and avoid leaking file existence

loadGcsCredentialsContent passed the raw credentials string to os.ReadFile,
so a documented ~/path (as written by remote.configure
-gcs.appCredentialsFile=~/...) was rejected because os.ReadFile does not
expand ~. It also wrapped the os.ReadFile error, which includes the
file path, exposing file existence to a caller who planted a conf with
an arbitrary path.

Resolve the path with util.ResolvePath, matching the gcs client's own
behavior in MakeWithHTTPClient. Return a generic sentinel error on read
failure so the path is not reflected in the error message. The credential
type validation still runs on the file content, so a path that does not
contain valid gcs credentials is rejected before any client is built.
2026-09-13 14:43:45 -07:00
Chris LuandGitHub 5d8a463b3e test/ec: fix EC interruption matrix slot exhaustion (#11295)
* test/ec: fix EC interruption matrix slot exhaustion

The EC integration test cluster (test/erasure_coding/chaos_lifecycle_test.go)
configured each disk with -max 4 and the seedAndSpread spread loop fired
volume.grow -count 4 every 2 s with no per-server cap. Because the master
topology lags the volume.grow writes, the loop re-fired before the prior
grow was visible, over-filling disks to capacity. A full disk leaves zero
free EC shard slots (failing the cluster-wide capacity check with "no
free ec shard slots") and drops the source disk below the encode's
FreeVolumeCount >= 2 health check (failing with "no healthy replicas"),
which aborted ec.encode before any phase marker printed and made every
encode scenario in TestECInterruptionMatrix fail.

Three changes to the test cluster:

1. Raise -max from 4 to 8 per disk so the source disk always retains
   FreeVolumeCount >= 2 for ec.encode's 14-shard generation (2 volume-slot
   equivalents) even after the spread loop and multiple encodes.

2. Switch the spread loop from -count 4 to -count 1 so each grow lands
   exactly one volume on the volume server's least-loaded disk, giving
   deterministic cross-disk spreading instead of relying on a single
   multi-volume grow to fan out.

3. Cap grows per server at 4 so heartbeat lag cannot run away and
   over-fill disks before the master registers the prior grow.

4. Pass -minFreeSpace 0 so the test is not falsely gated by the physical
   disk's free-space percentage on the host running CI (the EC shard slot
   calculation separately enforces a 90 % disk-usage cap via
   balancer.DiskTooFullAfter, which already guards against an over-set
   maxVolumeCount on a physically full disk).

Verified locally by running TestECInterruptionMatrix twice (all encode,
decode, and balance scenarios pass, including the previously failing
encode@Deletingoriginalvolumes).

* test/ec: only count successful grows toward the spread cap

A failed volume.grow (e.g. a transient collectTopologyInfo or VolumeGrow
RPC error) would otherwise consume one of the four permitted attempts
without creating any volume, exhausting the retry budget and leaving the
loop to only poll until the Eventually timeout. Increment the per-server
counter only when commandGrow.Do returns nil.
2026-09-13 14:41:45 -07:00
Chris LuandGitHub bea10e269f iceberg/s3tables: confine stored metadataLocation to the authorized table bucket (#11292)
* iceberg: confine commit/transaction/view-update write paths to authorized bucket

The create, register, and createView handlers already confine the client-
supplied metadata location to the caller table bucket and reject ".."
segments. The commit, create-on-commit, transaction, and view-update paths
read the stored metadataLocation back from the catalog and skipped the same
guard, so a location poisoned via the raw S3Tables UpdateTable API (which
persists metadataLocation verbatim) could escape the caller bucket through
a ".." segment that path.Join collapses in saveMetadataBlob.

Add confineMetadataLocation and apply it after parseS3Location on every
commit/update/transaction/view write path, mirroring the create/register/
createView check. Reject with 400 so a poisoned stored location fails the
commit instead of writing into another tenant bucket tree.

* s3tables: validate metadataLocation at the store layer

The raw S3Tables API (CreateTable, RegisterTable, UpdateTable, CreateView,
UpdateView) persisted the client-supplied metadataLocation verbatim with no
bucket-confinement or traversal check, so a caller could store a location
pointing outside its own bucket. The Iceberg REST gateway commit paths then
read that stored value back and wrote through it.

Add ValidateMetadataLocation and call it in every s3tables store handler
that accepts a metadataLocation, rejecting locations whose bucket differs
from the caller table bucket or whose path contains traversal segments. This
prevents a poisoned location from ever being persisted, complementing the
per-write-path guard added to the Iceberg commit handlers.

* iceberg/s3tables: validate location before repair and after idempotency check

Address review feedback:
- Move the commit-path confinement check ahead of repairManifests so a
  poisoned stored location cannot reach manifest repair I/O before the
  commit is rejected.
- Move ValidateMetadataLocation in CreateTable/CreateView to after the
  existing-resource check so idempotent retries that do not consume the
  requested location are not rejected for an unused bad location.
- Assert HTTP 400 in the cross-tenant reproduction tests so an unrelated
  failure cannot satisfy them.

* iceberg: confine staged metadata location before load in create-on-commit

The create-on-commit path parsed the staged metadata location from the
stage-create marker and called loadMetadataFile before validating that the
staged bucket/path stay within the authorized bucket. Add the same
confineMetadataLocation guard before the read so a tampered marker cannot
direct a cross-tenant metadata read.

* iceberg/s3tables: reject bucket-only metadata locations

ValidateMetadataLocation and confineMetadataLocation accepted s3://bucket
with an empty table path. metadataDirPath then maps every such table to the
shared <TablesPath>/<bucket>/metadata directory, so tables could overwrite
or read each other's metadata files. Require a non-empty table path in both
validators; the empty-location case (where the catalog derives one) is
unaffected.

* iceberg/s3tables: reject slash-only table paths in location validation

s3://bkt/// parses to tablePath="/" which passed the empty-string check
but path.Join cleans it away, mapping to the bucket-level metadata
directory shared across tables. Update isValidTablePath to require at
least one non-empty segment and mirror the same check in
ValidateMetadataLocation, closing the gap in all callers.
2026-09-13 13:48:13 -07:00
Chris LuandGitHub 10c0857476 s3: gate internal LifecycleDelete gRPC behind admin Bearer auth (#11291)
* s3/lifecycle: attach admin Bearer token on internal LifecycleDelete clients

Export credential.WithS3InternalAdminAuth (renamed from withIamCacheAdminAuth)
and use it in the worker and shell lifecycle RPC adapters so lifecycle calls
carry the same admin token the IAM-cache propagation already attaches. No-op
when jwt.filer_signing.key is unset, matching the server-side checkAdminAuth.

Prepares the internal clients for the server-side auth gate that follows.

* s3/lifecycle: gate LifecycleDelete behind admin Bearer auth

Add checkAdminAuth to LifecycleDelete, matching the SeaweedS3IamCache handlers on the same internal gRPC listener (PR #11190). No-op when jwt.filer_signing.key is unset; rejects unauthenticated callers when it is. The internal worker/shell clients already attach the token in the previous commit.
2026-09-13 13:07:10 -07:00
c462fffce6 master: name the unlabeled disk layout plainly in assign errors (#11290)
* master: name the unlabeled disk layout plainly in assign errors

When no volume server serves the layout an assign targets, the error
named the empty disk type as "hdd" (HardDriveType is the empty string),
sending operators looking for servers labeled hdd when the actual
mismatch is labeled (e.g. -disk=ssd) servers versus unlabeled clients.

- describe the layout as "default (unlabeled)" when the disk type is
  empty, keep %q naming for labeled types
- log the unserved-layout condition once per option instead of letting
  every failing write repeat an unactionable line

Observed in production: volume servers started with -disk=ssd while CSI
mounts assign with the unlabeled layout; the per-write error stream
pointed at a nonexistent hdd fleet.

* master: bound and expire the unserved-layout warning dedupe

The dedupe map retained every distinct option key permanently. Option
keys embed request-derived fields (collection, disk type), so repeated
assignments with distinct options would grow master memory without
bound, and a retained key suppressed the warning if the same option
went unserved again after the topology recovered.

Remember last-warned timestamps instead, expiring after an hour, with a
hard cap that resets the set when a client-driven key flood fills it.

* master: silence per-retry unserved-layout log and name explicit hdd

Addresses Devin Review comments on #11290.

- The unserved-layout branch already rate-limits its warning via
  assignUnservedLayoutWarning.Do, but the common epilogue still logged
  lastErr at V(0) on every retry, so the flood the dedup was meant to
  stop continued. Skip the epilogue log when the unserved-layout branch
  owns the logging; the error is still returned to the client.
- describeDiskLayout took the canonicalized option.DiskType, but
  ToDiskType folds both "" and "hdd" into HardDriveType, so an explicit
  disk=hdd request was mislabeled "default (unlabeled)". Pass the
  original request disk type instead: only an empty request is the
  unlabeled default; an explicit hdd is named "hdd".

Adds TestAssignFailsFastNamesExplicitHdd covering the explicit-hdd
wording.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-13 11:56:32 -07:00
Chris LuandGitHub 99d2479528 fix(vacuum): batch fsync in makeupDiff to prevent test timeout (#11289)
makeupDiff called dstDatBackend.Sync() (fsync) per needle in the loop
over incrementedHasUpdatedIndexEntry. With 20000 entries in
TestLDBIndexCompaction this resulted in up to 20000 fsync calls, which
on slow CI disks exceeded the 10-minute test timeout.

Batch the sync: write all needles/tombstones first, then fsync the dat
file once in the defer alongside the existing idx fsync. The durability
guarantee is unchanged — both files are still synced before CommitCompact
writes the .cpc commit marker and swaps the files.
2026-09-13 00:05:30 -07:00
8db41d0217 [Mount] Cache Chunk Manifest Resolution for Repeated File Opens (#11266)
* cache resolved chunk manifests for Mount

* Address PR review: per-mount cache, singleflight, reuse ResolveOneChunkManifest

- Own the manifest cache per WFS mount instead of a process-global
  variable, so manifests from one filer backend are never served to
  another (Devin/CodeRabbit major bug).
- Coalesce concurrent cold misses via singleflight so only one fetch
  runs during a cold burst (Greptile P2).
- Copy cached data after releasing the mutex so a large copy does not
  block concurrent hits, inserts, and evictions (CodeRabbit nitpick).
- Reuse the existing ResolveOneChunkManifest function name instead of
  introducing a new resolveOneChunkManifest wrapper.
- Validate (unmarshal) manifest bytes before caching so malformed
  manifests do not poison the cache.
- Add TestChunkGroupManifestResolutionCoalescesColdMisses covering
  the singleflight cold-miss path.

* Address round 2 review: coalesced-miss cancellation, test overlap

- Use singleflight.DoChan in fetchOrLoad and select on ctx.Done() so a
  caller whose context is canceled while waiting for an in-flight fetch
  returns ctx.Err() promptly instead of blocking for the leader's
  result (Devin BUG).
- Add TestResolveOneChunkManifestCanceledWaiterReturnsDuringCoalescedMiss
  covering the canceled-waiter path.
- Delay the cold-miss fixture response so the leader's fetch is still
  in flight when concurrent opens join the singleflight, making the
  one-fetch assertions reliable (CodeRabbit Minor).

* Address review: keep ResolveOneChunkManifest four-argument

Restore the exported ResolveOneChunkManifest to its original
four-argument signature so external callers keep compiling. Move the
cache-aware resolution into an unexported resolveOneChunkManifest
helper that accepts the per-mount ChunkManifestCache. The exported
function delegates to the helper with a nil cache, preserving the
historical uncached behavior for every non-Mount caller. The Mount
path (ChunkGroup.SetChunks) now calls the unexported helper with the
mount-owned cache. Tests and benchmarks that exercise the cache path
call the unexported helper directly.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-12 20:06:25 -07:00
github-actions[bot] bd6bcd47e3 docs: regenerate star history chart 2026-09-13 00:54:31 +00:00
Bruce ZouandGitHub eb6a7e93ca Fix mount eio on manifest resolve failure (#11287)
* mount: fail reads with error when chunk manifest resolution fails

When SetChunks fails to resolve a chunk manifest (e.g. the volume is on a
remote tier with reads disabled), the sections map stays empty and
readDataAtSequential/readDataAtParallel zero-fill every missing section as
if it were a sparse hole. Reads then return all-zero data with no error,
so a plain cp of a large manifest-based file silently produces a
completely zero-filled file.

Remember the resolve error in ChunkGroup (guarded by sectionsLock) and
return it from ReadDataAt. A later successful SetChunks clears it.

Fixes the mount path of #11286.

* filer: propagate manifest resolve errors in streaming read paths

ViewFromChunks discards the chunk manifest resolve error returned by
NonOverlappingVisibleIntervals. On failure the chunk views come back
empty, and the streaming paths zero-fill the entire requested range,
serving HTTP 200 / WebDAV 200 responses whose body is all zeros.

Propagate the error in PrepareStreamContentWithThrottler,
PrepareStreamContentWithPrefetch and the WebDAV read path so these
requests fail with 500 instead.

Fixes the filer HTTP and WebDAV paths of #11286.

* mount: fail lseek with EIO when chunk manifest resolution fails

SearchChunks still consulted the stale section map after SetChunks
recorded a manifest resolution failure, so SEEK_DATA/SEEK_HOLE would
describe the unresolved regions as sparse holes or return ENXIO.
Return the recorded error from SearchChunks and map it to EIO in
Lseek.

Also add regression tests for the stream preparation error paths.

Addresses review feedback on #11287.
2026-09-12 14:37:36 -07:00
5b2fe374fc [Volume] Scrub every disk's EC shards for a volume id, not just the first (#11258)
* storage: add Store::find_all_ec_volumes for split-disk EC lookups

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: add merge_ec_runtimes to resolve a vid's per-disk shard set

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: replace dead slots.get(14) assertion with a width-14 pin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: build the checksum scrub plan from every per-disk runtime

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: build the local scrub plan from every per-disk runtime

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: prove the local scrub plan reaches every runtime's slots

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: make the scrub plan tests falsifiable

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: report unverifiable protection when the sidecar predates the scrubbed encode

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: commit sidecar provenance with the sidecar it describes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: scrub every disk's EC shards for CHECKSUM and LOCAL

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: run the FULL/READS parity check across split-disk shards

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: report fenced-out runtimes in FULL/READS scrubs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: tighten verify_ec_shards ordering and missing-shard coverage

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: visit each EC volume id once in node-wide scrubs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: cover split-disk scrub aggregation end to end

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: pin fenced-out disks and sibling-disk shards in EC scrubs

Three scrub behaviors shipped without a test at the RPC seam. Task 8
showed the seam exists, so close them here.

FULL/READS (mode 2|5) now marks a volume broken when the identity fence
excludes a runtime, where it previously reported clean. Pinned against a
control fixture whose two disks AGREE and scrub clean, so the test fails
on the clean->broken transition, not only on the message text. That needs
a structurally valid, tombstone-only .ecx (so the needle walk finds
nothing to complain about) and a seeded shard-location cache (so the
absent master does not short-circuit the scrub with an error of its own).

LOCAL (mode 3) and CHECKSUM (mode 4) now build their plans from every
per-disk runtime. Made observable by moving shard 0 -- the shard the
volume's single needle spans and the one the checksum sidecar is checked
against -- to the SIBLING disk, leaving shard 5 on the disk the singular
find_ec_volume lookup returns. Built from that disk alone, neither scrub
ever looks at shard 0.

The split-disk fixture grows a config struct rather than more positional
arguments; its defaults reproduce the existing layout byte for byte, so
the node-wide dedupe test is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: report fenced-out disks on a malformed sidecar too

`errors.extend(self.skipped)` sat below the whole status match, so only
`(Some(p), On)` ever reached it. The Invalid arm already returns a
non-empty error vector of its own, so the Go-parity contract that
silences the Off arm (`case BitrotOff: return 0, nil, nil`) does not
reach it -- appending the fence lines there costs nothing that contract
protects. A volume with BOTH a malformed sidecar and a disk the identity
fence excluded reported only the sidecar, hiding the unscanned disk
behind an unrelated integrity error.

Off stays byte-identical, and so does the `(None, On)` arm that is
documented as treating a missing payload defensively as protection off.
Off is now the ONLY status that drops the report, and the comment at the
On-path copy says so: that is the one place the parity constraint costs
us coverage.

Also corrects a false claim in the FULL/READS test's doc comment. It
said a fenced-out disk "is a disk this scrub did NOT read", which is true
only of the merge-driven parity half. The per-needle walk still resolves
`store.find_ec_volume` (store_ec.rs:281) and binds
`expected_encode_ts_ns` to that runtime (:311) -- position 0, the
EXCLUDED one on that fixture -- so `read_local_intervals`' generation
filter (:1204) makes it read the excluded disk and treat the anchor's
shards as non-local, the inverse of what `skipped` reports. The fixture's
tombstone-only .ecx walks nothing, so the test cannot tell the two apart;
the comment now says that rather than implying coverage it does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: take CHECKSUM's bitrot protection from the disk that has the sidecar

`EcChecksumScrubPlan::for_volumes` read `(prot, status)` off the ANCHOR.
The anchor is the first shard-bearing runtime at the maximum `encode_ts_ns`,
chosen with no regard for which disk holds the `.ecsum`.

That sidecar is deliberately NOT mirrored across disks -- `ec_metadata_dirs()`
exists so one authoritative copy stays reachable rather than being duplicated
-- and at mount `EcVolume::new` resolves it via `load_active_bitrot_sidecar(&[])`
with no sibling directories at all; only the `VolumeEcShardsMount` RPC ever
passes `ec_metadata_dirs()`. So after EVERY volume-server restart, the
split-disk runtime that does not physically hold the sidecar mounts
`BitrotStatus::Off`. When the one copy lives on disk 1 and the anchor is disk 0,
`run()` hit `case BitrotOff` and returned `(0, [], [])`: the whole volume
scrubbed clean, silently. That is the steady state for roughly half of all
mirrored split-disk layouts, and it is the exact failure this branch exists to
remove.

Source protection from the first MERGED runtime that has any -- `On` if one
does, else `Invalid`, else the anchor's `Off`. Two facts make that safe, and
both are load-bearing:

  - Every runtime that mounted `On` already passed the `geometry_matches` gate
    in `load_bitrot_for_generation`, so its manifest agrees with the volume's
    layout. A sidecar that contradicted it would have failed the mount.
  - All merged runtimes share the same `encode_ts_ns` by construction of the
    identity fence, so a sidecar from any of them describes the same encode run.

The `unverifiable_sidecar` provenance rule four lines down read
`anchor.bitrot_source_dir`; it now reads the SAME runtime `prot` came from.
Otherwise the two would describe different sidecars and the rule would vouch
for a manifest nobody is scanning against. One consequence worth naming: that
source dir is now non-empty by construction (a runtime with protection found a
file), where the anchor's was often "" and short-circuited the rule -- so on a
fenced volume whose anchor had no sidecar, an unverifiable-protection note now
surfaces where previously nothing was reported at all.

`run()` is untouched, and the `BitrotStatus::Off` arm still returns
`(0, [], [])` exactly, for Go parity with `case BitrotOff: return 0, nil, nil`.
`parity_shards` still comes from the anchor while `prot` may come from a
sibling; the geometry gate above makes them agree, and slot-width agreement is
handled separately.

The test drives mode 4 through the real RPC against a split-disk volume whose
sidecar exists only on dir1, and asserts up front that the anchor mounted `Off`
and the sibling `On` -- otherwise it would prove nothing. Reverting this commit's
one-line source change makes it report `[]` instead of `[0, 5]`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: pin the slot width, contain the shard-size fallback, and cover multi-disk FULL

Five findings from the whole-branch review, none of which changes what a
healthy volume reports.

Slot width was undefined and the two consumers disagreed (ec_volume.rs).
`merge_ec_runtimes` sizes `slots` to the WIDEST merged runtime, but the identity
fence keys on `encode_ts_ns` alone and never on geometry -- so two
same-generation runtimes whose `.vif`s disagree do merge. The mode 2|5 arm
truncates to the anchor's `data+parity` and silently drops the surplus slots,
while `EcChecksumScrubPlan::for_volumes` iterated the full width and emitted
"present but missing from sidecar manifest" for exactly those ids. Nothing in
the volume describes them -- the sidecar manifest and the Reed-Solomon matrix
are both the anchor's -- so that message was the width disagreement talking, not
a finding. The `slots` field doc now states the contract (the range is the
anchor's geometry; every consumer truncates to it) and CHECKSUM truncates.

The LOCAL `shard_size` fallback had grown a node-wide blast radius
(ec_volume.rs). `anchor.shard_file_size()` returns the anchor's FIRST held
shard, not a maximum. Before aggregation the plan read only that runtime's own
shards, so a truncated shard was contained to its disk; now that one value sizes
every merged sibling's shards, mis-offsetting `locate_data` and manufacturing
needle corruption across the node. Take the max over the merged slots, which is
how `verify_ec_shards` already answers the same question
(`if size > shard_size { shard_size = size }`). Only on the legacy
`dat_file_size == 0` path.

Multi-disk `all_local` had no end-to-end test (grpc_server.rs). The parity check
is gated on every shard being present, and the one all-local fixture keeps them
in a single directory, so every entry of `dirs` is the same string and a
permutation or off-by-one in the `slots` -> `dirs` mapping is invisible;
`test_verify_ec_shards_reads_shards_from_multiple_dirs` builds its `dirs` by
hand and never goes through `merge_ec_runtimes`. The new fixture is a real 10+4
encode split 0..=6 / 7..=13 across two store locations (the `.dat`/`.idx` stay
outside both, so `prune_incomplete_ec_with_sibling_dat` has nothing to act on),
driven through the real RPC: clean first, then a corrupted PARITY shard on the
SECOND disk -- which only the parity half can see, and only through a correct
mapping. Shifting that mapping by one, or computing `all_local` from the anchor
alone, both make it report `[]` instead of `[13]`.

Deleted `test_ec_volume_enumeration_is_deduped` (store_ec_reconcile.rs). It
built `raw` from `store.locations` and then applied its OWN inline
`filter(|v| seen.insert(*v))`, asserting on that -- a property of
`HashSet::insert`, never reaching the production dedupe. That path is covered by
`test_scrub_ec_volume_node_wide_dedupes_a_split_disk_volume`, which does fail
(2 != 1) when the dedupe is removed.

Corrected `test_verify_ec_shards_treats_a_none_dir_as_missing`'s docstring
(ec_encoder.rs). It claimed the unmounted shard "must not drag the shards that
ARE mounted down with it", but `dirs[5] = None` puts shard 5 in `broken_shards`
before the block loop, so every iteration takes the `read_failed` arm and the
parity comparison never runs: corrupting a mounted shard in that fixture changes
nothing about the result. The assertions are unchanged; the docstring now states
what they actually establish.

Also refreshed two comments that cited `shard_file_size() - 1` as the reason
`merge_ec_runtimes` prefers a shard-bearing anchor -- true before this commit,
stale after it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: correct the Fix 1 rationale and truncate the shard-size scan

The safety argument attached to `EcChecksumScrubPlan::for_volumes`'s protection
selection was false as written, and it is the argument a reviewer reads first.
`geometry_matches` compares a sidecar against the MOUNTING runtime's own
data/parity/block size, not the anchor's, and returns true vacuously when
`ec_shard_config` is `None` -- so it establishes agreement only when all merged
runtimes share one geometry, which an `encode_ts_ns`-only fence does not
guarantee and which `test_checksum_scrub_truncates_slots_to_the_anchors_geometry`
constructs a counterexample to. The second clause was weaker than stated too: a
`.ecsum` records no encode identity at all, so merged runtimes agreeing on
`encode_ts_ns` does not transfer to the sidecar.

Replace it with the property that is true, checkable from the selection itself,
and stronger for what actually matters. `anchor` is an element of `merged`, so
the `.unwrap_or(anchor)` fallback is reached only when no merged runtime is `On`
and none is `Invalid` -- in which case the anchor is necessarily `Off`. The
status can therefore only move `Off -> On`, `Off -> Invalid` or
`Invalid -> On`; never `On -> Off`, never `Invalid -> Off`. This selection
cannot stop a volume that was being scanned from being scanned, and cannot turn
a reported integrity error into silence: every change it makes is toward more
verification. The comment now also states what it does NOT establish -- geometry
agreement is not guaranteed -- and names geometry fencing as the follow-up that
would close it.

Second, `EcLocalScrubPlan::for_volumes`'s `shard_size` max scanned the FULL slot
width, violating the `slots` contract documented in the same commit that
introduced the max: the volume's shard-id range is the anchor's geometry and
every consumer must truncate to it. Pre-fix that input could not exist, because
`anchor.shard_file_size()` read only the anchor's own anchor-sized vector -- so
the max opened a new, narrow path to the same node-wide mis-sizing it exists to
close (same-generation runtimes with disagreeing `.vif`s, the wider one holding
an out-of-geometry shard larger than the in-geometry ones, `dat_file_size == 0`).
`.take(anchor.data_shards + anchor.parity_shards)` mirrors the truncation
already applied to the CHECKSUM shard scan.

The sibling `shards:` vector is left untruncated on purpose: every access in
`EcLocalScrubPlan::run` is `shards.get(sid)` with `sid < data_shards`, so the
surplus entries are inert.

No behavior change for any healthy volume, and no test added -- the suite is
unchanged at 575 passing, 0 failing, 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: aggregate split-disk runtimes in Go scrubs, mirroring Rust

Go volume scrubs previously used FindEcVolume (first runtime only), so a
volume whose EC shards are split across multiple disks was scrubbed against
just one disk's shards and the others were silently skipped. Node-wide
ScrubEcVolume also appended each disk's EcVolumeIds without deduplication,
scrubbing a split-disk volume once per disk.

Add MergedEcRuntimes/MergeEcRuntimes (Go counterpart to Rust's
merge_ec_runtimes): select the maximum EncodeTsNs as the anchor generation,
fence out runtimes whose encode generation or geometry (DataShards,
ParityShards, BlockSize) disagrees with the anchor, merge shard handles by
shard ID, and report excluded runtimes rather than dropping them. Wire it
into every scrub mode:

- INDEX: scrub the anchor's index, report skipped runtimes.
- LOCAL: aggregate local shards across all merged runtimes via a synthetic
  EcVolume built from the merged shard slots.
- FULL/READS: resolve the runtime matching the anchor's encode generation
  (not the first match) so the needle walk and parity phase inspect one
  encode run; report skipped runtimes.
- CHECKSUM: take bitrot protection from the first merged runtime that has a
  valid sidecar (On, else Invalid, else anchor's Off), preserve invalid
  sidecar errors from every other merged runtime, and report skipped
  runtimes.

Deduplicate EC volume IDs in node-wide ScrubEcVolume so each volume is
scrubbed exactly once.

Refactor ScrubEcVolume to share the per-needle walk via scrubEcVolumeWalk,
called by both the legacy first-runtime path and the new merged path.

Add Go regression tests covering split-disk deduplication, encode-generation
fencing, geometry fencing, sibling-disk LOCAL reach, and merge anchor
selection.

Rust: keep the previously-landed merge/fence/checksum changes intact; revert
incidental cargo-fmt drift from unrelated files so the diff stays focused.

* ec: fence merged CHECKSUM on sidecar encode generation and fix legacy shard size

Address two review findings on the Go merged-runtime scrub:

1. Sidecar provenance: a merged runtime can load a bitrot sidecar from a
   sibling metadata directory (ReloadBitrotSidecar), and the merge fence may
   then exclude the runtime owning that directory. Generation-0 sidecars do
   not identify the encode run, so geometry validation alone cannot prove the
   borrowed manifest describes the anchor shards. If the sidecar records a
   non-zero EncodeTsNs that disagrees with the anchor, refuse the scan
   instead of applying stale checksums to current shards and reporting false
   corruption.

2. Legacy shard size: for volumes without datFileSize in .vif,
   LocateEcShardNeedleInterval derives the shard size from Shards[0].ecdFileSize.
   The merged shard set is compacted in shard-ID order, so a truncated
   lowest-ID shard would shrink every interval and misread intact sibling
   shards. Synthesize a datFileSize from the maximum mounted shard size when
   the anchor lacks one, so the datFileSize>0 path uses the largest shard
   size across all merged runtimes.

* ec: fix copylocks, legacy shard boundary, and encode-aware Rust lookups

Address review findings from CodeRabbit and Devin:

Go (ec_volume_merge.go):
- Remove bitrotLock copy from the synthetic EcVolume: copying a sync.RWMutex
  is a go vet copylocks error. The synthetic volume uses its own zero-value
  mutex; bitrot/bitrotStatus are set directly before ChecksumScrub reads them
  via BitrotProtection(), so no concurrent access occurs.
- Fix legacy shard-size boundary: synthesize datFileSize from
  (maxShardSize - 1) * DataShards, not maxShardSize * DataShards, to match
  the legacy fallback in LocateEcShardNeedleInterval (ecdFileSize - 1). An
  exact large-block boundary is ambiguous; the unadjusted size would select
  an extra large row and misread intact sibling shards.

Rust (store_ec.rs):
- Add find_ec_volume_for_scrub helper that resolves by encode generation
  (not first-match find_ec_volume) and use it in scrub_snapshot_under_lock,
  write_back_shard_locations, and the post-refresh shard-location read.
  Previously the encode-aware lookup was only used for the initial runtime
  selection; the cache write-back and per-needle snapshot still used
  first-match, so a split-disk volume whose first runtime was from an older
  encode run would write to and read from the wrong runtime's shard-location
  cache and falsely abort with 'remounted as a different encode run'.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-12 14:19:18 -07:00
Chris LuandGitHub 3b4a681e53 test(fuse_failover): dump chunk list and hex on append corruption (#11285)
* test(fuse_failover): dump chunk list and hex on append corruption

The failover append test (TestAppendWhileVolumeServerRestarts) failed
in CI with an 8-byte NUL region at offset 632 that appeared in both
the writer mount and the filer own view, but the failure message
only showed a quoted-string window around the divergence. That is
not enough to tell which chunk covered the zeroed bytes or which
volume server held it, so the next recurrence would be just as
unattributable.

Add a FileChunkList helper that reads the filer resolved chunk
list, and on failure dump:
  - every chunk fid, offset, size, volume id, and current master
    holders, flagging the chunk that covers the first divergence;
  - a hex+ASCII dump of the writer mount around the divergence so
    the exact zero-filled region is visible byte-for-byte.

No production code is touched; this only makes the test fail louder.

* test(fuse_failover): preserve diagnostic collection errors

Address review feedback from CodeRabbit and Greptile on PR #11285:

- FileChunkList now returns the wrapped ParseUint error when
  fid.volume_id is zero and the file_id prefix is invalid, matching
  FileVolumeIds instead of silently keeping vid=0 (which would
  query /dir/lookup?volumeId=0 and report the wrong holders).
- dumpChunkList captures the VolumeHolders error and renders it as
  'lookup failed: ...' so a failed master request is distinguishable
  from a successful lookup with no holders (both previously showed
  'holders=[unknown]').
- runChaosAppend captures the writer-mount read error and includes
  it in the failure message so an unavailable writer view is not
  mistaken for corrupted content.
2026-09-11 23:08:39 -07:00
Chris LuandGitHub c46f82d29a fix(master): stop goraft server on shutdown and bump raft to v1.2.1 (#11284)
MasterServer.Shutdown only stopped the Hashicorp raft implementation;
when using the default goraft backend, the raft event-loop goroutine
(leaderLoop/followerLoop) kept running after the master shut down. In
the in-process test harness this leaked goroutines across sequential
test runs, and a stale event occasionally reached a leader at term 0
and tripped the goraft "leader.elected.at.same.term" assertion,
crashing the whole test binary (CI run 34670959967, PR 11279).

Stop the goraft server in Shutdown() so its goroutines exit cleanly,
and bump seaweedfs/raft to v1.2.1 which replaces that assertion with a
graceful step-down to Follower instead of a panic.
2026-09-11 22:25:05 -07:00
Chris LuandGitHub 5a0e017457 s3: reject virtual-host bucket retargeting via X-Forwarded-Host (#11281)
* s3: reject virtual-host bucket retargeting via X-Forwarded-Host

SigV4 verification tries the client-supplied X-Forwarded-Host as a
signed host candidate, while routing and IAM select the bucket from
the actual Host header.  A presigned URL for one virtual-host bucket
could therefore be retargeted to another bucket accessible to the same
signing identity by changing Host and adding X-Forwarded-Host.

After the signature matches a host candidate, extract the bucket that
the candidate implies (via the configured virtual-host domains) and
compare it with the bucket the router selected.  Reject when they
differ, before returning success.

* test(s3api): cover virtual-host presigned URL retargeting

Add unit tests for bucketFromVirtualHost and end-to-end tests that
reproduce the X-Forwarded-Host retargeting attack for both presigned
and signed requests, plus a negative test confirming the legitimate
same-bucket case still verifies.

* s3: harden bucketFromVirtualHost for case and overlapping domains

Compare host and domain suffixes case-insensitively so a mixed-case
X-Forwarded-Host cannot bypass the consistency check.  Only treat the
exact path-style domain as non-virtual-host; subdomains of a path-style
domain still match the virtual-host router pattern and must be checked.
2026-09-11 22:23:26 -07:00
Chris LuandGitHub 210afacd12 s3: close list-type / ownership-controls routing mismatch (#11280)
* s3: reject list-type paired with another operation subresource

?list-type=2&ownershipControls= routes to ListObjectsV2 (the list-type
route is registered first) while the IAM action resolver resolves the
ownershipControls selector to s3:GetBucketOwnershipControls. A principal
denied s3:ListBucket but allowed s3:GetBucketOwnershipControls would
therefore list the bucket. list-type selects an operation just like the
other keys in operationSubresources, so add it there and reject the
combination before routing, matching the fix for policy&tagging (#10987).

* s3: resolve list-type to s3:ListBucket ahead of bucket subresources

The router registers the ListObjectsV2 route ahead of the bucket
subresource routes, so the action resolver should resolve list-type the
same way. Without this, a request carrying list-type and another operation
selector resolves to the subresource action (e.g. s3:GetBucketOwnershipControls)
while being served by ListObjectsV2. The ambiguity guard rejects such
combinations before routing, but resolving list-type to s3:ListBucket keeps
the resolver aligned with the router, mirroring how versions is handled.

* s3: match list-type=2 exactly in action resolver

The router selects ListObjectsV2 only for list-type=2; other values fall
through to the subresource routes. Resolve the same way so the action
matches the handler for every list-type value, not just 2.
2026-09-11 22:21:58 -07:00
Chris LuandGitHub 9f6feef299 feat(s3api): add bucket quota S3 extension via ?seaweedfs-quota (#11279)
* feat(s3api): add bucket quota S3 extension via ?seaweedfs-quota

Add a SeaweedFS-specific S3 subresource for bucket quota management:

  PUT /{bucket}?seaweedfs-quota   — set bucket quota (s3:PutBucketQuota)
  GET /{bucket}?seaweedfs-quota   — get bucket quota (s3:GetBucketQuota)

The request/response body is JSON:
  {"quota_size": 100, "quota_unit": "GB", "quota_enabled": true}

Quota is stored on the bucket's filer entry (positive = enabled,
negative = disabled but retained, zero = no quota), matching the
existing admin REST API behavior. When quota is cleared, the bucket's
read-only flag is also lifted.

Authentication uses the existing S3 SigV4 flow — no new global secret
is needed. Authorization uses two new dedicated IAM permissions:
  s3:PutBucketQuota
  s3:GetBucketQuota

This allows integrations like Apache CloudStack to manage per-bucket
quotas through the S3 endpoint with a scoped credential, without
exposing the broad admin REST API or requiring a separate admin token.
The credential can be limited to s3:PutBucketQuota/s3:GetBucketQuota
only, preventing bucket deletion, user management, or cluster topology
changes.

The coarse-grained ACTION_PUT_BUCKET_QUOTA/ACTION_GET_BUCKET_QUOTA
constants are added to s3_constants, and the action resolver maps the
seaweedfs-quota query parameter to the fine-grained s3: actions for
policy evaluation.

* docs: update design for S3 ?seaweedfs-quota extension approach

Replace the broad admin REST API + bearer-token design with the narrow,
scoped S3 ?seaweedfs-quota extension. Update quota, usage reporting, and
SeaweedFS-side changes sections to reflect PR #11279.

* fix(s3api): address review comments on quota handler

Fix four issues identified by Devin, Greptile, and CodeRabbit reviews:

1. Integer overflow in convertQuotaToBytes: large quota_size values
   (e.g. 8388608 TB) could overflow int64, wrapping to negative and
   being silently treated as zero quota. Now returns an error when
   size * multiplier would exceed math.MaxInt64.

2. Disabled quotas returned negative sizes in GET: the GET handler
   returned entry.Quota directly, which is negative for disabled-but-
   retained quotas. Now returns the absolute magnitude as quota_size
   and derives quota_enabled from the sign, making the response
   round-trippable.

3. Missing buckets returned 500 instead of NoSuchBucket: the PUT
   handler treated all lookup failures as internal errors. Now
   distinguishes filer_pb.ErrNotFound and returns ErrNoSuchBucket.

4. Trailing JSON was silently accepted: the decoder read only the
   first JSON object without checking for trailing data. Now
   requires EOF after the object, rejecting malformed payloads.

Also add tests for overflow detection and trailing data rejection.

* fix(s3api): cast math.MaxInt64 to int64 for 32-bit vet

On 32-bit platforms, math.MaxInt64 is an untyped int constant that
overflows int (32-bit) when used directly in fmt.Errorf with %d.
Cast to int64 explicitly to fix Go Vet 32-bit.

* docs: reconcile design doc with implementation and add AWS tools note

- Resolve open question about IAM endpoint path: driver accepts optional
  iamUrl and defaults to <s3Url>/iam
- Add note explaining ?seaweedfs-quota is not callable by standard AWS tools
  (aws s3api, s3cmd, rclone), and how this compares to MinIO and Ceph quota
  APIs which also live outside the standard S3 API

* docs: fix IAM endpoint default — SeaweedFS IAM is at POST / on S3 endpoint

SeaweedFS registers its embedded IAM API at POST / on the same S3
endpoint (UnifiedPostHandler), not under /iam. The design doc
previously said the driver defaults iamUrl to <s3Url>/iam, which would
send IAM operations to an unregistered path. Correct the default to
s3Url.

Found by Greptile review on PR #11279.

* docs: fix credential model, signer, and GET response shape in design doc

Three issues found by CodeRabbit review on PR #11279:

1. Credential-scope contradiction: the doc claimed the service credential
   is scoped to only s3:PutBucketQuota/s3:GetBucketQuota, but the
   implementation uses it as the admin credential for all operations
   (bucket CRUD, IAM user provisioning, quota). Document the actual
   model.

2. S3Signer -> AWSS3V4Signer: the doc said 'S3Signer for SigV4 signing'
   but S3Signer is legacy SigV2. Correct to AWSS3V4Signer.

3. GET response shape: the doc showed a single JSON example with 'GB'
   for both PUT and GET, but GET always returns quota_unit 'B' and the
   absolute byte count. Document PUT input and GET response separately.
2026-09-11 22:17:30 -07:00
Chris LuandGitHub 79994b69af s3: fail closed on unsupported bucket-policy condition operators (#11283)
* s3: support StringEqualsIgnoreCase and related condition operators

The S3 bucket-policy condition engine rejected StringEqualsIgnoreCase
(and StringNotEqualsIgnoreCase, StringLikeIgnoreCase,
StringNotLikeIgnoreCase), which AWS and the IAM policy engine both
accept. Add evaluators and register them in GetConditionEvaluator so
valid policies using these operators evaluate correctly instead of
being skipped.

* s3: reject bucket policies with unsupported condition operators

validateStatement did not check Condition operators, so a policy with
an unknown operator (e.g. a typo or unsupported key) was accepted at
upload time and only surfaced at evaluation, where it was silently
skipped. Reuse GetConditionEvaluator to reject unknown operators when
a policy is parsed or stored, failing closed at the entry point
instead of relying on evaluation-time handling.

* s3: fail closed on unsupported condition operators at evaluation

EvaluateConditions skipped statements whose condition operator was
unsupported, logging a warning and continuing. With no remaining
conditions to fail, the function returned true, so an Allow statement
conditioned on an unrecognized operator became unconditional and
granted access to private objects. Return false instead so an
unrecognized operator fails the condition block and the statement does
not match, matching the fail-closed behavior of the IAM policy engine.

* s3: validate condition operators at upload time only, not load time

Validating condition operators in validateStatement rejected the whole
policy document from ParsePolicy, which SetBucketPolicy uses when loading
stored bucket policies. A legacy policy saved before this change could
contain an unsupported operator, and rejecting it at load time dropped
the entire policy - including unrelated explicit Deny statements - so
the bucket lost its protections. Move the operator check into
ValidateBucketPolicy, which only the PutBucketPolicy handler and admin
UI run at upload time, so legacy policies still load and EvaluateConditions
fails the unsupported statement closed instead.

* s3: drop non-AWS StringLikeIgnoreCase and StringNotLikeIgnoreCase operators

AWS defines StringEqualsIgnoreCase and StringNotEqualsIgnoreCase but
not StringLikeIgnoreCase or StringNotLikeIgnoreCase (StringLike and
StringNotLike are case-sensitive only). Registering the wildcard
IgnoreCase variants made the engine accept operators AWS rejects. Keep
only the two AWS-defined IgnoreCase operators and add a test asserting
the wildcard IgnoreCase names are unsupported.
2026-09-11 22:17:11 -07:00
42b0ca7850 s3 sink: report the source read error the SDK hides (#11277)
* s3 sink: report the source read error the SDK hides

filer.backup stops for good on an event whose chunks are gone from the
volume servers: the uploader reads the body, the read fails with the
volume's 404, and the AWS SDK returns "ContentLength=N with Body length 0"
without the cause. isIgnorable404 would skip such an event, but it never
sees the 404, so the event is retried forever and the checkpoint never
advances.

ChunkStreamReader keeps its first source failure and the s3 sink returns
it when the upload fails.

* s3 sink: trim verbose comments on source error propagation

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-11 19:44:30 -07:00
github-actions[bot] 9b902a7662 docs: regenerate star history chart 2026-09-12 00:46:11 +00:00
d8aa7ecf04 fix(vacuum): stop comparing compact size against the live needle map (#11263)
* fix(vacuum): stop comparing compact size against the live needle map

CompactByIndex's post-copy integrity check compared bytes written to
the .cpd against v.nm.ContentSize()-DeletedSize(), the live map that
keeps mutating for as long as the volume stays writable during the
copy. Any write landing after the point-in-time index snapshot was
loaded made the live map's tally exceed what got copied, aborting
compaction with "unexpected new data size" — even though
CommitCompact's makeupDiff exists specifically to reconcile writes
that land mid-copy. On a busy volume this can fail every vacuum cycle.

Tally the expected live size from oldNm, the same frozen snapshot the
copy loop reads from, instead of the live map. This keeps the check's
original protection (destination smaller than what should have been
copied signals real data loss) while removing the false positive from
ordinary concurrent traffic.

* fix(vacuum): stop double-subtracting skipped bytes from the size check

Unreadable needles return before reaching the expectedLiveBytes tally,
so it already excludes them. Subtracting skippedDataBytes again on top
loosened the integrity check's margin by that same amount, letting a
.cpd short of the true expected size slip past undetected — the exact
failure mode the check exists to catch. Flagged independently by three
automated PR reviewers (Devin, Greptile, CodeRabbit).

Extract the comparison into exceedsExpectedCompactedSize and drop the
subtraction entirely; add TestExceedsExpectedCompactedSize to pin the
threshold to expectedLiveBytes alone.

* fix(vacuum): trim verbose integrity-check comment

Reduce the 8-line block comment to a concise 3-line rationale. No
behavior change.

* fix(vacuum): mirror compact integrity check in Rust volume server

Mirror the Go fix in the Rust volume server's do_compact_by_index:
tally expected_live_bytes from the frozen index snapshot (not the live
needle map) and compare the compacted .dat against it after the copy.
Unreadable needles already return before the tally, so no skipped-byte
adjustment is needed. Adds exceeds_expected_compacted_size and two
regression tests.

* fix(vacuum): exercise makeup_diff in Rust concurrent-write test

Address CodeRabbit review: write a needle after compaction (before
commit), then call commit_compact() and assert the late write survives
via makeup_diff. This actually exercises the concurrent-write path
rather than just confirming the integrity check passes.

---------

Co-authored-by: chrislusf <chris.lu@gmail.com>
2026-09-11 17:08:37 -07:00
Chris LuandGitHub 2ebfeabfce mount: rebuild expired directory cache on entry lookup (#11268)
* test: reproduce expired directory cache degrading lookup to N RPCs

After cacheMetaTtlSec elapses the kernel can still serve a directory
listing from its page cache, so ReadDir never runs and EnsureVisited is
not called. Metadata lookups then fall through to one LookupEntry RPC
per entry instead of rebuilding the directory cache once.

Issue #11262

* mount: add expired-directory rebuild predicate with cooldown to InodeToPath

ShouldRebuildExpiredDir distinguishes a TTL-expired cached directory from
a never-cached, invalidated, evicted, or read-through one (those clear
isChildrenCached, while a plain TTL expiry keeps it set). It also gates
retries on a cooldown since the last failed rebuild attempt, recorded by
MarkRebuildAttempt, so a transient listing failure does not trigger a
full rebuild on every later lookup.

Issue #11262

* mount: rebuild expired directory cache on entry lookup

When the kernel still serves a directory listing from its page cache past
cacheMetaTtlSec, ReadDir never runs and EnsureVisited is not called, so
lookupEntry issues one LookupEntry RPC per entry. Rebuild the expired
directory once via ensureDirectoryVisited before the cache-hit check so
later lookups are served locally. The EnsureVisited singleflight
deduplicates concurrent rebuilds.

On a non-oversized rebuild failure, record the attempt so the cooldown
suppresses repeated rebuilds while the listing keeps failing; once it
elapses a later lookup retries, recovering without waiting for ReadDir.
Oversized dirs are already marked read-through by ensureDirectoryVisited.

Issue #11262

* test: cover concurrent rebuild dedup and rebuild-cooldown fallback

Add a test that runs concurrent lookups into the same expired directory
behind a gated listing, asserting they share one rebuild via the
EnsureVisited singleflight. Add a test that a failed rebuild records the
attempt so an immediate retry is suppressed (per-entry RPC fallback), and
that once the cooldown elapses and the filer recovers a later lookup
rebuilds the cache.

Issue #11262

* mount: wait for pending async flush before rebuilding parent cache

The rebuild lists the parent directory from the filer, so a pending
async flush of the target entry must land first; otherwise the rebuilt
cache captures pre-flush metadata and the cache-hit path returns it
without the wait that guards the filer-fallback path. waitForPendingAsync
Flush is a no-op when no flush is pending, so the common case is unaffected.

Issue #11262
2026-09-11 11:54:10 -07:00
Chris LuandGitHub a3638e479e fix(s3api/audit): surface OIDC identity claim in audit log for STS sessions (#11269)
* Add ResolveIdentityClaim helper for OIDC audit identity

ComputeParentUser derives a stable per-identity hash from (sub, iss) for
internal keying, but it is opaque and not human-readable. Audit logs for
STS-assumed OIDC sessions currently surface that opaque value (or the
random session id) as the requester, leaving no authoritative trace of the
federated user.

Add ResolveIdentityClaim next to ComputeParentUser to recover a
human-readable, server-asserted identity attribute from the STS request
context populated at federation time. It walks a priority list
(preferred_username, email, name, sub) so a federated session always
audits against a stable OIDC claim rather than a client-supplied role
session name.

For #11264

* Surface authoritative OIDC identity claim in S3 audit log

For STS-assumed sessions minted from an OIDC web identity, the audit log
requester field is the opaque session subject, which cannot be traced back
to the federated user who performed the operation. The OIDC identity claims
(preferred_username, email, sub) are already carried in the session request
context and reach the auth layer as identity.Claims, but they were never
surfaced to the audit log.

Add a requester_identity field to the S3 access audit log, populated from
the authoritative OIDC identity claim resolved via ResolveIdentityClaim.
The claim is propagated through the shared identity holder (the same
mechanism the requester name and principal ARN already use) so it survives
the request-context copy that hides auth-set values from the outer audit
middleware.

The existing requester field is left unchanged for backward compatibility;
requester_identity is empty for non-federated sessions, where requester
already carries the real username.

For #11264

* Gate OIDC audit identity on federation marker and harden resolver

Address review feedback (Devin Review, Greptile) on the initial
implementation:

- Non-federated STS sessions no longer gain a false requester_identity.
  ValidateJWTWithClaims merges the JWT registered sub claim (the opaque
  session id) into RequestContext for sessions without an explicit request
  context, so the previous ResolveIdentityClaim fallback to sub surfaced
  that session id as an authoritative identity. Resolution is now gated on
  SessionInfo.ParentUser, which is set only for OIDC-federated sessions in
  AssumeRoleWithWebIdentity. The claim is resolved from the original
  sessionInfo.RequestContext (not the local claims map, whose sub the bearer
  path overwrites with the session subject) so SigV4 and bearer sessions
  surface the same identity.

- ResolveIdentityClaim now trims whitespace and treats whitespace-only
  claims as absent, so a blank preferred_username no longer masks a usable
  email or sub.

The resolved claim is carried on Identity.IdentityClaim (and IAMIdentity for
the bearer path) rather than re-derived in recordIdentityInContext, making
the federation gate explicit at the auth boundary.

For #11264

* Resolve OIDC identity claim for external bearer tokens

The external OIDC bearer path (a raw OIDC JWT presented directly, not via
STS) populates Claims with preferred_username/email/name/sub from the
validated token but did not set IdentityClaim, so requester_identity stayed
blank for that authentication path. Resolve the claim there too — sub is the
real OIDC subject on this path (not an STS session id), so no federation
gate is needed.

Also drop an ineffectual ctx assignment flagged by ineffassign in the audit
test.

For #11264
2026-09-11 11:29:46 -07:00
Chris LuandGitHub 80dae68dbf fix: write the new key when a remote-synced file is renamed (#11270)
* refactor: extract update event handling into processUpdateEvent

Pull the OldEntry/NewEntry update branch of the remote sync event
processor into its own function so the rename skip logic can be
exercised by tests with stub clients. No behavior change.

* test: reproduce remote sync rename dropping the new key

A rename under a remote mount arrives as an update whose NewEntry
inherits the source RemoteEntry. shouldSendToRemote returns false
for it, so processUpdateEvent skipped the event without writing the
new key, while the filer had already deleted the old object. The test
runs such an event through processUpdateEvent and expects both a
delete of the old key and a write of the new one. Fails before the
fix. See #11261.

* fix: write the new key when a remote-synced file is renamed

A rename under a remote mount arrives as an update whose NewEntry
inherits the source RemoteEntry, so shouldSendToRemote returns false
(RemoteMtime >= Mtime) and processUpdateEvent skipped the event.
That skip is only valid when the destination key is unchanged; a path
change always needs a write, and the delete-old/write-new handling
below the early return is exactly what a rename needs. Guard the skip
with proto.Equal(oldDest, dest) so a rename falls through to it.

Fixes #11261.

* fix: skip empty upload when renaming a remote-only entry

A remote-only entry (no local chunks or content, data lives only on
the remote object) carries a positive RemoteSize but nothing for
NewFileReader to read. After the previous commit lets a rename fall
through to the delete-old/write-new path, such a rename would upload
EOF and create a zero-byte object at the new key, then stamp it as
synced. Guard the write so a path change on a remote-only entry skips
the upload instead of replacing the file with zero bytes. The filer
has already deleted the old object, so the data is gone regardless;
this avoids leaving a misleading empty object behind.

* fix: propagate old-key delete errors except already-deleted

When deleting the old key on a rename fails for a non-multipart entry,
the error was swallowed and the write proceeded, which could leave both
remote keys. Return the error so MetadataProcessor retries the event.

The filer deletes the source remote object synchronously during the
rename, so the sync delete is redundant and the object may already be
gone. GCS reports that as ErrRemoteObjectNotFound (unlike S3/Azure,
whose deletes are idempotent), so treat it as a successful deletion and
continue to retriedWriteFile rather than pinning the sync offset.
2026-09-11 10:47:37 -07:00
Chris LuandGitHub 5ff49909a0 fix(s3api/iam): avoid transient AccessDenied from full reloads on single IAM file changes (#11271)
* fix(s3api/iam): fail config snapshot on empty or malformed IAM files

A full IAM reload reads every identity/policy/service-account/group file
from the filer. When an external secrets tool rewrites a file, a reload
that reads it mid-rewrite sees empty or partially-written content. The
identity, policy and service-account loaders silently skipped such files
(``continue``), so the snapshot was missing entries that still existed
on disk. The atomic swap then installed an incomplete identity set while
``isAuthEnabled`` stayed on, denying unrelated clients mid-reload
(#11259).

The group loader and the read-error paths already fail the snapshot in
this situation (a skipped entry reads as deleted). Apply the same
behavior to empty content and unmarshal failures across the identity,
policy, service-account and group loaders, so a transient mid-rewrite
fails the reload (preserving the last known-good state) instead of
silently dropping entries.

* fix(s3api/iam): coalesce burst IAM config reloads through the reload queue

onIamConfigChange did a full synchronous reload for every identity/policy
file change event. When several independently-refreshing credentials
rewrite their files within the same second, that produced a burst of
dozens of back-to-back full reloads, each reading the whole store and
widening the window where a mid-rewrite file is observed (#11259).

Route every IAM config change through the existing coalescing reload
queue (scheduleReload/reloadRetryLoop) instead. A burst of N events now
collapses into a single reload (plus one tail reload for events that
arrived while one was in flight). scheduleReload gains a reason argument
for the existing log line; the reloadRetryLoop already retries failed
reloads, so the per-event failure handoff is no longer needed.

Tests that asserted on the synchronous reload now wire up the queue
(centralized in newTestS3ApiServerWithMemoryIAM) and poll via
waitForIdentity/waitForIdentityGone. Adds TestOnIamConfigChangeCoalescesBurstReloads
showing 50 events coalesce into <=3 reloads.

* fix(s3api/iam): skip non-JSON auxiliary files before failing IAM snapshot

Per review: the multi-file loaders unmarshal every entry in an IAM
directory, so a non-JSON auxiliary file (README, .DS_Store, a migration
backup such as identity.json.old) would hit the new empty/malformed
errors and reject the whole snapshot, blocking all later IAM reloads.

Only *.json files are IAM objects (SeaweedFS writes identities,
policies, service accounts and groups as <name>.json, and other call
sites already gate on the .json suffix). Skip non-.json entries at the
top of each loader loop, before reading content, so auxiliary files are
ignored while empty/malformed .json files still fail the snapshot.

Adds TestLoadConfigurationIgnoresNonJsonAuxiliaryFiles.

* fix(s3api/iam): reject IAM files with empty identifiers and skip aux in listing

Per review:

- ListPolicyNames listed every regular entry in the policies directory as a
  policy name, including non-JSON auxiliary files, but GetPolicy cannot
  retrieve them. Apply the same .json suffix filter used by the loader so
  the list only exposes retrievable policies.

- json.Unmarshal accepts `{}` and unknown fields. The identity and group
  loaders merge by the decoded Name (not the file name), so a `{}` file
  could install an empty-key record and displace a real one; the
  service-account loader accepted an empty Id. Validate Identity.Name,
  Group.Name and ServiceAccount.Id (via validateServiceAccountId) after
  unmarshal and fail the snapshot on empty identifiers.

Adds TestFilerEtcStoreListPolicyNamesSkipsNonJsonAuxiliary and
empty-identifier regression tests for identity, group and service-account
files.
2026-09-11 10:42:19 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bc0efa4d10 build(deps): bump github.com/rclone/rclone from 1.75.0 to 1.75.1 (#11274)
Bumps [github.com/rclone/rclone](https://github.com/rclone/rclone) from 1.75.0 to 1.75.1.
- [Release notes](https://github.com/rclone/rclone/releases)
- [Changelog](https://github.com/rclone/rclone/blob/master/RELEASE.md)
- [Commits](https://github.com/rclone/rclone/compare/v1.75.0...v1.75.1)

---
updated-dependencies:
- dependency-name: github.com/rclone/rclone
  dependency-version: 1.75.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-11 09:42:47 -07:00
3ae9e332ec rust volume: honour is_last in the tail sender instead of rescanning the whole volume (#11273)
* rust volume: honour is_last in the tail sender instead of rescanning

volume_tail_sender discarded the is_last flag from
binary_search_by_append_at_ns:

    Ok((offset, _is_last)) => {
        if offset.is_zero() { Ok(sb_size) } ...

is_last means the caller is already caught up. Go answers that with a
heartbeat and does not scan at all (volume_grpc_tail.go, `if isLastOne`).
Dropping it is expensive rather than untidy, because the branches interact:
when the search reports caught-up it returns Offset::default(), which is
zero, so the start offset falls back to sb_size -- the beginning of the
data -- and scan_raw_needles_from materialises every needle from there to
EOF into a Vec. The timestamp filter discards all of it, the loop sleeps
2s, and it happens again.

A volume being moved is marked read-only before the copy, so it is ALWAYS
caught up during the tail phase. Measured on one volume.move of a 2.15 GB
volume, sampling the source's cgroup anon every 2s against the move's own
phase output:

  copying   16 -> 37 MB          CopyFile streams correctly, stays bounded
  tailing   904 -> 2166 -> 629 -> 2166 -> 342 -> 2173 -> 2179 MB
  deleting  46 MB

Six full-volume allocate/free cycles in 35s, peak 2179 MB against a volume
of 2147 MiB. The destination never exceeded 35 MB, so this is entirely
source-side. Under a per-process memory cap it OOM-kills the source
whenever the volume exceeds the cap.

The ordering here is the whole fix and is easy to get wrong: resolve the
start offset and is_last under a brief lock, return the heartbeat
immediately when caught up, and only then reach the scan. An earlier cut
set the flag correctly but placed the early return after the block that
performs the scan -- the heartbeat fired and the destination received
nothing, yet every iteration still read the whole volume and discarded it.
Production showed no improvement (1770 MB across five cycles), which is
what caught it. The binary search is over the .idx and costs nothing; the
scan is the expensive part and must not run speculatively.

Three tests, and the last two matter as much as the first: a fix that
always reported "caught up" would make tailing silently lose needles, a
worse bug than the one being fixed. One asserts is_last for a caller at or
beyond the newest append_at_ns; one asserts NOT is_last for a caller that
is behind, so real tail data is still scanned and shipped; one asserts NOT
is_last when the only newer record is a delete, and that scanning from the
returned offset ships exactly that tombstone.

Left deliberately unfixed, and worth separate changes: the scan still
collects into a Vec rather than streaming through a visitor as Go's
ScanVolumeFileFrom does, and it runs while holding store.read(), the same
lock-across-a-large-read shape as #11235. Both are latent once the rescan
is gone, since remaining scans are bounded by genuinely new data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MFr2v4BUqrXdgj4LEUAwVF
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: resolve and scan the tail under one store guard

The tail sender took store.read() once for the binary search and again
for the scan. A vacuum commit takes the store write lock and swaps
.dat/.idx, so it could land between the two: the offset resolved against
the old files would then be applied to the new ones and start the scan
inside an unrelated record. The code before the is_last fix held a
single guard for both. Restore that, and scan only when the caller is not
caught up, so the caught-up heartbeat still skips the scan and is sent
outside the lock.

Also pin the compacted-volume boundary raised in review. Compaction
writes .idx in needle-id order in both Go and Rust, so the search can
report caught-up while an earlier row is newer; such a caller's since_ns
is the last row's timestamp, so those rows were in the files it copied.
A write made afterwards is appended as the final row, which the search
cannot step past. The new test asserts it still reaches the scan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: make the compaction tail test a genuine overwrite

The compaction regression test's second id=1 write reused the first
write's data, so write_needle's dedup short-circuit (is_file_unchanged)
returned without appending or updating append_at_ns. Compaction then
kept key 1's original (older) timestamp, so the test passed without
exercising the overwrite it describes -- key 2 was the final row only
because key 1 was never actually newer.

Give the overwrite distinct data so it appends a new record, and assert
key1_ns > key2_ns up front so a future dedup regression fails the test
instead of silently hollowing it out. Trim the verbose comments on the
tail sender and the binary-search tests to their essentials.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-11 09:42:20 -07:00
Chris LuandGitHub 0de9c1f231 fix(s3api/sts): respect MaxSessionLength config in DurationSeconds validation (#11267)
* Refactor parseDurationSeconds into a STSHandlers method

Convert the parseDurationSeconds wrapper from a package-level function
into a method on STSHandlers so it can reach the configured STS service.
No behavior change; the three AssumeRole* handlers now invoke it via
their receiver.

* Respect MaxSessionLength config in STS DurationSeconds validation

parseDurationSeconds validated DurationSeconds against a hardcoded
43200s (12h) ceiling, so raising maxSessionLength in iam.json above
12h had no effect on AssumeRole, AssumeRoleWithWebIdentity, or
AssumeRoleWithLDAPIdentity — requests were rejected at the handler
before reaching the service layer.

Derive the upper bound from the configured STS MaxSessionLength,
falling back to maxDurationSeconds (43200s) when unset. The service
layer (calculateSessionDuration) already caps the issued duration
at MaxSessionLength, so this only relaxes the input-validation gate.

* Add tests for STS DurationSeconds MaxSessionLength bound

Cover the configured MaxSessionLength upper bound, rejection above
it, fallback to the 43200s default when STS config is unset, the
900s minimum, and the empty-parameter nil path.

* Refactor validateSessionDurationSeconds into a STSService method

Convert validateSessionDurationSeconds from a package-level function
into a method on STSService so it can reach the configured STS config.
No behavior change; the three assume-role entry points in the service
(AssumeRoleForPrincipal, validateAssumeRoleWithWebIdentityRequest,
validateAssumeRoleWithCredentialsRequest) now invoke it via their
receiver.

* Respect MaxSessionLength config in STS service DurationSeconds validation

The STS service validateSessionDurationSeconds rejected DurationSeconds
above a hardcoded 43200s (12h) ceiling, so even after the handler
accepted a longer duration it was rejected again in the service layer
for AssumeRoleForPrincipal, AssumeRoleWithWebIdentity, and
AssumeRoleWithCredentials.

Derive the upper bound from the configured MaxSessionLength, falling
back to DefaultMaxSessionLength (43200s) when unset. The issued
duration is still capped at MaxSessionLength by calculateSessionDuration.

* Add tests for STS service DurationSeconds MaxSessionLength bound

Cover the configured MaxSessionLength upper bound, rejection above
it, fallback to the 43200s default when STS config is unset, the
900s minimum, and the nil DurationSeconds path.

* Preserve capping when MaxSessionLength is below the API minimum

Deriving the DurationSeconds upper bound directly from MaxSessionLength
created an empty valid range when MaxSessionLength is configured below
the 900s API minimum, rejecting every explicit DurationSeconds that the
old code silently capped via calculateSessionDuration.

Only apply the configured MaxSessionLength as the upper bound when it is
at least minDurationSeconds; otherwise keep the default bound and let
calculateSessionDuration enforce the shorter configured limit.

* Add tests for sub-minimum MaxSessionLength capping behavior

Verify that a MaxSessionLength below the 900s API minimum keeps the
default upper bound so explicit DurationSeconds within the default
range are still accepted (and later capped by calculateSessionDuration).
2026-09-10 23:03:32 -07:00
github-actions[bot] b3aace2a08 docs: regenerate star history chart 2026-09-11 00:47:23 +00:00
4f9bbd51cb rust volume: stop glibc retaining freed EC buffers as unreturnable heap (#11255)
* rust volume: stop glibc retaining freed EC buffers as unreturnable heap

A Rust volume server doing EC work accumulates hundreds of MB of resident
anonymous memory that it never gives back, and under a hard cgroup
MemoryMax that ends in an OOM kill while most of the resident set is
free-but-unreturned.

It is not a leak. glibc serves allocations >= M_MMAP_THRESHOLD with mmap
and munmaps them on free, but the threshold is ADAPTIVE: freeing an
mmap'd block raises it toward that block's size, up to 32 MiB. EC
reconstruction and needle reassembly allocate large short-lived buffers,
so the first few train the threshold upward and every later buffer is
carved from the heap instead. Heap pages only return to the OS from the
top of the arena, so they stay resident for the life of the process --
reusable, but anonymous, and anonymous pages cannot be reclaimed under
pressure the way page cache can. The retained footprint is exactly the
headroom a burst of maintenance work needs.

Measured on a 17-node cluster (EC 10+4, --index=redb), one node, two
identical `ec.scrub -mode full` rounds over 10912 EC files each, same
unit restarted with and without a pinned threshold:

                       baseline  round 1  round 2  60s idle
  default (adaptive)      10 MB    84 MB    88 MB     88 MB
  pinned threshold        10 MB    13 MB    14 MB     14 MB

78 MB retained versus 4 MB for identical work. On heavier mixed scrub
workloads the same effect reached ~600 MB per volume server against a
3 GiB cap, and restarting the process was the only way to release it.

Calling mallopt(M_MMAP_THRESHOLD, ...) sets the threshold and disables
the dynamic adjustment. Pin it to glibc's own default rather than
inventing a value: the goal is to stop the adaptation, not to second-guess
the default. MALLOC_MMAP_THRESHOLD_ still wins if an operator sets it,
glibc-only, and a failed mallopt is logged rather than fatal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MFr2v4BUqrXdgj4LEUAwVF

* Address PR review: validate env overrides, honour GLIBC_TUNABLES, fix non-glibc test compile

Three review-bot findings on seaweed-volume/src/malloc_tuning.rs:

1. (CodeRabbit) The test used cfg!(...), which keeps both branches in
   compilation. On non-glibc targets DEFAULT_MMAP_THRESHOLD is undefined,
   so the test failed to compile. Split into #[cfg]-gated tests so each
   branch only references items defined for that target.

2. (Greptile) MALLOC_MMAP_THRESHOLD_ was checked by presence only. An
   empty or non-numeric value makes glibc ignore the override while we still
   skipped mallopt, leaving the adaptive threshold enabled -- exactly the
   behaviour this module exists to prevent. Now we defer only when the value
   is non-empty and parses as an integer; otherwise we fall through to
   pinning.

3. (Codex) The modern GLIBC_TUNABLES=glibc.malloc.mmap_threshold=... tunable
   was missed, so mallopt could overwrite an operator's explicit tunable. Now
   we detect that tunable (with the same validation) and defer to it.

The override check moved into the glibc-gated inner function, so off glibc
pin_mmap_threshold() always reports NotApplicable regardless of any
allocator env vars that happen to be set. The startup log for DeferredToEnv
is reworded to cover both override sources. Added tests for the override
parsers and the off-glibc no-op.

* Address round-2 review: match glibc's actual override parsing

Three follow-up review-bot findings after the first round of fixes, all
rooted in our validation not matching how glibc actually parses the
overrides:

1. (Greptile, P1) parse::<i64>() accepted negative values like "-1" and
   returned DeferredToEnv, but glibc's threshold is unsigned and rejects
   negatives — so we skipped mallopt while glibc also ignored the override,
   leaving the adaptive threshold enabled. Now we reject negatives and
   zero.

2. (Devin, BUG) glibc parses thresholds as unsigned (strtoul for tunables,
   atoi for the legacy var). Values above i64::MAX are valid for glibc but
   were rejected by parse::<i64>(), so we pinned 128 KiB over the operator's
   explicit setting. Now we parse as u64, accepting the full unsigned range.

3. (CodeRabbit, Major) Two issues in usable_glibc_tunable_threshold:
   a. A malformed sibling entry (e.g. glibc.malloc.check=2=2:...) makes
      glibc reject the entire GLIBC_TUNABLES string, but our per-entry scan
      still returned true for the valid-looking mmap_threshold entry. Now
      we validate every entry (exactly one '=') before accepting any.
   b. Hex values (0x20000) are accepted by glibc's strtoul but were rejected
      by parse::<i64>(). Now parse_strtoul_threshold handles 0x-prefixed hex.
      MALLOC_MMAP_THRESHOLD_ stays decimal-only (atoi), matching glibc.

Added regression tests for negatives, zero, >i64::MAX, hex tunables, and
malformed mixed GLIBC_TUNABLES entries. Verified: clippy clean and tests
pass on macOS (non-glibc); glibc-gated code type-checks for
x86_64-unknown-linux-gnu.

* Address round-3 review: match glibc's actual override parsing

Three follow-up review-bot findings (Greptile P1, Devin BUG, CodeRabbit
Major) all on the same issue: the round-2 fix rejected negative and zero
override values, but glibc actually accepts them.

Verified against the glibc source (malloc/malloc.c, malloc/arena.c,
elf/dl-tunables.c, elf/dl-misc.c):

- do_set_mmap_threshold(size_t value) does NO clamping — it just sets
  mp_.mmap_threshold = value and mp_.no_dyn_threshold = 1.
- MALLOC_MMAP_THRESHOLD_: glibc calls atoi(value) then mallopt, which
  always sets the threshold and disables dynamic adjustment — even for
  empty, negative, or non-numeric values (atoi returns 0). So ANY
  presence of the variable means the operator's override is in effect.
  Reverted to presence-only check for the legacy variable. The round-1
  Greptile comment claiming glibc "cannot apply the override" for
  empty/malformed values was incorrect.
- GLIBC_TUNABLES: glibc parses values with _dl_strtoul (elf/dl-misc.c),
  which accepts decimal, 0x hex, 0 octal, an optional sign (negatives
  wrap to unsigned long), and requires the entire value consumed
  (tunable_parse_num checks endptr == strval + len). Replaced
  parse_strtoul_threshold with dl_strtoul_consumes_all that replicates
  _dl_strtoul's parsing and checks full consumption. Now accepts -1
  (wraps to SIZE_MAX), 0, 0x20000, 010 (octal), and values above
  i64::MAX.

The duplicate-= validation for GLIBC_TUNABLES (from round 1) is kept —
glibc's parse_tunables_string returns -1 if any entry's value contains
a duplicate =, rejecting the entire string.

Added dl_strtoul_consumes_all tests covering decimal, hex, octal,
negative, zero, empty, whitespace, trailing garbage, and sign-only
inputs. Updated usable_glibc_tunable_threshold tests to accept
negative, zero, and empty values. Verified: clippy clean and tests
pass on macOS (non-glibc); glibc-gated code type-checks and clippy
clean for x86_64-unknown-linux-gnu.

* Address round-4 review: add overflow detection, fix sign-only test assertions

Two Greptile P1 findings:

1. Overflowing tunables bypass threshold pinning: dl_strtoul_consumes_all
   consumed every digit and returned true for values like
   18446744073709551616 (u64::MAX + 1), but glibc's _dl_strtoul stops at
   the overflowing digit (sets endptr there, returns UINT64_MAX), so
   tunable_parse_num rejects the value (endptr != strval + len). Added
   overflow detection matching glibc's cutoff/cutlim logic — on overflow,
   the parser stops and returns false.

2. Sign-only parser assertions fail: the test asserted
   !dl_strtoul_consumes_all("-") and !dl_strtoul_consumes_all("+"), but
   _dl_strtoul skips the sign, finds no digit, sets endptr to the position
   after the sign (== end of string), and returns 0. tunable_parse_num
   sees endptr == strval + len → true. So glibc accepts sign-only strings
   as value 0. Fixed the test assertions to expect true.

Also fixed "0x" with no hex digits: _dl_strtoul parses "0" as octal, then
stops at "x" (not an octal digit), so endptr != end of string → rejected.
The base-detection now requires a hex digit after "0x" before switching
to hex; otherwise "0" is parsed as octal and "x" stops the parser.

Added overflow regression tests: 18446744073709551616 (u64::MAX + 1),
99999999999999999999 (20 nines), 0x10000000000000000 (2^64). Verified:
clippy clean and tests pass on macOS (non-glibc); glibc-gated code
type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Address round-5 review: accept bare 0x prefix, remove unused helper

Two review-bot findings (Devin BUG + CodeRabbit Major) on the same issue:
the round-4 fix required a hex digit after "0x" before switching to hex
base, but glibc's _dl_strtoul unconditionally advances past "0x"/"0X"
when the first char is '0' and the next is 'x'/'X' — even if no hex digit
follows. In that case the digit loop breaks immediately, endptr reaches
the end, and the value is 0. tunable_parse_num accepts it.

Removed the is_digit_in_base lookahead from the base-detection condition
and the now-unused is_digit_in_base helper. Updated the test assertions
for "0x" and "0X" to expect true (accepted as value 0).

The Greptile P1 overflow comment is invalid: glibc's _dl_strtoul rejects
18446744073709551616 (u64::MAX + 1) — on overflow it sets endptr to the
overflowing digit (not end of string) and returns UINT64_MAX, so
tunable_parse_num sees endptr != strval + len and rejects. My
implementation correctly returns false for this value, matching glibc.

Verified: clippy clean and tests pass on macOS (non-glibc); glibc-gated
code type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Address round-6 review: rewrite tunable parser to match glibc exactly

Two Greptile P1 comments (3975151906, 3975151911) both invalid, but
investigation revealed a real bug in the split(':')-based parser:

Bug: usable_glibc_tunable_threshold used split(':') which loses the
distinction between an entry terminated by ':' (glibc skips it) and one
terminated by '\0' with no '=' (glibc rejects the entire string). Examples:
  - "glibc.malloc.mmap_threshold=262144:glibc.cpu.x" (no '=' at end):
    glibc rejects entire string, old code accepted it.
  - "glibc.malloc.mmap_threshold=262144:" (trailing ':'):
    glibc rejects entire string, old code accepted it.

Fix: replaced split(':') with a character-by-character parser matching
glibc's parse_tunables_string exactly. The parser tracks position in the
original string and correctly handles all three terminators ('=', ':', '\0')
for both name and value scanning.

Comment 3975151906 (near-maximum values): Invalid. Verified against
_dl_strtoul: for 18446744073709551615 (u64::MAX), cutoff = u64::MAX/10,
cutlim = u64::MAX%10 = 5. After 19 digits result == cutoff. 20th digit 5:
overflow check (digval > cutlim) is 5 > 5 = false → no overflow. glibc
accepts u64::MAX. Added regression test asserting it's accepted.

Comment 3975151911 (later malformed entry): Invalid. Verified against
parse_tunables (elf/dl-tunables.c): when parse_tunables_string returns -1,
parse_tunables prints a warning and returns immediately without applying
ANY tunable — including ones already parsed into the array. Added
regression test for "threshold=262144:check=2=2" (threshold before
malformed sibling) asserting it's rejected.

Added regression tests: u64::MAX accepted, threshold-before-malformed
rejected, no-'=' at end rejected, trailing ':' rejected, leading ':'
accepted. Verified: clippy clean and tests pass on macOS; glibc-gated
code type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Fix CI: correct hex trailing-garbage test assertion

The test asserted !dl_strtoul_consumes_all("0x20000abc"), but in hex
mode a-f are valid digits — "0x20000abc" is a valid hex number
(0x20000abc = 536874044), not trailing garbage. _dl_strtoul consumes
the entire string and tunable_parse_num accepts it. The assertion
failed on Linux CI where the glibc-gated test actually runs.

Replaced with "0x20000g" — 'g' is not a hex digit, so _dl_strtoul
stops at 'g' and tunable_parse_num rejects the value.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-10 09:29:31 -07:00
e919bec9d1 fix(volume): Harden Volume Copy Validation and Failure Handling (#11252)
* fix(volume): harden volume copy validation

* fix(volume): use stream context for ReadVolumeFileStatus in VolumeCopy

ReadVolumeFileStatus ran on context.Background() while the adjacent
VolumeStatus call used stream.Context(), an inconsistency left over
from the context revert in #11252. Use stream.Context() consistently
so the source status check is cancelled with the VolumeCopy stream.

* fix(volume): reserve destination before deleting existing replica

FindFreeLocation now runs before DeleteVolume so a full target fails
without destroying the existing replica. Previously, when the initial
VolumeStatus check failed (advisory) but ReadVolumeFileStatus
succeeded, the existing replica was deleted before a destination was
reserved, risking data loss if no location had enough free space.

Add a regression test verifying the existing replica survives when
the destination is full and the initial status check fails.

* fix(volume): count replaced replica slot in FindFreeLocation

FindFreeLocation now accepts the volume being replaced so its slot is
treated as available. Without this, a location at its MaxVolumeCount
limit could not replace its sole replica even though deleting it would
free the slot. VolumeCopy passes the volume ID so destination selection
succeeds before the existing replica is deleted.

Add TestVolumeCopyReplacesReplicaAtSlotLimit covering a single-slot
location that must replace its only replica.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 23:33:05 -07:00
2cd6c36c54 filer: end local-only metadata subscriptions when remote peers appear (#11251)
* filer: end local-only metadata subscriptions when remote peers appear

SubscribeMetadata delegates to SubscribeLocalMetadata whenever the
MetaAggregator knows no remote peers at stream setup. Peer discovery is
asynchronous with the gRPC server accepting streams: the master announces
filers after Filer.Init, via ListExistingPeerUpdates and OnPeerUpdate.
A subscriber that connects inside that window is pinned to a filer-local
stream for its whole life, silently missing every other filer's writes.
For filer.remote.sync in a multi-filer cluster this means the remote tier
permanently stops receiving writes served by other filers (#11247).

End the delegated local stream when the first remote peer appears, so the
client reconnects into the aggregated stream. The end surfaces as an
error, not a clean EOF: RetryUntil-driven followers (mount meta cache,
s3api IAM) treat a clean end as following finished and stop
reconnecting. The arrival channel is armed under the same lock as the
peer check in RemotePeerArrivedChan, so a peer learned in between sends
the stream straight to the aggregated path instead of parking on a
channel that would never fire.

A standalone filer is unaffected: no peer ever appears, the channel
never fires, and the local stream serves indefinitely.

Fixes #11247

* filer: interrupt disk replay on peer arrival, trim comments

Check upgradeOnRemotePeer inside eachLogEntryFn and chunkDiskPass so a
peer arriving during a backlog replay stops the stream before the
cursor advances past older remote events. Wrap errAggregationUpgrade
with StopReadingError so LoopProcessLogData does not log it. Remove
issue references from comments and trim verbose commentary.

* filer: check upgrade signal between ref batches

Pass upgradeOnRemotePeer to sendRefsBatched so a peer arriving while
refs are shipped to a slow client is detected between batches, not
only after the full batch completes.

* filer: interrupt gap park on peer arrival

Pass upgradeOnRemotePeer through gapPass to parkOnGap so a peer
arriving during a gap park ends the stream immediately instead of
waiting for the retry timer (up to one minute).

---------

Co-authored-by: Tyagiquamar <Tyagiquamar@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 20:16:50 -07:00
Chris LuandGitHub 7fa2f75f30 s3: bucket-policy Allow must not override an identity explicit Deny (#11256)
* s3: add isActionExplicitlyDeniedByApplicablePolicies helper

Add a helper that reports whether any applicable identity-side policy
(attached IAM policies, enabled-group policies, or the IAM-integration
session policy) explicitly denies an action. It reuses the existing
evaluateAttachedIAMPolicies, resolveS3AuthTarget, buildPrincipalARN, and
isActionExplicitlyDeniedByIAM helpers, and fails closed on evaluation
errors. A nil identity has no identity-side policy plane, so the helper
returns false to keep the bucket policy authoritative for anonymous
access. No behavior change yet; the next commits apply it to the two
bucket-policy Allow short-circuits.

* s3: enforce identity explicit Deny before bucket-policy Allow

authRequestWithAuthType short-circuits on a matching bucket-policy Allow
and skips VerifyActionPermission, so an explicit Deny in an authenticated
identity attached, group, or session policy is bypassed. A non-admin
principal with s3:PutBucketPolicy can install a bucket-policy Allow for
itself and read an object its identity policy explicitly denies.

Before honoring a bucket-policy Allow, check the applicable identity-side
policies for a matching explicit Deny via the new
isActionExplicitlyDeniedByApplicablePolicies helper, and fail closed.
The cross-account behavior is preserved: a bucket Allow still supplies the
Allow an identity policy omits (implicit denial), and a nil identity keeps
the bucket policy authoritative for anonymous access.

Regression tests cover the explicit-Deny override, the implicit-deny
Allow preservation, and the unmatched-key fall-through control.

* s3: enforce identity explicit Deny in secondary object-key auth

authorizeObjectKeyAction authorizes keys the request URL does not name
(CopySource, DeleteObjects body keys, POST Object form keys) and shares
the same bucket-policy Allow short-circuit as the primary path, so an
explicit Deny in the identity, group, or session policy is bypassed the
same way when a bucket policy allows the secondary key.

Apply isActionExplicitlyDeniedByApplicablePolicies before accepting the
bucket-policy Allow, mirroring the primary path. A regression test covers
AuthorizeCopySource for both the explicit-Deny override and the
implicit-deny Allow preservation.
2026-09-09 20:15:43 -07:00
github-actions[bot] 5061a16b12 docs: regenerate star history chart 2026-09-10 02:39:04 +00:00
Chris LuandGitHub 3c9a4bbdda rust: prevent phantom volumes + validate collection hint in mount_volume_by_id (#11254)
* rust: prevent phantom volumes + validate collection hint in mount_volume_by_id

The collection-hint path (and the find_volume_file_base fallback) called
create_volume on any matching .vif/.idx sidecar. create_volume ->
Volume::new -> load(create_dat_if_missing=true) writes an empty .dat and
registers a phantom normal volume, which can shadow a real EC volume whose
.ecx lives on a sibling disk. This reintroduces the phantom-volume bug the
codebase explicitly guards against in load_existing_volumes.

Apply the same guard load_existing_volumes uses to both paths: only mount
when a real .dat is present or the .vif references a remote-tiered file;
otherwise skip the candidate (no phantom). Also reject path-bearing
collection hints ('/', '\\', '..') so the shortcut cannot route .dat
creation outside the storage directory, falling back to the safe scan.

Adds 3 regression tests; all 336 storage:: tests pass.

Addresses Devin + Greptile review comments on PR #11249.

* rust: address review — .note guard, multi-candidate scan, foo..bar hint

Address the four review comments on #11254:

1. Greptile (P1): contains("..") rejected valid collections like "foo..bar".
   Replaced with collection != ".." — volume_file_name joins with "_" so a
   ".." inside a name is part of the filename, not a parent reference. Only
   the exact ".." name is rejected. Added a test that "foo..bar" mounts.

2. Devin #0001 (bug): mount_volume_by_id did not check the .note marker, so
   an interrupted VolumeCopy could mount as a live (truncated) volume. Added
   a .note check before create_volume in both the collection-hint path and
   the fallback — a candidate with .note is skipped (matches
   load_existing_volumes). Added a test covering both paths.

3. Devin #0002 + CodeRabbit (major): find_volume_file_base returned only the
   first matching candidate, so a lone sidecar on disk 0 hid a real .dat on
   disk 1 (the split-disk EC layout the phantom guard protects against).
   Added find_volume_file_bases (plural) that collects all candidates; the
   fallback now iterates every candidate and mounts the first with a real
   .dat or remote .vif. find_volume_file_base delegates to it for
   configure_volume. Added a two-disk test: sidecar on disk 0, real .dat on
   disk 1 — mount succeeds from disk 1.

All 339 storage:: tests pass (6 mount_volume_by_id tests).

* rust: continue past create_volume failure in mount_volume_by_id

Address Devin review comment on #11254: when create_volume fails on an
earlier candidate (e.g. an unreadable .dat), mount_volume_by_id returned
the error immediately instead of trying later candidates. A valid volume
on another disk remained unmounted.

Both the collection-hint loop and the find_volume_file_bases fallback now
remember the last error and continue scanning. A successful mount returns
immediately; if no candidate succeeds, the last error (or NotFound) is
returned. Matches DiskLocation::open_volumes and Go Store.mountVolume.

Added test_mount_volume_by_id_continues_past_open_failure (chmod 000 .dat
on disk 0, real volume on disk 1, mounts from disk 1).

All 340 storage:: tests pass.
2026-09-09 14:54:09 -07:00
Feng ShaoandGitHub 13bf056a15 Mount req with collection (#11249)
* volume mount req support specify collection

* rust mirror change
2026-09-09 12:55:47 -07:00
Nguyễn Đăng Minh LựcandGitHub c968084b34 iceberg: fix OAuth token expiry handling (401 + token-exchange + configurable TTL) (#11242)
* iceberg: return 401 for invalid or expired Bearer tokens

BUG-0001: when the OAuth JWT expired, Server.Auth fell through to the S3
SigV4 authenticator, which rejects the "Authorization: Bearer" scheme
with NotImplemented — a 501. Iceberg clients (Java OAuth2Manager,
pyiceberg) only refresh tokens on 401, so they retried the dead token
forever: RisingWave sinks stalled and Doris catalog queries failed every
token TTL (1h) until the client process was restarted.

A request carrying a Bearer header is an Iceberg REST client: answer 401
(+ WWW-Authenticate: Bearer, RFC 6750) when the token fails, and only
fall through to the S3 authenticator when no Bearer header is present.

* iceberg: make OAuth token TTL configurable via ICEBERG_OAUTH_TOKEN_EXPIRY

BUG-0001 follow-up: production evidence shows Iceberg Java 1.10.x
clients (RisingWave connector node, Doris FE) never re-fetch tokens on
401 — the sink stalled again on token expiry even with the 501→401 fix,
and no POST /v1/oauth/tokens appeared in server logs across dozens of
retries. 401 is necessary but not sufficient for these clients.

The TTL was hardcoded to 3600 with no knob. Read the expiry (seconds)
from ICEBERG_OAUTH_TOKEN_EXPIRY, defaulting to 3600, so deployments can
issue longer-lived tokens (e.g. 86400) to survive client restart cycles.

* iceberg: support OAuth token exchange (RFC 8693) for client refresh

Decompiling the Iceberg Java 1.10.1 client bundled with Doris FE showed
the missing half of BUG-0001: OAuth2Manager refreshes via token-exchange
(AuthConfig.exchangeEnabled defaults to true — the client_credentials
re-fetch branch only runs with exchange disabled), so a server that only
accepts client_credentials leaves Iceberg clients unable to ever refresh
their token, regardless of 401 correctness.

Accept grant_type=urn:ietf:params:oauth:grant-type:token-exchange on
POST /v1/oauth/tokens: verify the subject_token signature against the
issuing credential, allow exchange within a recovery grace window
(max(2*TTL, 1h), capped 24h) so clients holding tokens that expired
while the grant was unsupported recover without a restart, and mint a
fresh access token with the configured TTL.

* iceberg: harden OAuth token exchange and Bearer matching per review

- match the Bearer scheme case-insensitively (RFC 7235), like
  authenticateBearer already does
- accept optional client authentication on the token-exchange grant
  (Basic or form credentials, bound to the subject token's client);
  expired subject tokens now require it. Iceberg Java's proactive
  refresh sends Bearer-only headers, so the grant cannot require it
- reject subject tokens without an exp claim, and re-check the issuer
  on the verified claims
- unauthenticated exchange cannot extend the lifetime past the
  subject token's own expiry (no chain-refresh from a leaked token)
- return 400 invalid_grant per RFC 6749 §5.2 (was 401)
- include issued_token_type on exchange responses (RFC 8693)
- clamp ICEBERG_OAUTH_TOKEN_EXPIRY to 365d so Duration math cannot
  overflow into already-expired tokens

* iceberg: give authenticated token exchanges a fresh full TTL

The remaining-lifetime cap only guards unauthenticated (Bearer-only)
exchanges; an authenticated client renewing a live token must get the
full configured TTL, matching client_credentials.

* iceberg: reject token exchange when no lifetime remains

A Bearer-only exchange with under a second of subject lifetime would
mint a token with expires_in: 0. Reject with invalid_grant instead.

* iceberg: pin near-expiry test token to the next second boundary

jwt/v5 serializes exp at one-second precision, so a 300 ms offset can
round into the current second and route the test through the expired
branch instead of the ttlSeconds<=0 guard. Mint the subject with the
next whole-second expiry: live at exchange time, deterministically
under a second of remaining lifetime.

* iceberg: drop internal ticket reference from comments

* iceberg: clamp oversized OAuth TTLs on 32-bit platforms

strconv.Atoi on an int-sized value fails with ErrRange on 386, so an
oversized ICEBERG_OAUTH_TOKEN_EXPIRY silently fell back to the default
instead of clamping. Parse in 64-bit space and clamp, then narrow.

* iceberg: make OAuth TTL narrowing explicit

* iceberg: disable legacy OAuth in PyIceberg integration tests
2026-09-09 10:54:39 -07:00
516e251f9e rust volume: move the crate to edition 2024 (#11244)
* rust volume: move the crate to edition 2024

Edition 2024 turns three things in this crate into hard errors, and changes
drop order in a further 34 places without changing compilation. The compiler
errors are fixed here; the silent changes were audited against
`RUSTFLAGS='-W rust-2024-compatibility' cargo check --all-targets` output
captured before the flip, since edition 2024 stops reporting them.

`std::env::set_var`/`remove_var` are unsafe as of 2024 because they race with
concurrent readers. All six call sites are safe by construction rather than by
assertion, and the SAFETY comments say why: the build script runs
single-threaded before anything else in the process, and every test reaching
the `config.rs` helpers holds `process_state_lock()` for the duration.

The two `ref` bindings in handlers.rs sit in patterns that already borrow
implicitly, so removing the modifier leaves both bindings at `&String`.

On the 34 drop-order sites: no lock guard's scope is extended anywhere, and
`volume.rs` has none. Most are moved-from `Option`/`Result` husks — `if let
Some(v) = map.remove(&k)`, `while let Some(m) = stream.next().await` — where
the value is moved into the binding and the temporary has nothing left to drop;
where closing order actually matters these paths already call `v.close()`,
`ec_vol.destroy()` or `drop(writer)` explicitly. Two sites get strictly better
ordering: the metrics read guard in `run_metrics_push_loop` shrinks to the end
of its initializer block (it never crossed an `.await` either way), and an EC
test now closes the volume's descriptors before the `TempDir` removes the
directory.

No `rust-version` is declared. Edition 2024 needs rustc 1.85, but that is not
the binding constraint — the dependency tree already requires 1.91.1 through
the `aws-sdk-s3`/`aws-smithy-*` family, so `cargo +1.85 check` fails on the
deps regardless. CI builds on `dtolnay/rust-toolchain@stable`.

`vendor/reed-solomon-erasure` is a separate package and keeps edition 2021.
Cargo.lock is unchanged despite edition 2024 implying resolver 3.

Verified: `cargo test` 551 passed / 0 failed, `cargo test
--no-default-features` 550 passed / 0 failed (the two feature sets produce an
identical migration site list), `cargo build --release` clean. No automated
test covers shutdown ordering, so the channel and runtime sites in `main.rs`,
`write_queue.rs` and `grpc_server.rs` were read individually.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust volume: address edition-2024 review feedback

Three fixes from review of the edition bump.

Serialize the two environment-reading tests. The SAFETY comments on the
`env::set_var`/`remove_var` helpers claim every test touching the environment
holds `process_state_lock()`, but `test_resolve_config_defaults_dir_to_platform_temp_dir`
and `test_resolve_config_index_accepts_redb_and_leveldb_aliases` called
`resolve_config` — which reads HOME/USERPROFILE, SEAWEED_WRITE_QUEUE and the
WEED_* set — without taking it. `set_var` is unsafe precisely because a
concurrent *reader* is UB, not only a concurrent writer, so the comment was
overclaiming. An audit of the module found exactly these two; every other
environment-touching test already held the lock. The race predates edition
2024, which only made the requirement explicit.

Declare `rust-version = "1.91.1"`. The edition needs 1.85, but that was never
the binding constraint: `cargo +1.90 check --all-targets` fails on the
`aws-sdk-s3`/`aws-smithy-*` family, and 1.91.1 checks clean. Declaring the
verified floor turns a wall of per-dependency errors into one clear message.
Cargo.lock is unchanged despite this making the resolver MSRV-aware.

Update the README, which advertised "Rust 1.75+ (2021 edition)". 1.75 was
already stale before this branch — the tree has needed 1.91 for a while.

Verified: `cargo test` 551 passed / 0 failed, `cargo test --no-default-features`
550 passed / 0 failed, `cargo build --release` clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust volume: state the exact MSRV patch release in the README

The README said "Rust 1.91+", which reads as 1.91.0 and is wrong by one patch
release: `cargo +1.91.0 check --all-targets` fails on the aws-sdk-s3 family,
`cargo +1.91.1` passes. Say 1.91.1+, matching `rust-version` in Cargo.toml, and
call out that the patch component is load-bearing so nobody installs 1.91.0 and
hits the same wall.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust worker: move the workspace to edition 2024

Moves the seaweed-worker workspace (core, lance, sort) from edition 2021 to
2024, the same migration seaweed-volume just got in this branch.

Edition 2024 turns exactly one thing in this workspace into a hard error. The
baseline came from RUSTFLAGS='-W rust-2024-compatibility' cargo check
--all-targets, run before the flip; unlike seaweed-volume's 34 silent +
8 hard sites, the worker reports only the one hard site and no
tail_expr_drop_order or if_let_rescope sites at all. The worker is a much
smaller crate and none of its expressions hold a guard or temporary whose drop
order the edition changes, so there is nothing to audit on the silent side.

Fixed (1 site):

std::env::set_var is unsafe as of 2024 because it races with concurrent
readers. The single call is in crates/core/build.rs, which sets PROTOC from
protoc_bin_vendored the way seaweed-volume's build script does. A build
script's main runs single-threaded before anything else in the process, so
no other thread can be reading the environment concurrently; the SAFETY comment
says so. There are no config.rs-style test helpers here -- the worker's tests
do not mutate the environment -- so unlike the volume crate there are no
process_state_lock() callers to audit.

No redundant ref bindings to clean up: a grep for ref across the three
crates finds none.

MSRV:

rust-version = "1.94.1", verified rather than inferred. Edition 2024 only
needs 1.85, but the dependency tree needs more: lance's aws feature pulls in
a newer cut of the same aws-sdk-*/aws-smithy-* family that sets
seaweed-volume's 1.91.1 floor, and that newer cut requires 1.94.1.
cargo +1.94.0 check --all-targets fails on that family; cargo +1.94.1
check --all-targets is clean. The worker's floor is therefore higher than
the volume's, and moves with lance and the AWS SDK rather than with the
edition. CI builds on dtolnay/rust-toolchain@stable, so nothing changes
there.

The edition is set once in [workspace.package] and inherited by each member
via edition.workspace = true; rust-version is added the same way. The
workspace keeps its explicit resolver = "2" -- edition 2024 would default to
resolver 3, but the pin is deliberate and Cargo.lock is unchanged by this
commit either way.

The README gains a "Requires Rust 1.94.1+ (2024 edition)" line in its Building
section, matching the one seaweed-volume's README now carries, and calling out
that the patch release is load-bearing (1.94.0 does not build) so nobody
installs 1.94.0 and hits the same wall.

Verification:

* cargo check --all-targets -- clean, zero warnings (default toolchain 1.97)
* cargo +1.94.1 check --all-targets -- clean
* cargo +1.94.0 check --all-targets -- fails on the AWS SDK, as claimed
* cargo test --all-targets -- 40 passed, 0 failed
  (core 13, sort 11, lance lib 3, lance bin 2, compaction 6, lifecycle 1,
  sort integration 4)
* Cargo.lock unchanged

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 10:54:06 -07:00
Chris LuandGitHub 01fc31cb71 fix(filer): use path.Split instead of filepath.Split for filer paths (#11246)
FullPath.DirAndName() and FullPath.Name() used filepath.Split, which is
OS-dependent: on Windows it treats backslash as a path separator,
corrupting filer paths that contain literal backslashes. Filer paths
always use "/" as the separator, so switch to path.Split and path.Join
which only split on "/" regardless of the host OS.

This fixes the backslash case from #11243 where a file saved as
/test/special\reverseslash4.jpg was stored with a corrupted path on
Windows filer builds. The #, ?, and % cases from the same issue are
client-side URL-encoding problems (the server never receives the raw
characters), but once the client properly percent-encodes them the
server now handles the decoded path correctly on all platforms.
2026-09-09 10:43:16 -07:00
ssshr-66andGitHub 966692fa23 [Volume] Validate record counts after volume copy (#11238)
* validate Volume Copy record counts

* Delete s3api_object_versioning_bench_test.go

* reply ai comments
2026-09-09 02:18:18 -07:00
github-actions[bot] 168b9c39f8 docs: regenerate star history chart 2026-09-09 05:10:20 +00:00
Chris LuandGitHub f1f6886d0e fix: rebase on latest master before pushing star history chart (#11240)
* fix: rebase on latest master before pushing star history chart

The daily star history workflow git push was rejected with a
non-fast-forward error because new commits landed on master between
the checkout and the push. Fetch full history (fetch-depth: 0) and
rebase the generated commit on top of the latest remote branch before
pushing so the workflow no longer fails when master has moved.

* fix: retry rebase-and-push to handle concurrent master updates

Address review feedback: a one-shot rebase still races if master advances
between the rebase and the push. Match the bounded retry loop used by
java_release.yml — push first, and on rejection rebase and retry up to
five times before failing.

* fix: serialize runs and ensure every rebase is followed by a push

Address review feedback:
- Devin (line 54-55): the old loop rebased after the 5th failed push but
  never pushed the rebased commit. Restructure so each rebase (attempts
  2-5) is followed by a push attempt, with a clear 5-attempt cap.
- Greptile: overlapping runs could conflict on the SVG during rebase.
  Add a concurrency group (cancel-in-progress: true, matching the repo
  convention) so only one chart regeneration runs at a time.

* fix: scope concurrency by ref and guard rebase against transient failures

Address review feedback:
- Devin (line 16): the global star-history concurrency group let a manual
  run on another branch cancel an in-flight daily master update. Scope
  the group by github.ref so only same-branch runs cancel each other.
- Greptile (line 58): git pull --rebase runs under the fail-fast shell,
  so a transient fetch error or conflict aborted the whole step before
  remaining attempts ran. Guard the rebase so a failure aborts the
  in-progress rebase and continues to the next attempt instead.
2026-09-08 21:43:24 -07:00
Chris LuandGitHub 2ffa696809 fix(volume): handle faulty storage media (Go + Rust) (#11233)
* fix(volume): track EC shard read errors and unmount on faulty media

Extract the volume EIO tracker into a reusable IoErrorTracker and add the
same tracking to EcVolume. Sustained EIO on .ecx lookups or .ecd shard
reads now unmounts the EC volume in the heartbeat (without deleting
files) so the master re-replicates from healthy peers, mirroring the
existing volume replica quarantine.

Closes #11227 (EC shard unmount).

* rust(volume): mirror EC shard read error tracking and unmount

Add EIO tracking to the Rust EcVolume mirroring Go: a streak counter
with IO_ERROR_TOLERANCE, a sticky quarantine flag, and unmount (not
file deletion) in the heartbeat so the master re-replicates from
healthy peers.

* feat(metrics): expose storage IO error counter and quarantine gauge

Add a storage_io_error_total counter incremented on every EIO recorded
by the volume or EC shard tracker, and an io_quarantine gauge labelled
by kind (volume/ec_shard) reflecting the count of replicas suppressed
in the heartbeat. Mirrored in Go and Rust.

* feat(healthz): report 503 when local replicas are IO-quarantined

Add Store.HasIoQuarantine (Go) / Store::has_io_quarantine (Rust) and
have /healthz return 503 when any local volume or EC shard is
quarantined due to sustained storage-media EIO, so a load balancer
can drain a server whose underlying media is faulty. Mirrored in Go
and Rust.

* fix(volume): keep quarantined EC volumes in memory and reset EIO on success

Address review feedback: instead of unloading quarantined EC volumes
(which discards the quarantine state healthz needs), keep them in
memory and just skip them from heartbeat reporting, mirroring the
regular volume quarantine. Also clear the EIO streak on successful
.ecx reads in Rust so a transient error does not accumulate, and add
an ec_shard label to the io_quarantine gauge in both Go and Rust.

* fix(volume): exclude quarantined EC shards from heartbeat and add Rust volume tolerance

Address review feedback:
- Filter quarantined EC volumes from CollectErasureCodingHeartbeat
  (Go) and collect_ec_shard_delta_messages / collect_live_ec_shards
  (Rust) so the master stops advertising faulty shards and
  re-replicates from healthy peers.
- Add consecutive EIO count and sticky quarantine to the Rust
  regular Volume, mirroring Go IoErrorTracker: a single EIO no
  longer deletes the replica; the heartbeat quarantines after the
  tolerance threshold and keeps the volume in memory.
- Use the quarantine flag (not last_io_error) in has_io_quarantine
  so /healthz reflects sustained, not transient, failures.

* fix(volume): make Rust quarantined volumes read-only and wire recovery

Address Devin review:
- Set no_write_or_delete on Rust volumes when quarantined in the
  heartbeat, so cached or direct clients cannot mutate a faulty
  replica after the master removes it (mirrors Go).
- Wire reset_io_error_state into Volume::set_writable so an operator
  making a volume writable again clears the sticky quarantine and
  the volume re-enters heartbeat rotation.

* fix(volume): clear EC quarantine on shard re-mount for operator recovery

Address Greptile review: re-mounting EC shards (Go loadEcShardWithIdxDir
/ Rust mount_ec_shards_with_idx_dir) now calls ResetIoErrorState on the
existing EcVolume, giving operators a documented recovery path that
clears the sticky quarantine and returns the EC volume to heartbeat
rotation. Mirrored in Go and Rust.

* fix(volume): do not clear EC quarantine on routine shard mounts

Address review feedback: clearing the EC IO quarantine on every mount
(including duplicate, retry, sibling-shard, and reconciliation mounts)
is too aggressive and can re-advertise known-bad shards before the
storage media has been validated. Remove the automatic reset from the
mount path; quarantine clears naturally on restart or full unmount
when a fresh EcVolume is created with clean state.

* test(volume): update Rust IO error test for quarantine semantics

The heartbeat now quarantines a volume with sustained EIO (keeps it
mounted, makes it read-only, omits it from heartbeat) instead of
deleting it. Update test_collect_heartbeat_deletes_io_error_volume to
assert the volume stays in the store with no_write_or_delete set, and
update set_last_io_error_for_test to set the consecutive error count
at the tolerance threshold so the test reflects a sustained error.

* fix(volume): reset EIO streak after full write and match Windows media errors

Move the success-side EIO reset from append_needle (after write_all only)
to the end of do_write_request, after flush_dat/flush_idx complete, so a
successful write_all followed by a failed fsync no longer resets the
counter before the EIO is recorded. Repeated fsync EIOs now accumulate
toward the quarantine threshold as intended.

Recognize Windows storage-media failure codes ERROR_CRC (23) and
ERROR_IO_DEVICE (1117) in addition to Unix EIO (errno 5), so quarantined
heartbeat behavior is preserved on Windows. Mirrors the change in both
Go and Rust volume servers.

* fix(volume): preserve checkpoint EIO and clear streak on successful delete

maybe_checkpoint_index now returns whether the checkpoint succeeded;
the success-side EIO reset in do_write_request and do_delete_request
only fires when it did, so a checkpoint media failure is no longer
erased by the unconditional reset that followed it. do_delete_request
also gains the success reset that was lost when append_needle stopped
clearing the streak, so a successful delete still clears an earlier
failure streak.

is_storage_io_error now uses libc::EIO on Unix instead of a hard-coded
5, and the ECX binary-search read path gains a Windows fallback
(seek + read_exact) so the buffer is no longer zeroed on non-Unix
targets.
2026-09-08 21:42:56 -07:00
Chris LuandGitHub 0ce5ca42ea helm: supply admin auth in CI renders that enable admin (#11239)
PR #11236 added a render-time guard that fails the chart when admin.ip
is non-loopback (default 0.0.0.0) and admin auth is not configured,
since weed admin 4.46 refuses to bind a non-loopback address without
authentication. Several pre-existing helm_ci.yml test cases enable
admin.enabled=true as part of "everything on" renders without a
password, so helm template now exits non-zero and the Verify template
rendering step fails.

Add admin.secret.adminPassword to the four render calls that turn on
admin without auth (IAM gRPC opt-in, NetworkPolicy EVERYTHING, egress
without kubeApiServer.cidrs, and the license ALL_ON dict), using the
same key ci/admin-values.yaml already uses.
2026-09-08 20:51:53 -07:00
9b12d13934 volume server: release the store lock before scrubbing EC volumes (#11235)
* volume server: release the store lock before scrubbing EC volumes

`ec.scrub` makes a Rust volume server stop serving for the duration of the
scrub, and then kills its own gRPC connection:

    error: rpc error: code = Unavailable desc = keepalive ping failed to
    receive ACK within timeout

Measured on a 4.46 cluster (17 Rust volume servers on one host, ~520 volumes
and 53 EC volumes, --index=redb, EC 10+4). It reproduces against a SINGLE
node in 30-70s, in checksum, index and local modes, at -maxParallelization 1.

## Cause

The CHECKSUM arm of scrub_ec_volume reads every byte of every local shard
while holding the caller's store.read() guard:

    let store = self.state.store.read().unwrap();
    let ecv = store.find_ec_volume(vid)...?;
    let (blocks, broken, errs) = ecv.checksum_scrub();   // GBs of I/O, lock held

VolumeServerState::store is a std::sync::RwLock, which is write-preferring.
The periodic heartbeat's collect_heartbeat_with_snapshot takes store.write()
and blocks; once that writer is pending, every later store.read() queues
behind it. Every HTTP handler takes store.read(), so the node serves nothing,
stops heart-beating, and cannot answer the scrub RPC's own keepalive - the
scrub kills the connection it is running on.

The INDEX and LOCAL arms have the same shape, and the node-wide scrub_volume
loop is worse: it held ONE guard across every volume on the node.

## Evidence

offcputime, off-CPU stacks >1s in a 30s window during a scrub:

    futex_wait
      seaweed_volume::server::heartbeat::collect_heartbeat_with_snapshot
      - tokio-rt-worker
        27967020        <- 27.97s blocked, of a 30s window

A single HTTP /status request issued 12s into a scrub, with 180s of patience,
was accepted and queued for 120 seconds, then served once the scrub released.
Thread states throughout: 1 D + 48 S. One thread working, 48 idle - not
executor starvation and no thread pileup, which is what a single lock holder
looks like.

Memory was tested and ruled out as the cause: the same scrub was run at
MemoryMax 3G, 8G and unlimited. With no limit there is no reclaim at all,
page cache grows freely to 22 GB, and the node still goes unresponsive at
t+30s. anon stays flat at 48-86 MB in every run.

## Fix

checksum_scrub, scrub_index and scrub_local gain plan types -
EcChecksumScrubPlan, EcIndexScrubPlan and EcLocalScrubPlan - snapshotted from
the volume under a brief guard. The handler builds a plan, drops the guard,
and runs the scan in spawn_blocking, off the async workers, since it is
synchronous CPU + file I/O either way.

A plan captures DESCRIPTORS, not paths. Resolving a path again after the
guard is dropped would let a writer that legitimately unlinks the files - the
heartbeat's delete_expired_ec_volumes, which reaches EcVolume::destroy(), or
volume_ec_shards_delete - surface an intentional removal as "scrub read
error: No such file or directory" and put the volume in broken_volume_ids. A
descriptor outlives the name.

For the shards it duplicates the handle the mounted EcVolumeShard already
holds (try_clone_file), which is what Go does: ChecksumScrub reads through
shard.ReadAt (weed/storage/erasure_coding/ec_volume_scrub.go:71), never
through a path. That also inherits open_volume_file's O_NOATIME and drops a
dead branch - the old code built {base}.ec{id}.v{gen} for a non-zero
generation, a name nothing in this tree writes. dup shares the kernel offset,
so shard reads stay positional; the .ecx gets a fresh open instead, since
check_index_file seeks.

FULL/READS is unchanged here: it already released the guard across the index
walk, and still re-takes it per needle in store_ec::scrub_snapshot_under_lock
for that needle's local shard intervals - short holds, many of them.

scrub_volume now takes the read guard PER VOLUME instead of across the whole
loop, so the heartbeat can land between volumes. Its per-volume work still
runs under the guard; Volume needs an equivalent plan to fix that properly,
left as a follow-up and noted in the code.

## A failed scrub task must not take the whole RPC down

Moving the scans into spawn_blocking changed where a panic lands. It no
longer unwinds inside the handler's own future; it comes back as a JoinError
at the .await, and all four join points sat behind a `?`. So one bad volume
out of six hundred returned Err from the entire handler: the
broken_volume_ids, broken_shard_infos and details already gathered for the
other 599 were dropped, and emit_scrub_metrics - the only writer of
SCRUB_LAST_TIME_SECONDS, SCRUB_VOLUME_FAILURES and SCRUB_SHARD_FAILURES - was
never reached, so the staleness alert kept firing while real corruption went
unreported.

And there is a reachable panic behind it. EcLocalScrubPlan::run() sized its
reassembly buffer with

    Vec::with_capacity(get_actual_size(size, version) as usize)

which for any negative size that is not the -1 tombstone skipped above is a
capacity-overflow abort. Mode 3 (LOCAL) is the default of `weed shell
ec.scrub`, and a scrub is what you point at an index you already suspect, so
an arbitrary i32 in a .ecx size field is in-scope input. The buffer is
Rust-only - Go appends to a nil slice and has no capacity hint here. Guard on
`want <= 0` and fall through with an empty buffer: locate_data returns no
intervals for a non-positive size, read stays 0, and the existing
`read != want` error reports the row exactly as Go does.

Each join point now records the failure against its own volume and continues.
A panic is evidence about the volume and counts as broken; a non-panic
JoinError is not - spawn_blocking only reports one when the runtime is going
down, the volume was never scanned, and counting it would put a false
corruption into SCRUB_VOLUME_FAILURES. total_volumes moves before the join in
modes 1, 3 and 4 (2|5 already counted there) so a failed join cannot silently
shrink it. Mode 2|5's verify_ec_shards join is the one that must not
`continue`: the needle walk above has already produced findings for that
volume.

The tombstone guard stays is_tombstone() on purpose. ScrubLocal in
ec_volume_scrub.go:228 skips only IsTombstone(), while the distributed walk
in store_ec.go:516 skips all IsDeleted() - the asymmetry is Go's, and both
Rust walks mirror their own counterpart.

## Both servers: a node-wide scrub skips a volume that vanished mid-run

Releasing the lock makes the volume set legitimately mutable during a scrub,
so a node-wide run can reach a volume that has since been unmounted. That is
not a scrub failure. A node-wide run now logs and skips it; an explicitly
requested volume id still returns NotFound. The Go server is changed the same
way, so both implementations answer the same shell command identically.
mark_broken_volumes_readonly tolerates the same teardown one step later,
instead of throwing away the whole scrub report.

## Test

test_scrub_plans_are_self_contained_and_match_direct_call drops the EcVolume
and runs both plans on another thread, asserting the results match the direct
calls. A plan that borrowed from EcVolume could do neither, so the test stops
compiling if the snapshot regresses to a borrow.

test_scrub_plans_survive_files_removed_after_snapshot unlinks every shard and
the .ecx after the plans are built, then asserts the results still equal the
direct call. Against a path-resolving version it fails with all 14 shards
reported as "No such file or directory".

test_local_scrub_plan_reports_negative_size_ecx_row rewrites a .ecx row's
size to -1000 and runs the local plan on another thread, so the join is the
assertion - that thread is the spawn_blocking whose panic used to fail the
RPC. Without the capacity guard it fails with "capacity overflow"; with it,
the row is reported.

The Go tests cover both halves of the vanished-volume rule for volumes and EC
volumes.

517 lib tests pass, plus 34 across the other targets (`cargo test`).
`go test ./weed/server -run Scrub` passes.

## Known remaining, not fixed here

`ec.scrub -volumeId=N` is still fanned out to every node, and a node that
holds no shard of N returns NotFound, so the shell command errors even when
the nodes that do hold shards scrub cleanly. That is a shell-side fan-out
question rather than a volume-server one, and both servers keep the existing
behaviour for an explicitly requested id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DvHoW85w6SNKNBPvrqLMmK

* scrub: discard checksum block count from total_files; capture .ecx fd for FULL walk

Two review fixes:

1. CHECKSUM arm: plan.run() returns blocks scanned, not a file count.
   Go discards it (_, shardInfos, serrs = v.ChecksumScrub()) so TotalFiles
   stays a needle/file count. The Rust arm was adding it to total_files,
   inflating the count. Discard it to match Go.

2. FULL/READS (scrub_ec_volume_distributed): the needle walk reopened the
   .ecx by PATH after the store guard was released, so a concurrent teardown
   that unlinks or replaces the .ecx (heartbeat delete_expired_ec_volumes,
   volume_ec_shards_delete) could surface an intentional removal as a scrub
   error or mix index generations within one scrub. Capture a second .ecx
   descriptor under the guard (the index plan handle is consumed by its own
   structural walk, and both seek) and read through it instead -- the same
   descriptor-outlives-name invariant the checksum plan shard handles use.

* scrub: bind FULL/READS walk to one encode generation

Address Devin review: after capturing the .ecx descriptor under the guard,
scrub_snapshot_under_lock still re-resolves the volume by id per needle, so
a teardown-and-remount of the same vid between two rows would apply the
captured .ecx offsets to a replacement volume's shards -- falsely reporting
corruption.

Capture the volume's encode_ts_ns (encode-run identity) in Phase A and pass
it to scrub_snapshot_under_lock. If the mounted volume's encode_ts_ns no
longer matches, abort the walk like a mid-scan unmount instead of mixing
generations within one scrub.

* scrub: run FULL/READS index scan in the blocking pool

Address CodeRabbit review (5147767192): index_plan.run() reads the whole
.ecx synchronously, so running it on the async executor worker could block
unrelated RPC work handled on the same executor. Move it into spawn_blocking,
matching the treatment the CHECKSUM/LOCAL arms already give their plans. A
join failure (panic/cancellation) is reported as a seed error so the
per-volume findings below are not silently dropped.

* scrub: move ecx walk to blocking pool, classify join errors, guard encode_ts_ns==0

Three CodeRabbit review fixes (5148034447):

1. Move the FULL/READS needle walk (walk_index_file over the captured ecx
   descriptor) into spawn_blocking. It reads the full .ecx synchronously and
   was still running on the async executor worker, the same blocker the
   index_plan.run() fix in the previous commit addressed.

2. Preserve JoinError classification in both spawn_blocking join points in
   scrub_ec_volume_distributed. A panic is evidence about the volume and
   counts as broken; a cancellation only happens at runtime shutdown, the
   volume was never scanned, and returning it as an error would put a false
   corruption into broken_volume_ids (the FULL/READS arm marks the volume
   broken on any non-empty errs). Panics return an error; cancellations
   return clean.

3. Do not treat encode_ts_ns == 0 as a verified generation match. The .vif
   assigns 0 when it carries no encode-run identity (legacy/pre-feature
   volumes), so 0 == 0 would accept a teardown-and-remount and apply the old
   .ecx offsets to the replacement volume's shards. Only enforce the
   generation check when the captured identity is non-zero; when it is zero,
   fall back to the pre-check behavior (no generation binding) rather than
   aborting a scrub that was already running without the guard.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-08 19:12:30 -07:00
Chris LuandGitHub 723f473f02 filer: widen metadata-subscription readahead buffers (#11237)
The metadata-subscription readahead channels were sized for a
low-throughput era and now bottleneck replay catch-up:

- ReadPersistedLogBuffer's readaheadSize was 1024 entries: the
  background visitor fills the channel, then blocks on the consumer's
  gRPC Send, so volume-server I/O for the next log file never overlaps
  with delivery of the current one. Each disk pass takes longer, and
  the subscribe loop re-lists log files (ListDirectoryEntries on the
  filer store) more often to drain the same backlog. Raised to 8192
  so the reader stays ahead of the consumer through a full log file's
  worth of entries.

- readFilersMerged's logEntryChannelSize was 512 entries per filer
  stream: the same serialization on the client side, where weed mount
  (chunk mode) reads persisted log chunks directly from volume
  servers. A small channel means the producer stalls on the merge
  consumer's processEventFn, and the next log file's chunks are never
  fetched ahead. Raised to 4096 so volume I/O overlaps with event
  delivery.

The wider buffers keep the producer goroutines reading through a full
log file while the consumer is still processing the previous one,
turning serial read→process→read into pipelined read∥process. This
cuts the per-pass wall time that drives filer store listings and
volume-server round-trips, reducing filer workload under backlog
catch-up (e.g. CSI deployments where ~200 mounts reconnect on filer
restart).
2026-09-08 16:58:09 -07:00
5f77a0b67e admin: allow insecurely binding to any IP if -allowInsecureNoAuth is set (#11228)
* admin: allow insecurely binding to any IP if -allowInsecureNoAuth is set

* admin: rename -allowInsecureNoAuth to -allowInsecureBind

The new flag name is shorter and clearer: it describes what is being
allowed (an insecure bind to a non-loopback address) without the
redundant "NoAuth" suffix.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-08 16:57:44 -07:00
Chris LuandGitHub fd4fa72289 helm: pass -ip to admin so StatefulSet becomes Ready on 4.46 (#11236)
* helm: pass -ip to admin so StatefulSet becomes Ready on 4.46

Since weed admin 4.46 changed its default listen address from all
interfaces to loopback (127.0.0.1), the admin StatefulSet template never
passed -ip, so the admin server bound to loopback only. The chart's httpGet
readiness/liveness probes dial the pod IP, not loopback, so the probes
never succeeded and the admin StatefulSet stayed 0/1 forever — breaking
upgrades with helm --wait or GitOps controllers.

Add an admin.ip value (default "0.0.0.0", restoring the pre-4.46
behaviour) and render it as -ip. A non-loopback bind requires
authentication, so fail at render time when admin.ip is non-loopback and
neither admin.secret.adminPassword nor admin.secret.existingSecret is
set, instead of letting the pod crash-loop. Document the value and add a
chart-testing CI values file.

Bumps chart to 4.46.1.

Fixes #11234

* helm: address review feedback on admin bind validation

Align the chart's loopback classification with weed admin's isLoopbackIp
(net.ParseIP + IsLoopback): the whole 127.0.0.0/8 range and ::1 are
loopback; localhost and wildcard addresses are non-loopback, matching
the binary. Previously the exact-string check rejected valid loopback
addresses like 127.0.0.2 while permitting localhost (which the binary
treats as non-loopback).

Recognize WEED_ADMIN_PASSWORD supplied via admin.extraEnvironmentVars /
admin.secretExtraEnvironmentVars as authentication, since weed admin
picks it up through viper's AutomaticEnv. Previously such deployments
were wrongly rejected at render time.

Remove [https.admin] mTLS from the validation message and docs: the
chart only generates [grpc.admin] (gRPC mTLS), not [https.admin] (HTTP
mTLS), so mentioning it as an alternative was misleading.

Document that the -ip flag requires SeaweedFS 4.46 or newer, so pinning
admin.imageOverride to an older image is not supported with this chart.

Extracted the loopback and auth checks into reusable helpers
(seaweedfs.admin.isLoopbackIp, seaweedfs.admin.authEnabled) following
the existing seaweedfs.filer.mysqlEnabled pattern.

* helm: tighten loopback classification to reject malformed 127.x addresses

Use regexMatch instead of hasPrefix for the IPv4 loopback check so
malformed values like "127.not-an-ip" are not accepted as loopback
(net.ParseIP returns nil for them, so weed admin treats them as
non-loopback). Also recognize the expanded IPv6 loopback form
"0:0:0:0:0:0:0:1" in addition to "::1", matching net.ParseIP behavior
for the two common representations.
2026-09-08 16:18:27 -07:00
Chris LuandGitHub cb9fcd39d2 filer/postgres: create filemeta table on startup via createTable config (#11229)
* filer/postgres: create default filemeta table on startup

The postgres filer store hardcoded CreateTableSqlTemplate to empty and
never created the filemeta table, unlike postgres2/mysql2/sqlite which
all create it during Initialize. Users had to create the table manually
or the filer would crash loop with "relation filemeta does not exist".

Read the createTable config option (same as postgres2), default to
DefaultCreateTableQuery when unset, and execute CREATE TABLE IF NOT EXISTS
on the default table after the connection pool is established.

SupportBucketTable stays false so per-bucket table creation remains a
no-op; only the shared filemeta table is created, via a direct ExecContext
since AbstractSqlStore.CreateTable short-circuits without bucket support.

* filer/postgres: accept boolean createTable = true/false

viper reads a TOML boolean as the string "true"/"false" via GetString, so
createTable = true was being used as a SQL template and failed. Add
ResolveCreateTableQuery to normalize the value: true and empty select the
default template, false disables table creation, anything else is a custom
template. Both postgres and postgres2 now use it, and both skip the CREATE
TABLE call when the resolved template is empty.

* scaffold: document createTable option for postgres filer store

Replace the commented-out CREATE TABLE SQL in the [postgres] scaffold with
a createTable config hint, matching the [postgres2] section. Users no longer
need to manually create the filemeta table before starting the filer.

* filer/postgres: make createTable opt-in for postgres, keep postgres2 default

The previous commit defaulted postgres to create the filemeta table even
when createTable was unset, which could break existing deployments whose
DB user lacks CREATE TABLE privileges. ResolveCreateTableQuery now returns
empty for an unset value so postgres only creates the table when
createTable is explicitly true or a custom template — preserving the
prior no-DDL behaviour for existing configurations.

postgres2 keeps its existing always-create default: it defaults an empty
resolved value to DefaultCreateTableQuery, and only skips when
createTable is explicitly false.

* filer/postgres2: simplify createTable handling, document all modes

Drop the false opt-out from postgres2 — it only skipped the default table
while per-bucket CreateTable still ran, leaving restricted DB roles broken
on bucket access. postgres2 now accepts true the same way (defaulting to
DefaultCreateTableQuery) and keeps its existing always-create behaviour
for every other value, matching the original semantics.

The scaffold comment now documents true/false/custom for the postgres
section so users know false (or unset) is the backward-compatible default.

* filer/postgres2: normalize false via ResolveCreateTableQuery

postgres2 only handled "" and "true", leaving createTable = false as the
literal string "false" which CreateTable then executed as invalid SQL.
Route it through ResolveCreateTableQuery (which maps false to empty) and
default the empty result to DefaultCreateTableQuery, so false is treated
the same as unset for the bucket-aware store.

* filer/postgres2: honor createTable = false for default table

postgres2 treated false the same as unset and always created the default
filemeta table, failing startup for restricted DB roles that explicitly
opted out. Track the original false value before ResolveCreateTableQuery
collapses it to empty, and skip the default CreateTable call when set.

Per-bucket table creation is unaffected — it is a runtime requirement of
the bucket-aware store. Users who need to suppress all DDL should use the
postgres (non-bucket) store with createTable unset.

* filer/postgres2: disable bucket tables when createTable = false

Setting SupportBucketTable = false when createTable is explicitly false
makes AbstractSqlStore.CreateTable a no-op (it already returns nil when
SupportBucketTable is false), so neither the default filemeta table nor
per-bucket tables are created. The template stays empty and no DDL runs,
honouring the opt-out for restricted DB roles. All data routes to the
pre-provisioned filemeta table, matching the postgres (non-bucket) store.

* filer: suppress DDL without disabling bucket routing

Setting SupportBucketTable = false when createTable = false also disabled
per-bucket routing, hiding objects in pre-provisioned per-bucket tables.
Keep SupportBucketTable true and instead skip the CREATE TABLE execution
when the resolved template is empty.

GetSqlCreateTable now returns empty for both postgres and mysql SQL
generators when CreateTableSqlTemplate is empty, and
AbstractSqlStore.CreateTable skips the ExecContext call when the SQL is
empty. This preserves bucket routing while suppressing all DDL for users
who explicitly set createTable = false and pre-provision their tables.

* filer: add SkipDDL to suppress CREATE and DROP without disabling routing

createTable = false with SupportBucketTable = true preserved bucket
routing but deleteTable still executed DROP TABLE on bucket deletion,
dropping externally managed tables. CanDropWholeBucket also returned
true, so the S3 layer tried whole-table drops instead of row-by-row
deletes.

Add a SkipDDL flag to AbstractSqlStore, independent of SupportBucketTable.
CreateTable and deleteTable both skip when SkipDDL is set, and
CanDropWholeBucket returns false so bucket deletion falls back to
row-by-row metadata deletes. postgres2 sets SkipDDL when createTable is
explicitly false — bucket routing is preserved, no DDL runs.

* filer: fall back to row-by-row delete when CanDropWholeBucket is false

DeleteFolderChildren took the whole-table drop path whenever the path was
a bucket root, even when SkipDDL made deleteTable a no-op. The no-op
returned nil, the caller returned early, and rows inserted after the
recursive enumeration survived the bucket deletion.

Gate the whole-table drop on CanDropWholeBucket so the row-by-row
DeleteFolderChildren SQL runs when SkipDDL is set, removing all metadata
without issuing DROP TABLE.
2026-09-08 16:17:02 -07:00
Chris LuandGitHub 8782749f26 admin: reject IAM policy deletion while still attached to a user/group (#11230)
* admin: add IsPolicyAttached helper to detect user/group attachments

Introduces AdminServer.IsPolicyAttached, which lists the users and groups
that still have a managed policy attached, reusing the existing credential
manager ListUsers / ListAttachedUserPolicies / ListGroups / GetGroup
methods. This is the building block for rejecting policy deletion while a
policy is still referenced, so deleted policy names stop lingering in a
user attached policy names list (issue #11225).

* admin: reject IAM policy deletion while still attached

Guards AdminServer.DeletePolicy with the new IsPolicyAttached check and
returns the typed ErrPolicyStillAttached error when the policy is still
referenced by a user or group. This matches AWS IAM and the existing IAM
API handler behavior, fixing the stale reference where a deleted policy
name kept showing up in a user attached policy names list (issue #11225).

* admin: return 409 Conflict when deleting an attached IAM policy

The admin UI DeletePolicy handler now maps ErrPolicyStillAttached to HTTP
409 Conflict instead of 500, so the dashboard can surface the attachment
conflict to the user rather than reporting a generic server error.

* admin: skip vanished groups when checking policy attachments

IsPolicyAttached now treats a group that disappears between ListGroups
and GetGroup (credential.ErrGroupNotFound) as no longer attached instead
of failing the whole deletion with HTTP 500, matching the IAM API handler
which skips vanished groups.

* test: assert policy state after deletion paths

Strengthen GetPolicy assertions in the policy deletion tests to check the
returned policy is non-nil after a rejected deletion and nil after a
successful one, not just that no lookup error occurred (GetPolicy returns
nil, nil when a policy is absent).
2026-09-08 16:16:01 -07:00
Chris LuandGitHub b88156fe6b fix(s3api): evaluate aws:SourceIp from the direct TCP peer, not forwarded headers (#11231)
* fix(s3api): use direct peer IP for aws:SourceIp in bucket policy engine

extractSourceIP in the bucket-policy engine trusted X-Forwarded-For and
X-Real-Ip whenever the TCP peer looked private (loopback/RFC1918/link-local),
with no configurable trusted-proxy allowlist. In containerized deployments
the gateway peer is almost always private, so any caller reaching it directly
or from a co-located workload could spoof aws:SourceIp and bypass
IpAddress/NotIpAddress bucket-policy restrictions.

Always return the direct peer address (r.RemoteAddr), matching AWS S3
semantics. Remove the now-unused isPrivateIP helper and header-trust branch.

Update TestExtractConditionValuesFromRequestSourceIPPrecedence to assert the
peer IP is used regardless of forwarding headers, and add regression tests
TestExtractSourceIP_IgnoresForwardedHeaders and
TestExtractSourceIP_EnforcesIPRestrictionPolicy.

* fix(s3api): use direct peer IP for aws:SourceIp in IAM role/session policies

The IAM middleware's extractSourceIP trusted X-Forwarded-For and X-Real-IP
whenever the TCP peer looked private (loopback/RFC1918/link-local), with no
configurable trusted-proxy allowlist. In containerized deployments the gateway
peer is almost always private, so any caller reaching it directly or from a
co-located workload could spoof aws:SourceIp and bypass IpAddress/NotIpAddress
conditions on role and session policies (IsPrincipalActionExplicitlyDenied).

Always return the direct peer address (r.RemoteAddr), matching AWS S3
semantics. Remove the now-unused isPrivateIP helper, privateNetworks table,
and its init().

Update TestRequestContextExtraction and TestIPBasedPolicyEnforcement to assert
the peer IP is enforced regardless of forwarding headers, and add regression
test TestUserInlinePolicySourceIpCondition_IgnoresForwardedHeaders.
2026-09-08 15:09:31 -07:00
Chris LuandGitHub 557fffa350 iam: preserve native Admin when IAM policies are attached (#11226) (#11232)
* iam: expose tri-state result from attached policy evaluation

evaluateIAMPolicies returned a bool that collapsed explicit Deny and
no-match into a single false, so the authorization path could not tell
"policies forbid this" from "policies say nothing". Introduce
evaluateAttachedIAMPolicies returning Allow/Deny/NoMatch and keep
evaluateIAMPolicies as a bool projection for existing callers. This is
preparation for unioning native permissions with attached policies while
preserving deny-always-wins.

* iam: preserve native Admin when IAM policies are attached

Attaching an IAM policy routed authorization exclusively to the attached
policies, dropping the identity native permissions. A user with native
Admin lost all access after attaching a non-granting policy, and stayed
locked out if that policy was deleted without being detached first
(#11226).

Treat a native bare Admin grant as a permission floor that survives
attached policies: when the attached policies do not explicitly allow,
fall back to isAdmin() on the attached-policy path, and on the IAM
integration path allow unless an attached policy explicitly denies.
Explicit Deny still wins on both paths.

Only bare Admin is consulted because inline policies flatten lossily into
Actions (dropping conditions), so scoped actions are not unambiguously
native and must keep flowing through the policy engine.

* iam: regression tests for native Admin surviving attached policies

Reproduces issue #11226:

- TestNativeAdminSurvivesAttachedPolicy: a user with native Admin keeps
  Write access after attaching a policy that does not grant it.
- TestNativeAdminSurvivesDeletedPolicy: the same user keeps Write access
  after the attached policy is deleted without being detached.
- TestAttachedPolicyExplicitDenyOverridesNativeAdmin: an explicit Deny in
  an attached policy still constrains a native admin (deny-always-wins).

* iam: apply native Admin floor before IAM principal validation

The native Admin floor in authorizeWithIAM ran after the auth-path
switch, which denies when no session principal or PrincipalArn is
present. An Admin identity without a PrincipalArn (no session token)
was therefore denied before the floor executed. Move the floor ahead of
the switch and derive the principal for its explicit-deny check with
buildPrincipalARN, which already handles identities without a
PrincipalArn. Adds a regression case for an Admin identity with an
empty PrincipalArn.

Addresses CodeRabbit review feedback on PR #11232.
2026-09-08 14:39:43 -07:00
Chris LuandGitHub 4a1d65939f fix(mount): bound reader cache memory across open files (#11220)
* fix(filer): bound retained reader cache buffers by bytes

* test(filer): keep in-flight downloads during cache trimming

* feat(mem): expose pooled allocation capacity for byte reservations

* fix(mount): share a configurable reader buffer budget across files

* fix(filer): release failed prefetch slots and memory reservations

* feat(mount): expose a soft Go runtime memory limit

* docs(filer): restore shared-download rationale in startCaching

The one-line comment replacing the original context.Background() explanation was too thin for readChunkAt to cross-reference shared resource semantics. Restore a concise note on why request cancellation must not abort a download shared by concurrent readers.

* test(filer): loosen reader cache test deadlines to 5s

Three tests used 1-second deadlines that can flake on CI under load:
TestReaderCacheBudgetInFlight, TestReaderCacheEvictionDoesNotHoldCacheLock,
and TestReaderCacheFailedPrefetchReleasesBudget. Increase to 5 seconds.

* test(filer): cover re-read after reader cache eviction

Add TestReaderCacheReReadAfterEviction: reads chunk 'a', reads chunk 'b'
(evicting 'a' via budget pressure), then re-reads 'a' and asserts a
fresh download returns correct data. Verifies the core correctness
property that eviction never exposes missing or stale data to readers.
2026-09-08 10:51:28 -07:00
ssshr-66andGitHub c6b330be2b [Mount] Add ChunkGroup seeking tests and fix boundary handling (#11223)
* fix issue 11221

* reply ai comments
2026-09-08 10:50:41 -07:00
213f4c5d5c release: wait for Go proxy propagation (#11219)
Allow normal post-tag proxy propagation before dispatching downstream releases, while preserving the check that prevents them from pinning the previous commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-07 21:38:49 -07:00
github-actions[bot] d997fba157 4.46 2026-09-08 03:45:21 +00:00
Chris LuandGitHub c0a7dbb2bb iam: bind CreateServiceAccount ParentUser to the caller (#11218)
* iam: bind CreateServiceAccount target to caller in AuthorizeIamAction

A non-admin holding iam:CreateServiceAccount could pass an arbitrary
ParentUser and mint a service account for any identity, inheriting that
identity permissions. Add a self-target category so a granted non-admin
may only target their own identity; admins remain unrestricted.

* iam: authorize CreateServiceAccount against its ParentUser target

AuthIamManagement passed UserName as the authorization target for every
action, so CreateServiceAccount was authorized with an empty target and
the self-target binding never saw the caller-supplied ParentUser. Pass
ParentUser for that action so the binding takes effect on the live path.

* iam: test CreateServiceAccount binds target to caller

Regression test: a non-admin holding iam:CreateServiceAccount may target
itself but is denied targeting another identity; admins remain
unrestricted.

* iam: authorize CreateServiceAccount against ParentUser on the S3 port

UnifiedPostHandler passed UserName as the authorization target for every
IAM action, so CreateServiceAccount was authorized with an empty target
on the S3-port route and the self-target binding never saw the caller
ParentUser. Extract iamTargetUserName (ParentUser for CreateServiceAccount,
UserName otherwise) and use it from both IAM dispatch surfaces so the
binding applies on the live S3-port path as well as the standalone iam
server.

* iam: test CreateServiceAccount ParentUser binding on the S3 port

End-to-end regression test through UnifiedPostHandler: a non-admin
holding iam:CreateServiceAccount is denied (403) when targeting another
identity and passes authorization when targeting itself.
2026-09-07 20:25:03 -07:00
361fd6b263 [Filer] Parallelize Chunk Manifest Resolution to Reduce Large File Read Latency (#11215)
* fix issue-11214

* fix(filer): cancel sibling manifest reads on failure

* fix(filer): scope manifest cancellation to read batch and propagate context to encrypted reads

Address PR review comments on #11215:

- Scope cancellation to each parallel read batch instead of the resolver-wide
  context, so a later manifest failure does not cancel recursive work for an
  earlier successful manifest (CodeRabbit #3952446554).
- Propagate the resolver context through GetAuthenticatedWithContext so
  encrypted sibling reads observe cancellation and stop promptly when another
  manifest fails (Greptile #3952422343).
- Use net.ListenConfig.Listen with an explicit context in the test fixture to
  satisfy the noctx linter (CodeRabbit #3952111696).
- Add regression tests for encrypted sibling cancellation and for preserving
  earlier manifest children on later failure.

* fix(filer): return known manifest errors without blocking on earlier children

Address Greptile review comment on #11215:

After all parallel reads complete, pre-scan slots for the first real
(non-internal-cancel) error before recursing into earlier manifests'
children. If a later manifest already failed, return its error promptly
with data chunks already in hand, instead of blocking on recursive
network reads of earlier manifests' children.

Updated the regression test to verify the error returns within 1 second
when an earlier manifest's child has a 2-second delay, and that the
child is never loaded.

* fix(filer): filter partial child chunks by requested range on error path

Address CodeRabbit review comment on #11215:

The pre-scan error path appended non-manifest child chunks from earlier
manifests without applying the [startOffset, stopOffset) overlap check
used for top-level chunks. A child outside the requested range could be
returned in dataChunks alongside the later manifest's error.

Apply the same range predicate before appending. Add regression test
with an out-of-range child chunk.

* fix(filer): buffer job channel and abort submission on batch cancellation

Address Greptile review comment on #11215:

- Use a buffered job channel (capacity 128) so submission does not block
  when all workers are busy. This ensures a promptly-failing manifest is
  always queued and can cancel stalled sibling reads once a worker picks
  it up, instead of blocking the caller on the unbuffered channel send.
- Add batchCtx.Done() to the submit select so submission aborts promptly
  when the batch is already cancelled by a sibling failure.
- Add regression test with 5 manifests (4 stalled + 1 failing) verifying
  the failing job is queued and picked up after a stalled worker is freed.

* fix(filer): avoid double WaitGroup decrement on batch cancellation in submit

Address Devin review comment on #11215:

When batchCtx.Done() fired in submit, it called job.done.Done() and
returned false. The caller in resolve also called reads.Done() on the
same WaitGroup, causing a double decrement that would panic with a
negative counter.

Fix: submit sets the result error but does not decrement the WaitGroup.
The caller always owns the decrement and skips overwriting the result
when submit already set it.

* fix(filer): overflow execution when job queue buffer is full

Address Greptile follow-up review comment on #11215:

With a 128-entry buffer, if more than 132 in-range manifests (4 workers +
128 buffer) stall at one level, a promptly-failing manifest beyond the
buffer cannot be submitted and cannot cancel the stalled reads.

Fix: when the buffer is full, run the job directly in a goroutine instead
of blocking on the channel send. This only triggers for >132 manifests at
one level (exceedingly rare), so the bounded concurrency guarantee (4
workers) holds for all normal workloads. Extracted executeJob method
shared by both workers and overflow goroutines.

* fix(filer): bound overflow execution with a semaphore

Address Devin review comment on #11215:

The unbounded overflow goroutines could create thousands of concurrent
reads for large files, defeating the four-worker resource bound.

Fix: add a semaphore (capacity = maxChunkManifestResolveWorkers) that
overflow goroutines must acquire before doing the read. While waiting for
the semaphore, they also watch batchCtx and r.ctx so they exit promptly
on cancellation. Total concurrency is now bounded to 2 * workers (4
workers + 4 overflow) in the degenerate case.

* refactor(http): add ctx to GetAuthenticated signature instead of new function

Reuse the existing GetAuthenticated name by adding ctx as the first
parameter, matching the pattern of ReadUrl, ReadUrlAsStream, and
RetriedFetchChunkData. Removes the GetAuthenticatedWithContext wrapper.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-07 18:44:47 -07:00
Chris LuandGitHub 75ec5ec193 admin: allow setting volume read-only and read/write modes (#11217)
* admin: support setting volume read-only and read/write modes

* admin: address PR review on volume access-mode persistence

Reject trailing JSON values in the SetVolumeReadOnly handler so
requests like {"read_only":true}{} no longer pass validation, and add
a trailing-value case to the invalid-request test.

Propagate .vif persistence failures through the access-mode chain.
PersistReadOnly now returns the SaveVolumeInfo error and rolls back
the in-memory volumeInfo on failure; Store.MarkVolumeReadonly and
Store.MarkVolumeWritable propagate that error and roll back their
noWrite flags, so the API reports failure instead of success while
restart would revert the mode.

* admin: make .vif persistence atomic and preserve error chain

SaveVolumeInfo now writes to a .vif.tmp file, syncs it, renames it
over the target, and fsyncs the directory. A write/sync/close failure
leaves the existing .vif intact, so the PersistReadOnly in-memory
rollback matches the durable state instead of diverging from a
partially written file that restart would apply.

Switch the error wrappers in PersistReadOnly, MarkVolumeReadonly, and
MarkVolumeWritable from %v to %w so callers can use errors.Is and
errors.As to classify persistence failures.

* admin: treat post-rename dir fsync failure as a warning

After os.Rename commits the new .vif, the on-disk file already holds
the requested mode. A directory fsync failure only risks losing the
rename across a crash; returning an error here would make
PersistReadOnly roll back in-memory state while the durable file keeps
the new mode, splitting the replica. Log the failure as a warning
instead, matching the best-effort nature of FsyncDir (already skipped
on Windows).

* admin: distinguish post-rename durability failures and use unique temp files

SaveVolumeInfo now uses os.CreateTemp for the staging file, preventing
concurrent saves for the same volume from colliding on a shared .tmp
path.

A directory fsync failure after os.Rename returns a
NotCrashDurableError instead of being silently swallowed. The rename
already committed the new metadata to disk, so PersistReadOnly,
MarkVolumeReadonly, and MarkVolumeWritable skip the in-memory rollback
for this error type (keeping state aligned with the durable file) while
still propagating the failure to the API. Pre-commit failures continue
to roll back as before.

* admin: continue post-commit work after NotCrashDurableError

MarkVolumeWritable now clears the EIO quarantine and the gRPC handlers
(makeVolumeReadonly step 3, makeVolumeWritable master notification)
proceed with their post-commit work when SaveVolumeInfo returns a
NotCrashDurableError, instead of aborting and leaving the volume
unavailable or the master unaware of the mode change. The durability
warning is still propagated to the API caller. Pre-commit failures
continue to abort early as before.

* admin: handle NotCrashDurableError in tier and EC callers

VolumeTierMoveDatFromRemote and VolumeEcShardsGenerate now check for
NotCrashDurableError from SaveVolumeInfo. When the rename has already
committed the new .vif, they continue with their post-commit work
(backend switch, remote deletion, keeping generated EC shards) instead
of aborting and leaving the on-disk metadata inconsistent with the
file layout. The durability warning is logged for the operator.
2026-09-07 18:40:37 -07:00
github-actions[bot] 53a18ecadd docs: regenerate star history chart 2026-09-08 00:54:44 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
2d4b730a2f build(deps): bump github.com/twmb/avro from 1.7.2 to 1.8.0 (#11210)
* build(deps): bump github.com/twmb/avro from 1.7.2 to 1.8.0

Bumps [github.com/twmb/avro](https://github.com/twmb/avro) from 1.7.2 to 1.8.0.
- [Commits](https://github.com/twmb/avro/compare/v1.7.2...v1.8.0)

---
updated-dependencies:
- dependency-name: github.com/twmb/avro
  dependency-version: 1.8.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* iceberg: adapt to twmb/avro v1.8.0 and iceberg-go defensive copies

avro v1.8.0 changes Schema.Root() to return *SchemaNode, which breaks
iceberg-go v0.6.0's internal avro_schemas.go. The fix (apache/iceberg-go#1843)
is only on iceberg-go's main branch, unreleased, so bump iceberg-go to
that commit (c210509) alongside the avro bump.

That iceberg-go revision also changes two behaviors seaweedfs worked
around:

- It now infers a manifest list's format version from the embedded
  writer schema, so a list missing the "format-version" header entry
  (DuckDB's shape is read as v2, not v1. ReadManifestList's header
  patching is now a redundant safety net; tests updated to expect v2.

- It returns defensive copies from DataFile.Partition(), so the
  ReadManifest shim's in-place partition normalization was silently
  discarded. Rebuild the entry through NewDataFileBuilder when any
  partition value is normalized, copying every other DataFile field so
  manifest round-trips are preserved.

- It converts day-transform partitions to iceberg.Date on read
  (applyDayTransformDates), so the day-partition cases the shim and
  tests guarded now convert without help; tests updated to expect
  iceberg.Date from the raw read.
EOF
)

* iceberg: accept assert-ref-snapshot-id without snapshot-id

iceberg-go's new nullableInt64 parser rejects an assert-ref-snapshot-id
requirement whose "snapshot-id" field is absent from the JSON, even
though the Iceberg REST spec makes it optional (null means the ref must
not already exist). v0.6.0 used a plain *int64, so absent was nil and
accepted. ClickHouse sends the requirement without snapshot-id when
asserting a branch does not yet exist, so its writes fail with
"missing required field \"snapshot-id\"".

normalizeRequirements splices an explicit null into any
assert-ref-snapshot-id requirement missing the field before handing
the JSON to iceberg-go's parser, restoring the v0.6.0 behavior across
both iceberg-go versions.

* iceberg: fix v1 block_size_in_bytes default in rebuilt manifest entries

rebuildManifestEntry set block_size_in_bytes to 0, but the v1 manifest
schema requires the default of 64 MiB ("Always write default in v1").
The original value is not exposed on the DataFile interface, so use the
spec default. Also clarify the fallback comment to note that empty
(zero-record / zero-byte) files also trigger it, not just a nil spec.

Add a round-trip test that writes a rebuilt entry as v1 and verifies
block_size_in_bytes is 64 MiB via Avro decoding.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-07 15:58:24 -07:00
005012edcf rust volume: quick-repair redb on durable checkpoints (#11203)
* rust volume: insert redb rebuilds in needle-id order

Unlink the .rdb before create (create does not truncate). Collapse
last-write-wins, then insert live keys sorted so 4.2.0 packs leaves.

* rust volume: rebuild redb from a BTreeMap and clear leftover keys

Peak rebuild memory is one ordered map instead of HashMap + Vec +
stable-sort scratch. Unlink stays best-effort: if it fails, retain
clears the leftover table before sorted insert. Compute idx metrics
before the write so a read error does not unlink a committed .rdb.

* rust volume: drop the extra redb read transaction on put/delete

put uses insert()'s previous value. delete gets then inserts the
tombstone in the same write transaction. Truncate the .idx row on
any failed redb write after the append.

* rust volume: unpack redb blobs through packed_to_needle_value

save_to_idx, ascending_visit, and collect_entries used the same
length-check copy as get. Route them through the helper so a
wrong-length value is absent everywhere, not a panic.

* rust volume: quick-repair redb on durable checkpoints

set_quick_repair(true) on the durable checkpoint transaction so an
OOM-killed volume server opens without a full-file repair scan.

* rust volume: reopen redb from .idx on non-poisoned commit error

redb 4.2.0 can make a Durability::None commit visible before returning
Err(CommitError::Storage(..)). In that state the database refuses further
write transactions, so truncating the .idx row (the old behavior) would
leave a redb-only put or tombstone that the stored idx_size makes the
reload skip.

Distinguish CommitError::TransactionPoisoned (txn rolled back, db still
usable -- truncate the orphan .idx row as before) from other commit errors
(change may be visible, db refuses writes -- keep the .idx row, close the
database, and reopen from .idx to repair redb's internal state).

db becomes Option<Database> so reopen_from_idx can drop the old file lock
before load_from_idx opens the same path. rdb_path, version, and
cache_bytes are stored so the reopen uses the same configuration.

* rust volume: truncate .idx row when redb is closed in put

put appends the .idx entry before acquiring the write transaction.
When db_or_err() fails (db is None after a failed reopen), the ?
returned without calling truncate_idx_to_offset, so a write reported
failed remained in the authoritative .idx and was replayed on restart.

Handle db_or_err() explicitly and truncate the orphan .idx row before
returning the error, matching the existing handling for begin_write,
open_table, and insert failures.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-07 13:50:49 -07:00
Chris LuandGitHub 5d5ea18287 topology: fix fatal concurrent map read/write on VolumeLayout.crowded (#11216)
SetVolumeCrowded mutated the crowded map under accessLock.RLock(), while
GetWritableVolumeCount reads the same map under RLock() on the Assign hot
path. Two concurrent RLock holders with one writing and one reading the
map triggers a fatal "concurrent map read and map write" that kills the
master process (unrecoverable, bypasses recover).

Take the write lock in SetVolumeCrowded instead. This event path is a
low-frequency single consumer driven by the crowded-volume event loop,
and every other mutation of crowded already holds Lock(); setVolumeCrowded
takes no nested locks, so there is no deadlock path. The hot readers
(GetWritableVolumeCount, CloneWritableVolumes) keep using RLock.

Adds a -race regression test that fails (race detected) on the old RLock
and passes with the write lock.

Fixes #11211
2026-09-07 13:47:22 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
34936c610d build(deps): bump helm/kind-action from 1.14.0 to 1.15.0 (#11212)
Bumps [helm/kind-action](https://github.com/helm/kind-action) from 1.14.0 to 1.15.0.
- [Release notes](https://github.com/helm/kind-action/releases)
- [Commits](https://github.com/helm/kind-action/compare/v1.14.0...v1.15.0)

---
updated-dependencies:
- dependency-name: helm/kind-action
  dependency-version: 1.15.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 13:27:09 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
5e3af51d91 build(deps): bump docker/setup-qemu-action from 4.2.0 to 4.3.0 (#11213)
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 4.2.0 to 4.3.0.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](https://github.com/docker/setup-qemu-action/compare/v4.2.0...v4.3.0)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: 4.3.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 13:27:01 -07:00
a4885b7975 log_buffer: stop closing a notification channel another reader still holds (#11177)
* fix(log_buffer): stop closing a notification channel another reader still holds - #10810

The report blames the polling loop for the busy spin, but that loop is
not what burns the core. LogBuffer keeps one notification channel per
subscriberID, and UnregisterSubscriber closes it. Two registrations that
share a subscriberID share that channel, which happens whenever a client
opens a second stream or an old stream has not yet noticed it was
replaced, so the first unregister closes a channel the other reader is
parked on. A closed channel makes every receive in
awaitNotificationOrTimeoutFor return instantly, and that reader then
spins at full speed for the rest of its life.

Subscriptions are now reference counted. Registering an existing
subscriberID hands back the same channel and raises the count; the
channel is only closed when the last holder unregisters.

* test: fail if the surviving reader stops instead of keeps reading

Review caught that the iteration count alone proves nothing: had
LoopProcessLogData returned when the duplicate reader unregistered, the
counter would sit at 0 and the assertion would pass without a reader
ever having been there to spin. Check the reader is still running before
trusting its low count.

---------

Co-authored-by: Junker der Provinz <jdp@braethoria.com>
2026-09-07 13:20:27 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e0f9e02761 build(deps): bump github.com/prometheus/client_model from 0.6.2 to 0.6.3 (#11208)
Bumps [github.com/prometheus/client_model](https://github.com/prometheus/client_model) from 0.6.2 to 0.6.3.
- [Release notes](https://github.com/prometheus/client_model/releases)
- [Commits](https://github.com/prometheus/client_model/compare/v0.6.2...v0.6.3)

---
updated-dependencies:
- dependency-name: github.com/prometheus/client_model
  dependency-version: 0.6.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 13:19:23 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3f8cc380fb build(deps): bump cloud.google.com/go/storage from 1.64.0 to 1.67.0 (#11206)
Bumps [cloud.google.com/go/storage](https://github.com/googleapis/google-cloud-go) from 1.64.0 to 1.67.0.
- [Release notes](https://github.com/googleapis/google-cloud-go/releases)
- [Changelog](https://github.com/googleapis/google-cloud-go/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-cloud-go/compare/compute/v1.64.0...compute/v1.67.0)

---
updated-dependencies:
- dependency-name: cloud.google.com/go/storage
  dependency-version: 1.67.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 13:19:12 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3eba8ebb9b build(deps): bump github.com/go-ldap/ldap/v3 from 3.4.13 to 3.4.14 (#11209)
Bumps [github.com/go-ldap/ldap/v3](https://github.com/go-ldap/ldap) from 3.4.13 to 3.4.14.
- [Release notes](https://github.com/go-ldap/ldap/releases)
- [Commits](https://github.com/go-ldap/ldap/compare/v3.4.13...v3.4.14)

---
updated-dependencies:
- dependency-name: github.com/go-ldap/ldap/v3
  dependency-version: 3.4.14
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 13:19:00 -07:00
Eliah RusinandGitHub ed5b342f0c rust volume: optional redb insert_before bulk load (#11205)
* rust volume: quick-repair redb on durable checkpoints

set_quick_repair(true) on the durable checkpoint transaction so an
OOM-killed volume server opens without a full-file repair scan.

* rust volume: optional redb insert_before bulk load

Behind redb-experimental-cursor (default off). Production binary
stays on sorted insert(). CI unit tests run both feature settings.

* rust volume: exercise insert_before across leaf splits

Replace the 5-key cfg clone with a 4000-key reverse-order rebuild
so CursorMut::insert_before hits page splits. CI runs the feature
only on storage::needle_map unit tests.
2026-09-07 13:15:47 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
99b84eeb88 build(deps): bump github.com/getsentry/sentry-go from 0.48.0 to 0.49.0 (#11207)
Bumps [github.com/getsentry/sentry-go](https://github.com/getsentry/sentry-go) from 0.48.0 to 0.49.0.
- [Release notes](https://github.com/getsentry/sentry-go/releases)
- [Changelog](https://github.com/getsentry/sentry-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/getsentry/sentry-go/compare/v0.48.0...v0.49.0)

---
updated-dependencies:
- dependency-name: github.com/getsentry/sentry-go
  dependency-version: 0.49.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 09:37:41 -07:00
b6690cfbc8 rust volume: drop the extra redb read transaction on put/delete (#11204)
* rust volume: insert redb rebuilds in needle-id order

Unlink the .rdb before create (create does not truncate). Collapse
last-write-wins, then insert live keys sorted so 4.2.0 packs leaves.

* rust volume: rebuild redb from a BTreeMap and clear leftover keys

Peak rebuild memory is one ordered map instead of HashMap + Vec +
stable-sort scratch. Unlink stays best-effort: if it fails, retain
clears the leftover table before sorted insert. Compute idx metrics
before the write so a read error does not unlink a committed .rdb.

* rust volume: drop the extra redb read transaction on put/delete

put uses insert()'s previous value. delete gets then inserts the
tombstone in the same write transaction. Truncate the .idx row on
any failed redb write after the append.

* rust volume: unpack redb blobs through packed_to_needle_value

save_to_idx, ascending_visit, and collect_entries used the same
length-check copy as get. Route them through the helper so a
wrong-length value is absent everywhere, not a panic.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-07 09:31:00 -07:00
Eliah RusinandGitHub 94b10c006d rust volume: insert redb rebuilds in needle-id order (#11202)
* rust volume: insert redb rebuilds in needle-id order

Unlink the .rdb before create (create does not truncate). Collapse
last-write-wins, then insert live keys sorted so 4.2.0 packs leaves.

* rust volume: rebuild redb from a BTreeMap and clear leftover keys

Peak rebuild memory is one ordered map instead of HashMap + Vec +
stable-sort scratch. Unlink stays best-effort: if it fails, retain
clears the leftover table before sorted insert. Compute idx metrics
before the write so a read error does not unlink a committed .rdb.
2026-09-07 09:29:12 -07:00
Chris LuandGitHub 15e4da65f7 volume: avoid read-only replica write targets (#11195)
* master: carry replica read-only state in volume lookups

* volume: refresh writable replica targets

* volume: preserve read-only replicas for deletes

* master: propagate read-only delete capability

* volume: target delete-capable replicas

* volume: honor configured HTTPS for replica deletes

* volume: reject insecure delete authorization forwarding

* master: broadcast delete capability changes

* volume: align Rust replica routing

* http: protect credentialed replica redirects

* master: preserve digest compatibility for delete capability

* volume: propagate read-only state in short heartbeats

* volume: report changed short volume state

* http: guard TLS client redirects

* master: announce mounted volume read-only state

* volume: replace changed identity deltas

* master: replace incremental volume layouts in order

* master: keep moved volume lookup available

* volume: announce read-only mounts
2026-09-07 09:23:56 -07:00
331c6c3642 shell: volume.check.disk — actionable verdict for diverged vacuumed replicas (#11197)
* shell: volume.check.disk — actionable verdict for diverged vacuumed replicas

When -resurrectMissingNeedles is gated off because both replicas have been
vacuumed (compaction revision > 0) — the normal state of any production
cluster — check.disk previously stopped at 'cannot prove they are missing
writes vs vacuumed deletes' and did nothing, leaving a diverged replica with
no repair path. volume.fix.replication does not catch it either: it only
acts when the replica COUNT is below the expected replication, never when
two replicas are both present but hold different live data.

Classify the divergence instead of dead-ending:

  liveDivergence() counts live (non-deleted) needles present on one replica's
  index but entirely absent from the other, in both directions. Tombstones
  are excluded, so vacuum asymmetry (a compacted replica that dropped deleted
  entries) is not mistaken for divergence.

  reportDivergenceVerdict() turns the count into an operator action:
    - one-sided (one replica has all the live data, the other has no unique
      live needles) -> print the exact safe repair:
        volume.copy -source <complete> -target <lagging> -volumeId <id>
      Re-copying the complete replica is safe precisely because the lagging
      side holds no unique live data; VolumeCopy's verify-before-destroy gate
      independently confirms the source holds the volume before deleting the
      target.
    - two-sided (split-brain, both sides have unique live data) -> warn and
      do NOT emit an auto repair; point to volume.fsck -findMissingChunksInFiler
      to confirm the 'missing' needles are orphans before converging.

Report-only: no data is modified and the resurrection safety gate is
untouched. This is what lets a 13-volume diverged cluster be diagnosed and
repaired in minutes instead of by hand-diffing every index.

Observed motivating case: home SeaweedFS 4.45 cluster, 13 010 cross-rack
volumes diverged after failed replicate writes (ReplicatedWrite MaxAttempts=1
fire-and-forget), 11 one-sided + 2 two-sided, all repaired via volume.copy.

* shell: volume.check.disk — address review nits on divergence verdict

- Use pb.NewServerAddressFromDataNode (dialable ip:port, Address with Id
  fallback) for the advertised volume.copy -source/-target instead of the
  logical node Id, which may not be dialable.
- Make the one-sided verdict tombstone-aware: when the lagging replica has
  been vacuumed, absent live needles may be valid deletions whose tombstones
  were dropped, so a whole-volume re-copy would resurrect them. The command
  is only advertised as safe when the lagging side is proven never-vacuumed
  (compaction revision 0 read under -resurrectMissingNeedles); otherwise a
  caveat is printed pointing at fsck/needle-level repair.
- Fix reversed copy direction when the source replica is the lagging one
  (must copy complete -> lagging in both cases).
- Test: real tombstone (negative size) with the correct 0/0 expectation and
  || assertion; verdict test now covers dialable address, corrected
  direction, and caveat on/off.

* shell: volume.check.disk — per-replica revision knowledge, no copy command for vacuumed lagging side

- Track srcRevKnown/tgtRevKnown separately: in unidirectional mode the
  target revision IS read, so a proven never-vacuumed target no longer
  gets a false resurrection warning (regression: bidi=false, target rev 0).
- A one-sided verdict now only emits the volume.copy command when the
  lagging side is proven never-vacuumed; when it is vacuumed (or unproven)
  the verdict refuses to print the destructive command and points at
  fsck/needle-level repair instead — an appended caveat next to a ready-to-
  paste copy command was still inviting the resurrection.
- deletionCaveat now returns the boolean safety decision.

* shell: volume.check.disk — preserve gRPC port, proven two-sided is not split-brain

- Emit the raw ServerAddress string (host:port.grpcPort) instead of
  String()/ToHttpAddress(), which drops the custom gRPC port and would
  make the suggested volume.copy dial the default port and fail.
- Two-sided divergence with both replicas proven never-vacuumed under
  -resurrectMissingNeedles is mutually missed writes, not split-brain:
  recommend re-running with -apply (in-place resurrection both
  directions) instead of the split-brain no-auto-repair warning.
- Tests: case F (proven two-sided -> -apply, no split-brain warning),
  case G (custom gRPC ports preserved in emitted addresses).

---------

Co-authored-by: timolow <tim@timolow.dev>
2026-09-06 23:25:28 -07:00
Chris LuandGitHub 3225d2b0ce rust worker: install rustls CryptoProvider to fix TLS panic (#11194) (#11196)
* rust worker: add install_default_crypto_provider helper

lance's aws backend pulls aws-lc-rs and reqwest's rustls-tls pulls ring,
so rustls 0.23 cannot auto-select a CryptoProvider and tonic's client TLS
panics on first use. Add install_default_crypto_provider, pinning the
default to aws-lc-rs, mirroring the Rust volume server's helper of the
same name. Includes a regression test that builds a TLS channel and
panics without the install in this crate, where both providers link.

* rust worker: install the crypto provider at startup

Call install_default_crypto_provider before any TLS use, the way the
Rust volume server does in its main. Without this a worker started with
--tls-ca/--tls-cert/--tls-key panics on the first admin dial (#11194).
2026-09-06 22:05:40 -07:00
github-actions[bot] 8537d5bc08 docs: regenerate star history chart 2026-09-07 01:07:12 +00:00
Chris LuandGitHub 70a26cb5d2 s3: gate IAM-cache gRPC RPCs behind admin Bearer auth (#11190)
* s3: gate IAM-cache gRPC RPCs behind admin Bearer auth

The SeaweedS3IamCacheServer registered on the S3 gateway's internal gRPC
port (default 0.0.0.0:18333) accepted PutIdentity/RemoveIdentity/PutPolicy/
DeletePolicy/GetPolicy/ListPolicies/PutGroup/RemoveGroup with no per-RPC
authentication. An unauthenticated network peer could call PutIdentity with
Actions:[Admin] and write straight into the live accessKeyIdent map that
the SigV4 path reads, bypassing S3 authentication entirely.

Mirror the filer's IamGrpcServer.checkAdminAuth: require a Bearer token
signed with jwt.filer_signing.key (read from the existing s3a.filerGuard)
at the top of every IAM-cache RPC. With no key configured the check is a
no-op, matching the rest of SeaweedFS's gRPC surface.

* credential: attach admin Bearer token to S3 IAM-cache propagation

The filer's PropagatingCredentialStore fans IAM mutations out to peer S3
servers over the SeaweedS3IamCache gRPC service. Now that the S3 handlers
require a Bearer token signed with jwt.filer_signing.key, attach one to the
outgoing propagation context (mirroring shell/iamAdminAuthContext). With no
key configured it is a no-op, so deployments that run without the signing
key keep working.

* credential: mint IAM-cache admin token after master discovery

propagateChange attached the admin Bearer token before ListClusterNodes,
so master-client retries could run down the (default 10s) token lifetime
before the peer S3 fan-out began, leaving peers to reject an expired token
and IAM caches stale. Move withIamCacheAdminAuth to after discovery
succeeds, immediately before the propagation timeout is derived.

* credential: cap IAM-cache propagation timeout below JWT lifetime

The propagation fan-out used a fixed 10s timeout. If an operator
configures jwt.filer_signing.expires_after_seconds below 10, the admin
token can expire while slower S3 peers are still being contacted, leaving
their IAM caches stale. Derive the propagation deadline as
min(10s, tokenTTL) so it never outlives the token. withIamCacheAdminAuth
now returns the token's lifetime (0 = no expiry) for this purpose.
2026-09-06 12:21:51 -07:00
ade4bdf9e6 rust volume: stop a tier move whose caller has gone (#11192)
Both tier-move handlers run in a detached tokio::spawn and report
progress through a closure that returns (), with the send result
discarded. Nothing observes the caller leaving, so an abandoned move
uploads or downloads the whole .dat anyway and then commits the
transition.

Go aborts both. Its progress callback returns `stream.Send`'s error,
which surfaces out of the reader in s3_upload.go:99 and the writer in
s3_download.go:84 and fails the transfer, so the volume info is never
rewritten. The Rust port dropped that by typing the callback as
FnMut(i64, f32) with no result.

Give the callback Go's signature -- FnMut(i64, f32) -> Result<(), String>
-- and abort when the caller's channel is closed. Checked on every part
rather than only where progress is reported, since the report is
rate-limited to one a second and would miss a caller that left in
between. A merely full channel is a slow reader, not a departed one, so
only TrySendError::Closed counts as cancellation.

Two consequences of aborting mid-transfer that the old code never had to
handle:

- upload_file now aborts the multipart upload when the transfer fails.
  An abandoned multipart upload does not show up in an ordinary object
  listing but still accrues storage charges until a lifecycle rule reaps
  it, and cancellation makes that a routine path rather than a rare one.
- The tier-down handler removes the partial .dat. download_file
  pre-allocates the destination to the object's full size, so an aborted
  download leaves a file of the right length and the wrong content --
  and this handler refuses to run at all when a local .dat exists, so
  leaving one wedges every retry on "already on local disk" and a
  restart would load the sparse file as the volume's data.

There is deliberately no check between a finished transfer and the
bookkeeping that follows. Once the object is in S3, or the .dat is on
disk, that bookkeeping is what makes the state consistent; stopping
there would leave an object paid for and referenced by nothing, or a
complete local .dat the volume still calls remote. Go does not gate
there either -- its callback only runs during the transfer.


Claude-Session: https://claude.ai/code/session_0122W3eqt6gmLUMxmRoZdPAb

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-06 12:18:47 -07:00
4b41329e12 rust volume: stop a VolumeCopy whose caller has gone (#11188)
* rust volume: stop a VolumeCopy whose caller has gone

VolumeCopy runs its copy in a detached tokio::spawn and reports progress
with the send error discarded, so nothing observes the client leaving.
When the caller cancels the RPC -- which weed-admin's batch balance does
routinely, starting far more copies than it finishes -- the server
streamed the whole volume from the source, wrote it to disk, and mounted
it. The destination is then left holding a volume nobody took delivery
of: its index cache is never reclaimed, and under replication=000 one
volume id ends up on two servers, both writable, which concurrent writes
can diverge.

Three checks now reach the task:

- Every chunk in copy_file_from_source, via the sender's is_closed().
  This is the one that matters in practice. The first progress report is
  128MB in, so for a smaller volume -- the ordinary balance move -- no
  send ever happens and its result says nothing; only the closed channel
  does. The sender is passed for the .idx and .vif copies too, with
  reporting gated separately, so those phases notice as well.
- The throttle sleep, which for a throttled copy runs for seconds at a
  time, now races the sender's closed() instead of being slept through.
- Immediately before mount_volume, and once at the top of the task.

Cancellation surfaces as an ordinary Err(Status::cancelled), so it lands
in the existing error branch that already removes the partial .dat/.idx/
.vif and the .note. That branch also logs now: the error otherwise went
to a channel nobody was reading, leaving the operator with the
balancer's "delete that copy, then re-run the move" and no cause.

This also clears the stranded read-only sources reported on the issue.
They are downstream of the orphan mount, not a separate defect:
LiveMoveVolume's cleanup probes the target before undoing the freeze
(volume_move.go:95, "the server can finish the copy and mount the target
even when the client loses the stream"), and when it finds a mounted
copy it cannot attribute, or cannot delete, it deliberately keeps the
source readonly rather than risk two writable replicas -- the messages
at volume_move.go:110 and :123. With nothing mounted on the target the
probe reports clean and the freeze is undone.

On Go parity: the progress send result is honoured here too, matching
`return false` in volume_grpc_copy.go. But that report is Go's only
abort signal, and measured against a 120MiB volume -- above the
throttler's activation threshold, below the 128MiB report interval -- a
Go destination mounts an abandoned copy as well. The issue's premise
that Go aborts holds only above the report interval. The Rust side now
stops in both cases; the Go behaviour is worth its own issue.

Tests: the integration test runs against both implementations and is
green on Go, red on Rust before this change. Its 192MiB fixture is sized
for two separate constraints, documented at the fixture: IoBytePerSecond
is a no-op below ~100ms of wall clock (64MiB copies in ~110ms on a tmpfs
loopback cluster), and the payload must exceed the 128MiB report
interval for the Go leg to pass at all. The two Rust unit tests cover
what the integration test cannot reach: cancellation detected with no
progress report at all, and the cleanup of the partial files plus the
.note.

Fixes #11186

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0122W3eqt6gmLUMxmRoZdPAb

* rust volume: surface VolumeCopy cancellation as Status::cancelled

copy_file_from_source returned Result<_, String>, so the per-chunk
cancellation path -- the one that matters in practice for volumes
below the 128MB report interval -- was wrapped to Status::internal at
the call sites. The spawn logging branch then classified it as a
generic failure instead of the intended "abandoned by caller",
defeating the logging change in the same PR for the case that occurs
most often.

Return Result<_, Status> from copy_file_from_source: Status::cancelled
for caller-gone, Status::internal for the existing errors. Drop the
.map_err(|e| Status::internal(e)) at the three call sites.

The unit test now asserts the code is Cancelled, not just the message
text, so the classification is locked in.

* rust volume: close cancellation gaps in VolumeCopy

Address two review findings on the same PR:

1. Roll back a mount that races a departing caller. The pre-mount
is_closed() check cannot close the window between the check and
mount_volume: if the receiver drops in that gap, the volume mounts
and the final tx.send(Ok(...)) fails, but its error was discarded
(let _ =), so the task returned Ok(()) and the error branch never
ran. The destination then held an orphaned mounted replica — the
exact defect this PR prevents.

   Fix: track a mounted flag. The final send now checks its result;
   on failure it returns Status::cancelled, and the error branch
   calls store.delete_volume (which unmounts AND removes the files)
   when mounted is true, instead of only unlinking.

2. Observe cancellation while awaiting the source stream. The
per-chunk is_closed() check only runs after stream.message().await
returns. A stalled source (slow disk, partition, GC pause) never
delivers a chunk, so a caller that has already left cannot preempt
the read: the task, the source connection, and the partial files
(including the .note) all outlive the caller indefinitely.

   Fix: race stream.message() against progress_tx.closed() in a
   tokio::select!, so a departing caller preempts a stalled source.

Adds test_volume_copy_after_mount_cancellation_rolls_back_mount to
cover the after-mount rollback path. cargo test --release green
(497 + 5 + 1 + 28).

* rust volume: keep remote data on after-mount rollback, race RPC startup

Two review findings on the after-mount rollback added in ff51f6a:

1. The rollback called delete_volume(vid, false, false), i.e.
   keep_remote_data=false. A remote-tier copy .vif points at the same
   cloud object the source replica references, so destroying the
   abandoned destination with keep_remote_data=false deletes the
   source remote data via Volume::destroy backend.delete_file_blocking.
   Use keep_remote_data=true, matching the pre-spawn delete_volume at
   the top of volume_copy.

2. client.copy_file(copy_req).await (the initial RPC establishment)
   was not raced against progress_tx.closed(). If the source stalls
   before sending response headers, the per-message select! added in
   ff51f6a is never reached, and the task, source connection, and
   preallocated files outlive a departed caller. Race the RPC
   establishment against progress_tx.closed() the same way.

cargo test --release green (497 + 5 + 1 + 28).

* rust volume: race master-configuration wait against caller cancellation

try_get_master_configuration().await was the last un-raced await in
the VolumeCopy task before copy_file_from_source. A stalled master
(or slow leader election) would hold the task and its .note past a
departing caller, since the per-chunk cancellation checks are never
reached. Race it against tx.closed() the same way the source stream
reads already are.

cargo test --release green (497 + 5 + 1 + 28).

* rust volume: fix after-mount rollback test to actually reach that path

The previous version of
test_volume_copy_after_mount_cancellation_rolls_back_mount dropped the
response immediately after volume_copy returned, so tx.is_closed() was
already true at the spawn first check and the task returned before
mount_volume. The test was green for the wrong reason: the mounted
flag, the rollback, and the keep_remote_data=true line were all
uncovered.

Use the store write lock as a seam: take it immediately after
volume_copy returns so the task runs the copy to completion with the
caller still attached, passes the pre-mount is_closed() check, then
parks entering the mount block. Drop the response (caller gone) and
release the guard: the task mounts, fails the final tx.send, and must
roll back via delete_volume.

Verified by setting the rollback guard to if false: the test fails
with "destination still holds a mounted volume". With the rollback
enabled, probe eprintlns confirmed the full path: about to mount ->
mounted = true -> final send failed -> rolling back mount -> rollback
done.

cargo test --release green (497 + 5 + 1 + 28).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-06 10:47:18 -07:00
b82cb05d71 rust volume: checkpoint the redb index durably every 1000 writes (#11182)
* rust volume: checkpoint the redb index durably every 1000 writes

Every put and delete on a redb-backed volume committed with
Durability::None and nothing ever committed durably, on the theory that
the .idx file is the source of truth. redb, however, keeps an entry in
its transaction tracker for every non-durable commit and cannot recycle
pages that were on disk at the last durable commit until a durable one
happens. With no durable commit for the life of the process, both grew
with every write, and .rdb files could bloat toward double size after a
restart (#11179, the hash-table rehash stacks in the memleak output).

The needle map now counts non-durable commits and reports when a
checkpoint is due; the volume takes it, data first: flush the .dat, then
the map fsyncs the .idx and commits redb durably, recording in the same
transaction how much of the .idx the table reflects. A checkpoint makes
the index durable, so the bytes it points at must be down before it, or
after a power loss the index would reference past the end of the .dat
and the volume would load read-only. A failed .dat flush skips the
checkpoint; it is retried on the next write.

Volume::close() now closes the needle map instead of only syncing it,
and the redb map's close() takes the same checkpoint. Before, a clean
shutdown left the table durable (redb flushes on drop) but the recorded
.idx size stale at its load-time value, so the next load replayed every
entry written since load on top of the counters.

On load, the redb map's counters now come from the whole .idx history,
the way Go's LevelDB map rebuilds them (newest entry first, with a bloom
filter of seen keys), instead of from the table's final state. Both the
reuse and the full-rebuild path use it, so overwritten and deleted bytes
keep counting as garbage across restarts, and the incremental replay of
the .idx tail only touches the table, which makes it idempotent whether
or not the table is ahead of the recorded .idx size.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019x36FiSeyePh77YXao15kK

* rust volume: skip redundant .idx fsync on checkpoint after flush_idx

On the fsync=true write path, flush_idx() already fsyncs the .idx before
maybe_checkpoint_index() runs, so the checkpoint's own sync() fsyncs the
same file a second time for nothing. Thread an idx_already_synced flag
from the volume through maybe_checkpoint_index into checkpoint(sync_idx):
when it is true the checkpoint skips its .idx fsync and only does the
durable redb commit. The delete path and close() still sync (they have
not flushed the .idx beforehand).

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: saturate writes_since_checkpoint to prevent u32 overflow

If checkpoints keep failing (e.g. a persistent .dat flush failure whose
error is not EIO and so does not mark the volume read-only), the counter
increments on every write with no upper bound and wraps at ~4.3 billion.
Use saturating_add so it pins at u32::MAX instead, which keeps
checkpoint_due() true and retries on every subsequent write.

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: only update max_file_key on live entries in idx metric rebuild

metrics_from_idx called maybe_set_max_file_key on every entry including
tombstones, but the live on_put path only calls it for puts and on_delete
never does. A tombstone always has a preceding put for the same key that
already set max_file_key, so the result is the same today; restricting it
to live entries makes the parity with the live path exact and self-evident.

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: advance idx_file_offset only after redb commit succeeds

put() and delete() appended to the .idx file and advanced idx_file_offset
before committing to redb. If the redb commit failed, the offset included
the orphan row that redb doesn't reflect. A later checkpoint would record
that offset as "the table reflects up to here," and the reload would skip
the orphan row entirely — the entry becomes permanently unindexed.

Move the idx_file_offset increment to after the successful redb commit.
The .idx file still has the orphan row (append-only), but idx_file_offset
stays behind it, so the next checkpoint records the smaller offset and
the reload replays the orphan row back into redb.

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: skip index checkpoint on close when .dat sync fails

Volume::close() discarded the .dat sync_all() result and always
checkpointed the redb index. If the .dat sync failed, the checkpoint
made the index durable with entries that may point past the unflushed
.dat tail, and a power loss would leave the volume read-only on reload
(the max_needle_end check fires).

Check the .dat sync result: on success, checkpoint as before; on
failure, call close_without_checkpoint() — sync the .idx and drop the
writer without a durable redb commit. META_IDX_SIZE stays at the last
successful checkpoint, so the reload replays the uncheckpointed tail
(redb still flushes on drop, but without recording idx_size).

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: schedule checkpoints on every index mutation path

maybe_checkpoint_index was only called from do_write_request and
do_delete_request. put_needle_index and write_needle_blob_and_index
also call NeedleMap::put, which increments writes_since_checkpoint,
but neither triggered the checkpoint. Through those paths the counter
could grow past the interval without ever being satisfied, leaving
non-durable redb transaction state until close().

Add maybe_checkpoint_index(false) after the successful nm.put in both
methods. The .dat flush inside maybe_checkpoint_index covers the blob
write in write_needle_blob_and_index; put_needle_index pairs with a
prior write_needle_blob, so the flush covers that too.

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

* rust volume: truncate orphan .idx row on failed redb commit

Commit 947ee28 moved the idx_file_offset increment after the redb commit
so a failed commit doesn't advance the watermark. But the .idx file is
append-only: the orphan row stays in the file, and the next successful
write appends after it. That write's idx_file_offset += entry_size
advances past the orphan, so a later checkpoint records an offset that
makes the reload skip the orphan row — hiding a persisted put or
restoring a deleted needle.

On a failed redb commit, truncate the .idx file back to idx_file_offset
before returning the error. This removes the orphan row, so the next
write appends at the correct position and idx_file_offset stays a
contiguous replay watermark. Add a truncate_to method to IdxFileWriter
(set_len for std::fs::File) and a truncate_idx_to_offset helper.

Co-Authored-By: Chris Lu <chris.lu@gmail.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: chrislusf <chris@chrislusf.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-05 21:46:34 -07:00
Chris LuandGitHub d8a40ef750 ci: pin actions/setup-python to v7 in star_history workflow (#11191)
The star_history workflow referenced actions/setup-python@v8, which
does not exist, causing the workflow to fail at the "Set up job" step.
Pin to v7, matching the version used across the other workflows.
2026-09-05 21:44:12 -07:00
Chris LuandGitHub e6f2386a0f admin: redact S3 secret keys for read-only sessions (#11189)
The Admin UI documents its read-only account as view-only and blocks its
write requests, but the authenticated read routes returned object-store
users with plaintext access and secret keys. A read-only admin user could
retrieve another user's live S3 credential pair from GET /api/users and
GET /api/users/{username} and use it directly against the S3 endpoint,
converting view-only access into the victim identity's object-store
authority.

Redact the reusable secret_key in GetUsers, GetUserDetails, and the
rendered users page whenever the requesting session has the read-only
role. The public access_key identifier is retained so identities remain
browsable; only the reusable secret is stripped. Admin and no-auth
sessions are unaffected.
2026-09-05 13:16:56 -07:00
Junker der ProvinzandGitHub 78f79a3919 master: honour -volume.fileSizeLimitMB on the master's /submit (#11176)
* fix(master): honour -volume.fileSizeLimitMB on the master's /submit - #6748

`weed server -volume.fileSizeLimitMB=2048` still refused anything over
256MB, and the reason is not the one the report assumes: the option does
reach the volume server. The master does not use it. Uploads through the
master's /submit are buffered by submitForClientHandler, which passed a
hardcoded 256MB to needle.ParseUpload, so the master rejected what the
volume server it started would have accepted.

The limit is now passed in. `weed master` gains its own -fileSizeLimitMB
with the same 256 default, so a standalone master behaves exactly as
before, and `weed server` and `weed mini` hand it the value their volume
server already got.

* master.follower: take the same upload limit, and say which flag to match

Review found the follower left behind. It serves /submit like the leader
and buffers uploads under the same limit, but kept the fixed 256MB, so a
cluster raised above that would accept an upload through the leader and
refuse the identical one through a follower.

Two smaller points from the same review: the master's flag description
named only the standalone volume server's spelling, and now names the
weed server and weed mini form too; and the under-limit test asserted on
the error message alone, so it would have passed had the limit rejected
that payload with different wording. It now requires the request to get
past parsing.
2026-09-05 12:58:17 -07:00
yanglongweiandGitHub eb717199d0 master: delete replica_placement_mismatch labels when volumes leave topology (#11062)
* master: delete replica_placement_mismatch labels when volumes leave topology

Fixes #10804. Setting the gauge to 0 left stale Prometheus time series
that grew unbounded with volume churn; remove the label set on unregister
instead.

* master: delete replica_placement_mismatch only after last placement leaves

Unconditional DeleteLabelValues on UnRegisterVolumeLayout dropped the series
while other data nodes still held the volume, hiding under-replication until
the next collect cycle. Delete only when Lookup is empty, and cover the
two-copy case in a regression test.
2026-09-05 12:52:27 -07:00
Chris LuandGitHub 3e85d9ec8e admin: bind to loopback by default, guard public unauthenticated bind (#11185)
admin: bind to loopback by default, refuse public unauthenticated bind

The admin HTTP server (port 23646) defaulted to binding 0.0.0.0 with
authentication disabled when -adminPassword was not supplied, exposing
the full admin REST API (user creation, credential issuance, bucket
deletion, filer deletion) unauthenticated on the network. This is the
footgun described in GHSA-m3m8-mrgq-hf9h.

Keep the no-auth mode for local dev, but remove the network exposure:

- Add -ip flag (default 127.0.0.1) so the server binds loopback only
  unless the operator explicitly chooses a public address.
- Refuse to start when binding a non-loopback address with no
  -adminPassword and no [https.admin] mTLS. The operator must enable
  auth or use loopback.
- weed mini sets -ip from its existing -ip.bind; the guard does not
  apply because mini calls startAdminServer directly, not runAdmin.

Addresses GHSA-m3m8-mrgq-hf9h.
2026-09-05 12:48:50 -07:00
Chris LuandGitHub f35e2ccf21 s3: warn when a lifecycle change leaves fast-path-stamped objects on their old TTL (#11184)
* s3: warn when a lifecycle change leaves fast-path-stamped objects on their old TTL

The per-write TTL fast path (opt-in via s3.bucket.lifecycle.fastpath)
stamps a volume TTL at PutObject time that can't be taken back. When an
operator lengthens or removes an Expiration.Days rule (or deletes the
bucket lifecycle) on a fast-path-enabled bucket, objects already written
keep their baked-in TTL and won't be rescued by the change — unlike the
default worker-driven path, which re-evaluates the current rules each
pass. This is the data-loss direction described in #11183.

Surface it: Put/DeleteBucketLifecycle now emit a glog warning and set
X-Seaweed-Lifecycle-Fastpath-Warning on the response when the change
removes, disables, lengthens, or re-scopes a fast-path-eligible rule.
Shortening a rule does not warn (old objects simply expire later, not
data loss). Tag-only and overflow-day rules are never on the fast path
and never warn.

Addresses the warning half of option 2 in #11183.

* s3: address review — emit warning after mutation succeeds, fix ID-rename false positive

Two issues raised by CodeRabbit, Greptile, and Devin reviews:

1. Failed mutations retained the warning header. The warning was set on
   the ResponseWriter before storeBucketLifecycleConfiguration /
   clearStoredBucketLifecycleConfiguration was called; if that failed,
   the error response carried a warning for a change that was never
   applied. Now the reason is computed before the mutation but the log
   and header are emitted only after it succeeds.

2. Rule renames produced false "removed" warnings. fastpathRuleKey used
   Rule.ID as the sole identity when present, so renaming a rule (same
   prefix/size/days, different ID) treated the old rule as removed.
   Replaced with two-pass matching: first by ID, then by fast-path
   predicates (prefix + size). An ID-only rename with unchanged
   predicates and days no longer warns. Greedy matching ensures each
   new rule is consumed by at most one old rule.

Added regression tests: ID-only rename (no warn), rename + lengthen
(warn), rename + shorten (no warn).
2026-09-05 12:13:37 -07:00
Junker der ProvinzGitHubChris Ludevin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris Lu
f99c4a1f14 s3: make RenameObject idempotent for a retried request (#11178)
* feat(s3): make RenameObject idempotent for a retried request - #10661

A rename that succeeds but whose response is lost leaves the client with
no safe move: retrying returned 404, because the source is already gone,
so a retry was indistinguishable from a rename that never happened.

The destination now carries what the rename that created it was, under
x-seaweedfs-rename-token: the client's token, the source key and the
time. A retry that names the same token and the same source and
destination is answered 200 without touching anything. The same token
sent for a different rename is refused with 409 rather than silently
answered, and a token older than 24 hours is treated as unrelated so a
key cannot answer for a request indefinitely.

Requests without the header behave exactly as before.

* s3: answer a reused rename token with 409, not 400

The PR promised Conflict and the code returned Bad Request. 400 tells a
client its request was malformed and invites it to give up; this request
is well formed and resending it unchanged will not help, because what it
collides with is a rename the same token already stands for.

The status code is now asserted in a test, since it is the part of this
behaviour a client actually acts on.

* Update weed/s3api/s3err/s3api_errors.go

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: fix rename token review notes

- ErrIdempotentParameterMismatch returns 409 Conflict, not 400. The
  comment and TestRenameTokenReuseAnswersConflict both expect 409; the
  code regressed to 400 in a later commit.
- stampRenameToken: clarify that markRenameToken mutates srcEntry in
  place, so the token reaches the destination via the move regardless
  of whether the UpdateEntry succeeds. The precondition only guards the
  pre-move write, not the move itself.
- Extract the handler retry branch into retryRenameDecision and add
  TestRetryRenameDecision, covering the source-still-exists fallthrough
  that was previously reasoned about but not tested.

* s3: IdempotentParameterMismatch returns 400, matching AWS docs

The AWS S3 RenameObject API documentation specifies HTTP Status Code: 400
for IdempotencyParameterMismatch. Revert the previous 409 change and align
the comment and test with the documented behavior.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-05 11:48:42 -07:00
3f9b05946b rust volume: bound the redb index cache per volume by --index tier (#11180)
The Rust volume server opens one redb database per volume and built
each with redb's defaults, which give every database a 1 GiB page
cache (0.9 GiB read cache + 0.1 GiB write buffer). With hundreds of
volumes behind one disk the process-wide ceiling was volumes x 1 GiB:
memory grew in proportion to the pages traffic touched, never shrank
when traffic stopped, and hosts running many instances were OOM-killed
under bulk ingest. redb, redbMedium and redbLarge were also treated
identically, so the "memory~performance" tiers did nothing.

Size the cache per tier instead: 4, 8 and 16 MiB per volume, mirroring
the Go server's 3/6/12 MiB LevelDB block cache + write buffer. Thread
the budget through RedbNeedleMap::new/load_from_idx so every open path
(create, reuse, full rebuild) uses Database::builder().set_cache_size.

Fixes #11179


Claude-Session: https://claude.ai/code/session_019x36FiSeyePh77YXao15kK

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 10:56:42 -07:00
adbee9452a rust volume: bump redb 3.1.3 -> 4.2.0 (#11181)
No source changes: the API surface the needle map uses (Database
create/open/builder, set_cache_size, set_durability, tables, iterators)
is unchanged and the on-disk format is still v3, so existing .rdb files
open as-is. The 4.0.0 breaking changes (Drop on AccessGuardMut, removal
of the Legacy type) do not touch this crate.

Relevant to the redb-backed index (#11179):
- 4.1.0: optimizes cache usage and memory usage; ~1.5x faster writes.
- 4.2.0: Durability::None commits ~2x faster; pages freed by a durable
  transaction are reused by the very next one; a crash-recovery fix for
  a crash during repair of an earlier crash.


Claude-Session: https://claude.ai/code/session_019x36FiSeyePh77YXao15kK

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 10:55:23 -07:00
Chris LuandGitHub 8a68337256 filer: pack SSE chunks into manifests (#11175)
* filer: pack SSE chunks into manifests

* s3: resolve encrypted manifests before reads

* s3: scope encrypted manifest resolution to ranges
2026-09-05 10:16:17 -07:00
Chris LuandGitHub 97154802c5 docs(star-history): make the chart taller (#11173)
Change the matplotlib figure size from (10, 4) to (10, 6) so the
star history chart renders vertically longer in the README. The
regenerated note/star_history.svg reflects the new 5:3 aspect ratio
(720x432pt) instead of the previous flat 2.5:1 (720x288pt).
2026-09-05 00:11:09 -07:00
Chris LuandGitHub ff0d5a9adf ci(mount-windows): clean up processes and don't let Logs step fail job (#11174)
The "Mount and exercise" step left the weed.exe mini server and the final
WinFsp mount running when it exited. The next step's pwsh.exe then failed
with STATUS_DLL_INIT_FAILED (0xC0000142), failing a job whose actual test
step had passed. The same code passed on both the PR branch and the next
master run, so this was a transient launch failure — but it was caused by
an unclean environment and made fatal by a diagnostic step.

Tear down all weed.exe processes at the end of the test step so subsequent
steps launch into a clean environment, and mark the Logs step
continue-on-error so a diagnostic step can never fail the job on its own.
2026-09-05 00:10:40 -07:00
Chris LuGitHubdevin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
811b8b5734 make the remote-mount cache wait configurable per mount (#11168)
* add a per-mount cache_wait_ms to the remote storage mount mapping

A read of an uncached remote-only object waits on a hardcoded size tier
before it can fall back to the origin, so every ranged read of a large
remote-only object pays that wait. Carry the wait in the mount mapping so
it can be tuned, or set to zero, per mount.

* resolve the cache wait of an uncached remote-only read from its mount

The wait came only from the object size, so an operator could not trade
cache hits for time to first byte. Both read paths now resolve the mount
covering the object and let its cache_wait_ms replace the size tiers.

* read straight from the remote when a mount waits zero for its cache

A mount used as a streaming source pays the cache wait on every ranged
read of an object too large to finish caching, and the caching itself is
wasted work. A zero wait now skips the cache call, so both read paths go
to the origin immediately.

* let remote.mount set the cache wait of a mount

remote.mount -cacheWait=0 turns a mount into a streaming source, and any
other duration trades cache hits against time to first byte.

* keep the size based wait for a version-specific read

A read pinned to a version cannot fall back to the origin, since the
mounted remote only holds the current key, so a mount that opts out of
caching would leave it on the 503 retry loop forever.

* let the operator allow a remote-only read to dial an internal endpoint

The remote-mount read paths in the filer and the S3 gateway always refused
an endpoint resolving to a loopback or private host, so a mount backed by
an internal S3 could never be read from its origin, only through the local
cache. Both now take the allowance the volume server already has, still
off by default.

* skip the background cache of a mount that waits zero for its cache

GetObjectHandler kicks off caching for every remote-only read, so a mount
serving as a streaming source kept downloading whole objects even though no
read ever waited for them.

* cover a zero cache wait end to end

The read has to reach a real origin, so the harness also opts the filer and
the S3 gateway into dialing the loopback remote it already allows for the
volume server.

* resolve the S3 cache wait once so the background cache follows it too

The background cache that GetObjectHandler starts read the mount on its
own, so it skipped a version-specific read that the foreground path still
waits for. Both now ask the same resolver.

* answer 404 when the origin of a zero-wait read is gone

Metadata can outlive the object it points at, and with no cache to fill
the read would sit on the 503 retry path forever. The remote backends
already report a missing object as ErrRemoteObjectNotFound.

* open the origin at write time for a multipart range

Every part of a multipart Range is prepared before any is written, so
opening eagerly would hold one origin connection per part and leak the
ones already opened when a later part fails to open.

* reject a cache wait shorter than a millisecond

The mapping stores milliseconds, so -cacheWait=500us truncated to zero
and silently turned caching off instead of waiting.

* restore the doc comment of cacheRemoteObjectForStreamingWithShortTimeout

Extracting the wait resolver left its comment on the new function.

* stat the origin before committing a multipart range

Opening at write time keeps no connection through the preparation, but it
also moved a failure past the point where the multipart body picks the
response status, so a gone origin truncated a 206 instead of answering
404. One stat up front puts the status back.

* stat the origin once per request

Every part of a multipart Range is prepared on its own, so the preflight
ran once per range instead of once per read.

* map Azure and GCS stream not-found to ErrRemoteObjectNotFound

ReadFileAsStream on Azure and GCS returned provider-specific not-found
errors instead of ErrRemoteObjectNotFound, so a zero-wait read of a
deleted object was misclassified as a transient cache failure and
retried indefinitely. Map BlobNotFound and ErrObjectNotExist the same
way StatFile already does.

* Update weed/remote_storage/gcs/gcs_storage_client.go

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: devin-ai-integration[bot] <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-04 23:50:11 -07:00
Chris LuandGitHub f79d83abf4 volume: expire TTL volumes whose only traffic is deletes (#11167)
* volume: count a TTL volume's age from its last write, not the .dat mtime

A delete appends a tombstone needle and vacuum rewrites the .dat wholesale,
so the file's mtime moves without any write ever landing. The loader read
lastModifiedTsSeconds back from that mtime, so every restart of a volume
taking delete traffic re-armed expired() for another full TTL: an
overwrite-heavy collection kept growing until it hit the max-volume cap.

Recover the clock from the newest .idx entry that is not a tombstone and
read that needle's append timestamp, falling back to the mtime when no
write is recoverable. Only TTL volumes pay for the scan.

Fixes #11160

* volume: count the .vif destroy time from the last write too

ExpireAtSec is what an EC volume is reclaimed on, and it was recomputed as
now+TTL every time the .vif was written. A read-only mark, a tier upload or
an EC encode therefore handed an already expiring volume another full TTL,
the same way the .dat mtime did.

Derive it from the volume's last write, falling back to now for a volume
that has not taken one yet so a fresh volume is not born expired.

* volume: mirror the last-write TTL clock in the Rust volume server

Same recovery as the Go loader: scan the .idx backwards for the newest
entry that is not a tombstone and take that needle's append timestamp,
leaving the clock on the .dat mtime when no write is recoverable.

* volume: mirror the last-write destroy time in the Rust volume server

Both .vif writers and the EC encode computed ExpireAtSec as now+TTL, the
same way Go did, so the destroy time moved every time the sidecar was
rewritten. Route all three through the volume's last write.

* volume: report the .dat mtime in the Rust heartbeat, like Go does

The Rust server reported its TTL clock as ModifiedAtSecond while Go
reports the .dat mtime. The shell's quiet-period gates (volume.tier.move,
volume.delete_empty) read that field as "last touched", which a delete
has to count towards even though the TTL clock deliberately ignores it --
and with the clock now recovered from the last write, the two drift
further apart.

* volume: take the newest write by timestamp on a vacuumed volume

The reverse .idx scan trusted position, which holds only while the .dat is
append ordered. Vacuum rewrites it in key order, and since an overwrite
keeps its original key, the highest-key survivor is not necessarily the
newest write -- the recovered clock could land up to a TTL early and take
the volume with data still inside its TTL.

A volume that has been vacuumed (CompactionRevision > 0) now takes the
maximum append timestamp over a bounded window of write entries instead.
An append-ordered volume still answers in one read.

* volume: never guess a vacuumed volume's last write, and resolve wrapped offsets

Two holes in the reverse scan, both from review:

A vacuumed volume's writes are ordered by key, so any of them can hold the
newest timestamp. Reading a capped window sampled the highest keys, which
could still miss a recently overwritten low-key needle and expire data
inside its TTL. The scan now covers every write a vacuumed volume indexes,
and a volume too large to scan keeps the .dat mtime rather than report a
partial maximum -- late is recoverable, early is not.

A .dat past MaxPossibleVolumeSize wraps the offsets in its .idx, so reading
a timestamp at the unwrapped offset picks up an unrelated needle. Resolve
the entry against the needle header first and retry one volume size in,
the way doCheckAndFixVolumeData already does.

* volume: drop GitHub issue references from TTL comments
2026-09-04 23:48:40 -07:00
Chris LuandGitHub 567578f08d docs(readme): replace star-history.com with self-generated chart (#11171)
* docs(readme): replace star-history.com with self-generated chart

The star-history.com SVG is a third-party dependency that can rate
limit or go down. Replace it with a GitHub Action that fetches
stargazers via the REST API and renders an SVG with matplotlib,
committing note/star_history.svg weekly. The README references the
committed file directly, so the chart has no runtime dependency on
any external service.

* ci(star-history): run daily instead of weekly
2026-09-04 23:45:50 -07:00
Chris LuandGitHub 070174726d docs(readme): replace rate-limited starchart with star-history (#11170)
starchart.cc is rate-limiting the SVG endpoint, so the Stargazers
chart renders blank. Switch to star-history.com, which serves a live
SVG for this repo and links to the interactive chart.
2026-09-04 23:28:48 -07:00
Chris LuandGitHub e4dc66c66b docs(readme): move sponsor section to the end (#11169)
The Patreon CTA and Gold Sponsors logos sat between the logo and the
project intro, pushing the actual description below the fold. Move the
whole block to a dedicated `# Sponsors #` section after `# License #`,
add it to the TOC, and give it a real markdown heading so the anchor
works on GitHub.
2026-09-04 23:17:11 -07:00
Chris LuandGitHub 1ca19ea2e2 mount: add -volumeName to name the disk explicitly (#11165)
* mount: let volumeName take an explicit override

volumeName only ever derived the disk's label from -filer.path, -dir,
or the filer address, so a name that happened to collide with
something else - e.g. a UNC share's own name - could not be changed
without moving what was mounted. Give it an override parameter that
wins over all three; nothing passes one yet.

* mount: add -volumeName to name the disk explicitly

Windows has no equivalent of the "weed fuse" -o passthrough that lets
a Linux or macOS mount override its derived volname, so a name picked
up from -dir - e.g. a UNC share's own name - could not be changed
short of moving what was mounted. -volumeName overrides it on every
platform.

* mount: document -volumeName

* mount: scope -volumeName's help text to macOS and Windows

Linux has no volume-label mount option for -volumeName to feed, so
the flag's own description says where it applies instead of leaving
that unstated.

* mount: forward -volumeName through the weed fuse option parser

weed fuse (the /etc/fstab helper) turns -o key=value into the same
MountOptions weed mount takes, but volumeName had no case, so it fell
through to being forwarded as a literal, unrecognized FUSE option
instead of ever reaching mountOptions.volumeName.

* mount: apply -volumeName to FsName on Linux and FreeBSD

FsName only ever took the filer address and -filer.path, so
-volumeName had nothing to override there and silently did nothing;
the skipAutofs case still forces "fuse", since that name is what
util-linux/mount requires to recognize the pseudo filesystem.
2026-09-04 23:14:57 -07:00
cfa8afec92 filer: guard FoundationDB 100KB value limit and pack earlier (#11161)
* filer: guard the FoundationDB value size limit, not the transaction limit

An entry's whole chunk list is one FoundationDB value, and FDB caps a value at
100,000 bytes while a transaction may reach 10MB. UpdateEntry checked the
transaction limit, so every entry between the two limits passed the guard and
was rejected by FDB itself with error 2103 (Value length exceeds limit). The
failure surfaced inside the store rather than at the guard, so the S3 layer
dropped the connection and clients saw a network fault instead of an error.

Check the value limit in UpdateEntry and KvPut instead, after gzip and before
the transaction, with an error that names the limit it hit. The removed
transaction-size constant guarded nothing else: DeleteFolderChildren batches by
entry count.

Refs #11158

* filer: fold at 500 chunks in the foundationdb build

Manifest packing is what keeps a large file's entry small, but it only ran once
a flat chunk list reached 10000 chunks. A FoundationDB value stops at 100,000
bytes and an entry's whole chunk list is one value, which at ~100 bytes per
chunk record is about 1000 chunks -- so on FDB the write always failed before
packing could help: a 3.3 GiB PutObject at the default -maxMB=4 was already
past the limit.

FoundationDB support is its own build (`go build -tags foundationdb`, shipped
as its own image), so the batch is a build-time choice and needs no negotiation
at run time. The tagged build folds at 500, every other build keeps 10000 and
is untouched.

500 is not arbitrary: a single fold level leaves (chunks/batch) manifest
pointers plus up to (batch-1) unfolded chunks in the entry, so the reachable
chunk count is highest when the two terms are near equal. For a 100,000-byte
budget that optimum is 500, which holds an entry inside the limit up to
~250,000 chunks -- ~1 TB at -maxMB=4, against ~4 GB before. Larger files need
nested packing, which no batch size substitutes for.

One binary serves every role in that image, so the filer and each client that
folds -- S3, mount, WebDAV, weed shell, filer.copy -- agree on the batch by
construction. A binary built with the tag but pointed at another store folds
earlier than that store requires, costing one manifest blob per 500 chunks and
one read to resolve it.

Fixes #11158

* filer: fold with rollback inside MaybeManifestize, not beside it

A fold that fails midway has already uploaded manifest blobs for its earlier
batches, and returns only the data chunks -- dropping the manifests it had
separated out of the caller's list. Both were wrong in ways that mattered:

  - AppendToEntry assigned that shortened list straight to entry.Chunks and
    created the entry, so an append to an already-folded file whose fold
    failed lost every previously folded chunk. weed mount had the same shape.
  - cleanupChunks logged the error as "not good, but should be ok" and then
    returned it through a named result, failing the whole CreateEntry or
    UpdateEntry, while the blobs it had written stayed behind referenced by
    nothing.

The S3 path was alone in handling this, through a private helper beside
MaybeManifestize. A second entry point next to the one everything else calls
just means the wrong one gets used, so the behaviour moves inside
MaybeManifestize: on failure it returns inputChunks as it received them, and
hands the blobs it saved to a deleteChunks callback. The filer, S3 and
filer.copy pass their existing deleters -- filer.copy already cleans up this
way after a failed upload -- and mount, WebDAV and weed shell pass nil, which
reports the blobs rather than collecting them, as before. Each caller keeps its
own error policy: the filer HTTP PUT path and filer.copy still fail the request,
the rest still continue with the flat list, which is a correct entry.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-04 23:12:27 -07:00
Chris LuandGitHub 5a515adab2 s3: HeadObject with partNumber returns the part's size and 206 (#11166)
* s3: HEAD with partNumber reports the part's size and range

HeadObject set its headers from the total object size and then only
validated the partNumber, so a client probing part 1 with HEAD got the
whole object's Content-Length and a 200 while the same GET returned the
part's size, a Content-Range and a 206.

Resolve the part's byte range before the headers are written, through the
range logic GetObject already used, and answer a partNumber HEAD as the
ranged HEAD that AWS documents.

* s3: answer an unsatisfiable partNumber with 416 InvalidPartNumber

GET and HEAD rejected a partNumber past the number of parts with 400
InvalidPart, the code for a missing part in CompleteMultipartUpload. AWS
answers a read of a part that does not exist with 416 InvalidPartNumber,
which lets a client probing for the part count tell the two apart.

The ceph suite pins RGW's 400 InvalidPart here, so the s3tests jobs patch
that expectation the way they already patch prefix ordering.

* s3: keep the whole-object checksum off a partNumber response

The stored checksum covers the whole object, so it is already withheld
from a ranged read. A partNumber HEAD now describes one part while the
request carries no Range header, so exclude it there too rather than
handing a client a checksum that does not match the bytes described.

* s3: resolve a partNumber against the parts the object records

Completion accepts ascending, not consecutive, part numbers, so the part
count is not the highest part number. Comparing the two rejected an
uploaded part 3 of a two-part object, and let a request for the absent
part 2 fall through to the positional chunk lookup and serve part 3's
bytes. Ask the recorded boundaries for the part instead, and keep the
count comparison for objects written before boundaries were stored.

* s3: apply a client Range within the part on HEAD too

GET narrowed the part by a Range sent alongside partNumber; HEAD reported
the whole part, so the two disagreed again for a request that carries
both. Move the narrowing into the shared range lookup so either verb
describes the same bytes.
2026-09-04 23:10:41 -07:00
27b2411cdd fix(chart): add missing [grpc.s3] TLS section — S3 internal gRPC served plaintext while peers dial mTLS (#11157)
* fix(chart): serve S3 internal gRPC with mTLS when security enabled

The security.toml generated by the chart has no [grpc.s3] section, so
security.LoadServerTLS(viper, "grpc.s3") returns nil in weed/command/s3.go
and the S3 server listens plaintext on its gRPC port (httpPort+10000 = 18333
by default). Workers dial that port with mTLS credentials (grpc.worker),
producing:

  walker dispatch ...: rpc error: code = Unavailable desc = connection
  error: desc = "transport: authentication handshake failed: tls: first
  record does not look like a TLS handshake"

This breaks the s3_lifecycle worker's LifecycleDelete RPC path (recovery
walk, daily replay) and any S3->S3 IAM cache propagation would fail the
same way if clients enforced TLS.

Add [grpc.s3] reusing the client cert already mounted on s3 pods (or
s3.tlsSecret when set, mirroring the seaweedfs.s3.tlsArgs helper for the
HTTPS listener).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf)

* fix(chart): always use internal client cert for grpc.s3 identity

s3.tlsSecret is the public HTTPS listener certificate (possibly issued by
a public CA); internal gRPC peers only trust grpc.ca, so presenting it on
the internal gRPC port would break lifecycle/IAM RPC verification. Keep
the two trust domains separate.

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)
2026-09-04 19:41:25 -07:00
98115ec2de deps: update golang.org/x/image to v0.45.0 for CVE-2026-46603 (#11164)
golang.org/x/image v0.44.0 is affected by CVE-2026-46603 (GO-2026-6222):
a denial of service via excessive memory allocation when decoding
malformed VP8L (lossless WebP) data. It is fixed in v0.45.0, released
2026-08-11.

The decoder is reachable from SeaweedFS: weed/images/resizing.go
blank-imports golang.org/x/image/webp, which registers the VP8L decoder
with image.Decode, so the filer image resizing path decodes attacker
supplied WebP data with the affected version.

This is a go.mod/go.sum only change produced by
`go get golang.org/x/image@v0.45.0 && go mod tidy`; no other dependency
moved. `go build ./weed/`, `go vet ./weed/images/...`,
`go test ./weed/images/...` and `go mod verify` all pass.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 19:38:01 -07:00
Chris LuandGitHub a0b1272cc3 filer: authorize the chunk proxy and the root listing like the rest of the filer port (#11152)
* filer: require a read token for the root listing

maybeCheckJwtAuthorization waved through every GET/HEAD on "/", so a filer
with jwt.filer_signing.read.key set still served its root directory listing --
entry names, sizes and chunks[].file_id -- to a caller holding no token at
all, and served the same listing to a token restricted by allowed_prefixes.

The exemption was added for health checks before the filer had /healthz and
/readyz. Both are registered on the default and read-only muxes ahead of the
"/" handler and answer without a token, so drop it.

Point the mTLS harness at /healthz, which is what it was probing for.

* filer: keep the jwt query parameter out of a proxied chunk request

The proxy stripped "jwt" from the forwarded query on reads only, on the
grounds that a writer's own credential travels there. It does not: an
uploader carries its AssignVolume token in the Authorization header, and the
query parameter on this path holds a filer credential.

Strip it for every method. A volume server has no business seeing a filer
token, and because security.GetJwt reads the query before the header,
relaying one would hide the writer's own token behind it.

* filer: dispatch the chunk proxy after the JWT gate

The ?proxyChunkId= branch returned before maybeCheckJwtAuthorization ran, so
GET, PUT, POST and DELETE against any needle in the cluster were reachable on
the filer's HTTP port with no filer credential, on a filer where every other
request answered 401. An anonymous caller read a stored object, replaced its
bytes, or deleted the needle, which the master's next vacuum makes permanent.

#10434 stopped the filer from minting a volume write token for that caller,
which closes the write half only where the volume server has a jwt.signing.key
of its own -- not the shipped default, and not what scaffold/security.toml
recommends for a filer deployment. The read half stayed open in every
configuration, because the filer mints the read token itself.

Move the dispatch below the gate. A file id carries no path, so a token
restricted by allowed_prefixes cannot be scoped against one and is refused
here; every consumer of this endpoint holds an unrestricted token.

* filer: mint the volume credential for a proxied write too

The proxy minted a volume token on reads and forwarded whatever the caller
sent on writes. #10434 made it that way because the branch ran ahead of the
JWT gate, so a token minted here would have been signed for an unauthenticated
caller; the branch now runs behind the gate, and the credential the caller
presents there is a filer one, which a volume server cannot validate and has
no business seeing.

Mint at the access level the request needs, and drop the caller's
Authorization when there is no key to mint from. A proxied uploader then needs
only the filer credential, instead of holding one for each hop with a single
header to put them in.

* mount, mq, filer.sync: send the filer credential for a proxied chunk

Every in-tree consumer of ?proxyChunkId= reached the filer anonymously: mount
and the broker put the AssignVolume token in the Authorization header, which
is a volume credential, and filer.sync sent nothing at all. That was enough
only while the branch ran ahead of the filer's JWT gate.

Build the URL through one helper, and pick the credential from the URL it
returns: a chunk proxied through a filer is a request to the filer, which
authorizes it and attaches the volume credential itself, so the token there is
a filer one at the access level the request needs.

* filer: honor -exposeDirectoryData

The flag was declared on all three commands that start a filer and read by
none of them: FilerOption.ExposeDirectoryData was only ever assigned from
filer.expose_directory_metadata in security.toml, so -exposeDirectoryData=false
silently left the listing exposed. Only the TOML key had any effect.

Plumb the flag through and let either switch turn the listing off.

* filer: count a proxied chunk request once

Moving the dispatch below the gate put it after the deferred request
observation, so every proxied chunk now landed in FilerRequestHistogram twice,
once under its HTTP method and once under chunkProxy. Name the deferred one
after the proxy instead, the way the unsupported-method branch already does,
which also gives the endpoint the status codes FilerRequestCounter records.
2026-09-04 16:39:36 -07:00
Chris LuandGitHub cda43f1976 filer: do not 404 a TUS session on a transient chunk-load failure (#11153)
* filer: do not 404 a TUS session on a transient chunk-load failure

readTusSessionInfo already proved the session exists before
loadTusSessionChunks is called, so a failure there is a read failure,
not evidence the session is gone: a volume-server timeout or a
canceled request context surfaces through ListDirectoryEntries the
same way a missing session would.

Every such error was mapped to writeTusSessionNotFound, answering 404
to HEAD/PATCH and 204 to DELETE. A spec-compliant TUS client trusts
that and discards the session, orphaning every chunk it had committed
until the 24h expiry sweep, or forever if it never issues a DELETE.

Only an error matching filer_pb.ErrNotFound is now reported as not
found; anything else answers 500 so the client retries against the
same session instead of abandoning it.

* test: cover a TUS session's transient chunk-load failure

Adds a listErr hook to the in-memory test store, alongside the
existing commitErr/deleteErr, to simulate a store or RPC failure from
ListDirectoryEntries.

HEAD, PATCH and DELETE against a live session all answer with a
server error instead of a not-found status when the chunk listing
fails transiently, and the session is left on disk untouched. A
listing failure that genuinely means not found, filer_pb.ErrNotFound,
still answers 404 (204 for DELETE).
2026-09-04 16:39:24 -07:00
Chris LuandGitHub ed9d58873e filer.remote.sync: skip an upload whose source entry was deleted or rewritten (#11149)
* filer.remote.sync: skip an upload whose source entry was deleted or rewritten

A replay from an earlier offset (-timeAgo) re-emits create and update
events for entries the filer has since deleted or rewritten. Their chunks
are gone from the volume servers, so the upload can never succeed, and
failing the event holds the sync offset before it: every restart of the
subscription replays it into the same dead chunks, and progress on
everything after it in the log is never persisted. One such entry stops
replication for the whole mount.

When the upload fails, look the entry up on the filer. Gone, or holding
other content than the event described, the event is superseded and is
skipped with an error log; the event that superseded it follows in the
log and brings the remote to the current state. Otherwise the failure
stands and the event is retried as before.

Fixes #11148

* filer.remote.sync: compare chunks by file id when deciding an event is superseded

filer.IsSameData compares chunk ETags, so a delete-and-recreate of
identical bytes, which stores the same content under new file ids and
drops the old ones, looked still as described and kept failing the event
on its dead chunks. Compare by file id with DoMinusChunks, the way the
filer itself decides which chunks an update leaves for deletion: the
event is superseded when the current entry no longer references every
chunk it named, and still as described when it does, including when more
chunks were appended after it.

* filer.remote.sync: ask the filer on the first failed upload attempt, not after the backoff

The superseded check ran after util.Retry had given up, so every dead
entry still cost the full retry cycle, about 13s, before it was skipped:
the SDK reports a missing chunk as "RequestError", which
IsTransientError takes as worth retrying. Move the check into the retry
loop with util.RetryOnError. Any failed attempt asks the filer, and the
loop stops at once when the entry is gone, surfacing errSuperseded for
the caller to skip. An entry the filer still holds keeps the retry policy
it had.

filer.remote.gateway shares retriedWriteFile and the same offset-pinning
processor, so its three call sites skip a superseded event the same way.
2026-09-04 00:18:37 -07:00
Chris LuandGitHub 06838e28b2 filer: serve "//" paths at the cleaned path instead of redirecting (#11150)
* filer: serve "//" paths at the cleaned path instead of redirecting

http.ServeMux redirects a non-canonical path ("//", "..") to its cleaned
form, but since Go 1.22 it builds the Location from the already-escaped
path, so it is percent-encoded twice (golang/go#79897). A client that
follows the redirect re-posts "/负极全景" as "/%25E8%25B4%259F...", and
the filer stores a directory literally named "%E8%B4%9F...".

Wrap the filer muxes in CleanPathHandler, which rewrites the request to
the same cleaned path ServeMux would have redirected to and dispatches
directly. The decoded name reaches the handler, the round trip goes
away, and clients that do not follow redirects work too.

Fixes #11125

* filer: keep RequestURI in step with the cleaned path

PostHandler derives storage rules, the bucket and the read-only check from
r.RequestURI while writing the entry at r.URL.Path. After CleanPathHandler
rewrote only the URL, a "//" or ".." request would be placed by the raw
path and written to the cleaned one. Rewrite RequestURI too, as the
redirect-following client used to.

* filer: match storage rules on the decoded write path

PostHandler resolved the storage rule from r.RequestURI, the raw
request-target. Clients percent-encode non-ASCII segments on the wire, so
a read-only or TTL rule configured on "/data/只读/" never matched a POST
to "/data/%E5%8F%AA%E8%AF%BB/" and the write went through. Use r.URL.Path,
the decoded path the entry is actually written to, as the header-based
destination check already does. The query string no longer reaches the
rule lookup, so the "?" trimming in the read-only error is gone.
2026-09-04 00:02:33 -07:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
cc281dabc9 master: keep new volumes and writes off servers in maintenance mode (#11147)
* master: keep new volumes and writes off servers in maintenance mode

The master recorded a volume server's maintenance flag from the heartbeat
but never consulted it. A server in maintenance (#7977) is being drained,
yet the master kept creating volumes on it whenever it had free slots and
kept handing out its volumes for writes. Nothing on the volume server
blocks plain HTTP uploads either, so "read-only mode" was only a name.

Volume growth: a data node in maintenance mode reports zero free slots
through AvailableSpaceFor, which takes it out of every candidate list,
feasibility count and capacity reservation. Its slots still roll up into
its rack and data center, so the random offset drawn from those totals for
an other-rack or other-DC replica could land in space the walk then skips
and fail with "No free volume slot found!" while siblings had room; the
walk now folds the offset into the space that is actually eligible. This
also covers the pre-existing case of an over-committed sibling.

Assignment: a replica on a server in maintenance mode is treated like a
read-only replica in isAllWritable, so its volume leaves the writable
list and returns when the flag clears. Topology.SetDataNodeMaintenanceMode
re-evaluates the node's volumes on every change, since heartbeats are
digest-based and a full volume list may not follow for a long time. Reads
and lookups are untouched. The flag moves to an atomic so the assign and
growth paths can read it without the node lock.

Heartbeat: the Go volume server sent its state only when it changed, so a
master elected while a server sat in maintenance never learned about it.
The state now rides along on every heartbeat, as the Rust server already
does; the master's compare is an atomic swap, and only a change does work.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* master: hold maintenance mode through vacuum commit and mark-writable

SetVolumeAvailable and SetVolumeWritable put a volume back on the writable
list on the replica count alone. A vacuum that started before the server
entered maintenance, or a vacuum worker's mark-writable arriving after it,
handed the volume back to assignment with a replica on the draining server.
Heartbeats carry only changed volumes, so nothing re-evaluated it until the
volume itself changed.

Apply isAllWritable on both paths, the same test EnsureCorrectWritables
uses. Also pin that re-evaluating a volume a concurrent disconnect already
removed from its layout is a no-op.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* master: record a server's read-only notification on its node before judging the volume

A volume server notifies the master the moment it flips a volume between
read-only and writable, ahead of the heartbeat that repeats the flag. The
layout only set its per-location flag, so isAllWritable, which reads the
node's heartbeat copy, still saw the old value: a mark-writable was
withheld until the next heartbeat, and a re-evaluation landing between a
mark-readonly and its heartbeat put the volume back on the writable list.

Record the flag on the node's volume first. AddOrUpdateVolume keeps the
digest and the active volume count in step, so the heartbeat that follows
finds nothing to change.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* master: a read-only mark does not confirm a provisional volume

DataNode.SetVolumeReadOnly went through Disk.AddOrUpdateVolume, which
treats its input as a server report and so ended the grace period that
keeps a just-grown volume safe from a full report collected before the
grow. A volume marked read-only before its first report could then be
removed by that stale report.

Give Disk a SetVolumeReadOnly that flips the flag and keeps the digest and
active volume count in step without touching volumeAddedAt.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 23:50:12 -07:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
9fef11526e filer.remote.sync: confirm a missing RemoteEntry against the filer before re-uploading (#11146)
* filer.remote.sync: confirm a missing RemoteEntry against the filer before re-uploading

The event is the entry as it was when the update was logged. A chmod or
utimes right after a write is logged while the sync is still uploading the
write, so it carries no RemoteEntry even though the object is on the remote
by the time it is processed. Gating on the event alone turned every such
update into a delete and a second upload of the same bytes; cp -p, rsync
and Django's FileSystemStorage all write that way.

Look up the filer's current entry when the event has no RemoteEntry: the
upload stamps it as soon as it completes, so the stamp is there for the
race and absent for a file that was never replicated. Skip the update when
the entry has since been deleted rather than upload from chunks that may be
gone; the delete event that follows removes the remote object.

Tests build entries from chunks, which is what IsSameData compares in
production, and cover both no-RemoteEntry cases through a stub filer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer.remote.sync: do not delete the remote object before overwriting it in place

The update write path deleted the old object and then wrote the new one,
even when both are the same key. S3, GCS and Azure all overwrite on write,
so the delete bought nothing and left the remote with no object between
the two calls, or at all if the write then failed and pinned the offset.
On a versioned remote bucket it also left a delete marker per rewrite.

Delete only when the key changes, which is what the delete was for.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer.remote.sync: trim comments

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 19:14:34 -07:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
24b8646ec3 volume: let evacuation proceed on a server in maintenance mode (#11145)
Maintenance mode exists to fence a volume server so it can be evacuated
without taking new writes (#7977), but the gate added in #8115 also
rejected the RPCs evacuation issues against the source: VolumeMarkReadonly
(the first step of every move, and the failure reported in #11066),
VolumeDelete (the last step), and VolumeEcShardsDelete (the last step for
EC shards). volumeServer.evacuate, volume.move and ec.balance therefore
all failed on exactly the server they were meant to drain.

Those three RPCs only remove data or restrict the server further, the same
class as DeleteCollection and the unmount RPCs that were never gated, so
they are exempted from the maintenance check in both the Go and Rust
volume servers. Everything that adds data or reopens the server for
writes (AllocateVolume, WriteNeedleBlob, BatchDelete, VolumeCopy,
ReceiveFile, EC generate/copy/rebuild, vacuum, tiering, VolumeMarkWritable)
stays blocked. A side effect is that scrub can now fence broken volumes
readonly on a server already in maintenance.

Fixes #11066

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 18:39:45 -07:00
Alex KandGitHub c9c6e6fb1d filer.remote.sync: upload files that were never replicated (#11140)
An entry whose content is rewritten unchanged before it first reached the
remote took the metadata-only branch, and UpdateFileMetadata returns early
when the extended attributes match without checking that the object is
there. shouldSendToRemote had already reported the entry as needing to be
sent, so the effect was that it stayed local for as long as its content did
not change, with the sync reporting healthy progress over it.

Require RemoteEntry to be set before treating an update as metadata-only.
Gating at the caller covers the S3, GCS and Azure clients, which share the
same early return.

Fixes #11139
2026-09-03 17:42:40 -07:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
57a5285020 ci: turn off cosign's signing config alongside the bundle format (#11144)
Every image-signing job has failed since signing was added (#11129):

  must provide --new-bundle-format or --bundle where applicable with
  --signing-config or --use-signing-config

Cosign 3 turned on two defaults, not one. The action only disabled
--new-bundle-format to keep the .sig tag layout, but --use-signing-config
is still on, and cosign refuses that pairing because the signing-config
path has nowhere to write its verification material without a bundle.
Disabling it too falls back to the default Fulcio and Rekor URLs, the
same services the .sig layout always used. The verify step needs no
change: cosign verify looks for a referrer bundle first and falls back
to the .sig tag when there is none.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-03 17:22:46 -07:00
Chris LuandGitHub 8112f2733a filer: batch exact lookup RPC, authoritative volume lookup, VolumeDelete status codes (#11122)
* storage: make DeleteVolume errors inspectable with errors.Is

An absent volume wraps ErrVolumeNotFound and an only-empty refusal now
wraps ErrVolumeNotEmpty with %w instead of %v, so callers no longer have
to match on the message.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* volume server: return NotFound and FailedPrecondition from VolumeDelete

An absent volume maps to codes.NotFound and a non-empty volume under
only_empty to codes.FailedPrecondition, so a caller retiring a volume can
treat NotFound as already done. The store message is kept in the status
description because the EC empty-replica sweep still matches on it.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* wdclient: add LookupVolumeIdsAuthoritative

Bypasses the vid map and asks the provider directly, for callers where a
stale positive location is unsafe.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: add LookupDirectoryEntries batch lookup RPC

Up to 4096 exact-path lookups in one call, resolved concurrently with
results in request order, plus one deduplicated location lookup for every
volume the returned entries reference and per-fid read tokens when the
filer signs reads. unavailable_volume_is_miss lets cache-style callers
take an entry whose volume has no live location as a miss, resolved
against the master rather than the filer's location cache.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: test that an expired file entry is deleted on read

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: test that AssignVolume and CreateEntry resolve the same TTL rule

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* master: refuse partial lookups while warming up

LookupVolume returned Unavailable during warm-up only when every requested
volume was missing. A batch mixing a reported volume with one whose server
has not reconnected yet came back as a partial answer with a per-volume
not-found, which a caller treating the master as authoritative reads as
gone. Any not-found during warm-up is now Unavailable, which callers
already retry.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: build batch test requests instead of copying a proto message

Copying a generated message copies its internal mutex, which go vet's
copylocks check rejects.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: match ErrNotFound with errors.Is and state the miss rule's contract

A wrapped not-found from the store would otherwise be reported as an
error rather than a miss. The comments now say why a nil location map is
the only sign of an unanswered lookup: the provider returns nil when it
got no answer and a populated map, with unserved volumes reported as
errors, when the master did answer.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* volume server: map absent and non-empty VolumeDelete errors in the Rust server

Matches the Go server: an absent volume is NotFound and an only_empty
refusal is FailedPrecondition instead of Internal, with the messages the
EC empty-replica sweep matches on.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm

* filer: test that a malformed entry keeps its error outside cache mode

Same test file as the enterprise tree, so the next sync sees one version.

Claude-Session: https://claude.ai/code/session_01T4MEV3ETqFFKN46Uu2ZrUm
2026-09-03 14:52:05 -07:00
Chris LuandGitHub eb9f54cebb telemetry: type MinDiskBytes so a 32-bit int cannot hold it by accident (#11142)
As an untyped constant it became int when passed to Infof, which
overflows on linux/386 and failed the 32-bit vet job. Every field it is
compared against is already uint64.

Claude-Session: https://claude.ai/code/session_015rYAmF8hV9yypb9yvy4A1z
2026-09-03 13:54:25 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
830bd7ad24 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.19.34 to 1.20.1 (#11132)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.19.34 to 1.20.1.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/credentials/v1.19.34...v1.20.1)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.20.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 13:52:58 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
581dcf6b36 build(deps): bump github.com/pierrec/lz4/v4 from 4.1.28 to 4.1.29 (#11133)
Bumps [github.com/pierrec/lz4/v4](https://github.com/pierrec/lz4) from 4.1.28 to 4.1.29.
- [Release notes](https://github.com/pierrec/lz4/releases)
- [Commits](https://github.com/pierrec/lz4/compare/v4.1.28...v4.1.29)

---
updated-dependencies:
- dependency-name: github.com/pierrec/lz4/v4
  dependency-version: 4.1.29
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 13:52:47 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
edbbf951f9 build(deps): bump github.com/prometheus/procfs from 0.21.1 to 0.22.0 (#11134)
Bumps [github.com/prometheus/procfs](https://github.com/prometheus/procfs) from 0.21.1 to 0.22.0.
- [Release notes](https://github.com/prometheus/procfs/releases)
- [Commits](https://github.com/prometheus/procfs/compare/v0.21.1...v0.22.0)

---
updated-dependencies:
- dependency-name: github.com/prometheus/procfs
  dependency-version: 0.22.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 13:52:38 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f8a8571886 build(deps): bump github.com/klauspost/reedsolomon from 1.14.1 to 1.14.2 (#11135)
Bumps [github.com/klauspost/reedsolomon](https://github.com/klauspost/reedsolomon) from 1.14.1 to 1.14.2.
- [Release notes](https://github.com/klauspost/reedsolomon/releases)
- [Commits](https://github.com/klauspost/reedsolomon/compare/v1.14.1...v1.14.2)

---
updated-dependencies:
- dependency-name: github.com/klauspost/reedsolomon
  dependency-version: 1.14.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 13:52:27 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0555d39c48 build(deps): bump google.golang.org/api from 0.294.0 to 0.296.0 (#11136)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.294.0 to 0.296.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.294.0...v0.296.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.296.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-03 13:52:16 -07:00
Chris LuandGitHub 1205630c43 ci: track the actions in the signing jobs by tag again (#11137)
The repository tracks actions by tag with dependabot; the signing jobs
follow the same convention.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa
2026-09-03 12:36:36 -07:00
Chris LuandGitHub 0973634fd4 telemetry: keep only clusters that store at least 10 GiB (#11138)
* telemetry: tidy the server module after the protobuf bump

Claude-Session: https://claude.ai/code/session_01VGiphDxpsKMwu9XpUFhybQ

* telemetry: keep only clusters that store at least 10 GiB

Fresh weed server runs, CI jobs and throwaway containers each mint their
own cluster id. They came in at tens of thousands a day, were most of
the counted clusters and held almost none of the bytes, and the state
file and the metrics page grew with every one of them. Reports under
the floor are counted and dropped, and a state file written before the
floor sheds them on the first restart.

Claude-Session: https://claude.ai/code/session_01VGiphDxpsKMwu9XpUFhybQ

* master: report telemetry only once the cluster stores 10 GiB

A throwaway cluster no longer registers itself with its first report a
minute after start; a real one begins reporting at the first daily tick
after it crosses the floor.

Claude-Session: https://claude.ai/code/session_01VGiphDxpsKMwu9XpUFhybQ
2026-09-03 12:36:00 -07:00
Chris LuandGitHub a8f763e717 Sign the published Docker images with cosign (#11129)
* ci: composite action that signs and verifies an image with cosign

Keyless, by digest, with a verification pass against the calling workflow's
own identity right after signing. Signatures use the .sig tag layout rather
than the OCI-referrer bundle cosign 3 writes by default, since that is what
the verifiers people run today read. Dependabot is pointed at the action so
the cosign-installer pin keeps moving.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker release: sign every variant on both registries

The merge job signs each variant's multi-arch index on GHCR and Docker Hub
once the tag exists, recursively so the platform images are covered too.
latest re-tags the same manifest and inherits the signature.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker dev: sign the dev image

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker latest: sign a latest rebuilt by hand

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker release: sign the foundationdb image

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker: sign the per-version foundationdb and rocksdb builds

They push to the same repository as the releases, so an admission policy
that verifies chrislusf/seaweedfs would otherwise reject them.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker: document image signature verification

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* ci: pin the actions the signing jobs newly run by commit

These run with registry credentials and the OIDC token that signs under
the repository's identity, so a retargeted tag upstream must not be able
to reach them.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker latest: pass the dispatch tag through env, not the script

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker: complete Kyverno policy, digest note, identity scope

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker latest: keep the dispatch tag out of the manifest script too

The step predates signing, but the job now holds the OIDC identity.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* ci: pin every action in the jobs that sign

The jobs that hold the OIDC identity run these with registry credentials,
so a retargeted tag upstream must not reach them.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker release: copy and sign the digest the run created, pin the rest

crane copy and the signature both resolved the tag, which another
publisher could move between the two steps. The index digest is read once,
right after it is created, and the Docker Hub copy and both signatures use
it. The manual latest rebuild gets the same treatment. The actions in these
jobs are pinned to commits, crane to v0.22.0 by checksum, and the sparse
checkout no longer keeps the token.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker latest: the signing job checks out the workflow's own commit

The job only assembles and signs manifests, so nothing there needs the
source_ref checkout; the local signing action now comes from the same
revision as the workflow file that calls it.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa

* docker release: take the index digest from the create result

imagetools create writes the descriptor it pushed with --metadata-file
(buildx 0.32+, the runners ship 0.36), so the digest no longer comes from
re-resolving the tag even within the same step.

Claude-Session: https://claude.ai/code/session_01A5zMqzaUg1Snur4Yg8xJGa
2026-09-03 11:52:10 -07:00
Chris LuandGitHub fe98520358 read: try a replica that stopped answering last, and relearn its volume's locations (#11130)
* http: try a volume server that failed to answer last

A cached location list is shuffled on every read, so once a replica dies
half the reads keep dialing it first and pay a connect failure or timeout
before the healthy replica answers. Remember, per host, when a request got
no answer at all and order such hosts last for the next half minute. Once
that passes, one read probes the host in its usual place while the others
keep it last until the probe settles, so a black-holed server costs one
stalled read per interval instead of one per read.

Nothing is ever skipped: a host that failed is still tried when the others
fail too. Any response, including an error status, counts as reachable.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs

* filer: refresh a chunk's locations after one of them fails

A mount's location cache is only relearned when every cached location
fails. When one replica dies and the other still answers, every read
succeeds and the dead replica stays in the cache, and in the shuffled
order it keeps being dialed first long after the master has dropped it.

When a read fails on one location and a later one answers, call the
refresh hook so the cached entry is dropped and looked up again. The read
that already paid for the failure returns its data; the reads after it
start from the locations the master knows now.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs

* http: claim the probe for every expired host, and try it first

The claim was only checked for the first url, so with two replicas whose
marks expired together the second was probed by every read at once. Claim
each expired host on its own and put the reads that won a claim ahead of
the reachable hosts, so a probe is always a real attempt and a lost claim
always means the host is tried last.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs

* filer: refresh a chunk's locations in the streaming read path too

The streaming loop had no refresh hook, so a manifest or streamed chunk
that failed on one cached location and was served by another kept the
stale entry until every location failed. Give it the same hook as the
buffered loop, built by one refreshUrls function shared by the reader
cache and the stream callers.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs

* http: probe at most one expired host per read

Claiming every expired host in one ordering left all but the first claim
without an attempt, since a read stops at its first answer, and a host that
had come back waited another interval for nothing. Claim only the first
expired host a read sees and leave the rest last and unclaimed, so each
following read probes one of them.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs

* test: start the live server before releasing the dead server's port

Closing the dead server first let the live server come up on the same
port, in which case the dead location answers and the partial failure
under test never happens.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs
2026-09-03 11:51:48 -07:00
Chris LuandGitHub 52db7a6aee test: write more files than the master grows volumes at a time (#11131)
The cached-location test needs two files on one volume, but it wrote six
files into the six volumes a 001 layout starts with, and every so often
each file landed on its own volume and the test had nothing to probe.
Seven files leave no way to spread them out.

Claude-Session: https://claude.ai/code/session_011NYXuzGttwrTMsfLYvmQFs
2026-09-03 10:29:35 -07:00
Chris LuandGitHub 59916d8978 helm: values-driven labels on every ingress (#11127)
* helm: values-driven labels on every ingress

Each ingress already takes annotations from values, but its labels were
a fixed block, so tools that select ingresses by label (ExternalDNS
label filters, for one) had nothing to key on. Every ingress block now
has a labels map rendered after the standard app.kubernetes.io labels,
including the Traefik IngressRouteTCP that shares the filer gRPC values.

Claude-Session: https://claude.ai/code/session_01L6eJGXtYkwe1W9QjGeUgr1

* helm ci: render check for ingress labels

Claude-Session: https://claude.ai/code/session_01L6eJGXtYkwe1W9QjGeUgr1
2026-09-03 09:17:27 -07:00
Chris LuandGitHub c3511e7c86 ci: let codespell past the sme variable in the mount tests (#11121)
weedfs_stream_mutate_error_test.go names its *streamMutateError local
sme, which codespell reads as a misspelling of same/some. It is an
identifier, so exempt it beside the other variable-name entries.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ
2026-09-03 08:56:43 -07:00
Chris LuandGitHub 31fb46f693 volume: rebuild a missing .idx from the .dat (#11115)
* volume: rebuild a missing .idx from the .dat

Pointing -dir.idx at a directory that holds no index aborted the whole
volume server: checkIdxFile found no .idx and load() called glog.Fatalf.
Every row of the index is derivable from the .dat, so walk it in append
order and write the index back, which reproduces byte for byte what the
server's own writes had left in the old directory.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: keep the index co-located with the data in the Rust server

Go's load() drops back to the data directory when an .idx already sits
beside the .dat, so naming a --dir.idx does not strand a pre-existing
index. Rust had no such adjustment: it opened the new directory with
create, and the volume came up on an empty index with every needle
invisible.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: rebuild a missing .idx from the .dat in the Rust server

Mirrors the Go side. Rust did not abort on a missing index the way
checkIdxFile did; it opened the new directory with create and mounted the
volume on an empty index, so every needle read as missing while the .dat
still held the data. Walk the .dat in append order and write the index
back, byte for byte what the server's own writes had left behind.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: stop the idx rebuild at a zero-padded .dat tail

An all-zero needle header is unwritten space, not a record. Go's .dat walk
keeps reading past it and would index a truncated data file's tail as
millions of needle 0 rows; the Rust walk already stops there. Stop the Go
rebuild at the same place.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: create the -dir.idx directory when it does not exist

Rust's DiskLocation creates the index directory as it takes it; Go only
resolved the path, so naming a directory that does not exist yet left every
volume unable to open or rebuild its index and took the server down.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: stop the idx rebuild at a torn .dat record

A crash between writing a needle's header and its body leaves a record
whose declared size runs past the end of .dat. Indexing it puts a row in
the .idx that points at bytes that do not exist, which fails every read of
that needle and trips the past-EOF check on the next load. Stop at the
first record that does not fit, in both servers.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: stop the idx rebuild at a negative-size header

A corrupt header whose size field is negative makes the .dat walk advance
backwards: NeedleBodyLength adds the negative size, so the next offset is
lower than the current one. The Go walk then reads at a negative offset and
the rebuild fails, which puts the volume server right back to exiting at
startup; the Rust walk seeks past EOF and truncates the index instead.
A negative size is never a record, so stop there.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: skip a volume whose index cannot be rebuilt, do not exit

glog.Fatalf calls os.Exit(255), so a rebuild that could not write -- a full
or read-only index directory -- put the server right back to dying at
startup for one bad volume. Return the error instead: loadExistingVolume
logs it and skips that volume, which is what the remote-volume branch just
above already does and what the Rust loader has always done.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* volume: create the index directory from the rebuild too

The rebuild is the first thing to write into a fresh -dir.idx, and it runs
before the loaders that create the directory on their way to opening .idx.
Create it in both rebuilds so the ordering does not matter.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ

* ci: let codespell past the sme variable in the mount tests

weedfs_stream_mutate_error_test.go names its *streamMutateError local
sme, which codespell reads as a misspelling of same/some. It is an
identifier, so exempt it beside the other variable-name entries.

Claude-Session: https://claude.ai/code/session_01BYrb2AdJSckq9FdHuqseDJ
2026-09-03 08:43:10 -07:00
Chris Lu e35b418693 readme: quick start commands that work on a laptop
-dir=/data fails on macOS, where the root filesystem is read-only, and
on any Linux box without root; -dir=./data is created on the spot.
go install of the weed package is refused because go.mod carries
replace directives, so the install script is the shortcut instead.

Claude-Session: https://claude.ai/code/session_014apMEkkquAtAYp89paTkAT
2026-09-03 00:14:50 -07:00
Chris Lu f877b99c90 readme: fix the quick start examples
The Helm values put filer metadata on a claim through filer.data, which
is what the chart reads; enablePVC was rendering a hostPath. Claims use
the cluster default storage class instead of local-path. The AWS CLI
test carries its own credentials, the compose download includes the
Prometheus config the compose file mounts, and the disk-read claim is
per blob, since large files are chunked.

Claude-Session: https://claude.ai/code/session_014apMEkkquAtAYp89paTkAT
2026-09-02 23:55:16 -07:00
Chris LuandGitHub 9f6efc8b53 filer: a listing over a hard link no longer deadlocks a bounded SQL pool (#11118)
* filer: give the SQL stores' key-value reads their own connections

A listing holds the connection its rows are on for the whole iteration, and
FilerStoreWrapper calls maybeReadHardLink -> KvGet from inside that iteration,
so a hard-linked entry needs a second connection while the first is still busy.
Out of one bounded pool that is a deadlock: the listings fill the pool and then
wait for a connection none of them will release, and the wrapper's
context.WithoutCancel leaves the waiters without a deadline, so the filer stays
wedged rather than erroring.

The sqlite store shows it at its sharpest -- it allows a single connection, so
one listing over one hard-linked entry never returns. On postgres with
connection_max_open = 50, 60 concurrent listings over hard-linked entries made
no progress at all.

Key-value reads now run on their own pool, carved out of connection_max_open
rather than added to it, so the operator's cap still bounds what the store opens
against the database. An unbounded pool keeps a single pool: nothing can wait
there. sqlite's single connection becomes two, one per pool, and its writes get
a busy timeout so a write that meets the reader waits instead of failing.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

* sqlite: keep both pools on one database, whatever the dbFile spells

A dbFile that already carries URI options got a second "?" appended, which the
driver reads as part of the preceding option value, and a bare :memory: is
private to each connection, so the key-value pool would open its own empty
database and every key-value operation would fail on a missing filemeta.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

* sqlite: assert the busy timeout on the in-memory DSN too

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t
2026-09-02 23:45:45 -07:00
Chris LuandGitHub 9b61289293 Remove the RDMA sidecar prototype and its mount client (#11119)
* rdma: drop the sidecar prototype

The Rust engine under it never touched a wire: rdma.rs fabricates pattern
bytes and the crate's default feature is mock-ucx, with real-ucx unimplemented
since the directory landed. Nothing builds it, no CI runs it, and its only
consumer is weed mount's RDMA client, removed next. Two 22MB binaries were
committed along with it.

Claude-Session: https://claude.ai/code/session_01X3zhqLYwQwEQCrRbuzQKvy

* mount: remove the RDMA client that spoke to the deleted sidecar

Its only server was the sidecar's HTTP API, and the path could never have
worked in production anyway: it served a single chunk per call, ignored the
buffer's chunk boundaries, and had no test. Removing it also removes the
per-handle cumulative-offset cache, which nothing else used.

The -rdma.* mount flags go with it. They defaulted to off and pointed at an
address no released build ever listened on.

Claude-Session: https://claude.ai/code/session_01X3zhqLYwQwEQCrRbuzQKvy
2026-09-02 23:44:06 -07:00
Chris Lu 87bead0ab2 readme: drop the streaming clause from the data warehouse bullet
Claude-Session: https://claude.ai/code/session_014apMEkkquAtAYp89paTkAT
2026-09-02 23:42:47 -07:00
Chris LuandGitHub 1fc80df187 Update README.md 2026-09-02 23:36:41 -07:00
Chris Lu 144a6c68d0 readme: keep the two objectives at the top
Claude-Session: https://claude.ai/code/session_014apMEkkquAtAYp89paTkAT
2026-09-02 23:33:44 -07:00
Chris LuandGitHub 8aea9c6ab2 readme: get to the point (#11120)
Lead with what SeaweedFS is and how to start it: one command, Docker,
Docker Compose, a production-shaped Helm values file, build from
source, scale out. Then why: fast, scalable, the S3 API surface with
operation counts, the lakehouse with S3 Tables and the engines that
share it, the cloud cache, and cross-cluster replication and the rest
of the feature list, each pointing at its wiki page.

The blob store walkthrough and the master and volume server internals
move to the Blob Store Architecture wiki page. The comparisons,
benchmark, enterprise and license sections stay. The dev plan is gone,
it was done.

Claude-Session: https://claude.ai/code/session_014apMEkkquAtAYp89paTkAT
2026-09-02 23:27:46 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
179b273350 build(deps): bump google.golang.org/grpc from 1.82.1 to 1.83.1 in /seaweedfs-rdma-sidecar (#11116)
build(deps): bump google.golang.org/grpc in /seaweedfs-rdma-sidecar

Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.82.1 to 1.83.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/compare/v1.82.1...v1.83.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.83.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-02 22:15:29 -07:00
9089a546fb shell: exit non-zero when a piped command fails (#11117)
* shell: non-interactive mode exits non-zero when a command fails

A failed command in a piped weed shell run printed 'error: ...' but the process
still exited 0, so a CronJob wrapping e.g.

  echo 's3.lifecycle.run-shard -shards 0-15' | weed shell -master=...

reported green while the run aborted partway (shards N+1..15 unwalked). An
unknown command likewise exited 0.

RunShell now returns the last command failure from the non-interactive stdin
path (unknown commands included), and the shell command exits 2 on it.
Interactive sessions are unchanged: errors are shown to the operator and the
session continues, exiting 0 as before.

* shell: route the piped-failure exit through main's shutdown path

Review follow-up: os.Exit(2) inside the shell command skipped main's shutdown
work. The command now records the status (SetCommandExitStatus) and returns
normally; main applies it via setExitStatus before exit(). exit() itself now
flushes sentry before os.Exit -- main's deferred sentry.Flush never ran on this
path (os.Exit skips defers), so the existing 'flush buffered events before the
program terminates' intent only worked for the autocomplete early-return.
Exit status 2 on a failed piped run is preserved (verified: piped success
exits 0, piped failing command exits 2).

* shell: test the registered-command failure path

Review follow-up: the error-propagation test only covered unknown commands.
A fake registered command now drives processEachCmd's real dispatch path:
a failing Do surfaces its exact error (errors.Is) and a succeeding one
returns nil. The non-interactive exit status itself is main-level plumbing,
verified end to end against the reproduction (piped failure exits 2).

* shell: trim the comments added with the exit status

Keep the non-obvious why -- why a piped run has to fail its wrapper, why the
status is recorded instead of os.Exit'ed -- and drop the narration.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

* shell: fail a piped run with the status weed already uses for that

weed.go spends 1 on a command that failed and 2 on a usage or syntax error, and
runShell returns true precisely so the usage dump is skipped. Exiting 2 there
told a wrapper the command line was wrong.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

---------

Co-authored-by: Carlos Leyva <carlos.leyva@idener.es>
2026-09-02 22:12:51 -07:00
241541c026 filer: SQL store pool defaults that survive a concurrent walk (#11110)
* filer: SQL store pool defaults survive concurrent walks (idle == open == 50, lifetime 300s)

The code defaults for the four SQL stores were connection_max_idle=2 with NO
default for connection_max_open (unlimited) or lifetime, while the scaffold
filer.toml documents 10/50/300 -- so an env-configured or minimal-toml filer got
the worst possible pool. Under a concurrent listing burst (s3.lifecycle.run-shard
walks 16 shards in parallel) every operation released above the 2 idle slots
closes its TCP connection, so the walk opens a fresh connection per operation
until the filer exhausts its ephemeral ports:

  list /buckets/... : failed to connect ... dial tcp ...:5432:
  connect: cannot assign requested address

Measured on a production filer: 0 -> 28k TIME_WAIT with only ~1.3k concurrent,
and in the minimal docker-compose reproduction (2000-dir bucket, port range
narrowed to 400): the whole range in TIME_WAIT with only ~12 ESTABLISHED.

Default all three knobs, with idle == open so released connections are kept and
reused: idle connections only accumulate up to the actual peak concurrency and
connection_max_lifetime_seconds recycles them, so a quiet deployment holds
nothing extra. An explicit 0 still disables the caps as before. The scaffold's
connection_max_idle moves 10 -> 50 to match.

With this change the same reproduction completes all 16 shards with the default
configuration (TIME_WAIT peak 19 vs the whole port range).

* filer: trim the SQL pool default comments

One line of the non-obvious why is enough; the rest narrated the code.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

* filer: leave the SQL stores' connection_max_open unset

A listing holds its connection for the whole row iteration while its callback
runs another query -- FilerStoreWrapper.maybeReadHardLink does a KvGet per
hard-linked entry -- so every concurrent listing needs two connections from the
same pool. With a default cap, listings past the cap wedge: 60 concurrent
listings over hard-linked entries made no progress at all against a 50
connection pool, and the wrapper's context.WithoutCancel leaves the waiters
without a deadline.

The idle pool is what fixes the connection churn: idle 50 with an unbounded
max_open holds the same 14 postgres sessions across a 16-way listing burst that
opened 455 with idle 2.

Claude-Session: https://claude.ai/code/session_018DWwctzD4T2DmnczPRM47t

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-02 22:11:44 -07:00
Chris LuandGitHub 292145303f mount: name the disk without changing what is mounted (#11114)
mount: name the disk after the mount point when the whole tree is mounted

The mounted path was the only thing that named the disk, so a mount of the
whole tree was labelled with the filer address and the only way to give it
a name was to mount a subtree under that name — which hides everything
outside it. Fall back to the mount point's own name first, so
-dir=\\seaweedfs\Images labels the disk while -filer.path stays "/".

Claude-Session: https://claude.ai/code/session_01Q9f8pWBXu1ceJvcQfYRQ7x
2026-09-02 21:36:40 -07:00
7465b6a80f fix(mount): make a concurrent duplicate mkdir fail with EEXIST instead of both succeeding (#11079)
* fix(mount): make Mkdir exclusive so a concurrent duplicate fails with EEXIST

Mkdir sent CreateEntryRequest without OExcl, so the filer treated a
concurrent duplicate as an update and reported success to both callers;
the kernel's pre-mkdir lookup only masks this when the winner's create
is already visible. Set OExcl, map the entry-already-exists sentinel to
EEXIST instead of EIO, and drop the parent's children cache on the
losing side so the next lookup fetches the winner's entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mount): route exclusive creates to the path's owner filer

The filer's per-path lock is filer-local and the store insert has
upsert semantics, so two mounts streaming to different filers can both
create the same path even with OExcl (measured 18/30 both-success on a
3-filer cluster). Hash the path over the sorted filer list so every
mount sends the same path's exclusive create to the same owner filer:
keep the mutation stream when it already targets the owner, fall back
to it when the owner is unreachable. Also let doUnary hand failed
creates to CreateEntry so the structured error code survives as EEXIST
instead of collapsing into the stream's generic EIO. Same race after
the change: 30/30 exactly one winner, every loser fails with EEXIST.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mount): review fixes — pin exclusive creates to the owner filer

An OExcl create now goes only to the path's owner filer: retrying on a
different filer would race the owner's possibly still-in-flight create
through a separate per-path lock, the very hole this routing closes. A
broken mutation stream retries the same owner over unary, and an
unreachable owner fails the create instead of degrading.

Pick the owner by rendezvous hashing so the choice is independent of the
configured filer order, and mounts configured with different but
overlapping lists still agree wherever the winning filer appears in
both.

Reject a stream create wrapper whose nested response is nil instead of
handing it to CreateEntry, which would dereference it.

Add ownerFilerAddress unit tests: order independence, subset agreement,
spread.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* mount: drop the client-side owner ring, the filer routes exclusive creates now

Exclusive creates are arbitrated cluster-wide on the server: the filer resolves
an OExcl create's ring owner and forwards one hop, so one filer's per-path lock
binds every creator — mount, S3, the HTTP surface and the Java client alike, not
only the ones that opted into a client-side ring.

That makes exclusiveCreateEntry redundant. It hashed the mount's configured
-filer list, which names a different owner than the master-maintained ring, and
failed the mkdir outright when its chosen owner was unreachable rather than
letting the ring reassign. Mkdir goes back to streamCreateEntry; OExcl and the
EEXIST mapping stay, and now mean what they say.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ

* mount: cover the create error plumbing that turns a lost race into EEXIST

Letting a failed create's structured code survive doUnary is what makes a lost
mkdir race report EEXIST instead of EIO, and it had no test. Pull the two steps
out so they can be exercised without a live stream: hasCreateResponse decides
whether a response still carries a code to unwrap, createEntryFromResponse does
the unwrapping.

Reading the guard the other way round also says what it means — consume the
response only when there is no nested code left to recover — rather than
negating a type assertion inline.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ

* mount: do not trust a create reply's shape before reading it

createEntryFromResponse read cr.ErrorCode without checking the nested response
was there. Nothing our filer sends is shaped that way, but the mount reads this
off the wire and a nil there panics the whole mount, so report it instead.

A top-level failure whose nested response carries no code was also returned as
success, silently losing the error. Fall back to the top-level errno when the
nested response explains nothing.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-02 21:26:04 -07:00
Chris LuandGitHub 9e34426a56 lance: a maintenance job that sorts a table by its declared fields (#11113)
* lance: a maintenance job that sorts a table by its declared fields

Lance appends fragments in write order and has no notion of a sorted table, so
nothing but a rewrite establishes one, and nothing but another rewrite restores
it once rows have been appended. lance_sort reads the order from the dataset's
own configuration, falls back to the worker's, and rewrites the table in it.

The spec and the marker live in crates/sort rather than in the job, because
weed/worker/tasks/iceberg sorts too: two jobs that disagreed about what
"id desc nulls-first" means would be two features wearing one name.

The sort spills. lance builds its DataFusion runtime with a FairSpillPool and a
disk manager, but only when LanceExecutionOptions::use_spilling is set, and that
struct derives Default over a plain bool — so Scanner::try_into_stream, which
fills its options with ..Default::default(), is precisely the path that does not
spill. The job builds the plan with create_plan and executes it with spilling on
and the operator's memory budget.

The marker rides in the same commit as the data: Operation::Overwrite is the one
operation carrying config values alongside fragments, so a sorted table and the
record of its sorting cannot disagree. It records the version the sort read, not
the one it wrote, which is not knowable while the marker is being assembled.
Detection treats anything committed after the sort's own commit as data the sort
did not produce — row counts alone cannot see a rewrite that leaves the count
where it was, and such a table would look sorted forever.

Claude-Session: https://claude.ai/code/session_015SZkLTUvd1svDu4xdr6Q3y

* lance: identify a sorted table by the files it wrote, not its version

Review found three ways the version-based marker misjudges a table, and they
share a cause: the version a sort produces is not knowable while the marker is
being assembled, so the marker recorded the version it read and detection
inferred the rest. A commit that rebases past a conflict lands on a different
number, and the inference then reads a rewrite into an ordinary table — a full
re-sort, and its indices, for nothing.

Data file names do not have that problem. They are chosen before the commit, so
the commit can carry them, and they do not change with the version it lands on.
The marker now records how many files the sort wrote and a digest of their
names, and detection asks whether the table still holds them: the same files
means untouched, the same files followed by more means appended, anything else
means the data was replaced.

That also closes the hole the row threshold left. A replacement that grew the
table by fewer rows than min_unsorted_rows read as sorted, however many rows had
actually moved; the threshold now applies only where the sorted files are still
in place, which is what it was for. A marker without a row count is stale rather
than a zero to compare against, and deletes stop forcing a re-sort — they write
a deletion file beside the data rather than rewriting it, and removing rows does
not unsort the ones that remain.

Sort fields are also compared exactly rather than case-folded. Arrow schemas are
case-sensitive, so `id` and `ID` are two columns, and folding them together
rejected a valid order.

Claude-Session: https://claude.ai/code/session_015SZkLTUvd1svDu4xdr6Q3y

* lance: count the rows appended after a sort, not the table's net growth

Review found that rows deleted from the sorted fragments hide appended rows one
for one: the threshold compared the live row count against the count recorded at
sort time, so 800 deletions and 300 appends read as a table that shrank, and a
table where deletions keep pace with appends stays "sorted" with an unsorted
tail forever.

The fragments say it directly. The marker already records how many fragments the
sort wrote, so the ones after that prefix are exactly what arrived since, and
the manifest carries each fragment's live row count — physical rows less its
deletions. Counting those is the arithmetic the threshold was always meant to
do, and it needs no row count from the marker at all.

A fragment whose length the manifest does not record cannot be counted, and a
table that cannot be judged is one to sort rather than one to leave alone
forever, so an uncountable appended fragment reads as stale.

Claude-Session: https://claude.ai/code/session_015SZkLTUvd1svDu4xdr6Q3y
2026-09-02 21:25:50 -07:00
Chris LuandGitHub 97b54adcf6 iceberg: sort compaction bins on disk instead of in memory (#11112)
A sorted rewrite collected every row of a bin into one slice and sorted it
there, so a bin larger than the worker's heap could not be sorted at all.
sort_max_input_mb existed for that reason and skipped the bins it capped.

parquet-go's SortingWriter buffers sort_buffer_rows rows, encodes each buffer
as a sorted run, and merges the runs at close; backing those runs with a
FileBufferPool keeps them in files rather than on the heap. sort_spill_dir says
where, defaulting to the system temp directory — NewFileBufferPool resolves an
empty path to the working directory, which is not what an unset setting means.

The output now also declares its sorting columns, which the plain writer the
sorted path used never did.

Claude-Session: https://claude.ai/code/session_015SZkLTUvd1svDu4xdr6Q3y
2026-09-02 20:17:30 -07:00
Chris LuandGitHub 8398af3572 filer: route exclusive and conditional creates to the entry's ring owner (#11109)
* proto: resync the java copy of filer.proto

The Makefile keeps other/java/client/src/main/proto/filer.proto a verbatim copy,
but AssignVolumeResponse.fsync and SubscribeMetadataResponse.flushed_ts_ns
landed without it. Copy them over; no behaviour change.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ

* filer: route exclusive and conditional creates to the entry's ring owner

CreateEntry with o_excl is a FindEntry-then-Insert. The per-path lock added for
it makes that atomic only on the filer running it, and the store's insert is an
upsert on every backend, so two filers both pass the existence check and both
report success. mkdir(2) then succeeds twice for the same path. The same hole
sits under the condition precondition, whose comment already told callers to
route the key's writes to the owner filer themselves.

Do it on the server instead, with the mechanism ObjectTransaction already uses:
resolve the entry's ring owner and forward one hop, bounded by is_moved. The
ring's membership comes from the master, so it tolerates a stale view and
reassigns when a filer dies, neither of which a client's configured filer list
can do. Every creator gets this — mount, S3, the filer's own HTTP surface, the
Java client — not only the ones that opted in.

Plain creates are upserts whoever applies them, so they stay local and pay
nothing. The route key shares the S3 gateway's namespace so an object's
ObjectTransaction and its CreateEntry land on the same filer's per-path lock.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ
2026-09-02 19:56:11 -07:00
Chris LuandGitHub 938a15eb98 filer: keep a moved key on its prior owner while the ring settles (#11108)
ObjectTransaction forwards to the ring owner so one filer's per-path lock
arbitrates every writer of a key. But a ring change hands the key over before
the new owner has rebuilt the locks the prior owner still holds, so for the
cooling-off window both can grant it. LockRing.PriorOwner exists for exactly
this and nothing consulted it.

Route to the prior owner while that window is open. LockRing.WriteOwner
resolves prior-else-current under one read lock, so the pair cannot come from
different rings and name the same filer twice.

An unreachable owner fails the request rather than falling back to the current
one. gRPC reports a response lost in transit as Unavailable, indistinguishable
from a request the owner never saw, so re-sending elsewhere could re-apply what
the owner already committed; and an owner unreachable from here may be
partitioned rather than down, still serving the key to everyone else — which is
the split brain the routing exists to prevent. The window is bounded: once it
closes the ring hands the key to its new owner.

The owner resolution and the forward move into writeOwner/forwardToWriteOwner
so the next routed RPC reuses them rather than copying the block.

Claude-Session: https://claude.ai/code/session_01Fx1Hx8RqsJqHpbfbgTf4WJ
2026-09-02 19:40:53 -07:00
0f05957bc4 filer: self-heal chunk manifest reads when volume locations go stale (#11107)
* filer: self-heal fetchWholeChunk on stale volume locations

Upstream #10156/#10800 wired cache invalidation into the buffer-based
read paths, but manifest resolution still goes through fetchWholeChunk,
which returns the raw error on failure. When cached volume locations
are stale (volume tiered to remote storage, server rolled), resolving
a large multipart file fails permanently even though other locations
are healthy.

Thread the ChunkGroup's cacheInvalidator through ResolveChunkManifest /
ResolveOneChunkManifest / fetchWholeChunk, and on failure invalidate,
re-lookup and retry once via the existing retryFetchWithFreshLocations
helper. The streaming bytesBuffer is reset before the retry so partial
bytes from the failed attempt cannot corrupt the manifest
proto.Unmarshal. Non-mount callers pass nil and keep their semantics.

* filer: move the manifest self-heal tests in with the other manifest tests

Also make the stale server stream a prefix and then abort mid-body, which is
what actually leaves partial bytes in the buffer: an HTTP error status returns
before ReadUrlAsStream ever calls the writer, so a 500 never exercised the
Reset the tests claimed to cover.

Claude-Session: https://claude.ai/code/session_01FK3oGC5ZVeJYvNBWgb9JUD

* filer: keep the cached volume locations when a manifest read is cancelled

A cancelled or timed-out read says nothing about where the volume lives, so
dropping the location and going back to the master only costs the next reader
a round trip. PrepareStreamContentWithThrottler already guards its self-heal
this way. The guard also goes inside retryFetchWithFreshLocations, since the
caller can be cancelled between its own check and the invalidation, and that
covers the reader cache and prefetch paths too.

fetchWholeChunk returns the context error rather than the stream failure it
provoked, and ResolveOneChunkManifest wraps with %w so errors.Is still sees it.
That matters even where no invalidator is passed: volume.fsck resolves
manifests with nil and tells its own abort from a corrupt manifest that way,
so the cancellation check sits ahead of the nil-invalidator return.

Claude-Session: https://claude.ai/code/session_01FK3oGC5ZVeJYvNBWgb9JUD

* filer: self-heal manifest reads on the filer and s3 paths too

Every caller that already holds the location cache backing its lookup function
can hand it over: the filer's read, copy and deletion paths and the log cache
have the MasterClient right there, and s3api has the FilerClient. MinusChunks
takes one for the same reason, since the deletion path resolves manifests
through it. Only the shell tools and the replication sinks, whose lookup
functions cache privately with nothing to invalidate, keep passing nil.

Claude-Session: https://claude.ai/code/session_01FK3oGC5ZVeJYvNBWgb9JUD

---------

Co-authored-by: bruce-zzz <bruce.zou@hhy-data.com>
2026-09-02 17:43:46 -07:00
Chris LuandGitHub cd064f6eef shell: clean up the target copy when a merge upload fails (#11104)
* shell: clean up the target copy when a merge upload fails

A replicated write commits the needle to the local volume before it fans out
to the other replicas, so an upload that reports failure can still have left a
copy on the target. fs.mergeVolumes printed "failed to move" and carried on,
so that copy stayed behind forever: the filer is never re-pointed at it, and
nothing else knows it exists.

One sick replica orphans roughly half the chunks of a merge, two thirds with
three copies, since the entry node is picked at random from the replicas and
the replica upload uses MaxAttempts 1. A volume with a single copy has no such
window: the write is one local append that either succeeds or leaves nothing.

Delete the needle we may have written before continuing. The source side
already did exactly this, so deleteMovedSourceNeedles is renamed to
deleteOrphanedNeedles and reused for both ends.

Claude-Session: https://claude.ai/code/session_01XjiMGK72F4Gs3yWNJhnVjy

* shell: verify the cookie before deleting a merge target needle

The target cleanup deletes a needle an upload may or may not have written, and
BatchDelete matches on the needle id alone. Needle ids come from one global
sequence, so normally nothing else can hold that id — but a volume restored
from elsewhere, or one written either side of a master sequence reset, can, and
then a failed move deletes a live needle out from under its filer entry.

Have the volume server verify the cookie for those. Source needles keep
deleting by id: they are the ones the filer just pointed at, matching every
other filer-driven delete.

Claude-Session: https://claude.ai/code/session_01XjiMGK72F4Gs3yWNJhnVjy

* shell: delete the target copies an abandoned manifest rewrite leaves

rewriteManifestChunk moves sub-chunks one at a time and only then uploads the
rewritten manifest. Every error after the first successful move — a nested
rewrite failing, the marshal, the manifest upload — returned without touching
the copies already written to the target volumes. The filer keeps pointing at
the old manifest, so those copies orphan, one per sub-chunk moved so far.

Track them alongside the sources and delete them on the way out. Nested
rewrites hand theirs up so an outer failure clears the whole subtree.

A failed UpdateEntry deliberately still leaks its copies: that error can also
mean the filer applied the update and lost the response, and deleting there
would turn a leak into data loss.

Claude-Session: https://claude.ai/code/session_01XjiMGK72F4Gs3yWNJhnVjy

* shell: give the plan its room back when a manifest rewrite is abandoned

allocate reserves plannedSize against the chosen target for every move a
multi-target source makes, and release hands it back when the move fails.
Abandoning a manifest rewrite now deletes the copies that did land, so those
reservations stopped matching anything on disk: the plan kept counting bytes
that are gone and refused later chunks with "no target volume has room".

Release them alongside the delete. Nested rewrites hand theirs up so an outer
failure unwinds the whole subtree's accounting.

Claude-Session: https://claude.ai/code/session_01XjiMGK72F4Gs3yWNJhnVjy
2026-09-02 17:32:11 -07:00
1d0b97f4c6 avro: field time.Time <> iceberg.date (#11091)
* iceberg: normalize foreign day partitions during manifest rewrite

* test: cover manifest rewrite with foreign day partitions

* iceberg: restore every foreign partition value, not just day transforms

iceberg-go takes a partition field's logical type from the last branch of
its Avro union, so a writer that spells an optional partition [<type>, null]
rather than [null, <type>] leaves the value as whatever the Avro decoder
produced. A day or date partition then arrives as a time.Time the manifest
writer cannot encode, and a time partition is worse: time.Duration converts
to int64 nanoseconds and silently records the wrong value.

ReadManifest sits next to ReadManifestList, the other shim for what foreign
writers put on the wire, and converts each partition value back to the
Iceberg representation for its field type.

Claude-Session: https://claude.ai/code/session_01FdQyRuWF9SuCnPn21iH9yR

* iceberg: read manifests that carry partition values through the shim

Compaction, delete rewrite and their detection passes read entries and write
the same partition values back into new manifests, so they fail on a foreign
day partition exactly as manifest rewrite does. Where filters see it too:
literalMatchesActual falls through to fmt.Sprint, so a time.Time renders as a
timestamp and never matches the day the user asked for.

The two remaining readers, orphan collection and the admin preview, only look
at file paths and stay on iceberg.ReadManifest.

Claude-Session: https://claude.ai/code/session_01FdQyRuWF9SuCnPn21iH9yR

* iceberg: convert partition values before the writer rebinds logical types

Dimonyga checked the manifests of a live Doris table: every input spells the
partition union null-first, with the date logical type present, so the union
ordering is not what breaks the merge.

The conversion is lazy. iceberg-go converts what the Avro decoder returned on
the first Partition() call, using the logical types read from the manifest
being parsed, and ManifestWriter.addEntry rebinds them to the manifest it is
about to write before it makes that call. A day partition is where the two
disagree -- iceberg-go's day transform reports an int32 result type, so the
manifest it writes carries no date logical type at all -- and an entry nobody
looked at in between converts against that and keeps its time.Time.

That is why only rewrite_manifests failed: compaction and delete rewrite group
entries by partitionKey(df.Partition()) first, which converts them, and a where
filter does the same. Reading every entry's partition here converts them all
while the manifest's own logical types are still in place.

Claude-Session: https://claude.ai/code/session_01FdQyRuWF9SuCnPn21iH9yR

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-02 17:24:23 -07:00
7620e96171 expose whether a volume replica is backed by remote storage, and prefer local replicas (#11105)
* expose whether a volume replica is backed by remote storage

Volume locations returned by lookups do not indicate whether a replica
has been tiered to remote storage. Readers cannot distinguish a local
replica from a remote-backed one, so they may hit a remote-backed
replica first even when a local replica is available.

Add DataInRemote to the lookup location message, populate it from the
master's volume info, and carry it through the wdclient vid map so
clients can prefer local replicas when resolving chunk locations.

* wdclient: prefer local volume replicas over remote-tier replicas on lookup

LookupFileIdWithFallback (and the publicUrl variant in FilerClient)
didn't honor the DataInRemote flag when shuffling URLs, so the
DataInRemote patch only took effect in LookupVolumeServerUrl. Apply
the same ReorderToFront(localUrls) to sameDcUrls/otherDcUrls so
non-remote replicas stay at the front, matching the existing vidMap
convention.

* wdclient: propagate DataInRemote across tier transitions on existing replicas

When a volume is tiered to remote storage or a remote-backed replica is
restored locally, the cached DataInRemote on the same volume-server URL
stayed at its old value because two pieces of state never updated:

* master_grpc_server.go only split newVolumes and (already-tracked) volumes
  into NewVids vs RemoteVids. ChangedVolumes went straight to NewVids, so
  the broadcast announced the re-classified volume as a fresh arrival and
  the client had no way to tell whether its existing cache was stale.

* vid_map.addLocationToMap early-returned when an entry already had the
  same URL. A tier transition reports the same URL with DataInRemote
  flipped, so the cached entry stayed at the old classification.

Wire both sides together: ChangedVolumes now go through the same IsRemote
split as newVolumes, and addLocationToMap replaces the existing entry in
place when the URL matches but DataInRemote has changed. The server
reference key only depends on URL/grpc port, so the refcount does not
move across the flip.

Adds vid_map_remote_transition_test.go covering the local->remote and
remote->local paths so the in-place update and the cache-key stability
are pinned by tests.

* wdclient: prefer local replicas across data-center boundaries

The previous local-first ordering hoisted local URLs to the front of each
data-center bucket separately, then concatenated same-DC before other-DC.
That meant a same-DC remote replica could still be tried before an
other-DC local replica even though the local one would answer cheaply.

Reorder once across the full candidate list: concatenate same-DC and
other-DC first, then ReorderToFront pulls every local replica to the very
front while preserving the DC preference inside each tier. Apply the same
ordering in all four lookup paths so the cached vidMap, the
LookupFileIdWithFallback provider path, FilerClient.GetLookupFileIdFunction
(PublicUrl-preferred variant), and the deprecated filer.LookupFn all agree:
- weed/wdclient/vid_map.go (LookupVolumeServerUrl)
- weed/wdclient/vidmap_client.go (LookupFileIdWithFallback)
- weed/wdclient/filer_client.go (LookupFileId)
- weed/filer/reader_at.go (LookupFn)

Strengthen the existing local-first tests: vidmap_client_localfirst_test
now asserts both endpoints are present (not just the local one is first),
and slice_test asserts an exact match instead of accepting two orderings.

Add TestLookupFileIdWithFallbackGlobalLocalFirst to pin the cross-DC
ordering invariant: any local replica (same or other DC) precedes every
remote-tier replica; within each tier DC1 precedes DC2.

Add docstrings to ToVolumeLocations, ReorderToFront, LookupVolumeServerUrl,
LookupFileId, GetVidLocations, GetLocations, LookupFileIdWithFallback, and
updateVidMap so the touched lookup paths are described in one place.

* topology: broadcast tier transitions on existing replicas

When a volume replica is tiered to remote storage or restored locally, the
wdclient's cached DataInRemote went stale: every connected client kept
preferring a remote-backed replica over a freshly restored local one, or
demoted a freshly tiered remote replica. The fix in commit 116982595 routed
ChangedVolumes to NewVids/RemoteVids on the master, but ApplyVolumeChanges
returned only fresh arrivals and previously servable replicas. An existing
replica whose IsRemote() classification flipped was neither, so it never
reached the broadcast loop and the wdclient never learned.

Make Disk.doAddOrUpdateVolume return a third signal -- tierTransition --
true exactly when an existing replica's IsRemote() flips. ApplyVolumeChanges
treats that as an arrival so the existing SendHeartbeat routing loop now
sees it. Add a master-side end-to-end test covering local->remote,
remote->local, no-op re-reports, and a mixed heartbeat that only announces
the tier transition.

Also add docstrings to LookupFileId, wdclientLocationsToPb, and
LookupVolume where the prior change touched their bodies.

* topology: broadcast tier transitions received through full reconciliation

The previous commit added tier-transition routing on the ChangedVolumes
delta path, but that is not the only way a re-tiered replica reaches the
master. After a digest mismatch the volume server resends a full Volumes
list, and SyncDataNodeRegistration applies the new IsRemote() classification
silently -- the changedVolumes return value was being thrown away. The
master therefore never broadcast NewVids/RemoteVids, and a wdclient connected
during the recovery kept the stale DataInRemote until it lost contact with
the master.

Surface the changed set through UpdateVolumes.changedVolumes (now covering
both ReadOnly flips and tier flips) and SyncDataNodeRegistration, then route
it through NewVids/RemoteVids in SendHeartbeat the same way the delta path
already does. Add an end-to-end test for the full reconciliation path.

* master: keep an EC volume's locations in the volume lookup

The nodes that answer for an EC volume hold shards, not a volume record,
so asking them for one fails. Dropping the location on that failure
emptied the result and turned every EC read through the master's HTTP
lookup and fid redirect into a 404.

Treat an absent volume record as a local read and keep the node in the
answer. The per-node conversion moves into topologyLocation so the EC
case is covered by a test.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

* wdclient: replace a tier-flipped location without writing under a reader

GetLocations hands back the entry's own slice and the caller walks it
after the read lock is dropped, which is why every other mutation here
builds a new slice. Writing the flipped replica into the array in place
raced LookupVolumeServerUrl, reported by -race.

Copy the slice, swap the one element, and publish it.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

* master: keep a remote volume on NewVids for older clients

Moving remote-tier volumes out of NewVids and into RemoteVids alone is a
wire break in the wrong direction. A master upgraded ahead of its filers
and mounts -- the usual order -- announces a tiered volume only on a
field the older client ignores, so the volume drops out of that client's
vid map entirely and reads for it fail.

Announce every volume on NewVids and repeat the remote-tier subset on
RemoteVids, so a new client still learns the tier and an old one keeps
the location. The routing moves into announceVolume, which the heartbeat
paths and their tests now share instead of each restating it.

On the client, RemoteVids no longer needs a second write per volume: the
tier is settled before anything is added.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

* topology: split the volume snapshot by tier without copying the records

ToVolumeLocations runs on every KeepConnected, so a filer or mount
connecting made the master allocate a full VolumeInfo per volume per node
just to read four bytes of id off each one. AppendVolumeIds exists to
avoid exactly that.

Extend it to fill the remote-tier list alongside the full one, and use it
again in the snapshot.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

* wdclient: keep the data-center preference ahead of the local-first ordering

Hoisting every local replica to the very front puts an other-DC local
read ahead of a same-DC remote one. When the remote tier sits in the same
region as the replicas -- the common arrangement -- that trades an
in-region GET for a WAN round trip and costs more than the remote read it
avoids.

Reorder inside each data-center bucket instead, so local still wins among
equals and the data-center preference still wins overall.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

* operation: pick the read replica from one list

The local-preferring lookup built a list of local URLs and then branched
on whether it was empty, duplicating the random pick. Fall back by
filling the same list with every replica instead.

Claude-Session: https://claude.ai/code/session_01FcSp6quCaxQ1cb3Vx7o9fW

---------

Co-authored-by: Bruce Zou <gift_secondst@msn.com>
Co-authored-by: bruce-zzz <bruce.zou@hhy-data.com>
2026-09-02 16:12:46 -07:00
Chris LuandGitHub 1d335357d6 helm: name both replication keys in the enableReplication comment (#11103)
The comment said enableReplication overrides "master & filer
defaultReplicaPlacement", but the two components take different keys:
master.defaultReplication and filer.defaultReplicaPlacement.

Claude-Session: https://claude.ai/code/session_014Yr1Asxq3qTjJo2r9G4cLX
2026-09-02 13:42:06 -07:00
96242a2be2 s3: ListParts on a completed or unknown upload answers NoSuchUpload (#11081)
* s3: ListParts on a completed or unknown upload answers NoSuchUpload

complete/abort delete the .uploads/<id> directory, but most filer stores list
a missing directory as empty rather than erroring, so listObjectParts answered
200 with an empty Parts list for an upload that no longer exists -- the same
response an open upload with no parts yet gets. AWS (and Ceph/RGW, MinIO)
answer NoSuchUpload, and clients lean on that: tusd derives the resumable
upload offset from the ListParts part sizes, so every completed upload read
back as zero bytes received.

Probe the upload record before listing, the way completeMultipartUpload
already does: not found, or a directory a late part write resurrected without
the destination key, answers NoSuchUpload. An open upload with no parts keeps
answering 200 with an empty list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GGQEfVUvoATtwRCR8oC2jG

* s3: have the ListParts test filer refuse a directory it was not asked about

The fake answered the upload lookup on the name alone and the part listing
regardless of directory, so a wrong genUploadsFolder or upload-id suffix
would still have passed. Both calls now refuse any other directory with an
Internal error, which surfaces as ErrInternalError rather than the
NoSuchUpload the tests expect.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StsRz9wbu5dUCMGbFgPoRM

---------

Co-authored-by: tomislavcivcija <9787657+tomislavcivcija@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 12:05:16 -07:00
Chris LuandGitHub eed5e8cdf6 s3: return the multipart object checksum in the CompleteMultipartUpload response (#11101)
* s3: return the multipart object checksum in the CompleteMultipartUpload body

S3 carries the flexible-checksum members of CompleteMultipartUploadResult in
the XML body, not in response headers, so every SDK read back an empty
checksum from an upload that asked for one.

Claude-Session: https://claude.ai/code/session_01Huux1uh7JxAbf8yypMYrMk

* s3: echo the checksum algorithm and type from CreateMultipartUpload

The upload directory already records both, but the response dropped them, so a
client could not confirm which checksum its parts had to carry.

Claude-Session: https://claude.ai/code/session_01Huux1uh7JxAbf8yypMYrMk

* test: multipart upload reports the object checksum it was asked for

Covers every algorithm end to end: the create response echoes the algorithm and
type, the complete response carries the checksum, and it matches what a later
HEAD reports.

Claude-Session: https://claude.ai/code/session_01Huux1uh7JxAbf8yypMYrMk
2026-09-02 11:57:45 -07:00
Chris LuandGitHub 5f787a25c3 master: survive a volume layout deleted twice (#11098)
* master: survive a layout deleted twice

Two volume servers dropping the last replica of volumes that share a layout
both find it empty and both delete it. The loser's lookup misses, and the
single-value type assertion on the result crashed the master before the
caller could look at the found flag.

Claude-Session: https://claude.ai/code/session_01WmX6Rchx298NQksHDXg7sk

* master: remove a layout and read it back in one step

DeleteVolumeLayout looked the layout up and then deleted it, so two deleters
could each release the lookup ownership of the same layout, or one could find
nothing to release at all. Have the map hand back what it removed.

Claude-Session: https://claude.ai/code/session_01WmX6Rchx298NQksHDXg7sk
2026-09-02 11:50:48 -07:00
Chris LuandGitHub 86761cc7d5 filer: keep empty folders that are s3tables catalog entries (#11102)
* s3tables: build the catalog attribute keys from one shared prefix

Every attribute the catalog stores on a bucket, namespace, table or view
entry is spelled out with the same literal prefix. Name it once in
s3_constants so code outside the package can recognize a catalog entry
without repeating the string.

Claude-Session: https://claude.ai/code/session_01GfZsc4cyNB2yr6KYLRv9q1

* filer: keep empty folders that are s3tables catalog entries

A namespace, table or view is a directory whose extended attributes are
the catalog record. Its files can live elsewhere - a rename moves only
the catalog pointer and leaves the data at the old path, and a view has
no files at all - so an empty one is still a live entry.

Drop a table, then rename another table onto that name: the drop queues
the old table's folders, the rename recreates the name path, and two
minutes later the cleaner deletes it and cascades into the namespace,
losing a table the catalog still lists.

Claude-Session: https://claude.ai/code/session_01GfZsc4cyNB2yr6KYLRv9q1

* filer: drop a queued cleanup when the folder is created again

A cleanup is queued against the folder that was found empty. If that
folder is deleted and a new one takes its name, the queue entry outlives
the folder it was about and the next pass deletes the replacement. A
drop followed by a rename onto the dropped name does exactly this: the
name path comes back as a live catalog entry two minutes before the
queue is read.

Claude-Session: https://claude.ai/code/session_01GfZsc4cyNB2yr6KYLRv9q1
2026-09-02 11:49:08 -07:00
Chris LuandGitHub 23a6b8feb5 filer_pb: walk the whole tree when the BFS start path ends in a slash (#11099)
* filer_pb: build BFS child paths with FullPath.Child

A start path with a trailing slash produced "/dir//sub" for every
subdirectory, and the filer only trims a trailing slash, so those
listings came back empty and the walk stopped after the first level.

Claude-Session: https://claude.ai/code/session_01Jp9tXRpBv9gvh8fkaVFqxQ

* filer_pb: normalize the BFS start path

Entries directly under the start path were reported with the caller's
trailing slash, so filer.meta.backup wrote them under a directory the
incremental stream never names again.

Claude-Session: https://claude.ai/code/session_01Jp9tXRpBv9gvh8fkaVFqxQ
2026-09-02 11:45:22 -07:00
Chris LuandGitHub 9ea52db219 s3: validate the version-id header used as a filer path segment (#11097)
* s3: reject a version-id header that is not a valid path segment

putToFiler stored the client-supplied Seaweed-X-Amz-Version-Id header
verbatim into object metadata. That value is later read back and used
as a filer path component when building the .versions/v_<id> path, so a
value containing "/", "\" or ".." could steer retention/legal-hold
writes and remote-cache reads outside the object's own bucket tree.

Validate the header with isValidVersionID before storing it, the same
check the versioned read paths already apply, and reject the request
otherwise. Server-set version ids ("null" and generated hex) pass.

Claude-Session: https://claude.ai/code/session_011QqNaxZwnpHgMoAZNp3RkY

* s3: validate a stored version-id before using it as a path

The retention and legal-hold sinks build a .versions/v_<id> path from a
version id read back out of object metadata, and the remote-cache path
builder does the same from either the request or the stored id, without
the isValidVersionID check the other version-id consumers apply. Guard
these so a value that is not a valid path segment falls back to the
regular / unversioned path instead of steering the write or read out of
the bucket tree.

Claude-Session: https://claude.ai/code/session_011QqNaxZwnpHgMoAZNp3RkY
2026-09-02 11:36:58 -07:00
Chris LuandGitHub 23adeb37e2 s3: check Object Lock on directory-marker keys before bucket deletion (#11096)
recursivelyCheckLocksWithClient tested EntryHasActiveLock only on
non-directory entries, so a directory-marker object (an S3 key ending
in "/") that carries retention or a legal hold was recursed into but
never lock-checked. DeleteBucket then saw no locks and removed the
bucket, destroying an object under active Object Lock along with the
rest of the bucket. DeleteObject already enforces the lock on the same
key, so the two paths disagreed.

Check the directory entry for an active lock before recursing.

Claude-Session: https://claude.ai/code/session_011QqNaxZwnpHgMoAZNp3RkY
2026-09-02 11:35:19 -07:00
Chris LuandGitHub 3cdfe648eb sftp: reject an empty password (#11095)
ValidatePassword compared the stored and supplied passwords with
subtle.ConstantTimeCompare, which returns 1 for two zero-length slices.
A user provisioned for public-key-only auth has an empty stored password,
so an empty supplied password authenticated as that user whenever
"password" was among the enabled auth methods (the default).

Treat an empty stored or supplied password as a non-match.

Claude-Session: https://claude.ai/code/session_011QqNaxZwnpHgMoAZNp3RkY
2026-09-02 11:33:25 -07:00
Chris LuandGitHub 398277a15d mini: expose -volume.max (#11100)
mini hardcoded the per-directory volume limit to 0, so the volume server
always auto-sized it as free disk space divided by the volume size. That
sizing reserves a whole volume size for every writable volume, so a
workload spreading small objects over many buckets runs out of slots long
before the disk fills and assign starts failing with "no free volumes
left". Same name and semantics as the flag weed server already carries,
and still 0 (auto) by default.

Claude-Session: https://claude.ai/code/session_01BiQLeBvZLzG8XitjiypDKu
2026-09-02 09:46:27 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1c57734e58 build(deps): bump google.golang.org/grpc from 1.81.1 to 1.83.1 in /test/kafka (#11078)
build(deps): bump google.golang.org/grpc in /test/kafka

Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.81.1 to 1.83.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/compare/v1.81.1...v1.83.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.83.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-02 09:02:46 -07:00
Chris LuandGitHub 68a175ef6f deps: drop the apache/thrift replace, v0.24.0 carries the 32-bit fix (#11077)
The replace pinned thrift to a post-v0.23.0 master commit so 32-bit
GOARCHes would compile. That fix shipped in v0.24.0, so the replace was
only overriding the v0.24.0 require back to the pseudo-version and
holding the build below the CVE-2026-43871 fix.

Claude-Session: https://claude.ai/code/session_01C5BpSeYD3yULWmfVXwPRmB
2026-09-01 14:37:32 -07:00
Chris LuandGitHub 8f2daad338 topology: mirror the writable volume list in a set (#11076)
Membership was a linear scan over a slice, and ensureCorrectWritables runs it
for every volume on every heartbeat, so the master's steady-state cost per
volume server is quadratic in that server's volume count.

BenchmarkSyncDataNodeRegistration, median of 3:

  1000 volumes     565.7us -> 535.0us    -5.4%
  100000 volumes    1.665s -> 55.3ms     -96.7%

Allocations are unchanged at both sizes, so the difference is the scan.

Claude-Session: https://claude.ai/code/session_01P3pE6J2UPFp6G3ksfMV4s1
2026-09-01 14:21:26 -07:00
Chris LuandGitHub 40b3d32fe5 test: metadata operations on unlinked open files and removed open directories (#11075)
The POSIX suites skirt this: pjdfstest's unlink/14.t covers only fstat and
pread on an unlinked descriptor — its driver has no fchmod at all and never
opens a directory. Pin the full rule in the FUSE integration suite:
ftruncate, fchmod, futimes, fstat, and the f*xattr calls keep working
between the removal of the last name and the final close, for a file after
unlink and a directory after rmdir, with nlink 0 and the changes visible to
a following fstat.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK
2026-09-01 14:13:16 -07:00
Chris LuandGitHub 86a189ff80 mount: keep metadata operations working on a removed open directory (#11073)
* mount: remember the entry of a directory removed while still referenced

A directory removed while a descriptor is open on it keeps its inode until
the kernel's final forget, but unlike a file it has no handle to live on
through: OpenDir hands out only a listing cursor. Keep the last-known entry
in memory, keyed by inode, from rmdir until that forget.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK

* mount: serve metadata ops on a removed open directory from its remembered entry

fchmod, futimens, and the f*xattr calls on a descriptor whose directory was
removed failed with ENOENT: maybeReadEntry resolved the inode to a path, and
rmdir had already dropped it. Fall back to the remembered entry the same way
an unlinked file falls back to its open handle. Mutations publish a changed
copy back rather than editing in place, so a concurrent reader never sees a
half-applied change, and the empty path keeps nlink 0 in every reply.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK

* mount: stash the entry the delete itself returned, not an earlier snapshot

A chmod landing between Rmdir's entry load and the delete RPC would be
resurrected pre-change: the remembered entry was the earlier local snapshot.
The filer serializes the delete against updates under the path lock and hands
the entry back in the delete event, so prefer that, keeping the local load
for the sticky-bit check and as fallback when no event comes back.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK

* mount: drop a remembered entry whose insert lost to the final forget

The forget's cleanup runs between RemovePath and the insert when the kernel
evicts the inode concurrently, finds nothing, and the entry would sit in the
map for the life of the mount. Re-check the inode after inserting and take
the entry back out; every interleaving now ends with the map empty.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK

* mount: insert the remembered entry under the inode table lock

The post-insert HasInode re-check could be fooled by inode number reuse: a
lookup landing between the forget and the check makes the number look alive
and the stale entry stays, keyed to someone else's inode. Do not check after
the fact — RemovePath now runs the retention callback inside its critical
section, where the forget that releases under the same lock cannot have run
and cannot be missed. Publishes need no such fence: their open descriptor
keeps the kernel from issuing the final forget in the first place.

Claude-Session: https://claude.ai/code/session_01GYqLENjZzbV5hgt4L8cSAK
2026-09-01 13:21:02 -07:00
Chris LuandGitHub 77a9dd4b9e s3: route per-key object authorization through a shared helper (#11072)
* s3: share the per-key object authorization across copy and delete

AuthorizeCopySource and AuthorizeObjectDelete both authorize a key the request
URL does not name by evaluating the bucket policy and IAM against a synthetic
per-key request; only the method and action differed. Extract that into
authorizeObjectKeyAction and make the two callers thin wrappers. No behavior
change.

Claude-Session: https://claude.ai/code/session_01Qo7p6VsoWxMo8816ogJFk5

* s3: route POST Object uploads through the shared object authorization

POST Object uploads (presigned-POST / HTML form) authorized the write with only
the coarse per-identity Write action, unlike the other write paths which also
check the resolved object against the bucket policy and IAM. Route POST through
authorizeObjectKeyAction via a new AuthorizeObjectWrite so it is authorized like
the equivalent PUT.

Claude-Session: https://claude.ai/code/session_01Qo7p6VsoWxMo8816ogJFk5

* s3: test POST Object per-key authorization

Drives a signed POST upload and checks the per-key authorization decision for a
denied, permitted, and admin caller.

Claude-Session: https://claude.ai/code/session_01Qo7p6VsoWxMo8816ogJFk5
2026-09-01 13:20:20 -07:00
Chris LuandGitHub 34f5442e9b s3api: push the listing prefix down to the filer in ListObjectVersions (#11070)
The version walk listed every directory with no prefix, transferring
all 1024-entry batches over gRPC and filtering gateway-side - and kept
paging past the point where names can no longer match. On wide
directories (many sibling orgs/jobs next to the requested prefix) that
is most of the transfer, decode, and CPU cost of every page.

Derive the next path component of the requested prefix per directory
level and hand it to the filer listing. A name holds no slash, so a
directory whose name does not start with the component cannot contain
a matching key and a file that does not cannot be one; stores with
native prefixed listing (sql, leveldb) turn this into a range scan and
stop the stream at the end of the prefix zone.

Claude-Session: https://claude.ai/code/session_01FquvGtTD2zA3uMZGQHAuV4
2026-09-01 10:37:52 -07:00
Chris LuandGitHub 0ba21174bf volume: an already-deleted EC needle is not a delete failure (#11071)
* volume: an already-deleted EC needle is not a delete failure

Deleting a needle that is already gone is what the caller asked for, and the
non-EC paths have always said so: BatchDelete reports StatusNotModified when
DeleteVolumeNeedle finds nothing to do, and DeleteHandler answers 404 from its
ReadVolumeNeedle pre-check. The EC branches had no such case, so ErrorDeleted
fell through to a generic failure -- 500 from both, and DeleteHandler also
counted it in VolumeServerFileWriteFailures, inflating a failure metric on a
replayed or duplicated delete.

The filer already tolerates this by string-matching "already deleted" on the
result, which leaves an error message load-bearing; the status is now right at
the source instead.

Claude-Session: https://claude.ai/code/session_01P3pE6J2UPFp6G3ksfMV4s1

* volume: close the EC fixture's disk location

Close stops the location's disk-space goroutine and releases the mounted EC
volume's file handles, which otherwise live until the test binary exits.

Claude-Session: https://claude.ai/code/session_01P3pE6J2UPFp6G3ksfMV4s1
2026-09-01 10:37:25 -07:00
Chris LuandGitHub 81ca5cb6c6 s3api: drop two redundant filer round-trips per listed version entry (#11068)
* s3api: drop two redundant filer round-trips per listed version entry

ListObjectVersions paid two avoidable getEntry calls while walking a
bucket, both re-fetching data the walk already held:

- getObjectVersionList re-read the .versions directory entry that every
  caller had just received from listing the parent directory (or from
  its own sibling probe). Pass the entry down instead: one RPC saved
  per object listed.

- getObjectOwnerFromVersion, on a version with no stamped owner,
  re-fetched the same version entry its OwnerID had been extracted
  from. The refetch cannot answer differently, so data written before
  owners were stamped cost one futile RPC per listed version, forever.

All round-trips on this path are sequential, so on large versioned
buckets (Veeam-style workloads) they add up to a visible share of
per-page latency and gateway CPU.

Claude-Session: https://claude.ai/code/session_01FquvGtTD2zA3uMZGQHAuV4

* s3api: treat a nil .versions entry as an empty version list

filer_pb.GetEntry's contract permits (nil, nil) for an absent entry, and
the old internal lookup answered that case with an empty list. Keep that
answer now that the entry arrives from the caller.

Claude-Session: https://claude.ai/code/session_01FquvGtTD2zA3uMZGQHAuV4
2026-09-01 10:29:52 -07:00
Chris LuandGitHub 4c9cbf72bc s3api: stop retrying a definitive NotFound in getLatestObjectVersion (#11067)
The .versions lookup retried every error through the full backoff
ladder, so a missing key spent 12.7s (8 attempts, 100ms..6.4s) before
the pre-versioning fallback could answer. NotFound is an answer, not a
transient failure: gate the retries on isRetryableFilerErr, the same
classifier retryFilerOp already uses, which also stops retrying for
callers whose context is canceled or past its deadline.

GetObject already treats NotFound on .versions/ as definitive; this
brings the retention/tagging/ACL/attributes/delete/copy paths that go
through getLatestObjectVersion in line with it.

Claude-Session: https://claude.ai/code/session_01FquvGtTD2zA3uMZGQHAuV4
2026-09-01 10:11:09 -07:00
Chris LuandGitHub 2ef0e60aeb filer.sync: export replication lag, event counters, and in-flight jobs (#11069)
* filer.sync: count received, processed, and failed events and export in-flight jobs

The metadata processor admits at most -concurrency jobs and blocks the
subscription stream past that, so the backlog lives in the source filer's
metadata log and cannot be counted here. What can be measured honestly:
events read off the stream, replication outcomes, and worker saturation.
in_flight_jobs pinned at the concurrency limit means the sync itself is
the bottleneck; near zero means it is caught up or starved by the source.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw

* filer.sync: export replication lag in seconds

Lag is now minus the freshest of the processed watermark and the last
idle heartbeat: the watermark stops at the last real event, so a quiet
caught-up stream would otherwise show phantom lag. A ticker drives the
gauge because the offset callback only fires while events flow and
freezes exactly when the workers are saturated. Until the first event
or heartbeat the gauge stays unset rather than reporting lag against a
zero or stale resume offset.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw

* filer.sync: track replicated data sizes alongside event counts

An event count hides that 32 in-flight jobs can be 32 KB or 300 GB. Byte
counters mirror the event counters, and in_flight_bytes pairs with
in_flight_jobs. An event's size is the chunk delta - new chunks the old
entry does not already have - so deletes, renames, and attribute-only
updates count zero and byte rates reflect data movement, not metadata
churn.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw

* grafana: chart the new filer.sync metrics

The lag panel reads lag_seconds directly instead of deriving it from
sync_offset, and the sync row gains event rate, throughput, and the
in-flight jobs and bytes gauges.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw

* filer.sync: a pinned failure keeps showing as lag

An idle heartbeat means the stream is consumed, not that every event
replicated. While a permanent failure pins the watermark, letting the
heartbeat advance lag_seconds or the sync_offset gauge would report a
caught-up stream with an unreplicated event in it, so both now ignore
heartbeats until a restart replays the failure.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw

* filer.sync: in-flight gauges survive subscription retries

A subscription retry builds a new processor sharing the gauge children
while the old processor's jobs may still be draining, so setting the
gauge from either side's local count clobbers the other. Each job now
increments and decrements for itself, keeping the total truthful across
generations.

Claude-Session: https://claude.ai/code/session_01VGUXS7kdqiaXenTYVCg6vw
2026-09-01 10:10:30 -07:00
Chris LuandGitHub 9d5525e747 master: keep periodic volume growth to the data centers a layout lives in (#11060)
* master: keep the periodic growth scan to data centers hosting the layout

The rack-aware scan planned growth for every data center in the topology,
so a collection pinned to one DC (fs.configure -dataCenter) sprouted
volumes in all the others within one scan cycle. Plan only for data
centers already hosting the layout's volumes; an empty DC gets its
volumes from the DC-constrained assign that first asks for them. The
lastGrowCount divisor likewise counts only the racks the scan can plan
for.

Claude-Session: https://claude.ai/code/session_01J22TVTyoCMzdHyJirsLMG5

* master: pin periodic must-grow growth to a single-DC layout's data center

The must-grow and crowded paths of the periodic loop grow with no
DataCenter, so even with the scan fixed a pinned collection's volumes
could still land in any DC once lastGrowCount demands more writables.
Stamp the grow request with the layout's data center when its volumes
all live in one; layouts spanning DCs keep unconstrained growth.

Claude-Session: https://claude.ai/code/session_01J22TVTyoCMzdHyJirsLMG5

* master: never pin growth of a cross-DC-replicated layout

A layout whose replication spans data centers cannot legitimately live
in one DC; observing a single hosting DC there means the other DCs are
down. Do not encode that outage as a placement constraint.

Claude-Session: https://claude.ai/code/session_01J22TVTyoCMzdHyJirsLMG5

* master: bound the hosting-DC walk by the answer it needs

listVolumeDataCenters walked every location of the layout under
accessLock — ~190ms for a million volumes, twice per layout per cycle,
stalling assigns behind the read lock. Stop once enough distinct DCs
answer the caller's question: two for the single-DC check, the
topology's DC count for the scan. A spanning million-volume layout now
finishes in microseconds; only a layout truly confined to fewer DCs
still pays a full walk, the same cost class as the under-replication
count this loop already takes each cycle.

Claude-Session: https://claude.ai/code/session_01J22TVTyoCMzdHyJirsLMG5
2026-09-01 00:40:50 -07:00
Chris LuandGitHub 8873f9775c shell: reset noLock in the admin script dispatcher too (#11059)
Three dispatchers reuse one CommandEnv: the interactive shell, the master's
maintenance script runner, and the plugin worker's admin script handler. The
first two were fixed; this is the third.

It changes nothing today -- ForceNoLock already exempts this path -- so it is
here to keep the rule the same everywhere rather than resting on that exemption
staying in place.
2026-08-31 23:14:37 -07:00
Chris LuandGitHub 1996c6aec6 volume: open volume files with O_NOATIME (#11055)
* volume server: open volume files with O_NOATIME

Nothing reads the atime of .dat, .idx, .sdx, or EC files, but every
needle read still dirtied the inode: even relatime writes atime on the
first read after each write, so an actively written volume paid a
metadata write per read/write cycle, and strictatime mounts paid one
per read. Open the serving handles with O_NOATIME, falling back to a
plain open when the file belongs to another owner (EPERM).

Claude-Session: https://claude.ai/code/session_015uVY4diBgEn3VYQoc2eMuD

* seaweed-volume: mirror the O_NOATIME volume file opens

Same change as the Go volume server: serving handles for .dat, .idx,
.sdx, .ecx, .ecj, and shard files open with O_NOATIME on Linux, with a
plain-open fallback on EPERM.

Claude-Session: https://claude.ai/code/session_015uVY4diBgEn3VYQoc2eMuD

* route the tier-down and recreate .dat opens through the no-atime helper

Review caught the Rust tier-down swap opening the local .dat directly.
The Go swapToLocalDatBackend and the zero-length read-only .dat
recreate in maybeWriteSuperBlock had the same gap: all three install
long-lived serving handles.

Claude-Session: https://claude.ai/code/session_015uVY4diBgEn3VYQoc2eMuD
2026-08-31 21:41:50 -07:00
Chris LuandGitHub 0c59c0fb05 master: scope the startup capacity shed to a truly empty topology (#11058)
* master: scope the startup capacity shed to a truly empty topology

The retryable "no volume server capacity registered yet" shed checked
capacity for the requested disk type, so a cluster serving only other
media -- where that capacity will never register -- shed every assign
until the client's deadline instead of failing fast. An unsteered write
to such a cluster hung for its full HTTP deadline and surfaced "context
deadline exceeded" in place of "No writable volumes". Shed only while
no disk type has any registered capacity, and name the unserved medium
in the fast failure.

Claude-Session: https://claude.ai/code/session_01TF7FQghfDkpdoZgakTMX4R

* master: name the unserved medium for every fail-fast caller

The diagnostic sat in the growth-initiator block, so a follower joining
an in-flight growth and a growth-disabled master failed the same way
with only the generic pick error. Wrap at the fail-fast break instead,
which every caller reaches, and cover all three paths in the test.

Claude-Session: https://claude.ai/code/session_01TF7FQghfDkpdoZgakTMX4R
2026-08-31 21:29:27 -07:00
Chris LuandGitHub 721499a05a release: judge downstream releases by their run, excluding runner-queue time (#11057)
Claude-Session: https://claude.ai/code/session_01Y228KU8MLsGmfpcbGxgjwh
2026-08-31 20:56:39 -07:00
642 changed files with 43279 additions and 22540 deletions
+2 -1
View File
@@ -19,4 +19,5 @@ ignore-regex = \b[a-z]+[A-Z]\w*\b|\b[A-Z][a-z]+[A-Z]\w*\b
# thirdparty: literal Maven groupId `org.apache.hadoop.thirdparty` (external, cannot rename)
# unknwon: GitHub username / Go module path (`github.com/unknwon/goconfig`)
# atleast: CLI mode literal string in test/benchmark/fuse_db/bin/sqlite_verify.py
ignore-words-list = visibles,fo,te,ser,bject,unparseable,keep-alives,tread,anc,ue,auther,thirdparty,unknwon,atleast
# sme: local variable for a *streamMutateError in mount tests
ignore-words-list = visibles,fo,te,ser,bject,unparseable,keep-alives,tread,anc,ue,auther,thirdparty,unknwon,atleast,sme
+44
View File
@@ -0,0 +1,44 @@
name: Sign container images
description: >
Keyless cosign signature on each image, then a verification pass against the
identity the signature should carry, so a misconfigured job fails here and not
on someone's cluster. That identity is the calling workflow's own,
https://github.com/<owner>/<repo>/.github/workflows/<file>@<ref>.
The calling job needs `id-token: write` and a registry login for every image.
inputs:
images:
description: Image references by digest (name@sha256:...), whitespace separated.
required: true
runs:
using: composite
steps:
- name: Install cosign
uses: sigstore/cosign-installer@v4.1.2
- name: Sign
shell: bash
env:
IMAGES: ${{ inputs.images }}
run: |
set -euo pipefail
# The .sig tag layout: the OCI-referrer bundle cosign 3 writes by default
# is not read by the Kyverno and policy-controller releases in use today.
# Cosign 3 also defaults to --use-signing-config, which insists on a
# bundle for its output; turning it off falls back to the default
# Fulcio and Rekor URLs, which is all the .sig layout ever used.
cosign sign --yes --recursive \
--new-bundle-format=false --use-signing-config=false \
$IMAGES
- name: Verify
shell: bash
env:
IMAGES: ${{ inputs.images }}
run: |
set -euo pipefail
cosign verify \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
--certificate-identity "https://github.com/$GITHUB_WORKFLOW_REF" \
$IMAGES
+3 -1
View File
@@ -1,7 +1,9 @@
version: 2
updates:
- package-ecosystem: "github-actions"
directory: "/"
directories:
- "/"
- "/.github/actions/sign-image"
schedule:
interval: "weekly"
- package-ecosystem: gomod
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env python3
"""Render the repository's star history to note/star_history.svg.
Uses the GitHub REST stargazers endpoint with the starred-at accept header,
which caps at 40,000 entries (400 pages of 100). The chart is regenerated on
a schedule; if the repo grows past that cap the script stops at 40,000 and
logs a warning rather than under-reporting.
"""
import json
import os
import sys
import urllib.request
from datetime import datetime
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt # noqa: E402
from matplotlib.dates import AutoDateLocator, DateFormatter # noqa: E402
REPO = "seaweedfs/seaweedfs"
TOKEN = os.environ["GITHUB_TOKEN"]
OUT = os.environ.get("OUT", "note/star_history.svg")
PAGE_CAP = 400 # GitHub's hard limit on stargazer pagination
def fetch_stargazers():
stars = []
page = 1
while page <= PAGE_CAP:
url = f"https://api.github.com/repos/{REPO}/stargazers?per_page=100&page={page}"
req = urllib.request.Request(
url,
headers={
"Accept": "application/vnd.github.star+json",
"Authorization": f"Bearer {TOKEN}",
"X-GitHub-Api-Version": "2022-11-28",
"User-Agent": "seaweedfs-star-history",
},
)
with urllib.request.urlopen(req) as resp:
batch = json.load(resp)
if not batch:
break
for u in batch:
sa = u.get("starred_at")
if sa:
stars.append(datetime.fromisoformat(sa.replace("Z", "+00:00")))
if len(batch) < 100:
break
page += 1
if page > PAGE_CAP:
print(
f"::warning::Hit the {PAGE_CAP}-page stargazer pagination cap; "
"chart reflects the first 40,000 stars only."
)
return stars
def render(stars, out):
stars.sort()
counts = list(range(1, len(stars) + 1))
fig, ax = plt.subplots(figsize=(10, 6), dpi=130)
ax.plot(stars, counts, color="#0969da", linewidth=1.6)
ax.set_xlabel("Date")
ax.set_ylabel("Stars")
ax.set_title(f"{REPO} star history")
ax.grid(True, linestyle="--", alpha=0.3)
ax.xaxis.set_major_locator(AutoDateLocator())
ax.xaxis.set_major_formatter(DateFormatter("%Y-%m"))
fig.autofmt_xdate()
fig.tight_layout()
fig.savefig(out, format="svg", transparent=False)
plt.close(fig)
def main():
stars = fetch_stargazers()
if not stars:
print("::error::No stargazers fetched; not updating the chart.")
sys.exit(1)
render(stars, OUT)
print(f"Rendered {len(stars)} stars to {OUT}")
if __name__ == "__main__":
main()
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.9
uses: github/codeql-action/init@v4.38.0
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.9
uses: github/codeql-action/autobuild@v4.38.0
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.9
uses: github/codeql-action/analyze@v4.38.0
+13 -1
View File
@@ -101,6 +101,9 @@ jobs:
build-dev-containers:
needs: [build-rust-binaries]
runs-on: [ubuntu-latest]
permissions:
contents: read
id-token: write
steps:
- name: Checkout
@@ -149,7 +152,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.2.0
uses: docker/setup-qemu-action@v4.3.0
- name: Create BuildKit config
run: |
@@ -180,6 +183,7 @@ jobs:
password: ${{ secrets.GHCR_TOKEN }}
- name: Build
id: build
uses: docker/build-push-action@v7
with:
context: ./docker
@@ -188,3 +192,11 @@ jobs:
platforms: linux/amd64, linux/arm64
tags: ${{ steps.docker_meta.outputs.tags }}
labels: ${{ steps.docker_meta.outputs.labels }}
- name: Sign
if: github.event_name != 'pull_request'
uses: ./.github/actions/sign-image
with:
images: >-
chrislusf/seaweedfs@${{ steps.build.outputs.digest }}
ghcr.io/chrislusf/seaweedfs@${{ steps.build.outputs.digest }}
@@ -30,6 +30,9 @@ permissions:
jobs:
build-foundationdb-image:
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- name: Checkout
@@ -126,7 +129,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.2.0
uses: docker/setup-qemu-action@v4.3.0
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
@@ -154,6 +157,7 @@ jobs:
fi
- name: Build and push image
id: build
uses: docker/build-push-action@v7
with:
context: ./docker
@@ -171,3 +175,8 @@ jobs:
org.opencontainers.image.description=SeaweedFS is a distributed storage system for blobs, objects, files, and data lake, to store and serve billions of files fast!
org.opencontainers.image.vendor=Chris Lu
- name: Sign
if: github.event_name != 'pull_request'
uses: ./.github/actions/sign-image
with:
images: chrislusf/seaweedfs@${{ steps.build.outputs.digest }}
+28 -9
View File
@@ -236,7 +236,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
if: matrix.platform != 'amd64'
uses: docker/setup-qemu-action@v4.2.0
uses: docker/setup-qemu-action@v4.3.0
- name: Create BuildKit config
run: |
cat > /tmp/buildkitd.toml <<EOF
@@ -405,7 +405,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.0
if: always()
with:
sarif_file: trivy-results.sarif
@@ -441,15 +441,19 @@ jobs:
runs-on: ubuntu-latest
needs: [setup, build, trivy-scan]
if: needs.setup.outputs.publish == 'true' && github.event_name != 'pull_request'
permissions:
contents: read
id-token: write
strategy:
matrix:
variant: ${{ fromJSON(needs.setup.outputs.variants) }}
steps:
- name: Checkout
- name: Checkout the signing action
uses: actions/checkout@v7
with:
ref: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.source_ref || github.ref }}
sparse-checkout: .github/actions
persist-credentials: false
- name: Configure variant
id: config
run: |
@@ -482,21 +486,29 @@ jobs:
run: |
# Install crane for efficient multi-arch image copying
cd $(mktemp -d)
curl -sL "https://github.com/google/go-containerregistry/releases/latest/download/go-containerregistry_Linux_x86_64.tar.gz" | tar xz
curl -sLO https://github.com/google/go-containerregistry/releases/download/v0.22.0/go-containerregistry_Linux_x86_64.tar.gz
echo "edb74d53fad9a596860f59d1c5d04a43dfb5f441dc71f57060dd0bf39483c833 go-containerregistry_Linux_x86_64.tar.gz" | sha256sum -c -
tar xzf go-containerregistry_Linux_x86_64.tar.gz crane
sudo mv crane /usr/local/bin/
crane version
- name: Create and push manifest
id: manifest
env:
BASE_TAG: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.image_tag || 'latest' }}
run: |
SUFFIX="${{ steps.config.outputs.tag_suffix }}"
BASE_TAG="${{ github.event_name == 'workflow_dispatch' && github.event.inputs.image_tag || 'latest' }}"
# Create manifest on GHCR first (no rate limits)
echo "Creating GHCR manifest (no rate limits)..."
docker buildx imagetools create -t ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX} \
--metadata-file /tmp/manifest.json \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-amd64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
# The copy and the signature below use the digest this run pushed, not whatever the tag points at by then.
DIGEST=$(jq -er '."containerimage.descriptor".digest' /tmp/manifest.json)
echo "digest=${DIGEST}" >> "$GITHUB_OUTPUT"
# Copy the complete multi-arch image from GHCR to Docker Hub
# This only requires one pull from GHCR (no rate limit) and one push to Docker Hub
@@ -532,10 +544,10 @@ jobs:
# Use crane or skopeo to copy, fallback to docker if not available
if command -v crane &> /dev/null; then
echo "Using crane to copy..."
retry_with_backoff crane copy ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX} chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
retry_with_backoff crane copy ghcr.io/chrislusf/seaweedfs@${DIGEST} chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
elif command -v skopeo &> /dev/null; then
echo "Using skopeo to copy..."
retry_with_backoff skopeo copy --all docker://ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX} docker://chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
retry_with_backoff skopeo copy --all docker://ghcr.io/chrislusf/seaweedfs@${DIGEST} docker://chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
else
echo "Using docker buildx imagetools (pulling 4 images from Docker Hub)..."
# Fallback: create manifest directly on Docker Hub (pulls from Docker Hub - rate limited)
@@ -545,3 +557,10 @@ jobs:
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
fi
- name: Sign
uses: ./.github/actions/sign-image
with:
images: >-
ghcr.io/chrislusf/seaweedfs@${{ steps.manifest.outputs.digest }}
chrislusf/seaweedfs@${{ steps.manifest.outputs.digest }}
@@ -21,6 +21,9 @@ jobs:
build-large-release-container_foundationdb:
runs-on: [ubuntu-latest]
permissions:
contents: read
id-token: write
steps:
-
@@ -43,7 +46,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
-
name: Set up QEMU
uses: docker/setup-qemu-action@v4.2.0
uses: docker/setup-qemu-action@v4.3.0
-
name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
@@ -65,6 +68,7 @@ jobs:
fi
-
name: Build
id: build
uses: docker/build-push-action@v7
with:
context: ./docker
@@ -76,4 +80,10 @@ jobs:
platforms: linux/amd64
tags: ${{ steps.docker_meta.outputs.tags }}
labels: ${{ steps.docker_meta.outputs.labels }}
-
name: Sign
if: github.event_name != 'pull_request'
uses: ./.github/actions/sign-image
with:
images: chrislusf/seaweedfs@${{ steps.build.outputs.digest }}
@@ -231,7 +231,7 @@ jobs:
- name: Set up QEMU
if: (github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant) && matrix.qemu
uses: docker/setup-qemu-action@v4.2.0
uses: docker/setup-qemu-action@v4.3.0
- name: Create BuildKit config
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -295,10 +295,14 @@ jobs:
if-no-files-found: error
retention-days: 1
# Assemble each variant's per-platform digests into one tag, then mirror to Docker Hub.
# Assemble each variant's per-platform digests into one tag, mirror it to
# Docker Hub, and sign the result on both registries.
merge:
needs: [build]
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
strategy:
fail-fast: false
matrix:
@@ -309,6 +313,13 @@ jobs:
- { variant: large_disk_full, tag_suffix: _large_disk_full }
- { variant: rocksdb, tag_suffix: _large_disk_rocksdb }
steps:
- name: Checkout the signing action
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
uses: actions/checkout@v7
with:
sparse-checkout: .github/actions
persist-credentials: false
- name: Download digests
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
uses: actions/download-artifact@v8
@@ -331,12 +342,17 @@ jobs:
- name: Create multi-arch tag on GHCR
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
id: manifest
working-directory: /tmp/digests
run: |
docker buildx imagetools create \
-t ${{ env.IMAGE }}:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }} \
--metadata-file /tmp/manifest.json \
$(printf '${{ env.IMAGE }}@sha256:%s ' *)
docker buildx imagetools inspect ${{ env.IMAGE }}:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }}
# The copy and the signature below use the digest this run pushed, not whatever the tag points at by then.
digest=$(jq -er '."containerimage.descriptor".digest' /tmp/manifest.json)
echo "digest=${digest}" >> "$GITHUB_OUTPUT"
- name: Login to Docker Hub
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -349,7 +365,9 @@ jobs:
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
run: |
cd $(mktemp -d)
curl -sL "https://github.com/google/go-containerregistry/releases/latest/download/go-containerregistry_Linux_x86_64.tar.gz" | tar xz
curl -sLO https://github.com/google/go-containerregistry/releases/download/v0.22.0/go-containerregistry_Linux_x86_64.tar.gz
echo "edb74d53fad9a596860f59d1c5d04a43dfb5f441dc71f57060dd0bf39483c833 go-containerregistry_Linux_x86_64.tar.gz" | sha256sum -c -
tar xzf go-containerregistry_Linux_x86_64.tar.gz crane
sudo mv crane /usr/local/bin/
crane version
@@ -380,10 +398,18 @@ jobs:
echo "Copying ${{ matrix.variant }} from GHCR to Docker Hub..."
retry_with_backoff crane copy \
${{ env.IMAGE }}:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }} \
${{ env.IMAGE }}@${{ steps.manifest.outputs.digest }} \
chrislusf/seaweedfs:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }}
echo "Copied ${{ matrix.variant }} to Docker Hub"
- name: Sign ${{ matrix.variant }}
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
uses: ./.github/actions/sign-image
with:
images: >-
${{ env.IMAGE }}@${{ steps.manifest.outputs.digest }}
chrislusf/seaweedfs@${{ steps.manifest.outputs.digest }}
# Report-only trivy scan: uploads fixable HIGH/CRITICAL findings to GitHub
# Security for visibility, but never blocks the release. Releases (including
# `latest`) ship regardless — vulnerabilities are tracked, not gated, since
@@ -430,7 +456,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.0
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
@@ -439,7 +465,8 @@ jobs:
# image. crane tag adds an extra tag to an existing manifest — no rebuild,
# no QEMU, no separate workflow. Replaces the old container_latest.yml
# rebuild that often failed or lagged behind the release. Independent of
# trivy-scan: vuln findings are reported but do not block `latest`.
# trivy-scan: vuln findings are reported but do not block `latest`. The
# cosign signature is attached to the digest, so `latest` carries it too.
tag-latest:
runs-on: ubuntu-latest
needs: [merge]
@@ -468,7 +495,9 @@ jobs:
- name: Install crane
run: |
cd $(mktemp -d)
curl -sL "https://github.com/google/go-containerregistry/releases/latest/download/go-containerregistry_Linux_x86_64.tar.gz" | tar xz
curl -sLO https://github.com/google/go-containerregistry/releases/download/v0.22.0/go-containerregistry_Linux_x86_64.tar.gz
echo "edb74d53fad9a596860f59d1c5d04a43dfb5f441dc71f57060dd0bf39483c833 go-containerregistry_Linux_x86_64.tar.gz" | sha256sum -c -
tar xzf go-containerregistry_Linux_x86_64.tar.gz crane
sudo mv crane /usr/local/bin/
crane version
@@ -22,6 +22,9 @@ permissions:
jobs:
build-rocksdb-image:
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- name: Checkout
@@ -82,7 +85,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@96fe6ef7f33517b61c61be40b68a1882f3264fb8 # v1
uses: docker/setup-qemu-action@1f40c72289eff860ee54a304f1438e3cff362e0a # v1
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v1
@@ -94,6 +97,7 @@ jobs:
password: ${{ secrets.DOCKER_PASSWORD }}
- name: Build and push image
id: build
uses: docker/build-push-action@bcafcacb16a39f128d818304e6c9c0c18556b85f # v2
with:
context: ./docker
@@ -108,3 +112,8 @@ jobs:
org.opencontainers.image.title=seaweedfs
org.opencontainers.image.description=SeaweedFS is a distributed storage system for blobs, objects, files, and data lake, to store and serve billions of files fast!
org.opencontainers.image.vendor=Chris Lu
- name: Sign
uses: ./.github/actions/sign-image
with:
images: chrislusf/seaweedfs@${{ steps.build.outputs.digest }}
+17 -2
View File
@@ -90,6 +90,12 @@ jobs:
echo "S3 credentials reference the existing secret for $workload"
done
echo "=== Testing ingress labels ==="
helm template test $CHART_DIR --set s3.enabled=true,s3.ingress.enabled=true \
--set s3.ingress.labels.external-dns=s3 > /tmp/s3-ingress-labels.yaml
grep -A 12 "^kind: Ingress" /tmp/s3-ingress-labels.yaml | grep -q "^ external-dns: s3"
echo "S3 ingress renders custom labels"
echo "=== Testing with all-in-one mode ==="
helm template test $CHART_DIR --set allInOne.enabled=true > /tmp/allinone.yaml
grep -q "seaweedfs-all-in-one" /tmp/allinone.yaml
@@ -221,6 +227,9 @@ jobs:
out = render({
"global.seaweedfs.securityConfig.jwtSigning.filerWrite": "true",
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
})
cm = configmap(out, "test-seaweedfs-security-config")
if cm is None:
@@ -1135,6 +1144,9 @@ jobs:
"s3.enabled": "true",
"sftp.enabled": "true",
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
"worker.enabled": "true",
"cosi.enabled": "true",
"s3.createBuckets[0].name": "b",
@@ -1331,7 +1343,7 @@ jobs:
# Which means egress on its own must render for a release that runs
# neither COSI nor a resize: no component of it reaches the API server,
# so nothing may demand a CIDR for one.
for label, values in {"defaults": {}, "admin": {"admin.enabled": "true"}}.items():
for label, values in {"defaults": {}, "admin": {"admin.enabled": "true", "admin.secret.adminPassword": "ci-admin-password"}}.items():
try:
render(dict(values, **{"networkPolicy.enabled": "true",
"networkPolicy.egress.enabled": "true"}))
@@ -1392,6 +1404,9 @@ jobs:
ALL_ON = {
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
"s3.enabled": "true",
"sftp.enabled": "true",
"worker.enabled": "true",
@@ -1574,7 +1589,7 @@ jobs:
echo "IMAGE_TAG=$TAG" >> $GITHUB_ENV
- name: Create kind cluster
uses: helm/kind-action@v1.14.0
uses: helm/kind-action@v1.15.0
- name: Run chart-testing (install)
run: |
+11
View File
@@ -185,8 +185,19 @@ jobs:
if (Test-Path S:\walk) { throw "directory survived recursive delete" }
Write-Host "::endgroup::"
# Tear down everything we started so the runner's later steps (Logs,
# post-checkout) launch into a clean environment. A WinFsp mount left
# active and a weed.exe holding winfsp-x64.dll have been seen to make
# the next step's pwsh.exe fail with STATUS_DLL_INIT_FAILED
# (0xC0000142), which fails a job whose actual test step passed.
try { Stop-Mount } catch { Write-Host "no mount to stop at teardown: $_" }
Get-CimInstance Win32_Process -Filter "Name = 'weed.exe'" |
ForEach-Object { Stop-Process -Id $_.ProcessId -Force -ErrorAction SilentlyContinue }
Start-Sleep -Seconds 3
- name: Logs
if: always()
continue-on-error: true
shell: pwsh
run: |
foreach ($f in 'C:\seaweed-mount.log','C:\seaweed-mount.err.log','C:\seaweed-remount.log','C:\seaweed-remount.err.log','C:\seaweed-remount2.log','C:\seaweed-remount2.err.log','C:\seaweed-dirmount.log','C:\seaweed-dirmount.err.log','C:\seaweed-mini.log','C:\seaweed-mini.err.log') {
+52 -10
View File
@@ -178,6 +178,8 @@ jobs:
needs: release
if: ${{ inputs.downstream && !inputs.dry_run }}
runs-on: ubuntu-latest
# The wait below puts no bound of its own on runner-queue time; this does.
timeout-minutes: 120
permissions: {}
strategy:
fail-fast: false
@@ -204,8 +206,9 @@ jobs:
# The dispatched workflow pins seaweedfs with `go get -u ...@latest`, so
# wait until the proxy serves the release commit as the tip. Asking for
# the commit by name is what makes the proxy fetch it.
for _ in $(seq 30); do
# the commit by name is what makes the proxy fetch it. The proxy can
# take longer than five minutes to refresh @latest after a new tag.
for _ in $(seq 120); do
curl -sf "https://proxy.golang.org/${MODULE}/@v/${SHA}.info" >/dev/null || true
TIP=$(curl -sf "https://proxy.golang.org/${MODULE}/@latest" | jq -r '.Origin.Hash // ""' || true)
[ "$TIP" = "$SHA" ] && break
@@ -216,21 +219,60 @@ jobs:
exit 1
fi
# Wait on the release the dispatched workflow publishes, not on the run
# that publishes it: a dispatch cannot be told apart from a concurrent
# one through the API, and the release is what we are here for.
# The dispatched workflow publishes the release as its last step, so
# its conclusion decides success. Wait on the run, not on a wall
# clock: time it spends queued for a runner must not count against
# the budget. A dispatch does not return its run id, so take the
# newest workflow_dispatch run created since ours; a concurrent
# dispatch would be performing this same release, and waiting on it
# is just as good.
released() { gh api "repos/${REPO}/releases?per_page=30" --jq '[.[].tag_name]'; }
BEFORE=$(released)
# A minute early, so runner clock skew cannot hide the run.
DISPATCHED_AT=$(date -u -d '1 minute ago' '+%Y-%m-%dT%H:%M:%SZ')
gh workflow run -R "$REPO" "$WORKFLOW" --ref master -f bump=patch -f update_seaweedfs=true
for _ in $(seq 80); do
sleep 15
NEW=$(released | jq -c --argjson before "$BEFORE" '. - $before')
[ "$(jq length <<<"$NEW")" -gt 0 ] && break
RUN_ID=""
for _ in $(seq 12); do
sleep 10
RUN_ID=$(gh api -X GET "repos/${REPO}/actions/workflows/${WORKFLOW}/runs" \
-f event=workflow_dispatch -f "created=>=${DISPATCHED_AT}" \
--jq '(.workflow_runs | sort_by(.created_at) | last | .id) // empty' || true)
[ -n "$RUN_ID" ] && break
done
if [ -z "$RUN_ID" ]; then
echo "::error::the dispatch created no ${WORKFLOW} run in ${REPO}; see https://github.com/${REPO}/actions/workflows/${WORKFLOW}"
exit 1
fi
RUN_URL="https://github.com/${REPO}/actions/runs/${RUN_ID}"
echo "waiting on ${RUN_URL}"
# 20 minutes of execution; polls that find the run still queued do
# not consume it.
RUNNING=0
STATE=""
while :; do
sleep 15
STATE=$(gh api "repos/${REPO}/actions/runs/${RUN_ID}" \
--jq '.status + "/" + (.conclusion // "")' || true)
case "$STATE" in
completed/*) break ;;
in_progress/*) RUNNING=$((RUNNING + 1)) ;;
esac
if [ "$RUNNING" -gt 80 ]; then
echo "::error::${RUN_URL} has been executing for over 20 minutes; giving up on it"
exit 1
fi
done
if [ "$STATE" != "completed/success" ]; then
echo "::error::${RUN_URL} concluded '${STATE#completed/}'"
exit 1
fi
NEW=$(released | jq -c --argjson before "$BEFORE" '. - $before')
if [ "$(jq length <<<"$NEW")" -eq 0 ]; then
echo "::error::${REPO} published no release within 20 minutes; see https://github.com/${REPO}/actions/workflows/${WORKFLOW}"
echo "::error::${RUN_URL} succeeded but ${REPO} shows no new release"
exit 1
fi
echo "${REPO} released $(jq -r 'join(", ")' <<<"$NEW")"
@@ -59,9 +59,18 @@ jobs:
- name: Build Rust volume server
run: cd seaweed-volume && cargo build --release
# The crate is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[lints.clippy]` in
# seaweed-volume/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-volume && cargo clippy --all-targets -- -D warnings
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
- name: Run Rust unit tests (redb experimental cursor)
run: cd seaweed-volume && cargo test --features redb-experimental-cursor --lib storage::needle_map
rust-integration-tests:
name: Rust Integration Tests
runs-on: ubuntu-22.04
+6
View File
@@ -73,6 +73,12 @@ jobs:
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The workspace is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[workspace.lints.clippy]`
# in seaweed-worker/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-worker && cargo clippy --workspace --all-targets -- -D warnings
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
+6
View File
@@ -128,6 +128,9 @@ jobs:
echo "All SeaweedFS components are ready!"
cd ../s3-tests
sed -i "s/assert prefixes == \['foo%2B1\/', 'foo\/', 'quux%20ab\/'\]/assert prefixes == \['foo\/', 'foo%2B1\/', 'quux%20ab\/'\]/" s3tests/functional/test_s3.py
# The suite expects RGW's 400 InvalidPart for a partNumber past the last
# part; AWS answers 416 InvalidPartNumber, which is what we return.
sed -i "/# request PartNumber out of range/,+4{s/assert status == 400/assert status == 416/; s/assert error_code == 'InvalidPart'/assert error_code == 'InvalidPartNumber'/}" s3tests/functional/test_s3.py
# Debug: Show the config file contents
echo "=== S3 Config File Contents ==="
@@ -968,6 +971,9 @@ jobs:
echo "All SeaweedFS components are ready!"
cd ../s3-tests
sed -i "s/assert prefixes == \['foo%2B1\/', 'foo\/', 'quux%20ab\/'\]/assert prefixes == \['foo\/', 'foo%2B1\/', 'quux%20ab\/'\]/" s3tests/functional/test_s3.py
# The suite expects RGW's 400 InvalidPart for a partNumber past the last
# part; AWS answers 416 InvalidPartNumber, which is what we return.
sed -i "/# request PartNumber out of range/,+4{s/assert status == 400/assert status == 416/; s/assert error_code == 'InvalidPart'/assert error_code == 'InvalidPartNumber'/}" s3tests/functional/test_s3.py
# Create and update s3tests.conf to use port 8004
cp ../docker/compose/s3tests.conf ../docker/compose/s3tests-sql.conf
sed -i 's/port = 8000/port = 8004/g' ../docker/compose/s3tests-sql.conf
+74
View File
@@ -0,0 +1,74 @@
---
name: Star History
on:
schedule:
- cron: "0 0 * * *" # daily, 00:00 UTC
workflow_dispatch:
permissions:
contents: write
concurrency:
# Only one chart regeneration per branch at a time; a newer run on the same
# branch cancels an in-flight one so overlapping runs never conflict on
# note/star_history.svg during rebase. Scoped by ref so a manual run on
# another branch can't cancel the daily master update.
group: star-history-${{ github.ref }}
cancel-in-progress: true
jobs:
render:
name: Regenerate star history chart
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history so the chart commit can rebase onto a moved master.
fetch-depth: 0
- name: Set up Python
uses: actions/setup-python@v7
with:
python-version: "3.x"
cache: pip
cache-dependency-path: .github/scripts/star_history.py
- name: Install matplotlib
run: pip install matplotlib
- name: Render chart
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: python .github/scripts/star_history.py
- name: Commit if changed
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
if git diff --quiet -- note/star_history.svg; then
echo "No changes to the chart."
exit 0
fi
git add note/star_history.svg
git commit -m "docs: regenerate star history chart"
# Rebase and retry so a concurrent push to master doesn't lose the chart.
for attempt in 1 2 3 4 5; do
if [ "$attempt" -gt 1 ]; then
# Guard the rebase: a transient fetch error or conflict must not
# abort the fail-fast shell before the remaining attempts run.
if ! git pull --rebase origin "$GITHUB_REF_NAME"; then
echo "rebase failed (attempt ${attempt}); aborting and retrying"
git rebase --abort || true
continue
fi
fi
if git push origin HEAD:"$GITHUB_REF_NAME"; then
exit 0
fi
echo "push rejected (attempt ${attempt}); will rebase and retry"
done
echo "::error::could not push star history chart after retries"
exit 1
-1
View File
@@ -116,7 +116,6 @@ test/s3/versioning/weed-test.log
/docker/admin_integration/data
docker/agent_pub_record
docker/admin_integration/weed-local
/seaweedfs-rdma-sidecar/bin
/test/s3/encryption/filerldb2
/test/s3/sse/filerldb2
test/s3/sse/weed-test.log
+318 -488
View File
@@ -12,425 +12,257 @@
![SeaweedFS Logo](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/seaweedfs.png)
<h2 align="center"><a href="https://www.patreon.com/seaweedfs">Sponsor SeaweedFS via Patreon</a></h2>
SeaweedFS is a simple and highly scalable distributed file system. There are two objectives:
SeaweedFS is an independent Apache-licensed open source project with its ongoing development made
possible entirely thanks to the support of these awesome [backers](https://github.com/seaweedfs/seaweedfs/blob/master/backers.md).
If you'd like to grow SeaweedFS even stronger, please consider joining our
<a href="https://www.patreon.com/seaweedfs">sponsors on Patreon</a>.
1. to store billions of files!
2. to serve the files fast!
Your support will be really appreciated by me and other supporters!
<!--
<h4 align="center">Platinum</h4>
<p align="center">
<a href="" target="_blank">
Add your name or icon here
</a>
</p>
-->
### Gold Sponsors
[![nodion](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/sponsor_nodion.png)](https://www.nodion.com)
[![piknik](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/piknik.png)](https://www.piknik.com)
[![keepsec](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/keepsec.png)](https://www.keepsec.ca)
[![zyner](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/sponsor_zyner.png)](https://zyner.org)
---
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently.
- [Download Binaries for different platforms](https://github.com/seaweedfs/seaweedfs/releases/latest)
- [SeaweedFS on Slack](https://join.slack.com/t/seaweedfs/shared_invite/enQtMzI4MTMwMjU2MzA3LTEyYzZmZWYzOGQ3MDJlZWMzYmI0OTE4OTJiZjJjODBmMzUxNmYwODg0YjY3MTNlMjBmZDQ1NzQ5NDJhZWI2ZmY)
- [SeaweedFS on Twitter](https://twitter.com/SeaweedFS)
- [SeaweedFS on Telegram](https://t.me/Seaweedfs)
- [SeaweedFS on Reddit](https://www.reddit.com/r/SeaweedFS/)
- [SeaweedFS Mailing List](https://groups.google.com/d/forum/seaweedfs)
- [Wiki Documentation](https://github.com/seaweedfs/seaweedfs/wiki)
- [SeaweedFS White Paper](https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS_Architecture.pdf)
- [SeaweedFS Introduction Slides 2025.5](https://docs.google.com/presentation/d/1tdkp45J01oRV68dIm4yoTXKJDof-EhainlA0LMXexQE/edit?usp=sharing)
- [SeaweedFS Introduction Slides 2021.5](https://docs.google.com/presentation/d/1DcxKWlINc-HNCjhYeERkpGXXm6nTCES8mi2W5G0Z4Ts/edit?usp=sharing)
- [SeaweedFS Introduction Slides 2019.3](https://www.slideshare.net/chrislusf/seaweedfs-introduction)
- Community: [Slack](https://join.slack.com/t/seaweedfs/shared_invite/enQtMzI4MTMwMjU2MzA3LTEyYzZmZWYzOGQ3MDJlZWMzYmI0OTE4OTJiZjJjODBmMzUxNmYwODg0YjY3MTNlMjBmZDQ1NzQ5NDJhZWI2ZmY), [Twitter](https://twitter.com/SeaweedFS), [Telegram](https://t.me/Seaweedfs), [Reddit](https://www.reddit.com/r/SeaweedFS/), [Mailing List](https://groups.google.com/d/forum/seaweedfs)
- [SeaweedFS White Paper](https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS_Architecture.pdf) and introduction slides: [2025.5](https://docs.google.com/presentation/d/1tdkp45J01oRV68dIm4yoTXKJDof-EhainlA0LMXexQE/edit?usp=sharing), [2021.5](https://docs.google.com/presentation/d/1DcxKWlINc-HNCjhYeERkpGXXm6nTCES8mi2W5G0Z4Ts/edit?usp=sharing), [2019.3](https://www.slideshare.net/chrislusf/seaweedfs-introduction)
Table of Contents
=================
* [Quick Start](#quick-start)
* [Quick Start with weed mini](#quick-start-with-weed-mini)
* [Quick Start for S3 API on Docker](#quick-start-for-s3-api-on-docker)
* [Introduction](#introduction)
* [Features](#features)
* [Additional Features](#additional-features)
* [Filer Features](#filer-features)
* [Example: Using Seaweed Blob Store](#example-using-seaweed-blob-store)
* [Architecture](#object-store-architecture)
* [Compared to Other File Systems](#compared-to-other-file-systems)
* [One command](#one-command)
* [Docker](#docker)
* [Docker Compose](#docker-compose)
* [Kubernetes with Helm](#kubernetes-with-helm)
* [Build from source](#build-from-source)
* [Scale out](#scale-out)
* [Why SeaweedFS](#why-seaweedfs)
* [Fast](#fast)
* [Scalable](#scalable)
* [The most complete S3 API](#the-most-complete-s3-api)
* [A data warehouse with S3 Tables](#a-data-warehouse-with-s3-tables)
* [A fast cache for cloud storage](#a-fast-cache-for-cloud-storage)
* [Active-active replication and more](#active-active-replication-and-more)
* [Architecture](#architecture)
* [Compared to Other Systems](#compared-to-other-systems)
* [Compared to HDFS](#compared-to-hdfs)
* [Compared to GlusterFS, Ceph](#compared-to-glusterfs-ceph)
* [Compared to GlusterFS](#compared-to-glusterfs)
* [Compared to MooseFS](#compared-to-moosefs)
* [Compared to Ceph](#compared-to-ceph)
* [Compared to MinIO, RustFS](#compared-to-minio-rustfs)
* [Dev Plan](#dev-plan)
* [Installation Guide](#installation-guide)
* [Disk Related Topics](#disk-related-topics)
* [Benchmark](#benchmark)
* [Enterprise](#enterprise)
* [License](#license)
* [Sponsors](#sponsors)
# Quick Start #
## One command ##
## Quick Start with weed mini ##
Download the latest binary from the [releases](https://github.com/seaweedfs/seaweedfs/releases/latest) page and unzip the single `weed` (or `weed.exe`) file, or let the install script put it in `/usr/local/bin`:
Download the latest binary from https://github.com/seaweedfs/seaweedfs/releases and unzip the single `weed` (or `weed.exe`) file, or run `go install github.com/seaweedfs/seaweedfs/weed@latest`. Then start a ready-to-use S3 object store with credentials and a pre-created bucket in one command:
```bash
curl -fsSL https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/install.sh | bash
```
Then start a ready-to-use S3 object store:
```bash
AWS_ACCESS_KEY_ID=admin \
AWS_SECRET_ACCESS_KEY=secret \
S3_BUCKET=my-bucket \
./weed mini -dir=/data
./weed mini -dir=./data
```
That's it — the S3 endpoint is at http://localhost:8333, `my-bucket` already exists, and `admin`/`secret` are valid credentials. `S3_BUCKET` accepts a comma-separated list (e.g. `raw,processed`); use `S3_TABLE_BUCKET` for S3 Tables buckets, each `name` or `name:FORMAT` where the format is `ICEBERG` (the default) or `LANCE`. Drop any of the env vars to skip that piece (no AWS keys → S3 runs in unauthenticated "Allow All" mode for development).
That's it. The S3 endpoint is at http://localhost:8333, `my-bucket` exists, and `admin`/`secret` are valid credentials:
The same command starts everything else too:
- **S3 Endpoint**: http://localhost:8333
- **Master UI**: http://localhost:9333
- **Volume Server**: http://localhost:9340
- **Filer UI**: http://localhost:8888
- **WebDAV**: http://localhost:7333
- **Admin UI**: http://localhost:23646
```bash
AWS_ACCESS_KEY_ID=admin AWS_SECRET_ACCESS_KEY=secret \
aws --endpoint-url http://localhost:8333 s3 cp README.md s3://my-bucket/
```
The same process also runs the master, a volume server, the filer, WebDAV, the Iceberg REST catalog, and the Admin UI. Add `S3_TABLE_BUCKET=warehouse` to also create an Iceberg table bucket, or `warehouse:LANCE` for a Lance one. Drop the AWS keys to run without authentication for development.
> macOS: if the binary is quarantined, run `xattr -d com.apple.quarantine ./weed` first.
Perfect for development, testing, learning SeaweedFS, and single-node deployments. To scale out, add more volume servers by running `weed volume -dir="/some/data/dir2" -master="<master_host>:9333" -port=8081` locally, on another machine, or on thousands of machines.
`weed mini` is auto-tuned for one node and is fine for single-node production, such as an S3 gateway that issues presigned URLs. See [Quick Start with weed mini][WeedMini].
## Quick Start for S3 API on Docker ##
## Docker ##
```bash
docker run -p 8333:8333 \
docker run -p 8333:8333 -v weed-data:/data \
-e AWS_ACCESS_KEY_ID=admin \
-e AWS_SECRET_ACCESS_KEY=secret \
-e S3_BUCKET=my-bucket \
chrislusf/seaweedfs
```
Same behavior as the `weed mini` command above — the S3 endpoint is at http://localhost:8333 with `my-bucket` pre-created. Drop the env vars to run anonymously for development.
Same behavior as the `weed mini` command above.
# Introduction #
## Docker Compose ##
SeaweedFS is a simple and highly scalable distributed file system. There are two objectives:
To run master, volume server, filer, S3, and WebDAV as separate services:
1. to store billions of files!
2. to serve the files fast!
```bash
wget https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/docker/seaweedfs-compose.yml
wget -P prometheus https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/docker/prometheus/prometheus.yml
docker compose -f seaweedfs-compose.yml -p seaweedfs up
```
SeaweedFS started as a blob store to handle small files efficiently.
Instead of managing all file metadata in a central master,
the central master only manages volumes on volume servers,
and these volume servers manage files and their metadata.
This relieves concurrency pressure from the central master and spreads file metadata into volume servers,
allowing faster file access (O(1), usually just one disk read operation).
[Docker Compose for S3][DockerComposeS3] adds credentials, and the [docker/compose](docker/compose) folder has variants for replication, mounts, message queues, and more.
There is only 40 bytes of disk storage overhead for each file's metadata.
It is so simple with O(1) disk reads that you are welcome to challenge the performance with your actual use cases.
## Kubernetes with Helm ##
SeaweedFS started by implementing [Facebook's Haystack design paper](http://www.usenix.org/event/osdi10/tech/full_papers/Beaver.pdf).
Also, SeaweedFS implements erasure coding with ideas from
[f4: Facebook’s Warm BLOB Storage System](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-muralidhar.pdf), and has a lot of similarities with [Facebook’s Tectonic Filesystem](https://www.usenix.org/system/files/fast21-pan.pdf) and [Google's Colossus File System](https://cloud.google.com/blog/products/storage-data-transfer/a-peek-behind-colossus-googles-file-system)
```bash
helm repo add seaweedfs https://seaweedfs.github.io/seaweedfs/helm
helm install seaweedfs seaweedfs/seaweedfs -n seaweedfs --create-namespace -f values.yaml
```
On top of the blob store, optional [Filer] can support directories and POSIX attributes.
Filer is a separate linearly-scalable stateless server with customizable metadata stores,
e.g., MySql, Postgres, Redis, Cassandra, HBase, Mongodb, Elastic Search, LevelDB, RocksDB, Sqlite, MemSql, TiDB, Etcd, CockroachDB, YDB, etc.
A production-shaped `values.yaml` for a three-node cluster: two copies of every write, three masters, and an S3 endpoint with credentials and a bucket.
SeaweedFS can transparently integrate with the cloud.
With hot data on local cluster, and warm data on the cloud with O(1) access time,
SeaweedFS can achieve both fast local access time and elastic cloud storage capacity.
What's more, the cloud storage access API cost is minimized.
Faster and cheaper than direct cloud storage!
```yaml
global:
seaweedfs:
enableReplication: true
replicationPlacement: "001" # one extra copy on another server; "002" for two
SeaweedFS also ships a built-in **Iceberg REST Catalog**, turning the same cluster into a self-contained lakehouse.
Spark, Trino, Dremio, DuckDB, and RisingWave can query Iceberg tables directly — no Hive Metastore, Glue, or
external catalog service required. Storage and table metadata live in one system, simplifying on-prem and
small-team analytics stacks.
master:
replicas: 3
data:
type: persistentVolumeClaim # the cluster's default storage class; add storageClass to pick one
size: 1Gi
volume:
replicas: 3 # at least 1 + the sum of the replication digits
dataDirs:
- name: data
type: persistentVolumeClaim
size: 500Gi
maxVolumes: 0 # size the volume count from the disk
filer:
replicas: 2
data:
type: persistentVolumeClaim
size: 20Gi
s3:
enabled: true
replicas: 2
enableAuth: true
credentials:
admin:
accessKey: admin
secretKey: change-me
createBuckets:
- name: app-storage
```
The S3 endpoint is the `seaweedfs-s3` service on port 8333. [Helm Chart Recipes][HelmRecipes] has values for a development cluster, a lakehouse with the Iceberg catalog exposed, filer metadata on PostgreSQL, and node-local disks. The [SeaweedFS Operator][Operator] and the [CSI driver][SeaweedFsCsiDriver] are the other Kubernetes paths.
## Build from source ##
```bash
git clone https://github.com/seaweedfs/seaweedfs.git
cd seaweedfs/weed && make install
```
`weed` lands in `$GOPATH/bin`. [Getting Started][GettingStarted] covers running master, volume, filer, and S3 as separate processes.
## Scale out ##
Capacity is a volume server. Start one on any machine with disk and point it at the master:
```bash
weed volume -dir=/data -master=<master_host>:9333
```
Nothing rebalances until you ask it to. Throughput is a filer or S3 gateway; they are stateless, so run as many as you need behind a load balancer. [Production Setup][ProductionSetup] walks through a multi-node cluster.
[Back to TOC](#table-of-contents)
# Features #
## Additional Blob Store Features ##
* Support different replication levels, with rack and data center aware.
* Automatic master servers failover - no single point of failure (SPOF).
* Automatic compression depending on file MIME type.
* Automatic compaction to reclaim disk space after deletion or update.
* [Automatic entry TTL expiration][VolumeServerTTL].
* Flexible Capacity Expansion: Any server with some disk space can add to the total storage space.
* Adding/Removing servers does **not** cause any data re-balancing unless triggered by admin commands.
* Optional picture resizing.
* Support ETag, Accept-Range, Last-Modified, etc.
* Support in-memory/leveldb/readonly mode tuning for memory/performance balance.
* Support rebalancing the writable and readonly volumes.
* [Customizable Multiple Storage Tiers][TieredStorage]: Customizable storage disk types to balance performance and cost.
* [Transparent cloud integration][CloudTier]: unlimited capacity via tiered cloud storage for warm data.
* [Erasure Coding for warm storage][ErasureCoding] Rack-Aware 10.4 erasure coding reduces storage cost and increases availability. Enterprise version can customize EC ratio.
# Why SeaweedFS #
## Fast ##
* One disk read per blob. A small file is one blob; a large file is split into chunks of a few MB, each its own blob. A volume server keeps a 16-byte index entry per blob in memory and reads it in a single seek, also for erasure-coded data.
* The master is not in the read path. Clients cache the volume-to-server mapping and talk to volume servers directly.
* 40 bytes of metadata per file on disk. Small files are packed into append-only volume files, so there is no per-file inode, no per-file metadata file, no fragmentation, and writes are SSD friendly.
* Hot data is replicated; [erasure coding][ErasureCoding] is applied to warm data in the background, so writes never pay the encoding cost.
* The [Rust volume server][RustVolume] is a drop-in for higher throughput and lower tail latency on the same on-disk format.
On one laptop, [`weed benchmark`][Benchmarks] writes 1KB files at 15,700 per second and reads them back at 47,000 per second, and a mixed S3 [warp][S3Benchmark] run totals 3.2 GiB/s. Numbers are in the [Benchmark](#benchmark) section; throughput grows with volume servers and gateways.
## Scalable ##
* The master tracks volumes, not files. A cluster with billions of files has a few thousand volumes, so the master stays small. One master is enough for most clusters; run three for [Raft failover][FailoverMaster].
* Adding a server adds capacity with no data reshuffle. Balancing, vacuum, erasure coding, and repair run on demand from [`weed shell`][WeedShell] or the [maintenance worker][Worker].
* Filer and S3 gateways are stateless and scale linearly. Directory metadata lives in a [store you already run][FilerStores]: LevelDB, RocksDB, SQLite, MySQL, PostgreSQL, Cassandra, HBase, MongoDB, Redis, Elasticsearch, etcd, TiKV, FoundationDB, YDB, ArangoDB, Tarantool, and MySQL or PostgreSQL compatible databases such as TiDB, CockroachDB, and MemSQL.
* Rack and data center aware [replication][Replication], [tiered storage][TieredStorage] across disk types, and [transparent cloud tiering][CloudTier] for unlimited capacity.
* Files from a byte to [tens of TB][SuperLargeFiles]. Volumes up to 8TB with the large-disk build.
## The most complete S3 API ##
The S3 gateway implements the object, bucket, S3 Tables, IAM, and STS APIs on one endpoint, so the AWS SDKs and CLI, rclone, restic, Spark, and Trino work unchanged.
| API | Operations |
| --- | --- |
| S3 bucket and object | 73 |
| S3 Tables | 36 |
| IAM | 39 |
| STS | 5 |
* [Versioning][Versioning], [Object Lock][ObjectLock] with retention and legal hold, [lifecycle][Lifecycle] rules, tagging, [CORS][CORS], [conditional reads and writes][ConditionalOps], checksums, presigned URLs, browser POST uploads, multipart uploads, and an atomic [RenameObject][RenameObject].
* [Bucket policies][BucketPolicies] with [conditions][PolicyConditions] and [variables][PolicyVariables]; IAM users, groups, and policies; STS with [OIDC][OIDC], LDAP, and [Kubernetes service accounts][K8sSA].
* [SSE-S3, SSE-KMS, and SSE-C][SSE] server-side encryption, with OpenBao and Vault, AWS KMS, Azure Key Vault, and GCP KMS as key providers.
* [Audit log][AuditLog], [bucket quota][BucketQuota], and [rate limiting][RateLimiting].
* Each bucket is its own collection, so deleting a bucket is instant.
The full operation list is in [Amazon S3 API][AmazonS3API], and [Supported APIs vs MinIO][S3vsMinio] compares. The S3 compatibility suite and the SDK, IAM, SSE, policy, and Spark integration tests run in CI on every change.
## A data warehouse with S3 Tables ##
SeaweedFS is a lakehouse in one system. [S3 Table Buckets][S3TableBucket] hold Apache Iceberg tables by default, or [Lance][LanceCatalog] tables for vectors and multimodal data, and the built-in [Iceberg REST Catalog][IcebergCatalog] and Lance namespace serve them directly. There is no Hive Metastore, Glue, or separate catalog service to deploy, secure, and back up.
* Query engines operate on the same tables at the same time: [Spark][SparkIceberg], [Trino][TrinoIceberg], [Dremio][DremioIceberg], [DuckDB][DuckDBIceberg], [Apache Doris][DorisIceberg], [RisingWave][RisingWaveIceberg], ClickHouse, and [LanceDB][LanceDB]. Catalog commits are atomic compare-and-swap, so concurrent writers are safe. [Lakekeeper][Lakekeeper] can front the same storage with STS-vended credentials.
* [Automated table maintenance][IcebergMaintenance]: compaction, snapshot expiration, orphan file removal, and manifest rewriting, configured per bucket or table through the S3 Tables maintenance APIs, and the same for [Lance][LanceMaintenance].
* IAM at the bucket, namespace, and table level with standard bucket policies, see [S3 Tables Security][S3TablesSecurity].
* A [Hadoop compatible file system][Hadoop] for Spark, Flink, and HBase.
`S3_TABLE_BUCKET=warehouse ./weed mini -dir=./data` brings the whole stack up on a laptop.
## A fast cache for cloud storage ##
[Cloud Drive][CloudDrive] mounts a bucket from S3, Google Cloud Storage, Azure, Backblaze B2, Wasabi, Storj, or any S3-compatible store into SeaweedFS and serves it at local speed:
* Metadata is pulled once, so listing, stat, and directory walks cost no cloud API calls.
* File content is downloaded once, on first read or [warmed][CacheRemote] by folder, name pattern, size, or age, and cached with the capacity of the whole cluster: cache everything, no churn.
* Local writes complete at local latency and are written back to the cloud asynchronously in the cloud's native layout, so other tools keep reading the bucket directly.
* Uncache by the same rules to free local disk while keeping the metadata.
[Cloud Tier][CloudTier] goes the other direction, moving whole warm volumes to cloud storage while keeping one-read access, and the [Gateway to Remote Object Storage][GatewayToRemoteObjectStore] mirrors every bucket to a remote store. Faster and cheaper than reading the cloud directly.
## Active-active replication and more ##
* [Active-active or active-passive replication][ActiveActiveAsyncReplication] between clusters, continuous and resumable, for the whole tree or chosen folders, across data centers.
* [Filer store replication][FilerStoreReplication] for metadata HA, [async backup][AsyncBackup] to cloud storage, [metadata backup][MetaBackup], and [change data capture][CDC] with [webhooks][Webhook] on every metadata event.
* The same data as a [FUSE mount][Mount] on Linux, macOS, and [Windows][MountWindows], over [WebDAV][WebDAV], [SFTP][SFTP], HDFS, HTTP, and [TUS resumable uploads][TUS]; on Kubernetes through the [CSI driver][SeaweedFsCsiDriver] and [Operator][Operator].
* [AES256-GCM encryption at rest][FilerDataEncryption], TLS and mTLS between components, JWT-signed volume access, and [FIPS][FIPS] builds.
* [Admin UI][AdminUI], Prometheus [metrics][Metrics], [TTL][VolumeServerTTL] per file or volume, automatic compression and compaction, and [seaweed-up][SeaweedUp] for bare-metal clusters.
[Back to TOC](#table-of-contents)
## Filer Features ##
* [Filer server][Filer] provides "normal" directories and files via HTTP.
* [File TTL][FilerTTL] automatically expires file metadata and actual file data.
* [Mount filer][Mount] reads and writes files directly as a local directory via FUSE.
* [Filer Store Replication][FilerStoreReplication] enables HA for filer meta data stores.
* [Active-Active Replication][ActiveActiveAsyncReplication] enables asynchronous one-way or two-way cross cluster continuous replication.
* [Amazon S3 compatible API][AmazonS3API] accesses files with S3 tooling.
* [Hadoop Compatible File System][Hadoop] accesses files from Hadoop/Spark/Flink/etc or even runs HBase.
* [Async Replication To Cloud][BackupToCloud] has extremely fast local access and backups to Amazon S3, Google Cloud Storage, Azure, BackBlaze.
* [WebDAV] accesses as a mapped drive on Mac and Windows, or from mobile devices.
* [AES256-GCM Encrypted Storage][FilerDataEncryption] safely stores the encrypted data.
* [Super Large Files][SuperLargeFiles] stores large or super large files in tens of TB.
* [Cloud Drive][CloudDrive] mounts cloud storage to local cluster, cached for fast read and write with asynchronous write back.
* [Gateway to Remote Object Store][GatewayToRemoteObjectStore] mirrors bucket operations to remote object storage, in addition to [Cloud Drive][CloudDrive]
# Architecture #
## Data Lakehouse Features ##
* [S3 Table Buckets][S3TableBucket] expose a dedicated namespace for Iceberg tables with strict layout validation.
* Built-in [Iceberg REST Catalog][IcebergCatalog] runs alongside the S3 endpoint — no external metastore needed.
* Native integrations with [Apache Spark][SparkIceberg], [Trino][TrinoIceberg], [Dremio][DremioIceberg], [DuckDB][DuckDBIceberg], and [RisingWave][RisingWaveIceberg].
* [Automated table maintenance][IcebergMaintenance]: compaction, snapshot expiration, orphan removal, manifest rewriting.
* Granular IAM at the bucket, namespace, and table level via standard S3 bucket policies.
![SeaweedFS Architecture](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/SeaweedFS_Architecture.png)
## Kubernetes ##
* [Kubernetes CSI Driver][SeaweedFsCsiDriver] A Container Storage Interface (CSI) Driver. [![Docker Pulls](https://img.shields.io/docker/pulls/chrislusf/seaweedfs-csi-driver.svg?maxAge=4800)](https://hub.docker.com/r/chrislusf/seaweedfs-csi-driver/)
* [SeaweedFS Operator](https://github.com/seaweedfs/seaweedfs-operator)
[Filer]: https://github.com/seaweedfs/seaweedfs/wiki/Directories-and-Files
[SuperLargeFiles]: https://github.com/seaweedfs/seaweedfs/wiki/Data-Structure-for-Large-Files
[Mount]: https://github.com/seaweedfs/seaweedfs/wiki/FUSE-Mount
[AmazonS3API]: https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API
[BackupToCloud]: https://github.com/seaweedfs/seaweedfs/wiki/Async-Replication-to-Cloud
[Hadoop]: https://github.com/seaweedfs/seaweedfs/wiki/Hadoop-Compatible-File-System
[WebDAV]: https://github.com/seaweedfs/seaweedfs/wiki/WebDAV
[ErasureCoding]: https://github.com/seaweedfs/seaweedfs/wiki/Erasure-coding-for-warm-storage
[TieredStorage]: https://github.com/seaweedfs/seaweedfs/wiki/Tiered-Storage
[CloudTier]: https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier
[FilerDataEncryption]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Data-Encryption
[FilerTTL]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Stores
[VolumeServerTTL]: https://github.com/seaweedfs/seaweedfs/wiki/Store-file-with-a-Time-To-Live
[SeaweedFsCsiDriver]: https://github.com/seaweedfs/seaweedfs-csi-driver
[ActiveActiveAsyncReplication]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Active-Active-cross-cluster-continuous-synchronization
[FilerStoreReplication]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Store-Replication
[KeyLargeValueStore]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-as-a-Key-Large-Value-Store
[CloudDrive]: https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Architecture
[GatewayToRemoteObjectStore]: https://github.com/seaweedfs/seaweedfs/wiki/Gateway-to-Remote-Object-Storage
[S3TableBucket]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Table-Bucket
[IcebergCatalog]: https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS-Iceberg-Catalog
[IcebergMaintenance]: https://github.com/seaweedfs/seaweedfs/wiki/Iceberg-Table-Maintenance
[SparkIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Spark-Iceberg-Integration
[TrinoIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Trino-Iceberg-Integration
[DremioIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Dremio-Iceberg-Integration
[DuckDBIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/DuckDB-Iceberg-Integration
[RisingWaveIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/RisingWave-Iceberg-Integration
* **Master** servers, one or a Raft group of three, track which volume lives on which volume server and hand out file ids. They are not in the read path.
* **Volume** servers store blobs in append-only volume files, keep a 16-byte in-memory index per blob, and replicate or erasure-code at the volume level.
* **Filer** servers add directories and files on top, with metadata in a store of your choice, and expose HTTP, S3, WebDAV, SFTP, FUSE, and the table catalogs.
The blob store started from [Facebook's Haystack](http://www.usenix.org/event/osdi10/tech/full_papers/Beaver.pdf), erasure coding takes ideas from [f4](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-muralidhar.pdf), and the whole has a lot in common with [Tectonic](https://www.usenix.org/system/files/fast21-pan.pdf) and [Colossus](https://cloud.google.com/blog/products/storage-data-transfer/a-peek-behind-colossus-googles-file-system). How file ids are assigned, written, and looked up, and why a master that tracks volumes scales, is in [Blob Store Architecture][BlobStoreArchitecture]; the services are in [Components][Components] and the [white paper][WhitePaper].
[Back to TOC](#table-of-contents)
## Example: Using Seaweed Blob Store ##
By default, the master node runs on port 9333, and the volume nodes run on port 8080.
Let's start one master node, and two volume nodes on port 8080 and 8081. Ideally, they should be started from different machines. We'll use localhost as an example.
SeaweedFS uses HTTP REST operations to read, write, and delete. The responses are in JSON or JSONP format.
### Start Master Server ###
```
> ./weed master
```
### Start Volume Servers ###
```
> weed volume -dir="/tmp/data1" -max=5 -master="localhost:9333" -port=8080 &
> weed volume -dir="/tmp/data2" -max=10 -master="localhost:9333" -port=8081 &
```
### Write A Blob ###
A blob, also referred as a needle, a chunk, or mistakenly as a file, is just a byte array. It can have attributes, such as name, mime type, create or update time, etc. But basically it is just a byte array of a relatively small size, such as 2 MB ~ 64 MB. The size is not fixed.
To upload a blob: first, send a HTTP POST, PUT, or GET request to `/dir/assign` to get an `fid` and a volume server URL:
```
> curl http://localhost:9333/dir/assign
{"count":1,"fid":"3,01637037d6","url":"127.0.0.1:8080","publicUrl":"localhost:8080"}
```
Second, to store the blob content, send a HTTP multi-part POST request to `url + '/' + fid` from the response:
```
> curl -F file=@/home/chris/myphoto.jpg http://127.0.0.1:8080/3,01637037d6
{"name":"myphoto.jpg","size":43234,"eTag":"1cc0118e"}
```
To update, send another POST request with updated blob content.
For deletion, send an HTTP DELETE request to the same `url + '/' + fid` URL:
```
> curl -X DELETE http://127.0.0.1:8080/3,01637037d6
```
### Save Blob Id ###
Now, you can save the `fid`, 3,01637037d6 in this case, to a database field.
The number 3 at the start represents a volume id. After the comma, it's one file key, 01, and a file cookie, 637037d6.
The volume id is an unsigned 32-bit integer. The file key is an unsigned 64-bit integer. The file cookie is an unsigned 32-bit integer, used to prevent URL guessing.
The file key and file cookie are both coded in hex. You can store the <volume id, file key, file cookie> tuple in your own format, or simply store the `fid` as a string.
If stored as a string, in theory, you would need 8+1+16+8=33 bytes. A char(33) would be enough, if not more than enough, since most uses will not need 2^32 volumes.
If space is really a concern, you can store the file id in the binary format. You would need one 4-byte integer for volume id, 8-byte long number for file key, and a 4-byte integer for the file cookie. So 16 bytes are more than enough.
### Read a Blob ###
Here is an example of how to render the URL.
First look up the volume server's URLs by the file's volumeId:
```
> curl http://localhost:9333/dir/lookup?volumeId=3
{"volumeId":"3","locations":[{"publicUrl":"localhost:8080","url":"localhost:8080"}]}
```
Since (usually) there are not too many volume servers, and volumes don't move often, you can cache the results most of the time. Depending on the replication type, one volume can have multiple replica locations. Just randomly pick one location to read.
Now you can take the public URL, render the URL or directly read from the volume server via URL:
```
http://localhost:8080/3,01637037d6.jpg
```
Notice we add a file extension ".jpg" here. It's optional and just one way for the client to specify the file content type.
If you want a nicer URL, you can use one of these alternative URL formats:
```
http://localhost:8080/3/01637037d6/my_preferred_name.jpg
http://localhost:8080/3/01637037d6.jpg
http://localhost:8080/3,01637037d6.jpg
http://localhost:8080/3/01637037d6
http://localhost:8080/3,01637037d6
```
If you want to get a scaled version of an image, you can add some params:
```
http://localhost:8080/3/01637037d6.jpg?height=200&width=200
http://localhost:8080/3/01637037d6.jpg?height=200&width=200&mode=fit
http://localhost:8080/3/01637037d6.jpg?height=200&width=200&mode=fill
```
### Rack-Aware and Data Center-Aware Replication ###
SeaweedFS applies the replication strategy at a volume level. So, when you are getting a blob id, you can specify the replication strategy. For example:
```
curl http://localhost:9333/dir/assign?replication=001
```
The replication parameter options are:
```
000: no replication
001: replicate once on the same rack
010: replicate once on a different rack, but same data center
100: replicate once on a different data center
200: replicate twice on two different data center
110: replicate once on a different rack, and once on a different data center
```
More details about replication can be found [on the wiki][Replication].
[Replication]: https://github.com/seaweedfs/seaweedfs/wiki/Replication
You can also set the default replication strategy when starting the master server.
### Allocate Blob Key on Specific Data Center ###
Volume servers can be started with a specific data center name:
```
weed volume -dir=/tmp/1 -port=8080 -dataCenter=dc1
weed volume -dir=/tmp/2 -port=8081 -dataCenter=dc2
```
When requesting a blob key, an optional "dataCenter" parameter can limit the assigned volume to the specific data center. For example, this specifies that the assigned volume should be limited to 'dc1':
```
http://localhost:9333/dir/assign?dataCenter=dc1
```
### Other Features ###
* [No Single Point of Failure][feat-1]
* [Insert with your own keys][feat-2]
* [Chunking large files][feat-3]
* [Collection as a Simple Name Space][feat-4]
[feat-1]: https://github.com/seaweedfs/seaweedfs/wiki/Failover-Master-Server
[feat-2]: https://github.com/seaweedfs/seaweedfs/wiki/Optimization#insert-with-your-own-keys
[feat-3]: https://github.com/seaweedfs/seaweedfs/wiki/Optimization#upload-large-files
[feat-4]: https://github.com/seaweedfs/seaweedfs/wiki/Optimization#collection-as-a-simple-name-space
[Back to TOC](#table-of-contents)
## Blob Store Architecture ##
Usually distributed file systems split each file into chunks. A central server keeps a mapping of filenames to chunks, and also which chunks each chunk server has.
The main drawback is that the central server can't handle many small files efficiently, and since all read requests need to go through the central master, so it might not scale well for many concurrent users.
Instead of managing chunks, SeaweedFS manages data volumes in the master server. Each data volume is 32GB in size, and can hold a lot of blobs. And each storage node can have many data volumes. So the master node only needs to store the metadata about the volumes, which is a fairly small amount of data and is generally stable.
The actual blob metadata, which are the blob volume, offset, and size, is stored in each volume on volume servers. Since each volume server only manages metadata of blobs on its own disk, with only 16 bytes for each blob, all access can read the metadata just from memory and only needs one disk operation to actually read file data.
For comparison, consider that an xfs inode structure in Linux is 536 bytes.
### Master Server and Volume Server ###
The architecture is fairly simple. The actual data is stored in volumes on storage nodes. One volume server can have multiple volumes, and can both support read and write access with basic authentication.
All volumes are managed by a master server. The master server contains the volume id to volume server mapping. This is fairly static information, and can be easily cached.
On each write request, the master server also generates a file key, which is a growing 64-bit unsigned integer. Since write requests are not generally as frequent as read requests, one master server should be able to handle the concurrency well.
### Write and Read files ###
When a client sends a write request, the master server returns (volume id, file key, file cookie, volume node URL) for the blob. The client then contacts the volume node and POSTs the blob content.
When a client needs to read a blob based on (volume id, file key, file cookie), it asks the master server by the volume id for the (volume node URL, volume node public URL), or retrieves this from a cache. Then the client can GET the content, or just render the URL on web pages and let browsers fetch the content.
### Saving memory ###
All blob metadata stored on a volume server is readable from memory without disk access. Each file takes just a 16-byte map entry of <64bit key, 32bit offset, 32bit size>. Of course, each map entry has its own space cost for the map. But usually the disk space runs out before the memory does.
### Tiered Storage to the cloud ###
The local volume servers are much faster, while cloud storages have elastic capacity and are actually more cost-efficient if not accessed often (usually free to upload, but relatively costly to access). With the append-only structure and O(1) access time, SeaweedFS can take advantage of both local and cloud storage by offloading the warm data to the cloud.
Usually hot data are fresh and warm data are old. SeaweedFS puts the newly created volumes on local servers, and optionally upload the older volumes on the cloud. If the older data are accessed less often, this literally gives you unlimited capacity with limited local servers, and still fast for new data.
With the O(1) access time, the network latency cost is kept at minimum.
If the hot/warm data is split as 20/80, with 20 servers, you can achieve storage capacity of 100 servers. That's a cost saving of 80%! Or you can repurpose the 80 servers to store new data also, and get 5X storage throughput.
[Back to TOC](#table-of-contents)
## SeaweedFS Filer ##
Built on top of the blob store, SeaweedFS Filer adds directory structure to create a file system. The directory structure is an interface that is implemented in many key-value stores or databases.
The content of a file is mapped to one or many blobs, distributed to multiple volumes on multiple volume servers.
## Compared to Other File Systems ##
# Compared to Other Systems #
Most other distributed file systems seem more complicated than necessary.
@@ -438,9 +270,7 @@ SeaweedFS is meant to be fast and simple, in both setup and operation. If you do
SeaweedFS is constantly moving forward. Same with other systems. These comparisons can be outdated quickly. Please help to keep them updated.
[Back to TOC](#table-of-contents)
### Compared to HDFS ###
## Compared to HDFS ##
HDFS uses the chunk approach for each file, and is ideal for storing large files.
@@ -448,9 +278,7 @@ SeaweedFS is ideal for serving relatively smaller files quickly and concurrently
SeaweedFS can also store extra large files by splitting them into manageable data chunks, and store the file ids of the data chunks into a meta chunk. This is managed by "weed upload/download" tool, and the weed master or volume servers are agnostic about it.
[Back to TOC](#table-of-contents)
### Compared to GlusterFS, Ceph ###
## Compared to GlusterFS, Ceph ##
The architectures are mostly the same. SeaweedFS aims to store and read files fast, with a simple and flat architecture. The main differences are
@@ -469,25 +297,15 @@ The architectures are mostly the same. SeaweedFS aims to store and read files fa
| MinIO | separate meta file per drive for each file | | | Yes | No |
| RustFS | separate meta file per drive for each file | | | Yes | No |
[Back to TOC](#table-of-contents)
GlusterFS stores files, both directories and content, in configurable volumes called "bricks". It hashes the path and filename into ids, and assigned to virtual volumes, and then mapped to "bricks".
### Compared to GlusterFS ###
GlusterFS stores files, both directories and content, in configurable volumes called "bricks".
GlusterFS hashes the path and filename into ids, and assigned to virtual volumes, and then mapped to "bricks".
[Back to TOC](#table-of-contents)
### Compared to MooseFS ###
## Compared to MooseFS ##
MooseFS chooses to neglect small file issue. From moosefs 3.0 manual, "even a small file will occupy 64KiB plus additionally 4KiB of checksums and 1KiB for the header", because it "was initially designed for keeping large amounts (like several thousands) of very big files"
MooseFS Master Server keeps all meta data in memory. Same issue as HDFS namenode.
MooseFS Master Server keeps all meta data in memory. Same issue as HDFS namenode.
[Back to TOC](#table-of-contents)
### Compared to Ceph ###
## Compared to Ceph ##
Ceph can be setup similar to SeaweedFS as a key->blob store. It is much more complicated, with the need to support layers on top of it. [Here is a more detailed comparison](https://github.com/seaweedfs/seaweedfs/issues/120)
@@ -507,13 +325,11 @@ SeaweedFS Filer uses off-the-shelf stores, such as MySql, Postgres, Sqlite, Mong
| Volume | OSD | optimized for small files |
| Filer | Ceph FS | linearly scalable, Customizable, O(1) or O(logN) |
[Back to TOC](#table-of-contents)
### Compared to MinIO, RustFS ###
## Compared to MinIO, RustFS ##
Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs. RustFS is a MinIO reimplementation in Rust, Apache 2.0 licensed and still developed, keeping MinIO's storage model down to a byte-compatible on-disk format. So the points below apply to both.
MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.
MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.
The metadata are in simple files. Each file write incurs extra writes to the corresponding meta file, on every drive of the erasure set. Changing only tags or retention rewrites that meta file on all of them, so the write amplification does not shrink with object size.
@@ -530,113 +346,18 @@ There are specific requirements on storage layout, which makes it hard to scale
[Back to TOC](#table-of-contents)
## Dev Plan ##
# Benchmark #
* More tools and documentation, on how to manage and scale the system.
* Read and write stream data.
* Support structured data.
Unscientific single-machine numbers from a MacBook with an SSD. [`weed benchmark`][Benchmarks], 1 million 1KB files, concurrency 16:
This is a super exciting project! And we need helpers and [support](https://www.patreon.com/seaweedfs)!
| | Requests per second | p50 | p99 |
| --- | --- | --- | --- |
| Write | 15,708 | 0.8 ms | 2.6 ms |
| Random read | 47,019 | 0.3 ms | 0.7 ms |
[Back to TOC](#table-of-contents)
## Installation Guide ##
> Installation guide for users who are not familiar with golang
Step 1: install go on your machine and setup the environment by following the instructions at:
https://golang.org/doc/install
make sure to define your $GOPATH
Step 2: checkout this repo:
```bash
git clone https://github.com/seaweedfs/seaweedfs.git
```
Step 3: download, compile, and install the project by executing the following command
```bash
cd seaweedfs/weed && make install
```
Once this is done, you will find the executable "weed" in your `$GOPATH/bin` directory
For more installation options, including how to run with Docker, see the [Getting Started guide](https://github.com/seaweedfs/seaweedfs/wiki/Getting-Started).
[Back to TOC](#table-of-contents)
## Disk Related Topics ##
### Hard Drive Performance ###
When testing read performance on SeaweedFS, it basically becomes a performance test of your hard drive's random read speed. Hard drives usually get 100MB/s~200MB/s.
### Solid State Disk ###
To modify or delete small files, SSD must delete a whole block at a time, and move content in existing blocks to a new block. SSD is fast when brand new, but will get fragmented over time and you have to garbage collect, compacting blocks. SeaweedFS is friendly to SSD since it is append-only. Deletion and compaction are done on volume level in the background, not slowing reading and not causing fragmentation.
[Back to TOC](#table-of-contents)
## Benchmark ##
My Own Unscientific Single Machine Results on Mac Book with Solid State Disk, CPU: 1 Intel Core i7 2.6GHz.
Write 1 million 1KB file:
```
Concurrency Level: 16
Time taken for tests: 66.753 seconds
Completed requests: 1048576
Failed requests: 0
Total transferred: 1106789009 bytes
Requests per second: 15708.23 [#/sec]
Transfer rate: 16191.69 [Kbytes/sec]
Connection Times (ms)
min avg max std
Total: 0.3 1.0 84.3 0.9
Percentage of the requests served within a certain time (ms)
50% 0.8 ms
66% 1.0 ms
75% 1.1 ms
80% 1.2 ms
90% 1.4 ms
95% 1.7 ms
98% 2.1 ms
99% 2.6 ms
100% 84.3 ms
```
Randomly read 1 million files:
```
Concurrency Level: 16
Time taken for tests: 22.301 seconds
Completed requests: 1048576
Failed requests: 0
Total transferred: 1106812873 bytes
Requests per second: 47019.38 [#/sec]
Transfer rate: 48467.57 [Kbytes/sec]
Connection Times (ms)
min avg max std
Total: 0.0 0.3 54.1 0.2
Percentage of the requests served within a certain time (ms)
50% 0.3 ms
90% 0.4 ms
98% 0.6 ms
99% 0.7 ms
100% 54.1 ms
```
### Run WARP and launch a mixed benchmark. ###
`make benchmark` runs [warp][S3Benchmark] mixed S3 traffic against a local `weed server`:
```
make benchmark
warp: Benchmark data written to "warp-mixed-2025-12-05[194844]-kBpU.csv.zst"
Mixed operations.
Operation: DELETE, 10%, Concurrency: 20, Ran 42s.
* Throughput: 55.13 obj/s
@@ -653,17 +374,19 @@ Operation: STAT, 30%, Concurrency: 20, Ran 42s.
Cluster Total: 3302.88 MiB/s, 550.51 obj/s over 43s.
```
Read throughput is bounded by the random read speed of the disks, and grows with every volume server added. More numbers, including multi-node, FUSE, and Hadoop, are in [Benchmarks][Benchmarks], [S3 API Benchmark][S3Benchmark], [FIO benchmark][FIO], and [Independent Benchmarks][IndependentBenchmarks].
[Back to TOC](#table-of-contents)
## Enterprise ##
# Enterprise #
For enterprise users, please visit [seaweedfs.com](https://seaweedfs.com) for the SeaweedFS Enterprise Edition,
which has advanced features, including data recovery, self-healing storage,
For enterprise users, please visit [seaweedfs.com](https://seaweedfs.com) for the SeaweedFS Enterprise Edition,
which has advanced features, including data recovery, self-healing storage,
customizable erasure coding, EC vacuum and repair, etc.
[Back to TOC](#table-of-contents)
## License ##
# License #
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
@@ -681,5 +404,112 @@ The text of this page is available for modification and reuse under the terms of
[Back to TOC](#table-of-contents)
## Stargazers over time
[![Stargazers over time](https://starchart.cc/seaweedfs/seaweedfs.svg?variant=adaptive)](https://starchart.cc/seaweedfs/seaweedfs)
# Sponsors #
<h3 align="center"><a href="https://www.patreon.com/seaweedfs">Sponsor SeaweedFS via Patreon</a></h3>
SeaweedFS is an independent Apache-licensed open source project with its ongoing development made
possible entirely thanks to the support of these awesome [backers](https://github.com/seaweedfs/seaweedfs/blob/master/backers.md).
If you'd like to grow SeaweedFS even stronger, please consider joining our
<a href="https://www.patreon.com/seaweedfs">sponsors on Patreon</a>.
Your support will be really appreciated by me and other supporters!
<!--
<h4 align="center">Platinum</h4>
<p align="center">
<a href="" target="_blank">
<img src="https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/sponsor_nodion.png" width="200" alt="nodion">
</a>
</p>
-->
### Gold Sponsors
[![nodion](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/sponsor_nodion.png)](https://www.nodion.com)
[![piknik](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/piknik.png)](https://www.piknik.com)
[![keepsec](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/keepsec.png)](https://www.keepsec.ca)
[![zyner](https://raw.githubusercontent.com/seaweedfs/seaweedfs/master/note/sponsor_zyner.png)](https://zyner.org)
[Back to TOC](#table-of-contents)
## Star History
![Star History](note/star_history.svg)
[WeedMini]: https://github.com/seaweedfs/seaweedfs/wiki/Quick-Start-with-weed-mini
[DockerComposeS3]: https://github.com/seaweedfs/seaweedfs/wiki/Docker-Compose-for-S3
[HelmRecipes]: https://github.com/seaweedfs/seaweedfs/wiki/Helm-Chart-Recipes
[Operator]: https://github.com/seaweedfs/seaweedfs-operator
[SeaweedFsCsiDriver]: https://github.com/seaweedfs/seaweedfs-csi-driver
[GettingStarted]: https://github.com/seaweedfs/seaweedfs/wiki/Getting-Started
[ProductionSetup]: https://github.com/seaweedfs/seaweedfs/wiki/Production-Setup
[ErasureCoding]: https://github.com/seaweedfs/seaweedfs/wiki/Erasure-Coding-for-warm-storage
[RustVolume]: https://github.com/seaweedfs/seaweedfs/wiki/Rust-Volume-Server
[Benchmarks]: https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks
[S3Benchmark]: https://github.com/seaweedfs/seaweedfs/wiki/S3-API-Benchmark
[FIO]: https://github.com/seaweedfs/seaweedfs/wiki/FIO-benchmark
[IndependentBenchmarks]: https://github.com/seaweedfs/seaweedfs/wiki/Independent-Benchmarks
[FailoverMaster]: https://github.com/seaweedfs/seaweedfs/wiki/Failover-Master-Server
[WeedShell]: https://github.com/seaweedfs/seaweedfs/wiki/weed-shell
[Worker]: https://github.com/seaweedfs/seaweedfs/wiki/Worker
[FilerStores]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Stores
[Replication]: https://github.com/seaweedfs/seaweedfs/wiki/Replication
[TieredStorage]: https://github.com/seaweedfs/seaweedfs/wiki/Tiered-Storage
[CloudTier]: https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier
[SuperLargeFiles]: https://github.com/seaweedfs/seaweedfs/wiki/Data-Structure-for-Large-Files
[Versioning]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Object-Versioning
[ObjectLock]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Object-Lock-and-Retention
[Lifecycle]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle
[CORS]: https://github.com/seaweedfs/seaweedfs/wiki/S3-CORS
[ConditionalOps]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Conditional-Operations
[RenameObject]: https://github.com/seaweedfs/seaweedfs/wiki/S3-RenameObject
[BucketPolicies]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Bucket-Policies
[PolicyConditions]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Policy-Conditions
[PolicyVariables]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Policy-Variables
[OIDC]: https://github.com/seaweedfs/seaweedfs/wiki/OIDC-Integration
[K8sSA]: https://github.com/seaweedfs/seaweedfs/wiki/Kubernetes-ServiceAccount-Authentication
[SSE]: https://github.com/seaweedfs/seaweedfs/wiki/Server-Side-Encryption
[AuditLog]: https://github.com/seaweedfs/seaweedfs/wiki/S3-API-Audit-log
[BucketQuota]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Bucket-Quota
[RateLimiting]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Rate-Limiting
[AmazonS3API]: https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API
[S3vsMinio]: https://github.com/seaweedfs/seaweedfs/wiki/Supported-APIs-vs-Minio
[S3TableBucket]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Table-Bucket
[LanceCatalog]: https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS-Lance-Catalog
[IcebergCatalog]: https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS-Iceberg-Catalog
[SparkIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Spark-Iceberg-Integration
[TrinoIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Trino-Iceberg-Integration
[DremioIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Dremio-Iceberg-Integration
[DuckDBIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/DuckDB-Iceberg-Integration
[DorisIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/Doris-Iceberg-Integration
[RisingWaveIceberg]: https://github.com/seaweedfs/seaweedfs/wiki/RisingWave-Iceberg-Integration
[LanceDB]: https://github.com/seaweedfs/seaweedfs/wiki/LanceDB-Integration
[Lakekeeper]: https://github.com/seaweedfs/seaweedfs/wiki/Lakekeeper-Iceberg-Integration
[IcebergMaintenance]: https://github.com/seaweedfs/seaweedfs/wiki/Iceberg-Table-Maintenance
[LanceMaintenance]: https://github.com/seaweedfs/seaweedfs/wiki/Lance-Maintenance-Worker
[S3TablesSecurity]: https://github.com/seaweedfs/seaweedfs/wiki/S3-Tables-Security
[Hadoop]: https://github.com/seaweedfs/seaweedfs/wiki/Hadoop-Compatible-File-System
[CloudDrive]: https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Architecture
[CacheRemote]: https://github.com/seaweedfs/seaweedfs/wiki/Cache-Remote-Storage
[GatewayToRemoteObjectStore]: https://github.com/seaweedfs/seaweedfs/wiki/Gateway-to-Remote-Object-Storage
[ActiveActiveAsyncReplication]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Active-Active-cross-cluster-continuous-synchronization
[FilerStoreReplication]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Store-Replication
[AsyncBackup]: https://github.com/seaweedfs/seaweedfs/wiki/Async-Backup
[MetaBackup]: https://github.com/seaweedfs/seaweedfs/wiki/Async-Filer-Metadata-Backup
[CDC]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Change-Data-Capture
[Webhook]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Notification-Webhook
[Mount]: https://github.com/seaweedfs/seaweedfs/wiki/FUSE-Mount
[MountWindows]: https://github.com/seaweedfs/seaweedfs/wiki/Mount-on-Windows
[WebDAV]: https://github.com/seaweedfs/seaweedfs/wiki/WebDAV
[SFTP]: https://github.com/seaweedfs/seaweedfs/wiki/SFTP-Server
[TUS]: https://github.com/seaweedfs/seaweedfs/wiki/TUS-Resumable-Uploads
[FilerDataEncryption]: https://github.com/seaweedfs/seaweedfs/wiki/Filer-Data-Encryption
[FIPS]: https://github.com/seaweedfs/seaweedfs/wiki/Cryptography-and-FIPS-Compliance
[AdminUI]: https://github.com/seaweedfs/seaweedfs/wiki/Admin-UI
[Metrics]: https://github.com/seaweedfs/seaweedfs/wiki/System-Metrics
[VolumeServerTTL]: https://github.com/seaweedfs/seaweedfs/wiki/Store-file-with-a-Time-To-Live
[SeaweedUp]: https://github.com/seaweedfs/seaweedfs/wiki/Deployment-with-seaweed-up
[BlobStoreArchitecture]: https://github.com/seaweedfs/seaweedfs/wiki/Blob-Store-Architecture
[Components]: https://github.com/seaweedfs/seaweedfs/wiki/Components
[WhitePaper]: https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS_Architecture.pdf
+418
View File
@@ -0,0 +1,418 @@
# SeaweedFS as an Apache CloudStack Object Storage Provider
A CloudStack ObjectStore plugin that makes SeaweedFS a first-class object storage
backend inside Apache CloudStack, alongside the existing MinIO and Ceph RGW
providers. This is a collaboration with proIO (Swen), who builds private clouds on
CloudStack and wants SeaweedFS as a storage option.
## The request
> We can only add MinIO and Ceph as object storage [in CloudStack] today. I want
> to get SeaweedFS into this project... What we need is to build a provider which
> does the communication between Cloudstack and SeaweedFS.
This is **not** a SeaweedFS-side feature. The work lives in the Apache CloudStack
repo (Java): a new plugin under `plugins/storage/object/seaweedfs/` that implements
CloudStack's ObjectStore plugin framework and talks to SeaweedFS over its S3 and
IAM APIs. SeaweedFS itself needs no changes for the core to work — its S3 API
already covers every bucket operation CloudStack requires, and its IAM API covers
user/credential management.
## How the CloudStack ObjectStore framework works
CloudStack 4.18+ introduced an Object Storage framework. An admin registers an
object storage pool via `addObjectStoragePool` (URL + provider + credentials);
tenants then create and manage buckets on it through CloudStack APIs. CloudStack
manages pool and bucket lifecycle; the underlying provider handles the actual
object protocol.
A provider is a plugin module implementing three interfaces:
### 1. `ObjectStoreProvider` — registration
`MinIOObjectStoreProviderImpl` is the reference. It is a Spring `@Component` that:
- Returns a provider name (`"MinIO"`)
- Returns `DataStoreProviderType.OBJECT`
- In `configure()`, injects the lifecycle and driver implementations and calls
`storeMgr.registerDriver(name, driver)`
### 2. `ObjectStoreLifeCycle` — pool add/remove
`MinIOObjectStoreLifeCycleImpl.initialize()` reads the URL, name, and
`accesskey`/`secretkey` details from the `addObjectStoragePool` call, tests the
connection by listing buckets, and persists an `ObjectStoreVO` via
`ObjectStoreHelper`. The other methods (attachCluster/Host/Zone, maintain,
deleteDataStore) are no-ops for object storage.
### 3. `ObjectStoreDriver` — bucket + user operations
`ObjectStoreDriver` (in `engine/storage/.../object/ObjectStoreDriver.java`) extends
`DataStoreDriver` and defines the bucket/user contract. Every provider must
implement:
| Method | Purpose |
| --- | --- |
| `createBucket(Bucket, boolean objectLock)` | Create a bucket |
| `listBuckets(long storeId)` | List all buckets |
| `deleteBucket(BucketTO, long storeId)` | Delete a bucket |
| `createUser(long accountId, long storeId)` | Provision a user + credentials for a CloudStack account |
| `setBucketPolicy` / `getBucketPolicy` / `deleteBucketPolicy` | Bucket policy CRUD |
| `setBucketEncryption` / `deleteBucketEncryption` | SSE config |
| `setBucketVersioning` / `deleteBucketVersioning` | Versioning enable/suspend |
| `setBucketQuota(BucketTO, long storeId, long size)` | Per-bucket quota |
| `getAllBucketsUsage(long storeId)` | Usage map for billing/accounting |
| `getBucketAcl` / `setBucketAcl` | ACLs (MinIO/Ceph return null / no-op) |
`BaseObjectStoreDriverImpl` provides no-op defaults for the `DataStoreDriver`
methods (`createAsync`, `deleteAsync`, `copyAsync`, `canCopy`, `resize`,
`getTO`, `getStoreTO`), so object-store providers only implement the bucket/user
methods above.
## How the four existing providers differ (and where SeaweedFS lands)
CloudStack ships four object-store providers. Three are relevant; the simulator
is a test stub.
| Concern | MinIO | Ceph RGW | Cloudian HyperStore | SeaweedFS |
| --- | --- | --- | --- | --- |
| Bucket CRUD | `MinioClient` (S3) | `AmazonS3` (AWS SDK v1) | `AmazonS3` (AWS SDK v1) | `AmazonS3` (AWS SDK v1) |
| Bucket policy | `MinioClient` | `AmazonS3` | `AmazonS3` | `AmazonS3` |
| Versioning | `MinioClient` | `AmazonS3` | `AmazonS3` | `AmazonS3` |
| Encryption | `MinioClient` | not implemented | `AmazonS3` | `AmazonS3` |
| **User creation** | `MinioAdminClient` | `RgwAdmin` | **`AmazonIdentityManagement`** | **`AmazonIdentityManagement`** |
| **Per-bucket quota** | `MinioAdminClient` | `RgwAdmin` | **not supported** (throws) | **S3 extension** (`PUT /{bucket}?seaweedfs-quota`, SigV4, `s3:PutBucketQuota`) |
| **Usage reporting** | `MinioAdminClient` | `RgwAdmin` | Cloudian admin API | S3 `ListObjectsV2` (MVP); Prometheus / SOSAPI `capacity.xml` (recommended) |
**Cloudian HyperStore is the direct precedent.** It is an S3-compatible store
that, like SeaweedFS, manages users via the **standard AWS IAM API** using the
AWS IAM Java SDK (`com.amazonaws.services.identitymanagement`). Its driver
(`CloudianHyperStoreObjectStoreDriverImpl`) and util
(`CloudianHyperStoreUtil`) are the template this design follows almost line for
line. Cloudian even validates the quota limitation the same way this design
proposes for the MVP: `setBucketQuota` throws for any non-zero size and only
accepts `0` (no quota).
The SeaweedFS plugin is therefore a **simpler Cloudian** — same AWS S3 + IAM SDK
clients, same store-details keys (`s3Url`, `iamUrl`, `accesskey`, `secretkey`),
same IAM-user-with-restricted-policy pattern, but with no proprietary admin
client at all (Cloudian has its own `CloudianClient` for its admin API; SeaweedFS
needs only S3 + IAM). For quota, the plugin uses a narrow SeaweedFS S3 extension
(see below); for usage reporting, it falls back to S3 `ListObjectsV2` in the MVP
and recommends Prometheus or SOSAPI `capacity.xml` for production scale.
### Quota via the S3 `?seaweedfs-quota` extension
SeaweedFS supports bucket quota natively (server-side enforcement via a
read-only flag when usage exceeds the limit). Rather than exposing the broad
admin REST API (which would require a global bearer token and grant cluster-wide
admin access), the integration uses a **narrow, scoped S3 subresource**:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
**PUT request body** (JSON):
```json
{"quota_size": 100, "quota_unit": "GB", "quota_enabled": true}
```
**GET response body** (JSON):
```json
{"quota_size": 107374182400, "quota_unit": "B", "quota_enabled": true}
```
Note: GET always returns `quota_unit: "B"` and the absolute byte count, not
the original unit. A disabled-but-retained quota returns a positive
`quota_size` with `quota_enabled: false`.
Quota is stored on the bucket's filer entry (positive = enabled, negative =
disabled but retained, zero = no quota), matching the existing admin REST API
behavior. When quota is cleared, the bucket's read-only flag is also lifted.
**Authentication** uses the existing S3 SigV4 flow — no new global secret is
needed. The CloudStack service credential (the `accesskey`/`secretkey` on the
object store) is the admin credential used for all driver operations: bucket
CRUD, IAM user provisioning, and quota management. It must have broad S3 and
IAM permissions. The per-account IAM users created by `createUser` are the
ones with restricted permissions (full S3 access except bucket
creation/deletion). A future hardening could split quota management onto a
separate credential scoped to only `s3:PutBucketQuota`/`s3:GetBucketQuota`,
but the MVP uses the single admin credential for simplicity, matching how
the MinIO and Ceph providers work.
The plugin's `setBucketQuota` signs and sends the `PUT /{bucket}?seaweedfs-quota`
request using the AWS SDK v1 `AWSS3V4Signer` for SigV4 signing, then sends the
signed request via `java.net.http.HttpClient` (the AWS S3 SDK doesn't natively
support custom subresources, so we sign manually and send the request
ourselves). The `seaweedfs-quota` query parameter is included in the signed
canonical query string.
### Usage reporting
`getAllBucketsUsage` must return a `Map<String, Long>` of bucket name → size.
MinIO uses `MinioAdminClient.getDataUsageInfo`; Ceph uses
`RgwAdmin.listBucketInfo`. SeaweedFS has no admin rollup endpoint, so the MVP
plugin computes it by listing buckets and summing object sizes via S3
`ListObjectsV2` — expensive for large stores.
For production scale, SeaweedFS already exposes per-bucket size in:
- **Prometheus metrics** (`bucket_size_bytes` gauge, refreshed every minute)
- **SOSAPI `capacity.xml`** (reports capacity, available space, and usage
through the S3 endpoint)
Operators should consume one of those instead of S3 list-based aggregation for
large deployments. The MVP's list-based approach is correct but slow; flag it as
a known limitation.
## SeaweedFS API surface (what the plugin relies on)
SeaweedFS exposes two relevant APIs, both AWS-compatible:
### S3 API (`weed s3`)
Full S3-compatible surface. Confirmed against the SeaweedFS S3 wiki and code:
- `CreateBucket`, `HeadBucket`, `ListBuckets`, `DeleteBucket`
- `PutBucketPolicy`, `GetBucketPolicy`, `DeleteBucketPolicy`
- `PutBucketVersioning` (Enabled / Suspended), `GetBucketVersioning`
- `PutBucketEncryption`, `GetBucketEncryption`, `DeleteBucketEncryption`
- `PutBucketAcl`, `GetBucketAcl`
- `ListObjectsV2`, `HeadObject`, `GetObject`, `PutObject`, `DeleteObject`
- Bucket quota via extended attributes / `s3.bucket.quota` (enforced server-side,
surfaced as a read-only state when exceeded — see PR #10224)
### IAM API (`weed iam` / `iamapi`)
AWS IAM-compatible REST endpoints, implemented in `weed/iamapi/`. Confirmed by
the test suite which uses the **AWS IAM SDK** (`aws-sdk-go/service/iam`) against
the same handlers CloudStack would call:
- `CreateUser`, `DeleteUser`, `ListUsers`, `GetUser`
- `CreateAccessKey`, `DeleteAccessKey`, `ListAccessKeys`
- `PutUserPolicy`, `GetUserPolicy`, `DeleteUserPolicy`
- `AttachUserPolicy`, `ListAttachedUserPolicies`
This means the CloudStack plugin can manage SeaweedFS users with the **AWS IAM
Java SDK** (`com.amazonaws.services.identitymanagement.AmazonIdentityManagement`),
exactly the way the AWS IAM Go SDK is used in SeaweedFS's own tests. No proprietary
admin client is needed. **Cloudian HyperStore already does exactly this** in the
CloudStack tree — the SeaweedFS plugin follows the same pattern.
## Design
### Module layout
New CloudStack plugin module, mirroring `plugins/storage/object/cloudian/`
(the closest precedent — same AWS S3 + IAM SDK approach):
```
plugins/storage/object/seaweedfs/
pom.xml
src/main/java/org/apache/cloudstack/storage/datastore/
driver/SeaweedFSObjectStoreDriverImpl.java
lifecycle/SeaweedFSObjectStoreLifeCycleImpl.java
provider/SeaweedFSObjectStoreProviderImpl.java
util/SeaweedFSObjectStoreUtil.java
src/test/java/org/apache/cloudstack/storage/datastore/
driver/SeaweedFSObjectStoreDriverImplTest.java
provider/SeaweedFSObjectStoreProviderImplTest.java
src/main/resources/META-INF/cloudstack/storage-object-seaweedfs/
module.properties
spring-storage-object-seaweedfs-context.xml
```
### `SeaweedFSObjectStoreProviderImpl`
Direct copy of `MinIOObjectStoreProviderImpl` with `providerName = "SeaweedFS"`,
injecting the SeaweedFS lifecycle and driver. Registers via
`storeMgr.registerDriver`.
### `SeaweedFSObjectStoreLifeCycleImpl`
Copy of `MinIOObjectStoreLifeCycleImpl`. `initialize()` reads `url`, `name`,
`accesskey`, `secretkey` from the `addObjectStoragePool` details map, tests the
connection by calling `AmazonS3.listBuckets()` against the SeaweedFS S3 endpoint,
and persists the `ObjectStoreVO`. No proprietary client needed — the AWS S3 SDK
is enough for the health check.
### `SeaweedFSObjectStoreDriverImpl`
The substantive class. Uses two AWS SDK v1 clients (same dependency Ceph already
pulls in, so no new CloudStack dependency):
- `AmazonS3` for bucket operations (path-style, endpoint-pinned, `us-east-1`
region placeholder — same as Ceph's `getS3Client`)
- `AmazonIdentityManagement` for user/credential operations, pointed at the
SeaweedFS IAM endpoint
#### Bucket operations — straightforward S3
| Interface method | Implementation |
| --- | --- |
| `createBucket` | `s3.createBucket(name)`; reject if `doesBucketExistV2`; persist access/secret key + URL on `BucketVO` (same as Ceph) |
| `listBuckets` | `s3.listBuckets()` → wrap as `BucketObject` (same as Ceph) |
| `deleteBucket` | `s3.deleteBucket(name)` (same as Ceph) |
| `setBucketPolicy` | `s3.setBucketPolicy(...)` with the same public/private JSON the MinIO/Ceph drivers build |
| `getBucketPolicy` / `deleteBucketPolicy` | `s3.getBucketPolicy` / `s3.deleteBucketPolicy` |
| `setBucketVersioning` | `s3.setBucketVersioningConfiguration(Enabled)` |
| `deleteBucketVersioning` | `s3.setBucketVersioningConfiguration(Suspended)` |
| `setBucketEncryption` | `s3.setBucketEncryptionConfiguration(SSE-S3 rule)` |
| `deleteBucketEncryption` | `s3.deleteBucketEncryptionConfiguration` |
| `getBucketAcl` / `setBucketAcl` | no-op / null (same as MinIO and Ceph) |
#### User creation — the key difference
MinIO calls `MinioAdminClient.addUser`; Ceph calls `RgwAdmin.createUser`. SeaweedFS
exposes the standard AWS IAM API, so the plugin calls:
```java
AmazonIdentityManagement iam = getIamClient(storeId);
String userName = "acs-" + account.getUuid();
// CreateUser (idempotent — check GetUser first, like Ceph does)
iam.createUser(new CreateUserRequest(userName));
// CreateAccessKey → returns the access key + secret key to persist
CreateAccessKeyResult result = iam.createAccessKey(
new CreateAccessKeyRequest().withUserName(userName));
AccessKey key = result.getAccessKey();
// Persist per-account, same pattern as Ceph's CEPH_ACCESS_KEY/CEPH_SECRET_KEY
details.put(SEAWEEDFS_ACCESS_KEY, key.getAccessKeyId());
details.put(SEAWEEDFS_SECRET_KEY, key.getSecretAccessKey());
_accountDetailsDao.persist(accountId, details);
```
This is the cleanest mapping of the three providers: no proprietary admin client,
just the AWS IAM SDK that CloudStack already has access to. The IAM endpoint URL
is provided as `iamUrl` in the store details. If `iamUrl` is omitted, the driver
defaults it to `s3Url` — SeaweedFS registers its embedded IAM API at `POST /` on
the same S3 endpoint (`UnifiedPostHandler` in `s3api_server.go`), so the IAM
endpoint is the same as the S3 endpoint unless the deployment runs a separate
`weed iam` server.
#### Bucket quota — S3 `?seaweedfs-quota` extension
This is the one genuine gap. MinIO and Ceph both have an admin API to set a
per-bucket quota that the backend enforces. SeaweedFS enforces bucket quota
server-side, but the configuration path was **not exposed over a standard S3 or
IAM API** — it was only set via the admin REST API or shell commands.
The integration adds a **narrow S3 subresource** to SeaweedFS:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
This is implemented in SeaweedFS PR #11279. It uses SigV4 authentication and
dedicated IAM permissions, so the CloudStack service credential can be scoped
to quota management only — no global admin token, no cluster-wide admin access.
The enforcement already exists (PR #10224); this PR only adds the HTTP
configuration surface.
An earlier approach (PR #11278, closed) added bearer-token auth to the broad
admin REST API. After review, that was unnecessary for this integration —
static S3 config plus standard S3 APIs plus one scoped quota mutation API is
sufficient and far safer.
> **Note on AWS tools compatibility.** `?seaweedfs-quota` is a SeaweedFS-specific
> S3 subresource, not part of the AWS S3 API. Standard AWS tools (`aws s3api`,
> `s3cmd`, `rclone`) cannot call it directly. This is the same limitation MinIO
> and Ceph have — MinIO quota lives behind a separate admin API (`mc admin
> bucket quota`), and Ceph quota lives behind the Admin Ops API
> (`radosgw-admin quota set`). Neither is callable via `aws s3api` either.
> SeaweedFS's approach is the closest to standard S3 because it uses the same
> endpoint and same SigV4 credentials, just with a custom query parameter.
> Interactive quota management remains available via `weed shell`; the S3
> extension exists for programmatic integration (CloudStack) where the
> integrator can sign SigV4 requests but cannot run shell commands.
#### Usage reporting
`getAllBucketsUsage` must return a `Map<String, Long>` of bucket name → size.
MinIO uses `MinioAdminClient.getDataUsageInfo`; Ceph uses
`RgwAdmin.listBucketInfo`. SeaweedFS has no admin rollup endpoint, so the MVP
plugin computes it by listing buckets and summing object sizes via S3
`ListObjectsV2` — expensive for large stores. Better options exist in
SeaweedFS already:
- **Prometheus metrics** (`bucket_size_bytes` gauge, refreshed every minute)
- **SOSAPI `capacity.xml`** (reports capacity, available space, and usage
through the S3 endpoint — note: the current "return zero on backend error"
behavior should be validated before using it for billing)
For the MVP, `listBuckets` + per-bucket size via the S3 API is correct but slow;
flag it as a known limitation. Operators should consume Prometheus or SOSAPI
for production-scale usage reporting.
### Spring wiring
`spring-storage-object-seaweedfs-context.xml` registers the provider bean,
identical to the MinIO one. `module.properties` sets
`name=storage-object-seaweedfs`, `parent=storage`.
### `pom.xml`
Depends on `aws-java-sdk-s3` and `aws-java-sdk-iam` — both already in the
CloudStack dependency tree (Ceph uses the S3 SDK; the IAM SDK is the standard AWS
bundle). No new third-party dependency, unlike MinIO which pulls in the MinIO
Java client.
## What changes on the SeaweedFS side
**One narrow S3 extension is required for quota management.** SeaweedFS PR #11279
adds the `?seaweedfs-quota` S3 subresource:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
This is authenticated via standard S3 SigV4 and authorized via dedicated IAM
permissions, so no global admin token is needed. The enforcement already exists
(PR #10224); this PR only adds the HTTP configuration surface.
One follow-up improvement on the SeaweedFS side would close the usage reporting
gap:
1. **Validate SOSAPI `capacity.xml` usage calculation** — the current "return
zero on backend error" behavior should be validated before using it for
billing. If reliable, CloudStack can consume it directly instead of
list-based aggregation.
## Open questions for proIO / Swen
1. **IAM endpoint path.** ~~Where does `weed iam` listen relative to the S3
endpoint in a typical proIO deployment?~~ **Resolved.** SeaweedFS registers
its embedded IAM API at `POST /` on the same S3 endpoint
(`UnifiedPostHandler`), so the driver defaults `iamUrl` to `s3Url`. A
separate `iamUrl` is only needed if the deployment runs a standalone
`weed iam` server on a different host/port.
2. **Quota requirements.** Do proIO's customers need server-enforced per-bucket
quotas, or is CloudStack-side accounting sufficient for the first release?
The `?seaweedfs-quota` S3 extension (PR #11279) provides server-enforced
quotas via a scoped credential; this is the recommended path.
3. **Object Lock.** `createBucket` takes an `objectLock` boolean. MinIO supports
it; Ceph ignores it. SeaweedFS has Object Lock support. Should the plugin pass
it through?
4. **Contribution model.** Does proIO want to submit the PR to
`apache/cloudstack` themselves (with SeaweedFS maintainers as reviewers), or
the reverse? Apache CloudStack requires an ICLA for non-trivial contributions.
## Files
All in the `apache/cloudstack` repo (new module):
| File | Purpose |
| --- | --- |
| `plugins/storage/object/seaweedfs/pom.xml` | Maven module |
| `.../datastore/util/SeaweedFSObjectStoreUtil.java` | S3 + IAM client builders, constants, URL validators |
| `.../datastore/provider/SeaweedFSObjectStoreProviderImpl.java` | Spring provider registration |
| `.../datastore/lifecycle/SeaweedFSObjectStoreLifeCycleImpl.java` | Pool add/health-check |
| `.../datastore/driver/SeaweedFSObjectStoreDriverImpl.java` | Bucket + user ops via S3 + IAM SDK |
| `.../resources/META-INF/cloudstack/storage-object-seaweedfs/module.properties` | Module name |
| `.../resources/META-INF/cloudstack/storage-object-seaweedfs/spring-storage-object-seaweedfs-context.xml` | Spring bean |
| `plugins/pom.xml` | Register `storage/object/seaweedfs` module |
No files in `seaweedfs/seaweedfs` for the MVP.
### SeaweedFS-side changes (PR #11279)
| File | Purpose |
| --- | --- |
| `weed/s3api/s3_constants/s3_action_strings.go` | Add `S3_ACTION_PUT_BUCKET_QUOTA` and `S3_ACTION_GET_BUCKET_QUOTA` |
| `weed/s3api/s3_constants/s3_actions.go` | Add coarse-grained `ACTION_PUT_BUCKET_QUOTA` and `ACTION_GET_BUCKET_QUOTA` |
| `weed/s3api/s3_action_resolver.go` | Map `seaweedfs-quota` query param to fine-grained s3: actions |
| `weed/s3api/s3api_bucket_quota_handlers.go` | New — `PutBucketQuotaHandler` and `GetBucketQuotaHandler` |
| `weed/s3api/s3api_bucket_quota_handlers_test.go` | New — tests for unit conversion, validation, and error paths |
| `weed/s3api/s3api_server.go` | Register the two routes in the bucket subrouter |
+43
View File
@@ -31,6 +31,49 @@ docker compose -f seaweedfs-dev-compose.yml -p seaweedfs up
```
## Verify an image signature
Every image CI pushes to `chrislusf/seaweedfs` and `ghcr.io/chrislusf/seaweedfs` is signed with [cosign](https://docs.sigstore.dev/cosign/verifying/verify/), keyless, by the GitHub Actions workflow that built it, so there is no key to fetch or pin. The signature is attached to the image digest and covers the multi-arch index and each platform image in it; `latest` is the release image under another tag and verifies the same way. Images published before September 2026 predate signing.
```bash
cosign verify \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
--certificate-identity-regexp '^https://github.com/seaweedfs/seaweedfs/\.github/workflows/container_release_unified\.yml@' \
chrislusf/seaweedfs:latest
```
cosign prints the digest it verified. Deploy by that digest, or let an admission controller resolve the tag, so what runs is what was checked.
The identity is `https://github.com/seaweedfs/seaweedfs/.github/workflows/<workflow>@<ref>`. The ref is `refs/tags/<version>` for a release and `refs/heads/master` when a variant was republished by hand. The workflow is `container_release_unified.yml` for the release images, `container_dev.yml` for `dev`, `container_latest.yml` for a `latest` rebuilt by hand, `container_release_foundationdb.yml` for the `_large_disk_foundationdb` release image, and `container_foundationdb_version.yml` or `container_rocksdb_version.yml` for the per-version builds. The regexp above accepts release images only; `container_[a-z_]+\.yml@` accepts everything this repository publishes, `dev` included.
The same check as a Kyverno policy, release images only:
```yaml
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: verify-seaweedfs-images
spec:
validationFailureAction: Enforce
webhookTimeoutSeconds: 30
rules:
- name: signed-by-the-release-workflow
match:
any:
- resources:
kinds:
- Pod
verifyImages:
- imageReferences:
- "docker.io/chrislusf/seaweedfs:*"
- "ghcr.io/chrislusf/seaweedfs:*"
attestors:
- entries:
- keyless:
issuer: https://token.actions.githubusercontent.com
subject: https://github.com/seaweedfs/seaweedfs/.github/workflows/container_release_unified.yml@refs/tags/*
```
## Local Development
```bash
+63 -66
View File
@@ -1,11 +1,11 @@
module github.com/seaweedfs/seaweedfs
go 1.26
go 1.26.0
require (
cloud.google.com/go v0.123.0 // indirect
cloud.google.com/go/pubsub v1.51.1
cloud.google.com/go/storage v1.64.0
cloud.google.com/go/storage v1.67.0
github.com/Shopify/sarama v1.38.1
github.com/aws/aws-sdk-go v1.55.8
github.com/beorn7/perks v1.0.1 // indirect
@@ -25,7 +25,7 @@ require (
github.com/facebookgo/subset v0.0.0-20200203212716-c811ad88dec4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-redsync/redsync/v4 v4.17.0
github.com/go-sql-driver/mysql v1.10.0
github.com/go-sql-driver/mysql v1.10.1
github.com/go-zookeeper/zk v1.0.4 // indirect
github.com/golang/protobuf v1.5.4
github.com/golang/snappy v1.0.0
@@ -44,8 +44,8 @@ require (
github.com/jmespath/go-jmespath v0.4.0 // indirect
github.com/json-iterator/go v1.1.12
github.com/karlseguin/ccache/v2 v2.0.8
github.com/klauspost/compress v1.19.1
github.com/klauspost/reedsolomon v1.14.1
github.com/klauspost/compress v1.19.2
github.com/klauspost/reedsolomon v1.14.2
github.com/kurin/blazer v0.5.3
github.com/linxGnu/grocksdb v1.10.8
github.com/mailru/easyjson v0.9.2 // indirect
@@ -59,13 +59,13 @@ require (
github.com/posener/complete v1.2.3
github.com/pquerna/cachecontrol v0.2.0
github.com/prometheus/client_golang v1.24.1
github.com/prometheus/client_model v0.6.2
github.com/prometheus/common v0.70.1 // indirect
github.com/prometheus/procfs v0.21.1
github.com/prometheus/client_model v0.6.3
github.com/prometheus/common v0.70.1
github.com/prometheus/procfs v0.22.0
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/seaweedfs/goexif v2.0.0+incompatible
github.com/seaweedfs/raft v1.2.0
github.com/seaweedfs/raft v1.2.1
github.com/sirupsen/logrus v1.9.4 // indirect
github.com/spf13/afero v1.15.0 // indirect
github.com/spf13/cast v1.10.0 // indirect
@@ -90,16 +90,16 @@ require (
gocloud.dev v0.46.0
gocloud.dev/pubsub/natspubsub v0.46.0
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.55.0
golang.org/x/crypto v0.56.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.44.0
golang.org/x/image v0.46.0
golang.org/x/net v0.58.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/sys v0.48.0
golang.org/x/text v0.42.0 // indirect
golang.org/x/tools v0.49.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.294.0
google.golang.org/api v0.297.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/protobuf v1.36.12
@@ -118,19 +118,19 @@ require (
github.com/ThreeDotsLabs/watermill v1.5.2
github.com/a-h/templ v0.3.1020
github.com/apache/cassandra-gocql-driver/v2 v2.1.2
github.com/apache/iceberg-go v0.6.0
github.com/apache/iceberg-go v0.6.1-0.20260817192109-c2105090c9e2
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2/config v1.32.33
github.com/aws/aws-sdk-go-v2/credentials v1.19.34
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2
github.com/aws/aws-sdk-go-v2 v1.47.0
github.com/aws/aws-sdk-go-v2/config v1.32.35
github.com/aws/aws-sdk-go-v2/credentials v1.20.4
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
github.com/fluent/fluent-logger-golang v1.10.1
github.com/getsentry/sentry-go v0.48.0
github.com/go-ldap/ldap/v3 v3.4.13
github.com/getsentry/sentry-go v0.49.0
github.com/go-ldap/ldap/v3 v3.4.14
github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/flatbuffers/go v0.0.0-20230108230133-3b8644d32c50
github.com/hashicorp/golang-lru/v2 v2.0.7
@@ -144,22 +144,23 @@ require (
github.com/parquet-go/parquet-go v0.32.0
github.com/pkg/sftp v1.13.11
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rclone/rclone v1.75.0
github.com/rclone/rclone v1.75.1
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/redis/go-redis/v9 v9.22.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.7
github.com/tarantool/go-option v1.1.0
github.com/tarantool/go-tarantool/v3 v3.0.1
github.com/testcontainers/testcontainers-go v0.43.0
github.com/testcontainers/testcontainers-go v0.44.0
github.com/tikv/client-go/v2 v2.0.7
github.com/twmb/avro v1.8.0
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
golang.org/x/sync v0.23.0
golang.org/x/tools/godoc v0.1.0-deprecated
google.golang.org/grpc/security/advancedtls v1.0.0
)
@@ -177,16 +178,18 @@ require (
github.com/Azure/go-ansiterm v0.0.0-20250102033503-faa5f7b0171c // indirect
github.com/FilenCloudDienste/filen-sdk-go v0.0.39 // indirect
github.com/ProtonMail/gopenpgp/v3 v3.4.1 // indirect
github.com/RoaringBitmap/roaring/v2 v2.24.0 // indirect
github.com/a1ex3/zstd-seekable-format-go/pkg v0.10.0 // indirect
github.com/adrg/xdg v0.5.3 // indirect
github.com/anchore/go-lzo v0.1.1 // indirect
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
github.com/bits-and-blooms/bitset v1.24.4 // indirect
github.com/blevesearch/snowballstem v0.9.0 // indirect
github.com/boombuler/barcode v1.1.0 // indirect
github.com/buger/jsonparser v1.2.0 // indirect
@@ -206,11 +209,12 @@ require (
github.com/dave/dst v0.27.2 // indirect
github.com/diskfs/go-diskfs v1.9.4 // indirect
github.com/distribution/reference v0.6.0 // indirect
github.com/docker/go-connections v0.6.0 // indirect
github.com/docker/go-connections v0.7.0 // indirect
github.com/docker/go-units v0.5.0 // indirect
github.com/dromara/dongle v1.0.1 // indirect
github.com/geoarrow/geoarrow-go v0.0.0-20260403143023-f54751c3e3a1 // indirect
github.com/gin-gonic/gin v1.11.0 // indirect
github.com/go-asn1-ber/asn1-ber v1.5.8-0.20250403174932-29230038a667 // indirect
github.com/go-asn1-ber/asn1-ber v1.5.8 // indirect
github.com/go-git/go-billy/v5 v5.9.0 // indirect
github.com/goccy/go-yaml v1.18.0 // indirect
github.com/golang/geo v0.0.0-20210211234256-740aa86cb551 // indirect
@@ -239,13 +243,14 @@ require (
github.com/magiconair/properties v1.8.10 // indirect
github.com/moby/docker-image-spec v1.3.1 // indirect
github.com/moby/go-archive v0.3.0 // indirect
github.com/moby/moby/api v1.54.2 // indirect
github.com/moby/moby/client v0.4.0 // indirect
github.com/moby/moby/api v1.55.0 // indirect
github.com/moby/moby/client v0.5.0 // indirect
github.com/moby/patternmatcher v0.6.1 // indirect
github.com/moby/sys/sequential v0.7.0 // indirect
github.com/moby/sys/user v0.4.1 // indirect
github.com/moby/sys/userns v0.1.0 // indirect
github.com/moby/term v0.5.2 // indirect
github.com/mschoch/smat v0.2.0 // indirect
github.com/oklog/ulid/v2 v2.1.1 // indirect
github.com/opencontainers/go-digest v1.0.0 // indirect
github.com/opencontainers/image-spec v1.1.1 // indirect
@@ -257,8 +262,8 @@ require (
github.com/pquerna/otp v1.5.0 // indirect
github.com/pterm/pterm v0.12.83 // indirect
github.com/quic-go/qpack v0.6.0 // indirect
github.com/rclone/Proton-API-Bridge v1.0.4 // indirect
github.com/rclone/go-proton-api v1.0.3 // indirect
github.com/rclone/Proton-API-Bridge v1.0.5 // indirect
github.com/rclone/go-proton-api v1.0.4 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd // indirect
github.com/ryanuber/go-glob v1.0.0 // indirect
@@ -269,7 +274,6 @@ require (
github.com/substrait-io/substrait v0.87.0 // indirect
github.com/substrait-io/substrait-go/v8 v8.1.1 // indirect
github.com/substrait-io/substrait-protobuf/go v0.85.0 // indirect
github.com/twmb/avro v1.7.2 // indirect
github.com/twpayne/go-geom v1.6.1 // indirect
github.com/twpayne/go-kml/v3 v3.2.1 // indirect
github.com/tyler-smith/go-bip39 v1.1.0 // indirect
@@ -279,14 +283,14 @@ require (
github.com/xeipuuv/gojsonreference v0.0.0-20180127040603-bd5ef7bd5415 // indirect
github.com/xo/terminfo v0.0.0-20220910002029-abceb7e1c41e // indirect
github.com/zeebo/xxh3 v1.1.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.43.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.42.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.uber.org/mock v0.5.2 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/mod v0.38.0 // indirect
golang.org/x/mod v0.41.0 // indirect
gonum.org/v1/gonum v0.17.0 // indirect
)
@@ -311,7 +315,7 @@ require (
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/resourcemapping v0.57.0 // indirect
github.com/IBM/go-sdk-core/v5 v5.23.1 // indirect
github.com/Max-Sum/base32768 v0.0.0-20230304063302-18e6ce5945fd // indirect
github.com/Microsoft/go-winio v0.6.2 // indirect
github.com/Microsoft/go-winio v0.6.3-0.20251027160822-ad3df93bed29 // indirect
github.com/ProtonMail/bcrypt v0.0.0-20211005172633-e235017c1baf // indirect
github.com/ProtonMail/gluon v0.17.1-0.20230724134000-308be39be96e // indirect
github.com/ProtonMail/go-crypto v1.4.1 // indirect
@@ -322,21 +326,21 @@ require (
github.com/andybalholm/cascadia v1.3.4 // indirect
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0
github.com/aws/smithy-go v1.28.1
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
@@ -369,7 +373,7 @@ require (
github.com/go-chi/chi/v5 v5.3.1 // indirect
github.com/go-darwin/apfs v0.0.0-20211011131704-f84b94dbf348 // indirect
github.com/go-jose/go-jose/v4 v4.1.4 // indirect
github.com/go-logr/logr v1.4.3 // indirect
github.com/go-logr/logr v1.4.4 // indirect
github.com/go-logr/stdr v1.2.2 // indirect
github.com/go-ole/go-ole v1.3.0 // indirect
github.com/go-openapi/errors v0.22.8 // indirect
@@ -434,10 +438,10 @@ require (
github.com/oracle/oci-go-sdk/v65 v65.121.0 // indirect
github.com/panjf2000/ants/v2 v2.12.1 // indirect
github.com/patrickmn/go-cache v2.1.0+incompatible // indirect
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/pelletier/go-toml/v2 v2.4.1 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.28
github.com/pierrec/lz4/v4 v4.1.29
github.com/pingcap/errors v0.11.5-0.20211224045212-9687c2b0f87c // indirect
github.com/pingcap/failpoint v0.0.0-20220801062533-2eaa32854a6c // indirect
github.com/pingcap/kvproto v0.0.0-20230403051650-e166ae588106 // indirect
@@ -486,13 +490,13 @@ require (
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.68.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/otel v1.44.0 // indirect
go.opentelemetry.io/otel/metric v1.44.0 // indirect
go.opentelemetry.io/otel/sdk v1.44.0 // indirect
go.opentelemetry.io/otel/sdk/metric v1.44.0 // indirect
go.opentelemetry.io/otel/trace v1.44.0 // indirect
go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.45.0 // indirect
go.opentelemetry.io/otel/sdk/metric v1.45.0 // indirect
go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
@@ -516,13 +520,6 @@ require (
// replace github.com/seaweedfs/raft => /Users/chrislu/go/src/github.com/seaweedfs/raft
// apache/thrift v0.23.0 fixes CVE-2026-41602 but compares int against the
// untyped math.MaxUint32 in lib/go/thrift/framed_transport.go, which overflows
// int on 32-bit GOARCHes (e.g. openbsd/arm, linux/arm) and fails to compile.
// Upstream fixed the range check post-release; pin to that commit until the
// next tagged release carries both the CVE fix and the 32-bit fix.
replace github.com/apache/thrift => github.com/apache/thrift v0.23.1-0.20260429145742-d2acd3c49e58
// tyler-smith/go-bip39 was deleted from GitHub, so `go mod download` fails for
// anyone resolving it directly (GOPROXY=direct). It only reaches us transitively
// through rclone's internxt backend, which calls IsMnemonicValid and NewSeed.
+177 -168
View File
@@ -485,8 +485,8 @@ cloud.google.com/go/storage v1.22.1/go.mod h1:S8N1cAStu7BOeFfE8KAQzmyyLkK8p/vmRq
cloud.google.com/go/storage v1.23.0/go.mod h1:vOEEDNFnciUMhBeT6hsJIn3ieU5cFRmzeLgDvXzfIXc=
cloud.google.com/go/storage v1.27.0/go.mod h1:x9DOL8TK/ygDUMieqwfhdpQryTeEkhGKMi80i/iqR2s=
cloud.google.com/go/storage v1.28.1/go.mod h1:Qnisd4CqDdo6BGs2AD5LLnEsmSQ80wQ5ogcBBKhU86Y=
cloud.google.com/go/storage v1.64.0 h1:KLpxI/oX9LxeRsNqn877d2WyeT3ryiEwnGt8pwcSPZg=
cloud.google.com/go/storage v1.64.0/go.mod h1:lWyAtwvDZHdL3k68WVKbESP6bmWaV23ZJJ/JEVw/ZaQ=
cloud.google.com/go/storage v1.67.0 h1:8xD3NvKuHE4c+b/kmSIVBdTc0BiOjplx+KecKGR9fr8=
cloud.google.com/go/storage v1.67.0/go.mod h1:UsS9OgFg/XHOSYakQ8ZtLWWeyGkk1WnmD/GsGfN0BHM=
cloud.google.com/go/storagetransfer v1.5.0/go.mod h1:dxNzUopWy7RQevYFHewchb29POFv3/AaBgnhqzqiK0w=
cloud.google.com/go/storagetransfer v1.6.0/go.mod h1:y77xm4CQV/ZhFZH75PLEXY0ROiS7Gh6pSKrM8dJyg6I=
cloud.google.com/go/storagetransfer v1.7.0/go.mod h1:8Giuj1QNb1kfLAiWM1bN6dHzfdlDAVC9rv9abHot2W4=
@@ -587,8 +587,8 @@ github.com/DATA-DOG/go-sqlmock v1.5.2 h1:OcvFkGmslmlZibjAjaHm3L//6LiuBgolP7Oputl
github.com/DATA-DOG/go-sqlmock v1.5.2/go.mod h1:88MAG/4G7SMwSE3CeA0ZKzrT5CiOU3OJ+JlNzwDqpNU=
github.com/DataDog/datadog-go v3.2.0+incompatible/go.mod h1:LButxg5PwREeZtORoXG3tL4fMGNddJ+vMq1mwgfaqoQ=
github.com/DataDog/zstd v1.5.2/go.mod h1:g4AWEaM3yOg3HYfnJ3YIawPnVdXJh9QME85blwSAmyw=
github.com/DefangLabs/secret-detector v0.0.0-20250403165618-22662109213e h1:rd4bOvKmDIx0WeTv9Qz+hghsgyjikFiPrseXHlKepO0=
github.com/DefangLabs/secret-detector v0.0.0-20250403165618-22662109213e/go.mod h1:blbwPQh4DTlCZEfk1BLU4oMIhLda2U+A840Uag9DsZw=
github.com/DefangLabs/secret-detector v0.0.0-20250811234530-d4b4214cd679 h1:qNT7R4qrN+5u5ajSbqSW1opHP4LA8lzA+ASyw5MQZjs=
github.com/DefangLabs/secret-detector v0.0.0-20250811234530-d4b4214cd679/go.mod h1:blbwPQh4DTlCZEfk1BLU4oMIhLda2U+A840Uag9DsZw=
github.com/FilenCloudDienste/filen-sdk-go v0.0.39 h1:tgV5jYL6dsXop9TpDTIQU6UwJjws122HrwskaEE/igY=
github.com/FilenCloudDienste/filen-sdk-go v0.0.39/go.mod h1:0cBhKXQg49XbKZZfk5TCDa3sVLP+xMxZTWL+7KY0XR0=
github.com/Files-com/files-sdk-go/v3 v3.3.194 h1:dtOFxSTWWRpkmvXa6ycNiw8dVDu1wkgzcXyVV1VafNc=
@@ -620,8 +620,8 @@ github.com/Masterminds/semver/v3 v3.2.0/go.mod h1:qvl/7zhW3nngYb5+80sSMF+FG2BjYr
github.com/Max-Sum/base32768 v0.0.0-20230304063302-18e6ce5945fd h1:nzE1YQBdx1bq9IlZinHa+HVffy+NmVRoKr+wHN8fpLE=
github.com/Max-Sum/base32768 v0.0.0-20230304063302-18e6ce5945fd/go.mod h1:C8yoIfvESpM3GD07OCHU7fqI7lhwyZ2Td1rbNbTAhnc=
github.com/Microsoft/go-winio v0.5.2/go.mod h1:WpS1mjBmmwHBEWmogvA2mj8546UReBk4v8QkMxJ6pZY=
github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERoyfY=
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
github.com/Microsoft/go-winio v0.6.3-0.20251027160822-ad3df93bed29 h1:0kQAzHq8vLs7Pptv+7TxjdETLf/nIqJpIB4oC6Ba4vY=
github.com/Microsoft/go-winio v0.6.3-0.20251027160822-ad3df93bed29/go.mod h1:ZWa7ssZJT30CCDGJ7fk/2SBTq9BIQrrVjrcss0UW2s0=
github.com/OneOfOne/xxhash v1.2.2/go.mod h1:HSdplMjZKSmBqAxg5vPj2TmRDmfkzw+cTzAElWljhcU=
github.com/ProtonMail/bcrypt v0.0.0-20210511135022-227b4adcab57/go.mod h1:HecWFHognK8GfRDGnFQbW/LiV7A3MX3gZVs45vk5h8I=
github.com/ProtonMail/bcrypt v0.0.0-20211005172633-e235017c1baf h1:yc9daCCYUefEs69zUkSzubzjBbL+cmOXgnmt9Fyd9ug=
@@ -637,6 +637,8 @@ github.com/ProtonMail/gopenpgp/v3 v3.4.1 h1:K7uUhSHSJxORZ+RuHpilTT6S4MA2whCRlXNw
github.com/ProtonMail/gopenpgp/v3 v3.4.1/go.mod h1:bGdV9f6edhmd581wzXsQCTKdH8bXBbyhkgDKPjwPc6U=
github.com/PuerkitoBio/goquery v1.12.0 h1:pAcL4g3WRXekcB9AU/y1mbKez2dbY2AajVhtkO8RIBo=
github.com/PuerkitoBio/goquery v1.12.0/go.mod h1:802ej+gV2y7bbIhOIoPY5sT183ZW0YFofScC4q/hIpQ=
github.com/RoaringBitmap/roaring/v2 v2.24.0 h1:zQkkBZtG3WRP4j+P3A5DO221SvL1Br88TJkhyqEQRZo=
github.com/RoaringBitmap/roaring/v2 v2.24.0/go.mod h1:SfT3of9nYh3vis1dIbCj4Yw6KQGujTN+f345nrN/0JA=
github.com/Sereal/Sereal/Go/sereal v0.0.0-20231009093132-b9187f1a92c6/go.mod h1:JwrycNnC8+sZPDyzM3MQ86LvaGzSpfxg885KOOwFRW4=
github.com/Shopify/sarama v1.38.1 h1:lqqPUPQZ7zPqYlWpTh+LQ9bhYNu2xJL6k1SJN4WVe2A=
github.com/Shopify/sarama v1.38.1/go.mod h1:iwv9a67Ha8VNa+TifujYoWGxWnu2kNVAQdSdZ4X2o5g=
@@ -690,10 +692,11 @@ github.com/apache/arrow-go/v18 v18.7.0/go.mod h1:PM6IigLJkdMwIpeHXnymo+xZ52f42a9
github.com/apache/arrow/go/v10 v10.0.1/go.mod h1:YvhnlEePVnBS4+0z3fhPfUy7W1Ikj0Ih0vcRo/gZ1M0=
github.com/apache/cassandra-gocql-driver/v2 v2.1.2 h1:lu/p0Db2av18enHJvWJQoChLssI0P+AR06STq4VdvCc=
github.com/apache/cassandra-gocql-driver/v2 v2.1.2/go.mod h1:QH/asJjB3mHvY6Dot6ZKMMpTcOrWJ8i9GhsvG1g0PK4=
github.com/apache/iceberg-go v0.6.0 h1:tOVhC5BhBOEgPTowo5AVrPnAgsDo00qEbUPprFcLd4s=
github.com/apache/iceberg-go v0.6.0/go.mod h1:kESfDlyaW/6hK0WW6TzA4EWps+0NMhhrDYE4qajVjlo=
github.com/apache/thrift v0.23.1-0.20260429145742-d2acd3c49e58 h1:rDLE+tSW60VzRD7v5I+DU22Mjhmm+mfLc5Xl5dHkx6w=
github.com/apache/thrift v0.23.1-0.20260429145742-d2acd3c49e58/go.mod h1:zPt6WxgvTOM6hF92y8C+MkEM5LMxZuk4JcQOiU4Esvs=
github.com/apache/iceberg-go v0.6.1-0.20260817192109-c2105090c9e2 h1:xRULj4L2wrlAAPAYNruyuP17lz2zq2DGLgIFDgqiBjo=
github.com/apache/iceberg-go v0.6.1-0.20260817192109-c2105090c9e2/go.mod h1:u6gs2aFRl7QOL0FqfHX6DOzvxOyev6mmq3XT+3VUS9M=
github.com/apache/thrift v0.16.0/go.mod h1:PHK3hniurgQaNMZYaCLEqXKsYK8upmhPbmdP2FXSqgU=
github.com/apache/thrift v0.24.0 h1:zy31L1a49QTNB2bG1BBfMXol3yJrTH975G3pPubQVLQ=
github.com/apache/thrift v0.24.0/go.mod h1:zPt6WxgvTOM6hF92y8C+MkEM5LMxZuk4JcQOiU4Esvs=
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3 h1:WZaTKNHCfcw7fWSR6/RKnCldVzvYZC+Y20Su4lffEIg=
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3/go.mod h1:OMVSB21p9+xQUIqlGizHPZfjK+SHws1ht+ZytVDoz9U=
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc h1:LoL75er+LKDHDUfU5tRvFwxH0LjPpZN8OoG8Ll+liGU=
@@ -707,48 +710,48 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 h1:3IZY0XAJquT3aHzbkHfPzy4ACPcEjVG0x87KOwtpqGY=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14/go.mod h1:zwM6veDkhGgQFqkBy+uT28AAYpLu+uFMlPl+rCg/73E=
github.com/aws/aws-sdk-go-v2/config v1.32.33 h1:M1m/Q6f0OKDEDGwhiNOqx1OjTdrewe3v+GDbHmKczWk=
github.com/aws/aws-sdk-go-v2/config v1.32.33/go.mod h1:fGj1iQj2QpIZzp7jE4aQQ+71TE8cd4z9K4+xCd6EqmE=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34 h1:y6GkSmcv5myd1ngrYbGmiLlwQqB6TQhOuN/tbSSuWDY=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34/go.mod h1:w3dTcnDVoQIewjo7JG45hduAToikiIFLC4FIO7fndvw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 h1:+S7kbJoLDDQ5tE+lHrUBgMkzC8NLgsaioS2F3dVoFAE=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35/go.mod h1:Ak7xXviIARfFdNUJ9Etb0bdVDt/KAvKjMGJVLWXDzik=
github.com/aws/aws-sdk-go-v2 v1.47.0 h1:0jsHallhJCeaU0Ko48c/3FK1ctOQ7NpzggxriJOQ8MQ=
github.com/aws/aws-sdk-go-v2 v1.47.0/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 h1:LAfOuhAH331fmOjTQpAaOlH+Ftn7RzSDJ2VFwjdMMy4=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18/go.mod h1:4e5xhuXHx1e4U9EthvbPP1r/DIMp5c2823OL8karzcM=
github.com/aws/aws-sdk-go-v2/config v1.32.35 h1:UEzXuET8E42lxBPijuACu/tEK7v5lFPlk0Q+GT5WD9E=
github.com/aws/aws-sdk-go-v2/config v1.32.35/go.mod h1:KaMtJpFa2JlL2BStjjHQVwQpzZEmw+ND/EgVrfFoo2g=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4 h1:hTvrJJseKbvw32kmiE0G+u/9ZqpqscjDrTigHIXP2qs=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4/go.mod h1:gWp9O1ZBWwpcIrgV+mVHk4gZUurAEDkgypu/OXOlIaw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 h1:AM4hHjww+PSFtt6E+UrBrPlZkWsePCLEt9AjkfQX+yM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0/go.mod h1:3x/yXezeQjpOvBb4jEMxrS8SXvpdvJ5abv6l5c1gWM8=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3 h1:w5OoDiMN6x53ROmiIImGzmVcxXv2q1GXY+aKV4WAJYM=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3/go.mod h1:dAhgYp776bX3LuWvnSCFwQEjNs6fuFg7YXIy5PXcP3Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 h1:kzVuGlatQtYinwBJEEyLAbggepCoavosiaHHX9+fD+c=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35/go.mod h1:0yLx0yEI+SfqeJMPvOtIEFoZbiQYXMGszBueiutQyaI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 h1:WK6CjihTuLisCjSKKbildJ79sGZZgbBz3iNa7VsKIhU=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35/go.mod h1:KYleN57luLoe97R7vTnx8PMcVrr9gAcRECtOjl91DNg=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 h1:jbGY4CXLzZElOXgGsexlC3Hi+3YM0rSmk4opFXKqg/k=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36/go.mod h1:uBu/9aKsS/UQGc72RAt3y54kjgYQxmhut8ZD2dXCDNE=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 h1:JJLBQxwY+AFwuPAi5ivGc1ChnTdUt4cXMv7e76m2c/Y=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15/go.mod h1:lQknBIe78MVL0cQOQDlag8KGflMbMEVFx9mB6O8ENvk=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 h1:9Fjh6fi/U5JEStVZijmaMpUwE/gvBJj7x2B/PjbO9To=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23/go.mod h1:iMoT2f1tClxrWAAnKCXjZQ6LOmfLrMG14wmnWpM+F14=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 h1:BBEElKh4a+rKshvjrfpajTe9CbpZvrbb4Jkg2PB7RzA=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35/go.mod h1:zaZk983w//8beSruBVec/mr4CmDwgZitW/qzGhAAX0g=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 h1:uao4A3QZ5UmB326V6KF+qRpv9Tjz7IlnlnTbbANntlU=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31/go.mod h1:I/1+z0VwL1GhQyLgkoHDlygpUZ+iTAwOQ/NsftiUL2I=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2 h1:5C00eQYpTrgQXnp6V3P6P7zPElna3AXvlukbANE6nJI=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2/go.mod h1:zdmCoFO/dSI7GlrwsPqFJI+WlFnSU4Tc8TJnlXrM1Do=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 h1:cOJELVNrq5Q3Udry2GLuHUM7MhwpeaQRdYaoa6GI/yI=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4/go.mod h1:f4LxzKBtaTxD7xh3PiVg3CE1tchQemfmghaJr+NbK2c=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11 h1:eBXB8KZgzQ8A9QB4iJS4aw/u6+4OY3i2hQXPABeAIOg=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11/go.mod h1:N9+5pG27Fy61GUL5YXVLXDTLmUudMrgwsuDbgBMNLxQ=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 h1:Hp/VgjP0BysR3OgLlR057Vz2LcbbVnoWeJ+3qWiS/fY=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3/go.mod h1:nwGV5qw7F1IZPgxCvA/ph8N2TAuz+BkRG/bXn808qMA=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 h1:MUaM4f+kj1ZIBPZfUS8cxP1GKXXZtHJjAthy93AN7SM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3/go.mod h1:6YmVmEVRI5ZZzRjCSsb9SryKH0hAlMRdgA7kG9aDvBU=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 h1:fuSCw4Z2qfRCztMPO3GXJNSiEp6Wee+WOLwrHHUMy9c=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3/go.mod h1:6SxcHheD1pPR5+kWm1wGvjlL/YqUsh267sAfEmN4K7A=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 h1:uZOinZb+h7lZw8IYzP1z1IuEnueB76/EFkcf/fEW4Ag=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31/go.mod h1:NRtwAM/p5VRt03TlEUs0pH3TeWamWdf4YyJpSrzPYLc=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 h1:bON1rJf67TSTDCKg816AAIE4xSTtoo9tl0XRkO72R+I=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3/go.mod h1:c5BBpjJcQXpfeq9iASyVKA3T6vX6B6LEXY4mL/gklDY=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 h1:HLPAVrlLDaN2boN0xJx7MgaQDNEO3Q+c9L6kl/8m47Q=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39/go.mod h1:Pg/dVfsNkm1hsIDK/gMvCKtmyNfNTV12mrgHqVE/6Oo=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3 h1:IKoCZqfWfZzSBi16QFQ+QcbQ3LRQ7QgB1S5tDAyPBQQ=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3/go.mod h1:RBpRcXiM4s2pOInVs32GsBonnje+fiAj4mcrStRmlCA=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 h1:ZD5qFpWcaOKdTuhBi431pIDkCgrMkMlMT6jlpSPoIRI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0/go.mod h1:8Nuuf+tR346PjJ3MvZPh9pekbLiLQFWJhzMXfwy7alA=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 h1:AMW7a7S8iQaHjBYZdU3PCq4GKRPijTPRAc7e6XtEThY=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4/go.mod h1:QQNsFV1DVXoXcZt18FS8lI8rtUrlDyAuWZLQ5shunv4=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 h1:AsbZcJAQPRmHDJG8K1N0pof/1zPWjVT8TFlTWuGLSvo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4/go.mod h1:6imqztH0//t0mKbl6yWl7swSEl7F/w32oAmqB3vP1ag=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4 h1:w/AryDYMjSUANSQ2uoZxJovUsMTwWJNTv3IMex30Y+4=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4/go.mod h1:WeBiAa67azG7Su9Vf+ChGDBLiAozJCXzdjXiPBUwtbc=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 h1:JGeeBcMlhg1xtOXYpeCaTQBZObtXMPQCUqBcmr65NRA=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0/go.mod h1:XwteswG9EOMRFm73UT0t+MbTwyLxMrEXkU6e+v92Lzo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 h1:obhahQXDEdVEv8y5bTKXR30LVaxYe1kyYM0L7l2Iq+k=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0/go.mod h1:6twZZ/aXHNy1vXUO8koUbp++MYzMASkOgEBdkbJYmO0=
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0 h1:khXV3+K5D3f4e8xtplaRdSFn1bEg3gj5EBHQvbCOZbQ=
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0/go.mod h1:/8JRcdTt//hG0Q4BTmGbuOplT7ABe+5rdtqUHqXvYIM=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
@@ -762,6 +765,8 @@ github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM=
github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw=
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f h1:+6okTAeUsUrdQr/qN7fIODzowrjjCrnJDg/gkYqcSXY=
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f/go.mod h1:z52shMwD6SGwRg2iYFjjDwX5Ene4ENTw6HfXraUy/08=
github.com/bits-and-blooms/bitset v1.24.4 h1:95H15Og1clikBrKr/DuzMXkQzECs1M6hhoGXLwLQOZE=
github.com/bits-and-blooms/bitset v1.24.4/go.mod h1:7hO7Gc7Pp1vODcmWvKMRA9BNmbv6a/7QIWpPxHddWR8=
github.com/blevesearch/snowballstem v0.9.0 h1:lMQ189YspGP6sXvZQ4WZ+MLawfV8wOmPoD/iWeNXm8s=
github.com/blevesearch/snowballstem v0.9.0/go.mod h1:PivSj3JMc8WuaFkTSRDW2SlrulNWPl4ABg1tC/hlgLs=
github.com/boltdb/bolt v1.3.1 h1:JQmyP4ZBrce+ZQu0dY660FMfatumYDLun9hBCUVIkF4=
@@ -860,17 +865,17 @@ github.com/cognusion/imaging v1.0.4 h1:hrmWVa4S9fiLJJFKEHYYRNBUnHEKZD8/oxnC77xTU
github.com/cognusion/imaging v1.0.4/go.mod h1:X1mPSSZpMyiva54chCO/1NROp+b6xG4uMcud8VLii+U=
github.com/colinmarc/hdfs/v2 v2.4.0 h1:v6R8oBx/Wu9fHpdPoJJjpGSUxo8NhHIwrwsfhFvU9W0=
github.com/colinmarc/hdfs/v2 v2.4.0/go.mod h1:0NAO+/3knbMx6+5pCv+Hcbaz4xn/Zzbn9+WIib2rKVI=
github.com/compose-spec/compose-go/v2 v2.10.2 h1:USa1NUbDcl/cjb8T9iwnuFsnO79H+2ho2L5SjFKz3uI=
github.com/compose-spec/compose-go/v2 v2.10.2/go.mod h1:ZU6zlcweCZKyiB7BVfCizQT9XmkEIMFE+PRZydVcsZg=
github.com/compose-spec/compose-go/v2 v2.12.1 h1:+xBZNxcgSus4atQJwXPEdhHRgCEyZmj/BuqN5m33Ou0=
github.com/compose-spec/compose-go/v2 v2.12.1/go.mod h1:ZU6zlcweCZKyiB7BVfCizQT9XmkEIMFE+PRZydVcsZg=
github.com/containerd/console v1.0.3/go.mod h1:7LqA/THxQ86k76b8c/EMSiaJ3h1eZkMkXar0TQ1gf3U=
github.com/containerd/console v1.0.5 h1:R0ymNeydRqH2DmakFNdmjR2k0t7UPuiOV/N/27/qqsc=
github.com/containerd/console v1.0.5/go.mod h1:YynlIjWYF8myEu6sdkwKIvGQq+cOckRm6So2avqoYAk=
github.com/containerd/containerd/api v1.10.0 h1:5n0oHYVBwN4VhoX9fFykCV9dF1/BvAXeg2F8W6UYq1o=
github.com/containerd/containerd/api v1.10.0/go.mod h1:NBm1OAk8ZL+LG8R0ceObGxT5hbUYj7CzTmR3xh0DlMM=
github.com/containerd/containerd/v2 v2.2.2 h1:mjVQdtfryzT7lOqs5EYUFZm8ioPVjOpkSoG1GJPxEMY=
github.com/containerd/containerd/v2 v2.2.2/go.mod h1:5Jhevmv6/2J+Iu/A2xXAdUIdI5Ah/hfyO7okJ4AFIdY=
github.com/containerd/continuity v0.4.5 h1:ZRoN1sXq9u7V6QoHMcVWGhOwDFqZ4B9i5H6un1Wh0x4=
github.com/containerd/continuity v0.4.5/go.mod h1:/lNJvtJKUQStBzpVQ1+rasXO1LAWtUQssk28EZvJ3nE=
github.com/containerd/containerd/api v1.11.1 h1:h8nfoDW9+fNsC/9TwiAHj8B1GzXKtR4eFtkhi/X5RLU=
github.com/containerd/containerd/api v1.11.1/go.mod h1:CaQFRu+N1MtbgL6JDOJLUB1hCKESU1lD6MuTJhgtdlw=
github.com/containerd/containerd/v2 v2.2.5 h1:KTFzB02LviYmmfRmz8r9UFd+n6YlddVFK+5lbgQXUTU=
github.com/containerd/containerd/v2 v2.2.5/go.mod h1:5t2+xFv2dGd/iDYp9Z8DXB4cmWrWQi1XqxGJPS2gBzU=
github.com/containerd/continuity v0.5.0 h1:7a85HZpCSs+1Zps0Ee3DPSuAWY+0SJM1JNM51nlEVDg=
github.com/containerd/continuity v0.5.0/go.mod h1:/lNJvtJKUQStBzpVQ1+rasXO1LAWtUQssk28EZvJ3nE=
github.com/containerd/errdefs v1.0.0 h1:tg5yIfIlQIrxYtu9ajqY42W3lpS19XqdxRQeEwYG8PI=
github.com/containerd/errdefs v1.0.0/go.mod h1:+YBYIdtsnF4Iw6nWZhJcqGSg/dwvV7tyJ/kCkyJ2k+M=
github.com/containerd/errdefs/pkg v0.3.0 h1:9IKJ06FvyNlexW690DXuQNx2KA2cUJXx151Xdx3ZPPE=
@@ -881,8 +886,8 @@ github.com/containerd/platforms v1.0.0-rc.4 h1:M42JrUT4zfZTqtkUwkr0GzmUWbfyO5VO0
github.com/containerd/platforms v1.0.0-rc.4/go.mod h1:lKlMXyLybmBedS/JJm11uDofzI8L2v0J2ZbYvNsbq1A=
github.com/containerd/ttrpc v1.2.8 h1:xbVu6D4qF2jihdh9rDVOKqUMiFBQk6YctTdo1zk087Y=
github.com/containerd/ttrpc v1.2.8/go.mod h1:wyZW2K79t4Hfcxl+GUvkZqRBzJlqFFvgEeeWXa42tyE=
github.com/containerd/typeurl/v2 v2.2.3 h1:yNA/94zxWdvYACdYO8zofhrTVuQY73fFU1y++dYSw40=
github.com/containerd/typeurl/v2 v2.2.3/go.mod h1:95ljDnPfD3bAbDJRugOiShd/DlAAsxGtUBhJxIn7SCk=
github.com/containerd/typeurl/v2 v2.3.0 h1:HZHPhRWo5XMy3QGQoPrUzbW/2ckwjfweHmOwlkIrPAQ=
github.com/containerd/typeurl/v2 v2.3.0/go.mod h1:Qk+PAdUYArVj41TnGi6rJ+48RF0PkcTc4i/taoBcK0w=
github.com/coreos/go-semver v0.3.1 h1:yi21YpKnrx1gt5R+la8n5WgS0kCrsPp33dmEyHReZr4=
github.com/coreos/go-semver v0.3.1/go.mod h1:irMmmIw/7yzSRPWryHsK7EYSg09caPQL03VsM8rvUec=
github.com/coreos/go-systemd/v22 v22.7.0 h1:LAEzFkke61DFROc7zNLX/WA2i5J8gYqe0rSj9KI28KA=
@@ -926,18 +931,18 @@ github.com/djherbis/times v1.6.0 h1:w2ctJ92J8fBvWPxugmXIv7Nz7Q3iDMKNx9v5ocVH20c=
github.com/djherbis/times v1.6.0/go.mod h1:gOHeRAz2h+VJNZ5Gmc/o7iD9k4wW7NMVqieYCY99oc0=
github.com/dnaeon/go-vcr v1.2.0 h1:zHCHvJYTMh1N7xnV7zf1m1GPBF9Ad0Jk/whtQ1663qI=
github.com/dnaeon/go-vcr v1.2.0/go.mod h1:R4UdLID7HZT3taECzJs4YgbbH6PIGXB6W/sc5OLb6RQ=
github.com/docker/buildx v0.33.0 h1:xuZeuQe/C/2tvLDgiIA6+Ynq3FFWSfsGNWIHM3q1hD8=
github.com/docker/buildx v0.33.0/go.mod h1:7JVma62htERKE5iy5YD1q64PKiAHUzXuhSBd4oq3I74=
github.com/docker/cli v29.4.0+incompatible h1:+IjXULMetlvWJiuSI0Nbor36lcJ5BTcVpUmB21KBoVM=
github.com/docker/cli v29.4.0+incompatible/go.mod h1:JLrzqnKDaYBop7H2jaqPtU4hHvMKP+vjCwu2uszcLI8=
github.com/docker/compose/v5 v5.1.2 h1:HxtfGZA0DESw/+hvrNJDjM8VKmCond7OkmgVJCHku48=
github.com/docker/compose/v5 v5.1.2/go.mod h1:JJ2H+lRSugOH50/zxCbkBwN8jU91qDtRCp7gEHWZdgo=
github.com/docker/buildx v0.35.0 h1:5r/ZcAC0s2oZ1XOJHvYFYbj44I4rxRsIdfsRBeXClks=
github.com/docker/buildx v0.35.0/go.mod h1:kO/9ZBoZmhi1OsZbeLJqWclckB2TKu/yeW6aj6xxdt0=
github.com/docker/cli v29.6.0+incompatible h1:nw9himxMMZ7eIeherJNlKQq+acnlzGgHd+4uf10QRSc=
github.com/docker/cli v29.6.0+incompatible/go.mod h1:JLrzqnKDaYBop7H2jaqPtU4hHvMKP+vjCwu2uszcLI8=
github.com/docker/compose/v5 v5.2.0 h1:WuZgtAM91BjFWy7xMKcM2Huq9OksVqRgkgBKFiEXs2Q=
github.com/docker/compose/v5 v5.2.0/go.mod h1:tjYsBMooQQZ8XGtVoUnc+/T9peXtkMmTdmMakW8/ZXE=
github.com/docker/docker v28.5.2+incompatible h1:DBX0Y0zAjZbSrm1uzOkdr1onVghKaftjlSWt4AFexzM=
github.com/docker/docker v28.5.2+incompatible/go.mod h1:eEKB0N0r5NX/I1kEveEz05bcu8tLC/8azJZsviup8Sk=
github.com/docker/docker-credential-helpers v0.9.5 h1:EFNN8DHvaiK8zVqFA2DT6BjXE0GzfLOZ38ggPTKePkY=
github.com/docker/docker-credential-helpers v0.9.5/go.mod h1:v1S+hepowrQXITkEfw6o4+BMbGot02wiKpzWhGUZK6c=
github.com/docker/go-connections v0.6.0 h1:LlMG9azAe1TqfR7sO+NJttz1gy6KO7VJBh+pMmjSD94=
github.com/docker/go-connections v0.6.0/go.mod h1:AahvXYshr6JgfUJGdDCs2b5EZG/vmaMAntpSFH5BFKE=
github.com/docker/docker-credential-helpers v0.9.8 h1:bIREROb7So6PRlq6KTtdS9MPEjC29OQRkFNlvK2OX8Q=
github.com/docker/docker-credential-helpers v0.9.8/go.mod h1:v1S+hepowrQXITkEfw6o4+BMbGot02wiKpzWhGUZK6c=
github.com/docker/go-connections v0.7.0 h1:6SsRfJddP22WMrCkj19x9WKjEDTB+ahsdiGYf0mN39c=
github.com/docker/go-connections v0.7.0/go.mod h1:no1qkHdjq7kLMGUXYAduOhYPSJxxvgWBh7ogVvptn3Q=
github.com/docker/go-units v0.5.0 h1:69rxXcBk27SvSaaxTtLh/8llcHD8vYHT7WSdRZ/jvr4=
github.com/docker/go-units v0.5.0/go.mod h1:fgPhTUdO+D/Jk86RDLlptpiXQzgHJF7gydDDbaIK4Dk=
github.com/docopt/docopt-go v0.0.0-20180111231733-ee0de3bc6815/go.mod h1:WwZ+bS3ebgob9U8Nd0kOddGdZWjyMGR8Wziv+TBNwSE=
@@ -1029,17 +1034,19 @@ github.com/fvbommel/sortorder v1.1.0 h1:fUmoe+HLsBTctBDoaBwpQo5N+nrCp8g/BjKb/6ZQ
github.com/fvbommel/sortorder v1.1.0/go.mod h1:uk88iVf1ovNn1iLfgUVU2F9o5eO30ui720w+kxuqRs0=
github.com/gabriel-vasile/mimetype v1.4.13 h1:46nXokslUBsAJE/wMsp5gtO500a4F3Nkz9Ufpk2AcUM=
github.com/gabriel-vasile/mimetype v1.4.13/go.mod h1:d+9Oxyo1wTzWdyVUPMmXFvp4F9tea18J8ufA774AB3s=
github.com/geoarrow/geoarrow-go v0.0.0-20260403143023-f54751c3e3a1 h1:8VcjuP5pKZ717K4zOLS9Lm4Zsn8JwD9X+CQGFCiRT2E=
github.com/geoarrow/geoarrow-go v0.0.0-20260403143023-f54751c3e3a1/go.mod h1:XMOIOA5J96jEzAOpBROPyz0n14sHsooeF+y2FtGdrCY=
github.com/geoffgarside/ber v1.2.0 h1:/loowoRcs/MWLYmGX9QtIAbA+V/FrnVLsMMPhwiRm64=
github.com/geoffgarside/ber v1.2.0/go.mod h1:jVPKeCbj6MvQZhwLYsGwaGI52oUorHoHKNecGT85ZCc=
github.com/getsentry/sentry-go v0.48.0 h1:FRZNr7Uk1C86ev1bSJmYlUkL9oyivQA6YOcdYfaaMmY=
github.com/getsentry/sentry-go v0.48.0/go.mod h1:E5UkA5wp1qR2+MDydNYlVeUiNN2xEdjYMidkgf0Qoss=
github.com/getsentry/sentry-go v0.49.0 h1:Ehejknu1l023Ub7QoRBVLAI7g3Jnhqku4oWx4B4Sh5s=
github.com/getsentry/sentry-go v0.49.0/go.mod h1:nuMJAoCfe1u0Bts2ocyNI+TW8HT84vRMqwA5Qq/SKUI=
github.com/ghodss/yaml v1.0.0/go.mod h1:4dBDuWmgqj2HViK6kFavaiC9ZROes6MMH2rRYeMEF04=
github.com/gin-contrib/sse v1.1.0 h1:n0w2GMuUpWDVp7qSpvze6fAu9iRxJY4Hmj6AmBOU05w=
github.com/gin-contrib/sse v1.1.0/go.mod h1:hxRZ5gVpWMT7Z0B0gSNYqqsSCNIJMjzvm6fqCz9vjwM=
github.com/gin-gonic/gin v1.11.0 h1:OW/6PLjyusp2PPXtyxKHU0RbX6I/l28FTdDlae5ueWk=
github.com/gin-gonic/gin v1.11.0/go.mod h1:+iq/FyxlGzII0KHiBGjuNn4UNENUlKbGlNmc+W50Dls=
github.com/go-asn1-ber/asn1-ber v1.5.8-0.20250403174932-29230038a667 h1:BP4M0CvQ4S3TGls2FvczZtj5Re/2ZzkV9VwqPHH/3Bo=
github.com/go-asn1-ber/asn1-ber v1.5.8-0.20250403174932-29230038a667/go.mod h1:hEBeB/ic+5LoWskz+yKT7vGhhPYkProFKoKdwZRWMe0=
github.com/go-asn1-ber/asn1-ber v1.5.8 h1:H9AZkK22UOmfX8J84ubyaZxKJZ3FMHVwn8swoMML7iQ=
github.com/go-asn1-ber/asn1-ber v1.5.8/go.mod h1:hEBeB/ic+5LoWskz+yKT7vGhhPYkProFKoKdwZRWMe0=
github.com/go-chi/chi/v5 v5.3.1 h1:3j4HZLGZQ3JpMCrPJF/Jl3mYJfWLKBfNJ6quurUGCf8=
github.com/go-chi/chi/v5 v5.3.1/go.mod h1:R+tYY2hNuVUUjxoPtqUdgBqevM9s9njzkTLutVsOCto=
github.com/go-darwin/apfs v0.0.0-20211011131704-f84b94dbf348 h1:JnrjqG5iR07/8k7NqrLNilRsl3s1EPRQEGvbPyOce68=
@@ -1063,14 +1070,14 @@ github.com/go-kit/kit v0.9.0/go.mod h1:xBxKIO96dXMWWy0MnWVtmwkA9/13aqxPnvrjFYMA2
github.com/go-kit/log v0.1.0/go.mod h1:zbhenjAZHb184qTLMA9ZjW7ThYL0H2mk7Q6pNt4vbaY=
github.com/go-latex/latex v0.0.0-20210118124228-b3d85cf34e07/go.mod h1:CO1AlKB2CSIqUrmQPqA0gdRIlnLEY0gK5JGjh37zN5U=
github.com/go-latex/latex v0.0.0-20210823091927-c0d11ff05a81/go.mod h1:SX0U8uGpxhq9o2S/CELCSUxEWWAuoCUcVCQWv7G2OCk=
github.com/go-ldap/ldap/v3 v3.4.13 h1:+x1nG9h+MZN7h/lUi5Q3UZ0fJ1GyDQYbPvbuH38baDQ=
github.com/go-ldap/ldap/v3 v3.4.13/go.mod h1:LxsGZV6vbaK0sIvYfsv47rfh4ca0JXokCoKjZxsszv0=
github.com/go-ldap/ldap/v3 v3.4.14 h1:D6PYdEgsaVzsXyr6w/yDC06Ria4uUhWm+Rb+er8lfAs=
github.com/go-ldap/ldap/v3 v3.4.14/go.mod h1:S4eJUMUNjDkE0ZJtIZdybwyb03sGGLW6gxXT1Hs8VKA=
github.com/go-logfmt/logfmt v0.3.0/go.mod h1:Qt1PoO58o5twSAckw1HlFXLmHsOX5/0LbT9GBnD5lWE=
github.com/go-logfmt/logfmt v0.4.0/go.mod h1:3RMwSq7FuexP4Kalkev3ejPJsZTpXXBr9+V4qmtdjCk=
github.com/go-logfmt/logfmt v0.5.0/go.mod h1:wCYkCAKZfumFQihp8CzCvQ3paCTfi41vtzG1KdI/P7A=
github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A=
github.com/go-logr/logr v1.4.3 h1:CjnDlHq8ikf6E492q6eKboGOC0T8CDaOvkHCIg8idEI=
github.com/go-logr/logr v1.4.3/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8=
github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag=
github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE=
github.com/go-ole/go-ole v1.2.6/go.mod h1:pprOEPIfldk/42T2oK7lQ4v4JSDwmV0As9GaiUsvbm0=
@@ -1102,8 +1109,8 @@ github.com/go-redsync/redsync/v4 v4.17.0 h1:FFJ+uxZs44y4Sq10//IFKic9T94AYl+u3Sog
github.com/go-redsync/redsync/v4 v4.17.0/go.mod h1:CKVA6qwT07S/916i+Yd9h1/8YFQhCCpPYTQhvvYytJo=
github.com/go-resty/resty/v2 v2.17.2 h1:FQW5oHYcIlkCNrMD2lloGScxcHJ0gkjshV3qcQAyHQk=
github.com/go-resty/resty/v2 v2.17.2/go.mod h1:kCKZ3wWmwJaNc7S29BRtUhJwy7iqmn+2mLtQrOyQlVA=
github.com/go-sql-driver/mysql v1.10.0 h1:Q+1LV8DkHJvSYAdR83XzuhDaTykuDx0l6fkXxoWCWfw=
github.com/go-sql-driver/mysql v1.10.0/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-stack/stack v1.8.0/go.mod h1:v0f6uXyyMGvRgIKkXu+yp6POWl0qKG85gN/melR3HDY=
github.com/go-task/slim-sprig v0.0.0-20230315185526-52ccab3ef572 h1:tfuBGBXKqDEevZMzYi5KSi8KkcZtzBcTgAUUtapy0OI=
github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI=
@@ -1360,8 +1367,8 @@ github.com/hpcloud/tail v1.0.0/go.mod h1:ab1qPbhIpdTxEkNHXyeSf5vhxWSCs/tWer42PpO
github.com/iancoleman/strcase v0.2.0/go.mod h1:iwCmte+B7n89clKwxIoIXy/HfoL7AsD47ZCWhYzw7ho=
github.com/ianlancetaylor/demangle v0.0.0-20181102032728-5e5cf60278f6/go.mod h1:aSSvb/t6k1mPoxDqO4vJh6VOCGPwU4O0C2/Eqndh1Sc=
github.com/ianlancetaylor/demangle v0.0.0-20200824232613-28f6c0f3b639/go.mod h1:aSSvb/t6k1mPoxDqO4vJh6VOCGPwU4O0C2/Eqndh1Sc=
github.com/in-toto/attestation v1.1.2 h1:MBFn6lsMq6dptQZJBhalXTcWMb/aJy3V+GX3VYj/V1E=
github.com/in-toto/attestation v1.1.2/go.mod h1:gYFddHMZj3DiQ0b62ltNi1Vj5rC879bTmBbrv9CRHpM=
github.com/in-toto/attestation v1.2.0 h1:aPRUZ3azbqD7yEBD5fP3TD8Dszf+YHo284SOcpahjQk=
github.com/in-toto/attestation v1.2.0/go.mod h1:r79G45gOmzPismgObLSL+rZTFxUgZLOQJI6LofTZgXk=
github.com/in-toto/in-toto-golang v0.11.0 h1:nfidMYBFx+E0lnmX5KUnN2Pdm8zdNKal1ayjJuzzRoA=
github.com/in-toto/in-toto-golang v0.11.0/go.mod h1:u3PjTnwFKjp5a1YCcw8SJg0G+tMeKfVoWsWeFMDCMtw=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
@@ -1446,15 +1453,15 @@ github.com/kisielk/gotool v1.0.0/go.mod h1:XhKaO+MFFWcvkIS/tQcRk01m1F5IRFswLeQ+o
github.com/kisielk/sqlstruct v0.0.0-20201105191214-5f3e10d3ab46/go.mod h1:yyMNCyc/Ib3bDTKd379tNMpB/7/H5TjM2Y9QJ5THLbE=
github.com/klauspost/asmfmt v1.3.2/go.mod h1:AG8TuvYojzulgDAMCnYn50l/5QV3Bs/tp6j0HLHbNSE=
github.com/klauspost/compress v1.15.9/go.mod h1:PhcZ0MbTNciWF3rruxRgKxI5NkcHHrHUDtV4Yw2GlzU=
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/compress v1.19.2 h1:hMRETovs/pu/dVWN7zIT1PGG8t509MwT6bO7XSi26R8=
github.com/klauspost/compress v1.19.2/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/cpuid/v2 v2.0.9/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa0213Md/qVLRg=
github.com/klauspost/cpuid/v2 v2.0.10/go.mod h1:g2LTdtYhdyuGPqyWyv7qRAmj1WBqxuObKfj5c0PQa7c=
github.com/klauspost/cpuid/v2 v2.0.12/go.mod h1:g2LTdtYhdyuGPqyWyv7qRAmj1WBqxuObKfj5c0PQa7c=
github.com/klauspost/cpuid/v2 v2.4.0 h1:S6Hrbc7+ywsr0r+RLapfGBHfyefhCTwEh3A0tV913Dw=
github.com/klauspost/cpuid/v2 v2.4.0/go.mod h1:19jmZ9mjzoF//ddRSUsv0zfBTJWh3QJh9FNxZTMrGxU=
github.com/klauspost/reedsolomon v1.14.1 h1:swE9kzyWXD/wVG+l5Pe8bWnQ0giIY7D1GjCBKk3kG2U=
github.com/klauspost/reedsolomon v1.14.1/go.mod h1:yjqqjgMTQkBUHSG97/rm4zipffCNbCiZcB3kTqr++sQ=
github.com/klauspost/reedsolomon v1.14.2 h1:SafJYwpBBQBI6amHUygcjxZjXeN2HpiENHQDwuPWCCQ=
github.com/klauspost/reedsolomon v1.14.2/go.mod h1:yjqqjgMTQkBUHSG97/rm4zipffCNbCiZcB3kTqr++sQ=
github.com/konsorten/go-windows-terminal-sequences v1.0.1/go.mod h1:T0+1ngSBFLxvqU3pZ+m/2kptfBszLMUkC4ZK/EgS/cQ=
github.com/konsorten/go-windows-terminal-sequences v1.0.3/go.mod h1:T0+1ngSBFLxvqU3pZ+m/2kptfBszLMUkC4ZK/EgS/cQ=
github.com/koofr/go-httpclient v0.0.0-20240520111329-e20f8f203988 h1:CjEMN21Xkr9+zwPmZPaJJw+apzVbjGL5uK/6g9Q2jGU=
@@ -1517,8 +1524,8 @@ github.com/mattn/go-runewidth v0.0.3/go.mod h1:LwmH8dsx7+W8Uxz3IHJYH5QSwggIsqBzp
github.com/mattn/go-runewidth v0.0.13/go.mod h1:Jdepj2loyihRzMpdS35Xk/zdY8IAYHsh153qUoGf23w=
github.com/mattn/go-runewidth v0.0.24 h1:cpokDiIn0MGnhdHwuWnJBITySJ20QyNGnY2kR/ay2DU=
github.com/mattn/go-runewidth v0.0.24/go.mod h1:XBkDxAl56ILZc9knddidhrOlY5R/pDhgLpndooCuJAs=
github.com/mattn/go-shellwords v1.0.12 h1:M2zGm7EW6UQJvDeQxo4T51eKPurbeFbe8WtebGE2xrk=
github.com/mattn/go-shellwords v1.0.12/go.mod h1:EZzvwXDESEeg03EKmM+RmDnNOPKG4lLtQsUlTZDWQ8Y=
github.com/mattn/go-shellwords v1.0.13 h1:DC0OMEpGjm6LfNFU4ckYcvbQKyp2vE8atyFGXNtDcf4=
github.com/mattn/go-shellwords v1.0.13/go.mod h1:EZzvwXDESEeg03EKmM+RmDnNOPKG4lLtQsUlTZDWQ8Y=
github.com/mattn/go-sqlite3 v1.14.14/go.mod h1:NyWgC/yNuGj7Q9rpYnZvas74GogHl5/Z4A/KQRfk6bU=
github.com/mattn/go-sqlite3 v1.14.34 h1:3NtcvcUnFBPsuRcno8pUtupspG/GM+9nZ88zgJcp6Zk=
github.com/mattn/go-sqlite3 v1.14.34/go.mod h1:Uh1q+B4BYcTPb+yiD3kU8Ct7aC0hY9fxUwlHK0RXw+Y=
@@ -1539,18 +1546,18 @@ github.com/mitchellh/mapstructure v1.5.1-0.20220423185008-bf980b35cac4 h1:BpfhmL
github.com/mitchellh/mapstructure v1.5.1-0.20220423185008-bf980b35cac4/go.mod h1:bFUtVrKA4DC2yAKiSyO/QUcy7e+RRV2QTWOzhPopBRo=
github.com/mmcloughlin/geohash v0.9.0 h1:FihR004p/aE1Sju6gcVq5OLDqGcMnpBY+8moBqIsVOs=
github.com/mmcloughlin/geohash v0.9.0/go.mod h1:oNZxQo5yWJh0eMQEP/8hwQuVx9Z9tjwFUqcTB1SmG0c=
github.com/moby/buildkit v0.29.0 h1:wxLEFbCOJntEDjSNNN2YWd8zxltZxT5muDQ0LzpbtpU=
github.com/moby/buildkit v0.29.0/go.mod h1:Dmv2FeDe34t75QuzeU87rBoZpAAkcpT5zeu4hXzmASc=
github.com/moby/buildkit v0.31.0 h1:hMUAbQGgjtzJDDOZ6o7MQk5XBZkBTyzLWEvnjguHHQI=
github.com/moby/buildkit v0.31.0/go.mod h1:YM5iNEbNCc6L1Zt3YWFB/aXNLufvf4Rcu0DPlc9HwQg=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
github.com/moby/docker-image-spec v1.3.1/go.mod h1:eKmb5VW8vQEh/BAr2yvVNvuiJuY6UIocYsFu/DxxRpo=
github.com/moby/go-archive v0.3.0 h1:nos4BtzzUIqB406BgQnWGMI4qib9BZ8XUHU+ucv/n1c=
github.com/moby/go-archive v0.3.0/go.mod h1:Npdv43fFqlhZW7Xo8fbm3ZMYFvAGNviUPqX21VERbcE=
github.com/moby/locker v1.0.1 h1:fOXqR41zeveg4fFODix+1Ch4mj/gT0NE1XJbp/epuBg=
github.com/moby/locker v1.0.1/go.mod h1:S7SDdo5zpBK84bzzVlKr2V0hz+7x9hWbYC/kq7oQppc=
github.com/moby/moby/api v1.54.2 h1:wiat9QAhnDQjA7wk1kh/TqHz2I1uUA7M7t9SAl/JNXg=
github.com/moby/moby/api v1.54.2/go.mod h1:+RQ6wluLwtYaTd1WnPLykIDPekkuyD/ROWQClE83pzs=
github.com/moby/moby/client v0.4.0 h1:S+2XegzHQrrvTCvF6s5HFzcrywWQmuVnhOXe2kiWjIw=
github.com/moby/moby/client v0.4.0/go.mod h1:QWPbvWchQbxBNdaLSpoKpCdf5E+WxFAgNHogCWDoa7g=
github.com/moby/moby/api v1.55.0 h1:2/sexvQyqIWS8pRSCFddBfpW2qE7vR7FCL+vN8pxwMc=
github.com/moby/moby/api v1.55.0/go.mod h1:+RQ6wluLwtYaTd1WnPLykIDPekkuyD/ROWQClE83pzs=
github.com/moby/moby/client v0.5.0 h1:5XhyPk2fuOWf6RlSFa3MkIIgDZkF25xToXW8Q/BH7cc=
github.com/moby/moby/client v0.5.0/go.mod h1:rcVpF8ncl9vo5gaIBdol6CnbEtSj1uxMvEV/UrykF/s=
github.com/moby/patternmatcher v0.6.1 h1:qlhtafmr6kgMIJjKJMDmMWq7WLkKIo23hsrpR3x084U=
github.com/moby/patternmatcher v0.6.1/go.mod h1:hDPoyOpDY7OrrMDLaYoY3hf52gNCR/YOUYxkhApJIxc=
github.com/moby/sys/atomicwriter v0.1.0 h1:kw5D/EqkBwsBFi0ss9v1VG3wIkVhzGvLklJ+w3A14Sw=
@@ -1582,6 +1589,8 @@ github.com/montanaflynn/stats v0.7.1 h1:etflOAAHORrCC44V+aR6Ftzort912ZU+YLiSTuV8
github.com/montanaflynn/stats v0.7.1/go.mod h1:etXPPgVO6n31NxCd9KQUMvCM+ve0ruNzt6R8Bnaayow=
github.com/morikuni/aec v1.1.0 h1:vBBl0pUnvi/Je71dsRrhMBtreIqNMYErSAbEeb8jrXQ=
github.com/morikuni/aec v1.1.0/go.mod h1:xDRgiq/iw5l+zkao76YTKzKttOp2cwPEne25HDkJnBw=
github.com/mschoch/smat v0.2.0 h1:8imxQsjDm8yFEAVBe7azKmKSgzSkZXDuKkSq9374khM=
github.com/mschoch/smat v0.2.0/go.mod h1:kc9mz7DoBKqDyiRL7VZN8KvXQMWeTaVnttLRXOlotKw=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ=
github.com/mwitkow/go-conntrack v0.0.0-20161129095857-cc309e4a2223/go.mod h1:qRWi+5nqEBWmkhHvq77mSJWrCKwh8bxhgT7d/eI7P4U=
@@ -1643,8 +1652,8 @@ github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144T
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
github.com/patrickmn/go-cache v2.1.0+incompatible/go.mod h1:3Qf8kWWT7OJRJbdiICTKqZju1ZixQ/KpMGzzAfe6+WQ=
github.com/pborman/getopt v0.0.0-20170112200414-7148bc3a4c30/go.mod h1:85jBQOZwpVEaDAr341tbn15RS4fCAsIst0qp7i8ex1o=
github.com/pelletier/go-toml/v2 v2.2.4 h1:mye9XuhQ6gvn5h28+VilKrrPoQVanw5PMw/TB0t5Ec4=
github.com/pelletier/go-toml/v2 v2.2.4/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pelletier/go-toml/v2 v2.4.1 h1:j5OMOImsH+j2k7GJ5YO+RxfWwohNiH6t5zB/+h3bagc=
github.com/pelletier/go-toml/v2 v2.4.1/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 h1:XeOYlK9W1uCmhjJSsY78Mcuh7MVkNjTzmHx1yBzizSU=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14/go.mod h1:jVblp62SafmidSkvWrXyxAme3gaTfEtWwRPGz5cpvHg=
github.com/peterh/liner v1.2.2 h1:aJ4AOodmL+JxOZZEL2u9iJf8omNRpqHc/EbrK+3mAXw=
@@ -1659,8 +1668,8 @@ github.com/phpdave11/gofpdf v1.4.2/go.mod h1:zpO6xFn9yxo3YLyMvW8HcKWVdbNqgIfOOp2
github.com/phpdave11/gofpdi v1.0.12/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/phpdave11/gofpdi v1.0.13/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/pierrec/lz4/v4 v4.1.15/go.mod h1:gZWDp/Ze/IJXGXf23ltt2EXimqmTUXEy0GFuRQyBid4=
github.com/pierrec/lz4/v4 v4.1.28 h1:pPEPwRJ4kybBTfGt28q7lQsRJQHhC08axprdLD5Ppio=
github.com/pierrec/lz4/v4 v4.1.28/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrec/lz4/v4 v4.1.29 h1:CDQY6qZOLI4DW0Nx6R1vRrifrCeQHnNXkMb0hZWXFjg=
github.com/pierrec/lz4/v4 v4.1.29/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrre/compare v1.0.2 h1:k4IUsHgh+dbcAOIWCfxVa/7G6STjADH2qmhomv+1quc=
github.com/pierrre/compare v1.0.2/go.mod h1:8UvyRHH+9HS8Pczdd2z5x/wvv67krDwVxoOndaIIDVU=
github.com/pierrre/geohash v1.0.0 h1:f/zfjdV4rVofTCz1FhP07T+EMQAvcMM2ioGZVt+zqjI=
@@ -1717,8 +1726,8 @@ github.com/prometheus/client_model v0.0.0-20190129233127-fd36f4220a90/go.mod h1:
github.com/prometheus/client_model v0.0.0-20190812154241-14fe0d1b01d4/go.mod h1:xMI15A0UPsDsEKsMN9yxemIoYk6Tm2C1GtYGdfGttqA=
github.com/prometheus/client_model v0.2.0/go.mod h1:xMI15A0UPsDsEKsMN9yxemIoYk6Tm2C1GtYGdfGttqA=
github.com/prometheus/client_model v0.3.0/go.mod h1:LDGWKZIo7rky3hgvBe+caln+Dr3dPggB5dvjtD7w9+w=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/client_model v0.6.3 h1:O0jaTVAYNxTHYInEPFJt5I3+sN8zqBtVMPTB1qyxiEo=
github.com/prometheus/client_model v0.6.3/go.mod h1:gpN5P9S7Rr6Yr92PiQ+Ixvhf6JZEkF1dnxsYL2aPBEM=
github.com/prometheus/common v0.4.1/go.mod h1:TNfzLD0ON7rHzMJeJkieUDPYmFC7Snx/y86RQel1bk4=
github.com/prometheus/common v0.9.1/go.mod h1:yhUN8i9wzaXS3w1O07YhxHEBxD+W35wd8bs7vj7HSQ4=
github.com/prometheus/common v0.10.0/go.mod h1:Tlit/dnDKsSWFlCLTWaA1cyBgKHSMdTB80sz/V91rCo=
@@ -1730,8 +1739,8 @@ github.com/prometheus/procfs v0.0.2/go.mod h1:TjEm7ze935MbeOT/UhFTIMYKhuLP4wbCsT
github.com/prometheus/procfs v0.0.8/go.mod h1:7Qr8sr6344vo1JqZ6HhLceV9o3AJ1Ff+GxbHq6oeK9A=
github.com/prometheus/procfs v0.1.3/go.mod h1:lV6e/gmhEcM9IjHGsFOCxxuZ+z1YqCvr4OA4YeYWdaU=
github.com/prometheus/procfs v0.6.0/go.mod h1:cz+aTbrPOrUb4q7XlbU9ygM+/jj0fzG6c1xBZuNvfVA=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/prometheus/procfs v0.22.0 h1:6q9+/JL9IKAPbCmBrv9n5O5Ty3NKnciV5X7YGw0oics=
github.com/prometheus/procfs v0.22.0/go.mod h1:CvmFr/GVhIjIvWJZW3tgkODBQMRIf0EyWMQLHCHab58=
github.com/pterm/pterm v0.12.27/go.mod h1:PhQ89w4i95rhgE+xedAoqous6K9X+r6aSOI2eFF7DZI=
github.com/pterm/pterm v0.12.29/go.mod h1:WI3qxgvoQFFGKGjGnJR849gU0TsEOvKn5Q8LlY1U7lg=
github.com/pterm/pterm v0.12.30/go.mod h1:MOqLIyMOgmTDz9yorcYbcw+HsgoZo3BQfg2wtl3HEFE=
@@ -1751,18 +1760,18 @@ github.com/quic-go/quic-go v0.59.0 h1:OLJkp1Mlm/aS7dpKgTc6cnpynnD2Xg7C1pwL6vy/SA
github.com/quic-go/quic-go v0.59.0/go.mod h1:upnsH4Ju1YkqpLXC305eW3yDZ4NfnNbmQRCMWS58IKU=
github.com/rabbitmq/amqp091-go v1.14.0 h1:RSaT7aOKt/OrkVUyswPDW29lnRz9psuGmfZFBmLqLek=
github.com/rabbitmq/amqp091-go v1.14.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rclone/Proton-API-Bridge v1.0.4 h1:uGQJRjQC1hVLd5kqLsXc6CWO6oqrVeLoKQYoHapEZDg=
github.com/rclone/Proton-API-Bridge v1.0.4/go.mod h1:VTPBYZotKAeDLlAzxU2O/s14NXk9FxUt9hn1jhH2iY8=
github.com/rclone/go-proton-api v1.0.3 h1:3gBTzR+j0dYiTwtj9yKIdN/aV3W2a8KIPKp0GArojyQ=
github.com/rclone/go-proton-api v1.0.3/go.mod h1:QAlkFfswzrBuxvCORWV8rZdddg52hahMN98CFWoFW1E=
github.com/rclone/rclone v1.75.0 h1:3ARHem4jXWltvl+b0PvDAG8s6J/inHd5BRzfwMRb3W8=
github.com/rclone/rclone v1.75.0/go.mod h1:PGLJUW/WSIJCysALqUcxmaCFyfMXUevf8CbuoOwsAdU=
github.com/rclone/Proton-API-Bridge v1.0.5 h1:K1++Qtk3PvgkiCCiv6Pahju1TMOzKY6VSwiwT7XLAVc=
github.com/rclone/Proton-API-Bridge v1.0.5/go.mod h1:vCeOPhlXzevN0AFojgh1zsjhetiShy/ArvJ/xkFUDWk=
github.com/rclone/go-proton-api v1.0.4 h1:AJW0e9pB4j0hVK4WqyGErFwaI+5MUQWPCtj5FYYxtPg=
github.com/rclone/go-proton-api v1.0.4/go.mod h1:QAlkFfswzrBuxvCORWV8rZdddg52hahMN98CFWoFW1E=
github.com/rclone/rclone v1.75.1 h1:kIxQcoDLj2Gke/gMSHK7OnxhX1Gu1cJBLP1kJZoaFp0=
github.com/rclone/rclone v1.75.1/go.mod h1:4zmMjGatCkSJPRZDpo+7y3xOl8S29EMUyKvZop5mHr4=
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 h1:N/ElC8H3+5XpJzTSTfLsJV/mx9Q9g7kxmchpfZyxgzM=
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475/go.mod h1:bCqnVzQkZxMG4s8nGwiZ5l3QUCyqpo9Y+/ZMZ9VjZe4=
github.com/rdleal/intervalst v1.5.0 h1:SEB9bCFz5IqD1yhfH1Wv8IBnY/JQxDplwkxHjT6hamU=
github.com/rdleal/intervalst v1.5.0/go.mod h1:xO89Z6BC+LQDH+IPQQw/OESt5UADgFD41tYMUINGpxQ=
github.com/redis/go-redis/v9 v9.21.0 h1:FPBE4hhbAke+TLmcY3WkpbDffJEomdqPn3HYiqAtL9E=
github.com/redis/go-redis/v9 v9.21.0/go.mod h1:v/M13XI1PVCDcm01VtPFOADfZtHf8YW3baQf57KlIkA=
github.com/redis/go-redis/v9 v9.22.0 h1:laDvpYXTJtZLloinw1fA5Kqd6HAEH2XKxOkG/PDq2F0=
github.com/redis/go-redis/v9 v9.22.0/go.mod h1:y2g0Wj8rQvuK0ELM+oxSudcLtC09JScs98I/X9gRWY4=
github.com/redis/rueidis v1.0.76 h1:RdDWuvlYBSp+bTrBvaXqJnNEL3VVzsnjo+0psPFgLc4=
github.com/redis/rueidis v1.0.76/go.mod h1:UsfHPSbomB6QAVMk4iiFkzRy0nh9o7scDGa+SitvBY4=
github.com/redis/rueidis/rueidiscompat v1.0.76 h1:7LikbiqCQqCsZXeZ+akgZMnjIV/J0VHih9PIX4gGZC4=
@@ -1800,8 +1809,8 @@ github.com/sagikazarmark/locafero v0.11.0 h1:1iurJgmM9G3PA/I+wWYIOw/5SyBtxapeHDc
github.com/sagikazarmark/locafero v0.11.0/go.mod h1:nVIGvgyzw595SUSUE6tvCp3YYTeHs15MvlmU87WwIik=
github.com/samber/lo v1.53.0 h1:t975lj2py4kJPQ6haz1QMgtId2gtmfktACxIXArw3HM=
github.com/samber/lo v1.53.0/go.mod h1:4+MXEGsJzbKGaUEQFKBq2xtfuznW9oz/WrgyzMzRoM0=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.1 h1:PKK9DyHxif4LZo+uQSgXNqs0jj5+xZwwfKHgph2lxBw=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.1/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU=
github.com/sasha-s/go-deadlock v0.3.1 h1:sqv7fDNShgjcaxkO0JNcOAlr8B9+cV5Ey/OB71efZx0=
github.com/sasha-s/go-deadlock v0.3.1/go.mod h1:F73l+cr82YSh10GxyRI6qZiCgK64VaZjwesgfQ1/iLM=
github.com/schollz/progressbar/v3 v3.19.1 h1:iv8BgwOvdML/S3p84uBpy/IMigv4U9594vPZYa2EdrU=
@@ -1812,10 +1821,10 @@ github.com/seaweedfs/go-fuse/v2 v2.9.4 h1:ACyloiuopdhRSjdLLeSWbsVaemMPskORaRF01T
github.com/seaweedfs/go-fuse/v2 v2.9.4/go.mod h1:zABdmWEa6A0bwaBeEOBUeUkGIZlxUhcdv+V1Dcc/U/I=
github.com/seaweedfs/goexif v2.0.0+incompatible h1:x8pckiT12QQhifwhDQpeISgDfsqmQ6VR4LFPQ64JRps=
github.com/seaweedfs/goexif v2.0.0+incompatible/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/raft v1.2.0 h1:Ez4Hw9ifBbTT7wg54DvGHBjw1vRlTb4roH0TKl0Oj9Y=
github.com/seaweedfs/raft v1.2.0/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/secure-systems-lab/go-securesystemslib v0.10.0 h1:l+H5ErcW0PAehBNrBxoGv1jjNpGYdZ9RcheFkB2WI14=
github.com/secure-systems-lab/go-securesystemslib v0.10.0/go.mod h1:MRKONWmRoFzPNQ9USRF9i1mc7MvAVvF1LlW8X5VWDvk=
github.com/seaweedfs/raft v1.2.1 h1:QgFl/aaPnagpUxYB6Bx+fFss1NyetVVcJmraMmnaQ5Q=
github.com/seaweedfs/raft v1.2.1/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/secure-systems-lab/go-securesystemslib v0.11.0 h1:iuCR9kcMFD4QurdKrGvPLoKZLv9YvwPYVr0473BdtFs=
github.com/secure-systems-lab/go-securesystemslib v0.11.0/go.mod h1:+PMOTjUGwHj2vcZ+TFKlb1tXRbrdWE1LYDT5i9JC80Q=
github.com/sergi/go-diff v1.0.0/go.mod h1:0CfEIISq7TuYL3j771MWULgwwjU+GofnZX9QAmXWZgo=
github.com/sergi/go-diff v1.1.0/go.mod h1:STckp+ISIX8hZLjrqAeVduY0gWCT9IjLuqbuNXdaHfM=
github.com/sergi/go-diff v1.2.0 h1:XU+rvMAioB0UC3q1MFrIQy4Vo5/4VsRDQQXHsEya6xQ=
@@ -1824,10 +1833,10 @@ github.com/shibumi/go-pathspec v1.3.0 h1:QUyMZhFo0Md5B8zV8x2tesohbb5kfbpTi9rBnKh
github.com/shibumi/go-pathspec v1.3.0/go.mod h1:Xutfslp817l2I1cZvgcfeMQJG5QnU2lh5tVaaMCl3jE=
github.com/shirou/gopsutil/v4 v4.26.7 h1:IXzpHz/dkMRYAhKkOXr1HB6SuzWU3eoyyeWe7g3bNZc=
github.com/shirou/gopsutil/v4 v4.26.7/go.mod h1:5O9FjBiXoTDFatIWjZZosqj4pV0DRtLx598xGbBehzM=
github.com/sigstore/sigstore v1.10.4 h1:ytOmxMgLdcUed3w1SbbZOgcxqwMG61lh1TmZLN+WeZE=
github.com/sigstore/sigstore v1.10.4/go.mod h1:tDiyrdOref3q6qJxm2G+JHghqfmvifB7hw+EReAfnbI=
github.com/sigstore/sigstore-go v1.1.4 h1:wTTsgCHOfqiEzVyBYA6mDczGtBkN7cM8mPpjJj5QvMg=
github.com/sigstore/sigstore-go v1.1.4/go.mod h1:2U/mQOT9cjjxrtIUeKDVhL+sHBKsnWddn8URlswdBsg=
github.com/sigstore/sigstore v1.10.8 h1:1Mgkxvkw4AXMfIP1DOjc6kw0GkUgA8pGVpveN/EfOq4=
github.com/sigstore/sigstore v1.10.8/go.mod h1:f9+B/4iaYimvUkySyb2mvc73n3RLqNn24grHZM/ET8M=
github.com/sigstore/sigstore-go v1.2.1 h1:YWP/rDbBaEBvtbkj6xtwsSj38ZCFEhTVVadNOXjVe3A=
github.com/sigstore/sigstore-go v1.2.1/go.mod h1:I8BqVwAb/SaQJ5pBu5IDFY+ksq8O/1/kCag8XUgrsko=
github.com/sirupsen/logrus v1.2.0/go.mod h1:LxeOpSwHxABJmUn/MG1IvRgCAasNZTLOkJPxbbu5VWo=
github.com/sirupsen/logrus v1.4.2/go.mod h1:tLMulIdttU9McNUspp0xgXVQah82FyeX6MwdIuYE2rE=
github.com/sirupsen/logrus v1.5.0/go.mod h1:+F7Ogzej0PZc/94MaYx/nvG9jOFMD2osvC3s+Squfpo=
@@ -1911,10 +1920,10 @@ github.com/tarantool/go-option v1.1.0 h1:ShoOhNsdL41sRpm4hXCRDjV8H0WzPkd4UnKhLKb
github.com/tarantool/go-option v1.1.0/go.mod h1:hMr9z2JXOWlgdCBpCPSL2nwp8718GKYvNBJ+ZuzJbCo=
github.com/tarantool/go-tarantool/v3 v3.0.1 h1:vaUX4xmVmXh2dIJ/LqlX1MXK3iYqAqV6YiE54Wwl/qg=
github.com/tarantool/go-tarantool/v3 v3.0.1/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/testcontainers/testcontainers-go v0.43.0 h1:oEQx5MW2DGd9z3AeEQfB2lPM0eLs7ztyaGRu75bFo5A=
github.com/testcontainers/testcontainers-go v0.43.0/go.mod h1:+VxkT2NQnKOZPKi6praMuMKYHYyOGXr0XSBSlSMCzFo=
github.com/testcontainers/testcontainers-go/modules/compose v0.42.0 h1:+t1ZN31TD36cwxmeLqGwe7wIdvblBm0Z+vlj4SX8Mv0=
github.com/testcontainers/testcontainers-go/modules/compose v0.42.0/go.mod h1:CfMpouDHqNTCHC8CijEURU2ZotTV3QhH6pXd48s6ofk=
github.com/testcontainers/testcontainers-go v0.44.0 h1:/Fwh6HY1mIikhnm9e7HwoxGycx0lzRAE0f5VQpjFxzI=
github.com/testcontainers/testcontainers-go v0.44.0/go.mod h1:IcnwQrYTO86xHXu5bvMaBH7ATlbS3Qn1M1QWW3c66rE=
github.com/testcontainers/testcontainers-go/modules/compose v0.44.0 h1:8YcW51jhgpkkiRVe10Wj9TCBthJmoNpU2fK5WSf7TQ8=
github.com/testcontainers/testcontainers-go/modules/compose v0.44.0/go.mod h1:0du5YaH5n8Of4bDM/RwIt421Q14kvqrXZ0mB8EjuMoo=
github.com/the42/cartconvert v0.0.0-20131203171324-aae784c392b8 h1:I4DY8wLxJXCrMYzDM6lKCGc3IQwJX0PlTLsd3nQqI3c=
github.com/the42/cartconvert v0.0.0-20131203171324-aae784c392b8/go.mod h1:fWO/msnJVhHqN1yX6OBoxSyfj7TEj1hHiL8bJSQsK30=
github.com/tiancaiamao/gp v0.0.0-20221230034425-4025bc8a4d4a h1:J/YdBZ46WKpXsxsW93SG+q0F8KI+yFrcIDT4c/RNoc4=
@@ -1943,8 +1952,8 @@ github.com/tmthrgd/go-hex v0.0.0-20190904060850-447a3041c3bc h1:9lRDQMhESg+zvGYm
github.com/tmthrgd/go-hex v0.0.0-20190904060850-447a3041c3bc/go.mod h1:bciPuU6GHm1iF1pBvUfxfsH0Wmnc2VbpgvbI9ZWuIRs=
github.com/tonistiigi/dchapes-mode v0.0.0-20250318174251-73d941a28323 h1:r0p7fK56l8WPequOaR3i9LBqfPtEdXIQbUTzT55iqT4=
github.com/tonistiigi/dchapes-mode v0.0.0-20250318174251-73d941a28323/go.mod h1:3Iuxbr0P7D3zUzBMAZB+ois3h/et0shEz0qApgHYGpY=
github.com/tonistiigi/fsutil v0.0.0-20251211185533-a2aa163d723f h1:Z4NEQ86qFl1mHuCu9gwcE+EYCwDKfXAYXZbdIXyxmEA=
github.com/tonistiigi/fsutil v0.0.0-20251211185533-a2aa163d723f/go.mod h1:BKdcez7BiVtBvIcef90ZPc6ebqIWr4JWD7+EvLm6J98=
github.com/tonistiigi/fsutil v0.0.0-20260609174605-b61e79c0c046 h1:j29MScUISj0S98EHdM6/pzKgqppfswl9OZ7OVpnQdzE=
github.com/tonistiigi/fsutil v0.0.0-20260609174605-b61e79c0c046/go.mod h1:K5zrLch9UaSGNiek5XHZeqZUf1zPWJHqDfLIcnpquQ4=
github.com/tonistiigi/go-csvvalue v0.0.0-20240814133006-030d3b2625d0 h1:2f304B10LaZdB8kkVEaoXvAMVan2tl9AiK4G0odjQtE=
github.com/tonistiigi/go-csvvalue v0.0.0-20240814133006-030d3b2625d0/go.mod h1:278M4p8WsNh3n4a1eqiFcV2FGk7wE5fwUpUom9mK9lE=
github.com/tonistiigi/units v0.0.0-20180711220420-6950e57a87ea h1:SXhTLE6pb6eld/v/cCndK0AMpt1wiVFb/YYmqB3/QG0=
@@ -1956,8 +1965,8 @@ github.com/tsuna/gohbase v0.0.0-20201125011725-348991136365/go.mod h1:zj0GJHGvyf
github.com/tv42/httpunix v0.0.0-20150427012821-b75d8614f926/go.mod h1:9ESjWnEqriFuLhtthL60Sar/7RFoluCcXsuvEwTV5KM=
github.com/twitchyliquid64/golang-asm v0.15.1 h1:SU5vSMR7hnwNxj24w34ZyCi/FmDZTkS4MhqMhdFk5YI=
github.com/twitchyliquid64/golang-asm v0.15.1/go.mod h1:a1lVb/DtPvCB8fslRZhAngC2+aY1QWCk3Cedj/Gdt08=
github.com/twmb/avro v1.7.2 h1:cmrEBRSbELRqsg/dRkQvVWuOaR2EfGifHIt/2iJ9lfI=
github.com/twmb/avro v1.7.2/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/avro v1.8.0 h1:UMWLg+nH4P3yad5Om7yFSohYLy2RG1s7BcFFiOvmK9Q=
github.com/twmb/avro v1.8.0/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/murmur3 v1.1.8 h1:8Yt9taO/WN3l08xErzjeschgZU2QSrwm1kclYq+0aRg=
github.com/twmb/murmur3 v1.1.8/go.mod h1:Qq/R7NUyOfr65zD+6Q5IHKsJLwP7exErjN6lyyq3OSQ=
github.com/twpayne/go-geom v1.6.1 h1:iLE+Opv0Ihm/ABIcvQFGIiFBXd76oBIar9drAwHFhR4=
@@ -2096,38 +2105,38 @@ go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 h1:NmLfL734pJhM0JKaYd2Y28+nY9dPRWYAAbxhRCrKXPw=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0/go.mod h1:tNAsgd8avTGke1+MndXlU5Cru4PQ9Ai/cCNWQv/ZJ/s=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.68.0 h1:0Qx7VGBacMm9ZENQ7TnNObTYI4ShC+lHI16seduaxZo=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.68.0/go.mod h1:Sje3i3MjSPKTSPvVWCaL8ugBzJwik3u4smCjUeuupqg=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.63.0 h1:2pn7OzMewmYRiNtv1doZnLo3gONcnMHlFnmOR8Vgt+8=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.63.0/go.mod h1:rjbQTDEPQymPE0YnRQp9/NuPwwtL0sesz/fnqRW/v84=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 h1:2yEATaop1/a1I4psnSLgWVPLWwCzkqWakgJy7xTDVy0=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0/go.mod h1:D7J12YRapIekYyPWgGPlA/23pRmpSEZC5xJC/TTLI9U=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0 h1:MCcYL7J6Vt/X0kjqbMZkekCmwsurbQRbL69vkiye2lk=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0/go.mod h1:3jnStNwSufK+f5ktjL4EPcwtig4rtd81NS70lqHuXl8=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/otel v1.44.0 h1:JjwHmHpA4iZ3wBxluu2fbbE7j4kqlE8jXyAyPXH7HqU=
go.opentelemetry.io/otel v1.44.0/go.mod h1:BMgjTHL9WPRlRjL2oZCBTL4whCGtXch2H4BhOPIAyYc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.42.0 h1:MdKucPl/HbzckWWEisiNqMPhRrAOQX8r4jTuGr636gk=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.42.0/go.mod h1:RolT8tWtfHcjajEH5wFIZ4Dgh5jpPdFXYV9pTAk/qjc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.43.0 h1:w1K+pCJoPpQifuVpsKamUdn9U0zM3xUziVOqsGksUrY=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.43.0/go.mod h1:HBy4BjzgVE8139ieRI75oXm3EcDN+6GhD88JT1Kjvxg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.43.0 h1:88Y4s2C8oTui1LGM6bTWkw0ICGcOLCAI5l6zsD1j20k=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.43.0/go.mod h1:Vl1/iaggsuRlrHf/hfPJPvVag77kKyvrLeD10kpMl+A=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.42.0 h1:zWWrB1U6nqhS/k6zYB74CjRpuiitRtLLi68VcgmOEto=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.42.0/go.mod h1:2qXPNBX1OVRC0IwOnfo1ljoid+RD0QK3443EaqVlsOU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.43.0 h1:3iZJKlCZufyRzPzlQhUIWVmfltrXuGyfjREgGP3UUjc=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.43.0/go.mod h1:/G+nUPfhq2e+qiXMGxMwumDrP5jtzU+mWN7/sjT2rak=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 h1:SUplec5dp06reu1zaXmOXdvqH398taqrDXqUl99jxSc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0/go.mod h1:ho2g4N+ane+swq5I/VBkKWnRDY4kUINH3FuqyZqX/Ug=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 h1:lgh3PiVrRUWMLOVSkQicxzZll5NjF1r+AtsX1XRIHw0=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0/go.mod h1:5Cnhth3m/AgOeTgE3ex12pPmiu/gGtZit03kSzx9X7s=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 h1:hqxVTu/GtBF+vJ8d1fzW7fRxZFvgoDjWcxwwCaFDYpU=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0/go.mod h1:z5fVEF4X5v0ESvlJqBrrFlBVoj5EQuefZpzsu7R+x5Q=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 h1:s0n95ya5tOG03exJ5JySOdJFtwGo4ZQ+KeY7Zro4CLI=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0/go.mod h1:m9wRxtKA2MZ1HcnNC4BKI+9aYe434qRZTCvI7QGUN7Y=
go.opentelemetry.io/otel/metric v1.44.0 h1:1w0gILTcHdr3YI+ixLyjemwrVnsMURbTZFrSYCdDdmc=
go.opentelemetry.io/otel/metric v1.44.0/go.mod h1:8O7hanEPBNgEMmybD3s2VBKcgWOCsA6tzHBPODAiquo=
go.opentelemetry.io/otel/metric/x v0.66.0 h1:YkCrx1zLOChi9ZcZ6euupOcsgzbVlec7D/xoEU1+cTA=
go.opentelemetry.io/otel/metric/x v0.66.0/go.mod h1:d1+BDj9t96do0/1LoU1ayfCv79ZgNE41qbhBvnMOBZk=
go.opentelemetry.io/otel/sdk v1.44.0 h1:nHYwb9lK+fJPU/dnT6s7W7Z8itMWyqrnVfbheVYrZ58=
go.opentelemetry.io/otel/sdk v1.44.0/go.mod h1:Osuydd3Se74nqjAKxid74N5eC+jfEqfTegHRnq58oK0=
go.opentelemetry.io/otel/sdk/metric v1.44.0 h1:3LlKgI+VjbVsjNRFZJZAJ30WjXC5VkNRks6si09iEfI=
go.opentelemetry.io/otel/sdk/metric v1.44.0/go.mod h1:5B5pMARnXxKhltooO4xUuCBorl65a4EpnTalObqOigA=
go.opentelemetry.io/otel/trace v1.44.0 h1:jxF5CsGYCe74MCRx2X4g7WsY/VBKRqqpNvXlX/6gtIk=
go.opentelemetry.io/otel/trace v1.44.0/go.mod h1:oLl1jrMQAVo6v3GAggN+1VH9VIz9iUSvW53sW1Q8PIE=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/metric/x v0.67.0 h1:PcicCNZFkZ4bXfSooXdo3WN7RBOVOtjVdo1wD358Uns=
go.opentelemetry.io/otel/metric/x v0.67.0/go.mod h1:FBjCWZe6wgcqxcMtjdGiClDKXb2YxxXii0CXftE4QtI=
go.opentelemetry.io/otel/sdk v1.45.0 h1:4VVSMgQ83dUgW2aoX5f6JgLvHwIvzcuLnF9lUdCSpCw=
go.opentelemetry.io/otel/sdk v1.45.0/go.mod h1:Sr40LgXV7DsKMMJMKOhUWOgMWTfAaqvm2kF0g7ilwuA=
go.opentelemetry.io/otel/sdk/metric v1.45.0 h1:oVFszMfyj1Am6s24Vtc7wBb8BKLcwepJjNEYILuiE3o=
go.opentelemetry.io/otel/sdk/metric v1.45.0/go.mod h1:vUWUxDZvu1WVRj8JA8S0AdhsPrZoDpA2DdZauIh4mDA=
go.opentelemetry.io/otel/trace v1.45.0 h1:l/mP6Uv7oNO7/TblbhpbgMidxhq1uO/rPsikOyVhxag=
go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbiv+jhs/CxI8bc=
go.opentelemetry.io/proto/otlp v0.7.0/go.mod h1:PqfVotwruBrMGOCsRd/89rSnXhoiJIqeYNgFYFoEGnI=
go.opentelemetry.io/proto/otlp v0.15.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v0.19.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
@@ -2156,8 +2165,8 @@ go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
go.yaml.in/yaml/v4 v4.0.0-rc.4 h1:UP4+v6fFrBIb1l934bDl//mmnoIZEDK0idg1+AIvX5U=
go.yaml.in/yaml/v4 v4.0.0-rc.4/go.mod h1:aZqd9kCMsGL7AuUv/m/PvWLdg5sjJsZ4oHDEnfPPfY0=
go.yaml.in/yaml/v4 v4.0.0-rc.6 h1:1h7H1ohdUh93/FyE4YaDa1Zh64K6VVbjF4K6WUxMtH4=
go.yaml.in/yaml/v4 v4.0.0-rc.6/go.mod h1:aZqd9kCMsGL7AuUv/m/PvWLdg5sjJsZ4oHDEnfPPfY0=
gocloud.dev v0.46.0 h1:niIuZwSjMtBx8K+ITB2s5kZullB13PGOS2ZoQPZxQ4Q=
gocloud.dev v0.46.0/go.mod h1:ACQe+2qO+hEO+pdcvvsM+RB63r8TyGD1W3ESCLFyzvM=
gocloud.dev/pubsub/natspubsub v0.46.0 h1:cTmJnWOmhb3dYRzX7PvNkq2m5mQdBLFFxZbl/21pfjk=
@@ -2184,8 +2193,8 @@ golang.org/x/crypto v0.6.0/go.mod h1:OFC/31mSvZgRz0V1QTNCzfAI1aIRzbiufJtkMIlEp58
golang.org/x/crypto v0.7.0/go.mod h1:pYwdfH91IfpZVANVyUOhSIPZaFoJGxTFbZhFTx+dXZU=
golang.org/x/crypto v0.13.0/go.mod h1:y6Z2r+Rw4iayiXXAIxJIDAJ1zMW4yaTpebo8fPOliYc=
golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf4=
golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M=
golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis=
golang.org/x/crypto v0.56.0 h1:GUh5Ii4J5jtcseSMiRqr1jXCNHoxjeV9Fmekc2oLy6Y=
golang.org/x/crypto v0.56.0/go.mod h1:OMW5y6CY9l38uPLmxU6l6pwcXp1obtLo3e6gT7gQR2I=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2216,8 +2225,8 @@ golang.org/x/image v0.0.0-20210607152325-775e3b0c77b9/go.mod h1:023OzeP/+EPmXeap
golang.org/x/image v0.0.0-20210628002857-a66eb6448b8d/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20211028202545-6944b10bf410/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20220302094943-723b81ca9867/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.44.0 h1:+tDekMZED9+LrtB3G5xzRggpVh9CARjZqROla3R3R+I=
golang.org/x/image v0.44.0/go.mod h1:V8K3KE9KKKE+pLpQDOeN18w9oacNSvy1tDOirTu4xtY=
golang.org/x/image v0.46.0 h1:b1+oYj0Jbp6K5MDT4i4/eZpYlk3V8SJhhDKh6LBHAyQ=
golang.org/x/image v0.46.0/go.mod h1:3B3W05VGVQyuXucLINLjXKrqISASfi4Xj+iCVkLMwew=
golang.org/x/lint v0.0.0-20181026193005-c67002cb31c3/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
golang.org/x/lint v0.0.0-20190227174305-5b3e6a55c961/go.mod h1:wehouNa3lNwaWXcvxsM5YxQ5yQlVC4a0KAMCusXpPoU=
golang.org/x/lint v0.0.0-20190301231843-5614ed5bae6f/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
@@ -2249,8 +2258,8 @@ golang.org/x/mod v0.8.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.9.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.12.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.13.0/go.mod h1:hTbmBsO62+eylJbnUtE2MGJUyE7QWk4xUqPFrRgJ+7c=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c=
golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o=
golang.org/x/net v0.0.0-20180724234803-3673e40ba225/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180826012351-8a410e7b638d/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
@@ -2365,8 +2374,8 @@ golang.org/x/sync v0.0.0-20220929204114-8fcdb60fdcc0/go.mod h1:RxMgew5VJxzue5/jJ
golang.org/x/sync v0.1.0/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.3.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.4.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.0.0-20180810173357-98c5dad5d1a0/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180830151530-49385e6e1522/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180905080454-ebe1bf3edb33/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
@@ -2468,8 +2477,8 @@ golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.8.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.12.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.13.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210220032956-6a3ed077a48d/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210615171337-6886f2dfbf5b/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
@@ -2502,8 +2511,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI=
golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2580,8 +2589,8 @@ golang.org/x/tools v0.6.0/go.mod h1:Xwgl3UAJ/d3gWutnCtw505GrjyAbvKui8lOU390QaIU=
golang.org/x/tools v0.7.0/go.mod h1:4pg6aUX35JBAogB10C9AtvVL+qowtN4pT3CGSQex14s=
golang.org/x/tools v0.13.0/go.mod h1:HvlwmtVNQAhOuCjW7xxvovg8wbNq7LwfXh/k7wXUl58=
golang.org/x/tools v0.14.0/go.mod h1:uYBEerGOWcJyEORxN+Ek8+TT266gXkNlHdJBwexUsBg=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
golang.org/x/tools/godoc v0.1.0-deprecated h1:o+aZ1BOj6Hsx/GBdJO/s815sqftjSnrZZwyYTHODvtk=
golang.org/x/tools/godoc v0.1.0-deprecated/go.mod h1:qM63CriJ961IHWmnWa9CjZnBndniPt4a3CK0PVB9bIg=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -2659,8 +2668,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.294.0 h1:8gASjJxdtcIieB3OqbkLcF0FfbXVNqKtU5iozD1ssvA=
google.golang.org/api v0.294.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/api v0.297.0 h1:WktxTsnnx0yZNnsR6j0q6hR21RnnK81FHTOPy/ux4OE=
google.golang.org/api v0.297.0/go.mod h1:S4m8x0M6OkQpkOzGk1y9JG2sm4fFQrMh6dxzjCTszhE=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.45"
appVersion: "4.47"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.45.0
version: 4.47.1
+22 -1
View File
@@ -286,7 +286,7 @@ metadata:
app.kubernetes.io/component: s3
stringData:
# this key must be an inline json config file
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read","List"]}]}'
```
#### Source S3 credentials from an existing Secret
@@ -363,6 +363,27 @@ If `adminPassword` is empty or not set, the admin interface runs without authent
As an alternative, a kubernetes Secret can be used (`admin.secret.existingSecret`).
### Admin listen address
Since SeaweedFS 4.46, `weed admin` defaults to listening on loopback (`127.0.0.1`).
The chart's httpGet readiness/liveness probes dial the pod IP, so the admin
server must bind a non-loopback address for the probes to succeed. The chart
therefore passes `-ip={{ .Values.admin.ip }}`, defaulting `admin.ip` to `0.0.0.0`
(the pre-4.46 behaviour of listening on all interfaces).
Binding a non-loopback address requires authentication: `weed admin` refuses to
start on a non-loopback address without `-adminPassword`, so the chart fails at
render time if `admin.ip` is non-loopback and authentication is not configured via
`admin.secret.adminPassword`, `admin.secret.existingSecret`, or
`WEED_ADMIN_PASSWORD` supplied through `admin.extraEnvironmentVars` /
`admin.secretExtraEnvironmentVars`. The whole `127.0.0.0/8` range and `::1` are
treated as loopback (matching `weed admin`); `localhost` is treated as
non-loopback. Set `admin.ip` to a loopback address only if you also replace the
httpGet probes (e.g. with an `exec` probe that checks `127.0.0.1`).
The `-ip` flag requires SeaweedFS 4.46 or newer; pinning `admin.imageOverride`
to an older image is not supported with this chart version.
### Admin Data Persistence
The admin component can store configuration and maintenance data. You can configure storage in several ways:
@@ -0,0 +1,7 @@
# Admin install: exercises the admin StatefulSet, which passes -ip (default
# 0.0.0.0) and therefore requires authentication to bind a non-loopback address.
admin:
enabled: true
secret:
adminUser: "admin"
adminPassword: "ci-admin-password"
@@ -23,6 +23,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: admin
{{- with .Values.admin.ingress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if and (semverCompare ">=1.18-0" .Capabilities.KubeVersion.GitVersion) .Values.admin.ingress.className }}
ingressClassName: {{ .Values.admin.ingress.className | quote }}
@@ -6,6 +6,11 @@
{{- if and (not .Values.admin.masters) (not .Values.global.seaweedfs.masterServer) (not .Values.master.enabled) }}
{{- fail "admin.masters or global.seaweedfs.masterServer must be set if master.enabled is false" -}}
{{- end }}
{{- $adminAuthEnabled := include "seaweedfs.admin.authEnabled" . }}
{{- $adminIp := .Values.admin.ip | default "0.0.0.0" }}
{{- if and (not (include "seaweedfs.admin.isLoopbackIp" $adminIp)) (ne $adminAuthEnabled "true") }}
{{- fail (printf "admin.ip is set to %q (non-loopback) but admin authentication is not configured. Since `weed admin` 4.46 refuses to bind a non-loopback address without authentication, the admin container would exit on startup. Set admin.secret.adminPassword or admin.secret.existingSecret, or supply WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars / admin.secretExtraEnvironmentVars, or set admin.ip to a loopback address such as 127.0.0.1 (note: a loopback bind makes the chart's httpGet readiness/liveness probes fail)." $adminIp) -}}
{{- end }}
apiVersion: apps/v1
kind: StatefulSet
metadata:
@@ -162,6 +167,7 @@ spec:
-v={{ .Values.global.seaweedfs.loggingLevel }} \
{{- end }}
admin \
-ip={{ .Values.admin.ip | default "0.0.0.0" }} \
-port={{ .Values.admin.port }} \
-port.grpc={{ .Values.admin.grpcPort }} \
{{- if or (eq .Values.admin.data.type "hostPath") (eq .Values.admin.data.type "persistentVolumeClaim") (eq .Values.admin.data.type "emptyDir") (eq .Values.admin.data.type "existingClaim") }}
@@ -44,6 +44,13 @@ spec:
{{- with .Values.allInOne.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $existingS3ConfigSecret := or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret }}
{{- if $existingS3ConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace $existingS3ConfigSecret) | default dict }}
checksum/s3config: {{ $configSecret | toYaml | sha256sum }}
{{- else }}
checksum/s3config: {{ include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum }}
{{- end }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.allInOne.restartPolicy }}
{{- if .Values.allInOne.affinity }}
@@ -24,6 +24,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- with .Values.filer.ingresses.grpc.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if $securityEnabled }}
tls:
@@ -29,6 +29,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- with .Values.filer.ingresses.http.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if and (semverCompare ">=1.18-0" .Capabilities.KubeVersion.GitVersion) .Values.filer.ingresses.http.className }}
ingressClassName: {{ .Values.filer.ingresses.http.className | quote }}
@@ -87,6 +90,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- with .Values.filer.ingresses.grpc.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if and (semverCompare ">=1.18-0" .Capabilities.KubeVersion.GitVersion) .Values.filer.ingresses.grpc.className }}
ingressClassName: {{ .Values.filer.ingresses.grpc.className | quote }}
@@ -21,6 +21,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: master
{{- with .Values.master.ingress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
ingressClassName: {{ .Values.master.ingress.className | quote }}
tls:
@@ -42,6 +42,12 @@ spec:
{{- with .Values.s3.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- if .Values.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.s3.existingConfigSecret) | default dict }}
checksum/s3config: {{ $configSecret | toYaml | sha256sum }}
{{- else }}
checksum/s3config: {{ include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum }}
{{- end }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.s3.restartPolicy }}
{{- if .Values.s3.affinity }}
@@ -41,6 +41,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3-iceberg
{{- with .Values.s3.icebergIngress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if .Values.s3.icebergIngress.className }}
ingressClassName: {{ .Values.s3.icebergIngress.className | quote }}
@@ -32,6 +32,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3
{{- with .Values.s3.ingress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
ingressClassName: {{ .Values.s3.ingress.className | quote }}
tls:
@@ -41,6 +41,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3-lance
{{- with .Values.s3.lanceIngress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if .Values.s3.lanceIngress.className }}
ingressClassName: {{ .Values.s3.lanceIngress.className | quote }}
@@ -60,7 +60,7 @@ stringData:
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read","List"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
{{ toJson .Values.filer.s3.auditLogConfig | nindent 4 }}
@@ -88,6 +88,43 @@ true
{{- end -}}
{{- end -}}
{{/* Classify an admin bind address as loopback, mirroring weed admin's
isLoopbackIp (net.ParseIP + IsLoopback). Helm templates cannot call
net.ParseIP, so we approximate: valid IPv4 addresses in 127.0.0.0/8
(validated via regex to reject malformed values like "127.not-an-ip")
and the IPv6 loopback "::1" / its expanded form "0:0:0:0:0:0:0:1" are
loopback. Hostnames (e.g. "localhost") and wildcard addresses
("0.0.0.0", "::") are non-loopback, matching the binary, which
treats unparseable hostnames as non-loopback to be safe. Other IPv6
loopback representations are not matched; the binary's own runtime
validation is the authoritative guard. */}}
{{- define "seaweedfs.admin.isLoopbackIp" -}}
{{- $ip := toString . -}}
{{- if or (regexMatch "^127\\.[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}$" $ip) (eq $ip "::1") (eq $ip "0:0:0:0:0:0:0:1") -}}
true
{{- end -}}
{{- end -}}
{{/* Whether admin authentication is enabled from any supported source:
admin.secret (adminPassword or existingSecret), or WEED_ADMIN_PASSWORD
supplied via extraEnvironmentVars / secretExtraEnvironmentVars (which
weed admin picks up through viper's AutomaticEnv). A secret-backed
entry counts as enabled even though the chart cannot read its value. */}}
{{- define "seaweedfs.admin.authEnabled" -}}
{{- if or .Values.admin.secret.existingSecret .Values.admin.secret.adminPassword -}}
true
{{- else -}}
{{- $merged := dict -}}
{{- $_ := include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.admin "target" $merged) -}}
{{- $envPassword := index $merged "WEED_ADMIN_PASSWORD" -}}
{{- if or (kindIs "map" $envPassword) (hasKey (.Values.admin.secretExtraEnvironmentVars | default dict) "WEED_ADMIN_PASSWORD") -}}
true
{{- else if and $envPassword (ne (toString $envPassword) "") -}}
true
{{- end -}}
{{- end -}}
{{- end -}}
{{/* Return the proper filer image */}}
{{- define "seaweedfs.filer.image" -}}
{{- if .Values.filer.imageOverride -}}
@@ -97,6 +97,19 @@ data:
cert = "/usr/local/share/ca-certificates/worker/tls.crt"
key = "/usr/local/share/ca-certificates/worker/tls.key"
{{- /* S3 gRPC server identity (weed s3 -port.grpc, default httpPort+10000).
Without this section, LoadServerTLS("grpc.s3") returns nil and the S3
server serves its internal gRPC (IAM cache propagation, s3 lifecycle
delete RPCs) as plaintext, while peers dial it with mTLS credentials:
"tls: first record does not look like a TLS handshake".
Always uses the internal CA-signed client certificate (already mounted
on s3 pods by s3-deployment.yaml). Deliberately NOT s3.tlsSecret: that
secret is for the public HTTPS listener and may be issued by a public
CA which internal gRPC peers (trusting only grpc.ca) would reject. */}}
[grpc.s3]
cert = "/usr/local/share/ca-certificates/client/tls.crt"
key = "/usr/local/share/ca-certificates/client/tls.key"
# use this for any place needs a grpc client
# i.e., "weed backup|benchmark|filer.copy|filer.replicate|mount|s3|upload"
[grpc.client]
@@ -27,6 +27,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: volume
{{- with .Values.volume.ingress.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if and (semverCompare ">=1.18-0" .Capabilities.KubeVersion.GitVersion) .Values.volume.ingress.className }}
ingressClassName: {{ .Values.volume.ingress.className | quote }}
+23 -1
View File
@@ -55,7 +55,7 @@ global:
gatewayHost: null
gatewayPort: null
additionalLabels: {}
# if enabled will use global.seaweedfs.replicationPlacement and override master & filer defaultReplicaPlacement config
# if enabled will use global.seaweedfs.replicationPlacement and override master.defaultReplication & filer.defaultReplicaPlacement config
enableReplication: false
# replication type is XYZ:
# X number of replica in other data centers
@@ -276,6 +276,7 @@ master:
host: "master.seaweedfs.local"
path: "/sw-master/?(.*)"
pathType: ImplementationSpecific
labels: {}
annotations: {}
# nginx.ingress.kubernetes.io/auth-type: "basic"
# nginx.ingress.kubernetes.io/auth-secret: "default/ingress-basic-auth-secret"
@@ -600,6 +601,7 @@ volume:
host: "volume.seaweedfs.local"
path: "/"
pathType: Prefix
labels: {}
annotations:
nginx.ingress.kubernetes.io/app-root: /ui/index.html
# nginx.ingress.kubernetes.io/use-regex: "true"
@@ -858,6 +860,7 @@ filer:
host: "seaweedfs.cluster.local"
path: "/sw-filer/?(.*)"
pathType: ImplementationSpecific
labels: {}
annotations: {}
# nginx.ingress.kubernetes.io/backend-protocol: GRPC
# nginx.ingress.kubernetes.io/auth-type: "basic"
@@ -884,6 +887,7 @@ filer:
# whole host, not the HTTP UI's regex path.
path: "/"
pathType: Prefix
labels: {}
annotations:
# Ingress terminates TLS and re-originates gRPC (HTTP/2) to the filer.
nginx.ingress.kubernetes.io/backend-protocol: "GRPC"
@@ -1180,6 +1184,7 @@ s3:
host: "seaweedfs.cluster.local"
path: "/"
pathType: Prefix
labels: {}
# additional ingress annotations for the s3 endpoint
annotations: {}
tls: []
@@ -1205,6 +1210,7 @@ s3:
host: "seaweedfs-iceberg.cluster.local"
path: "/"
pathType: Prefix
labels: {}
annotations: {}
tls: []
@@ -1214,6 +1220,7 @@ s3:
host: "seaweedfs-lance.cluster.local"
path: "/"
pathType: Prefix
labels: {}
annotations: {}
tls: []
@@ -1321,6 +1328,20 @@ admin:
replicas: 1
port: 23646 # Default admin port
grpcPort: 33646 # Default gRPC port for worker connections
# IP address the admin server listens on. Since `weed admin` 4.46 defaults to
# loopback (127.0.0.1), the chart must bind a non-loopback address for the
# kubelet's httpGet readiness/liveness probes (which dial the pod IP) to ever
# succeed. "0.0.0.0" restores the pre-4.46 behaviour of listening on all
# interfaces. A non-loopback address requires authentication: set
# admin.secret.adminPassword or admin.secret.existingSecret, or supply
# WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars /
# admin.secretExtraEnvironmentVars; otherwise the admin container will exit
# with a clear error rather than silently staying unready. The whole
# 127.0.0.0/8 range and ::1 are treated as loopback (matching weed admin).
# Set to a loopback address only if you also replace the httpGet probes.
# Note: the -ip flag requires SeaweedFS 4.46 or newer; pinning
# admin.imageOverride to an older image is not supported with this chart.
ip: "0.0.0.0"
loggingOverrideLevel: null
# Admin authentication
@@ -1445,6 +1466,7 @@ admin:
host: "admin.seaweedfs.local"
path: "/"
pathType: Prefix
labels: {}
annotations: {}
tls: []
File diff suppressed because it is too large Load Diff

After

Width:  |  Height:  |  Size: 53 KiB

+49 -2
View File
@@ -13,6 +13,9 @@ service SeaweedFiler {
rpc LookupDirectoryEntry (LookupDirectoryEntryRequest) returns (LookupDirectoryEntryResponse) {
}
rpc LookupDirectoryEntries (LookupDirectoryEntriesRequest) returns (LookupDirectoryEntriesResponse) {
}
rpc ListEntries (ListEntriesRequest) returns (stream ListEntriesResponse) {
}
@@ -250,9 +253,11 @@ message CreateEntryRequest {
repeated int32 signatures = 5;
bool skip_check_parent_directory = 6;
// Optional precondition evaluated against the current entry atomically with
// the write, under the filer's per-path lock. The caller must route the
// key's writes to this entry's owner filer for the check to be authoritative.
// the write, under the filer's per-path lock.
WriteCondition condition = 7;
// Set on a create a filer forwarded to the entry's ring owner, so the owner
// applies it locally instead of forwarding again. Clients leave it unset.
bool is_moved = 8;
}
// WriteCondition is the precondition the filer evaluates against the existing
@@ -561,6 +566,9 @@ message AssignVolumeResponse {
string error = 8;
Location location = 9;
repeated Location replicas = 10;
// fsync is the storage rule's fsync decision for the assigned path, so the
// client can carry it onto the volume server upload request.
bool fsync = 11;
}
message LookupVolumeRequest {
@@ -576,6 +584,7 @@ message Location {
string public_url = 2;
uint32 grpc_port = 3;
string data_center = 4;
bool data_in_remote = 5;
}
message LookupVolumeResponse {
map<string, Locations> locations_map = 1;
@@ -662,6 +671,7 @@ message SubscribeMetadataResponse {
int64 ts_ns = 3;
repeated SubscribeMetadataResponse events = 4; // batch of additional events (backlog catch-up)
repeated LogFileChunkRef log_file_refs = 5; // log file chunk refs for client direct-read
int64 flushed_ts_ns = 6; // local log-buffer flush watermark: everything at or below it is on disk
}
message ListMetadataSubscribersRequest {
repeated string client_types = 1; // optional filter by client type, e.g. "mount"; empty = all
@@ -906,3 +916,40 @@ message MountInfo {
int64 last_seen_ns = 3;
string data_center = 4;
}
// LookupDirectoryEntriesRequest batches independent exact-path lookups into
// one bounded RPC. Results preserve this order. Empty batches and batches over
// 4096 requests are rejected. Declared last to keep generated message indices
// stable.
message LookupDirectoryEntriesRequest {
repeated LookupDirectoryEntryRequest requests = 1;
// Cache callers may treat an Entry whose Volume the master no longer
// reports as a miss. A lookup the filer could not complete still marks the
// Entry with an error. Ordinary filesystem callers keep the fail-closed
// error in both cases.
bool unavailable_volume_is_miss = 2;
}
message LookupDirectoryEntryResult {
// found reports metadata presence. An error makes this item unusable even
// when entry is present, for example when a chunk volume cannot be located.
bool found = 1;
Entry entry = 2;
string error = 3;
// Per-entry read fence. Every filer event at or below this position is
// reflected in found/entry/error for this item.
int64 log_ts_ns = 4;
int32 log_signature = 5;
}
message LookupDirectoryEntriesResponse {
repeated LookupDirectoryEntryResult results = 1;
// One deduplicated location set for every volume referenced by returned
// Entry.chunks. A missing/unavailable volume has an empty Locations value
// and marks each affected result with an error.
map<string, Locations> locations_map = 2;
// Short-lived, exact-FID read capabilities. Direct data-plane clients must
// present the matching token to a Volume instead of treating a shared
// service credential as authority to read arbitrary needles.
map<string, string> read_auth = 3;
}
+421 -4
View File
@@ -5869,8 +5869,8 @@
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Filer Sync Offset Lag",
"description": "Seconds behind source (filer.sync); only present when filer sync runs",
"title": "Filer Sync Lag",
"description": "Seconds the replicated watermark trails the source filer (filer.sync); only present when filer sync runs",
"type": "timeseries",
"id": 155,
"gridPos": {
@@ -5952,10 +5952,10 @@
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "time() - (SeaweedFS_filerSync_sync_offset{cluster=~\"$cluster\"} > 0) / 1e9",
"expr": "SeaweedFS_filerSync_lag_seconds{cluster=~\"$cluster\"}",
"range": true,
"refId": "A",
"legendFormat": "{{clientName}} {{path}}"
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}}"
}
],
"pluginVersion": "10.3.1"
@@ -6055,6 +6055,423 @@
}
],
"pluginVersion": "10.3.1"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Filer Sync Event Rate",
"description": "Metadata events read off the source stream, replicated, and failed after retries (filer.sync)",
"type": "timeseries",
"id": 230,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "rate(SeaweedFS_filerSync_events_received_total{cluster=~\"$cluster\"}[$__rate_interval])",
"range": true,
"refId": "A",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}} received"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "rate(SeaweedFS_filerSync_events_processed_total{cluster=~\"$cluster\"}[$__rate_interval])",
"range": true,
"refId": "B",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}} processed"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "rate(SeaweedFS_filerSync_events_failed_total{cluster=~\"$cluster\"}[$__rate_interval])",
"range": true,
"refId": "C",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}} failed"
}
],
"pluginVersion": "10.3.1"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Filer Sync Throughput",
"description": "Chunk data bytes carried by sync events: the chunk delta, so deletes, renames, and attribute-only updates count zero (filer.sync)",
"type": "timeseries",
"id": 231,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "Bps",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "rate(SeaweedFS_filerSync_received_bytes_total{cluster=~\"$cluster\"}[$__rate_interval])",
"range": true,
"refId": "A",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}} received"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "rate(SeaweedFS_filerSync_processed_bytes_total{cluster=~\"$cluster\"}[$__rate_interval])",
"range": true,
"refId": "B",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}} processed"
}
],
"pluginVersion": "10.3.1"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Filer Sync In-flight Jobs",
"description": "Jobs currently replicating; pinned at -concurrency means the sync itself is the bottleneck, near zero means caught up or starved by the source (filer.sync)",
"type": "timeseries",
"id": 232,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "SeaweedFS_filerSync_in_flight_jobs{cluster=~\"$cluster\"}",
"range": true,
"refId": "A",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}}"
}
],
"pluginVersion": "10.3.1"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Filer Sync In-flight Bytes",
"description": "Chunk data carried by the jobs currently replicating; distinguishes workers stuck on a few large files from many small ones (filer.sync)",
"type": "timeseries",
"id": 233,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "bytes",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "SeaweedFS_filerSync_in_flight_bytes{cluster=~\"$cluster\"}",
"range": true,
"refId": "A",
"legendFormat": "{{sourceFiler}} -> {{targetFiler}} {{path}}"
}
],
"pluginVersion": "10.3.1"
}
]
},
+2 -3
View File
@@ -3160,9 +3160,9 @@ dependencies = [
[[package]]
name = "redb"
version = "3.1.3"
version = "4.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4ba239c1c1693315d3cc0e601db3b3965543afbf48c41730fdca2f069f510f4a"
checksum = "de6c3b63e007e90ce536ec2ae4690826136a20ec8dbbbb400daef1bb999d2e36"
dependencies = [
"libc",
]
@@ -4582,7 +4582,6 @@ dependencies = [
"image",
"jsonwebtoken",
"kamadak-exif",
"lazy_static",
"libc",
"md-5",
"memmap2",
+16 -3
View File
@@ -1,7 +1,10 @@
[package]
name = "weed-volume"
version = "0.1.0"
edition = "2021"
edition = "2024"
# The edition needs 1.85; the dependency tree needs more. Verified with
# `cargo +1.91.1 check --all-targets` (1.90 fails on the AWS SDK).
rust-version = "1.91.1"
description = "SeaweedFS Volume Server — Rust implementation"
[lib]
@@ -16,6 +19,17 @@ path = "src/main.rs"
# Disable with --no-default-features for 4-byte offsets (32GB max volume size).
default = ["5bytes"]
5bytes = []
# Unstable redb cursor bulk-load for full_rebuild. Off in production.
# Pulls redb's experimental_cursor (and therefore experimental-api-5).
redb-experimental-cursor = ["redb/experimental_cursor"]
[lints.clippy]
# Every RPC path returns tonic::Status (176 bytes). Boxing it would change
# every handler signature for no gain, so the large-Err lint is off.
result_large_err = "allow"
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[dependencies]
# Async runtime
@@ -42,7 +56,6 @@ clap = { version = "4", features = ["derive"] }
# Metrics
prometheus = { version = "0.13", default-features = false, features = ["process"] }
lazy_static = "1"
# JWT
jsonwebtoken = { version = "10", features = ["rust_crypto"] }
@@ -57,7 +70,7 @@ rustls-pemfile = "2"
rusty-leveldb = "3"
# Disk-backed needle map (alternative to in-memory HashMap)
redb = "3"
redb = "4"
# Reed-Solomon erasure coding
reed-solomon-erasure = "6"
+4 -1
View File
@@ -4,7 +4,10 @@ A drop-in replacement for the [SeaweedFS](https://github.com/seaweedfs/seaweedfs
## Building
Requires Rust 1.75+ (2021 edition).
Requires Rust 1.91.1+ (2024 edition), matching `rust-version` in `Cargo.toml`.
The patch release matters: 1.91.0 does not build. The edition itself only needs
1.85; the higher floor comes from the dependency tree — chiefly the AWS SDK — so
it moves with those crates. CI builds on the latest stable.
```bash
cd seaweed-volume
+6 -1
View File
@@ -3,7 +3,12 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
// one, so the build needs no package manager and always sees the same
// version. An explicit PROTOC still wins, for packagers supplying their own.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
// SAFETY: a build script's main runs single-threaded before anything
// else in this process, so no other thread can be reading the
// environment concurrently.
unsafe {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
+9
View File
@@ -145,6 +145,7 @@ message VolumeInformationMessage {
uint64 delete_count = 5;
uint64 deleted_byte_count = 6;
bool read_only = 7;
bool read_only_can_delete = 17;
uint32 replica_placement = 8;
uint32 version = 9;
uint32 ttl = 10;
@@ -164,6 +165,8 @@ message VolumeShortInformationMessage {
uint32 ttl = 10;
string disk_type = 15;
uint32 disk_id = 16;
bool read_only = 17;
bool read_only_can_delete = 18;
}
message VolumeEcShardInformationMessage {
@@ -222,6 +225,9 @@ message VolumeLocation {
uint32 grpc_port = 7;
repeated uint32 new_ec_vids = 8;
repeated uint32 deleted_ec_vids = 9;
repeated uint32 remote_vids = 10;
repeated uint32 read_only_vids = 11;
repeated uint32 read_only_can_delete_vids = 12;
}
message ClusterNodeUpdate {
@@ -267,6 +273,9 @@ message Location {
string public_url = 2;
uint32 grpc_port = 3;
string data_center = 4;
bool data_in_remote = 5;
bool read_only = 6;
bool read_only_can_delete = 7;
}
message AssignRequest {
+2
View File
@@ -75,4 +75,6 @@ message RemoteStorageLocation {
string name = 1;
string bucket = 2;
string path = 3;
int32 listing_cache_ttl_seconds = 4; // 0 = disabled; >0 enables on-demand directory listing with this TTL in seconds
optional int32 cache_wait_ms = 5; // unset = size based default; 0 = read straight from the remote without caching
}
+1
View File
@@ -236,6 +236,7 @@ message VolumeIncrementalCopyResponse {
message VolumeMountRequest {
uint32 volume_id = 1;
optional string collection = 2;
}
message VolumeMountResponse {
}
+108 -81
View File
@@ -64,6 +64,8 @@ pub struct Cli {
pub rack: String,
/// Choose [memory|redb|redbMedium|redbLarge] mode for memory~performance balance.
/// The redb tiers give each volume's on-disk index a 4, 8, or 16 MiB page
/// cache respectively; total index memory is roughly (volumes x cache).
/// `leveldb`/`leveldbMedium`/`leveldbLarge` are accepted as aliases for the
/// corresponding redb backends (Rust volume server uses redb under the hood).
#[arg(long = "index", default_value = "memory")]
@@ -369,17 +371,18 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
if arg == "--" {
break;
}
if arg.starts_with("--") {
let key = if let Some(eq) = arg.find('=') {
arg[2..eq].to_string()
if let Some(long) = arg.strip_prefix("--") {
let key = if let Some(eq) = long.find('=') {
long[..eq].to_string()
} else {
arg[2..].to_string()
long.to_string()
};
cli_flags.insert(key);
} else if arg.starts_with('-') && arg.len() > 2 {
} else if arg.len() > 2
&& let Some(without_dash) = arg.strip_prefix('-')
{
// Single-dash long option (already normalized to -- at this point,
// but handle both for safety)
let without_dash = &arg[1..];
let key = if let Some(eq) = without_dash.find('=') {
without_dash[..eq].to_string()
} else {
@@ -399,15 +402,14 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
}
// Split on first `=`, ` `, or `:`
let (name, value) =
if let Some(pos) = trimmed.find(|c: char| c == '=' || c == ' ' || c == ':') {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
let (name, value) = if let Some(pos) = trimmed.find(['=', ' ', ':']) {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
// Strip leading dashes from name
let name = name.trim_start_matches('-').to_string();
@@ -434,10 +436,8 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
/// Extract the options file path from args (looks for --options or -options).
fn find_options_arg(args: &[String]) -> String {
for i in 1..args.len() {
if args[i] == "--options" || args[i] == "-options" {
if i + 1 < args.len() {
return args[i + 1].clone();
}
if (args[i] == "--options" || args[i] == "-options") && i + 1 < args.len() {
return args[i + 1].clone();
}
if let Some(rest) = args[i].strip_prefix("--options=") {
return rest.to_string();
@@ -455,20 +455,22 @@ fn parse_duration(s: &str) -> std::time::Duration {
if s.is_empty() {
return std::time::Duration::from_secs(60);
}
if let Some(secs) = s.strip_suffix('s') {
if let Ok(v) = secs.parse::<u64>() {
return std::time::Duration::from_secs(v);
}
if let Some(secs) = s.strip_suffix('s')
&& let Ok(v) = secs.parse::<u64>()
{
return std::time::Duration::from_secs(v);
}
if let Some(mins) = s.strip_suffix('m') {
if let Ok(v) = mins.parse::<u64>() {
return std::time::Duration::from_secs(v * 60);
}
if let Some(mins) = s.strip_suffix('m')
&& let Ok(v) = mins.parse::<u64>()
&& let Some(seconds) = v.checked_mul(60)
{
return std::time::Duration::from_secs(seconds);
}
if let Some(hours) = s.strip_suffix('h') {
if let Ok(v) = hours.parse::<u64>() {
return std::time::Duration::from_secs(v * 3600);
}
if let Some(hours) = s.strip_suffix('h')
&& let Ok(v) = hours.parse::<u64>()
&& let Some(seconds) = v.checked_mul(3600)
{
return std::time::Duration::from_secs(seconds);
}
// Fallback: try parsing as raw seconds
if let Ok(v) = s.parse::<u64>() {
@@ -501,40 +503,40 @@ fn parse_min_free_spaces(min_free_space: &str, min_free_space_percent: &str) ->
}
// Try parsing human-readable bytes: e.g. "10GiB", "500MiB", "1TiB"
let s_upper = s.to_uppercase();
if let Some(rest) = s_upper.strip_suffix("TIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
// Default: 1%
MinFreeSpace::Percent(1.0)
@@ -1026,20 +1028,20 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
"cipher_suites" => cfg.tls_policy.cipher_suites = value.to_string(),
_ => {}
},
Section::Guard => match key {
"white_list" => {
Section::Guard => {
if key == "white_list" {
cfg.guard_white_list = value
.split(',')
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
.collect();
}
_ => {}
},
Section::Access => match key {
"ui" => cfg.access_ui = value.parse().unwrap_or(false),
_ => {}
},
}
Section::Access => {
if key == "ui" {
cfg.access_ui = value.parse().unwrap_or(false)
}
}
Section::None => {}
}
}
@@ -1186,12 +1188,11 @@ fn apply_env_overrides(cfg: &mut SecurityConfig) {
/// Mirrors Go's `util.DetectedHostAddress()`.
fn detect_host_address() -> String {
// Connect to a remote address to determine the local outbound IP
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0") {
if socket.connect("8.8.8.8:80").is_ok() {
if let Ok(addr) = socket.local_addr() {
return addr.ip().to_string();
}
}
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0")
&& socket.connect("8.8.8.8:80").is_ok()
&& let Ok(addr) = socket.local_addr()
{
return addr.ip().to_string();
}
"localhost".to_string()
}
@@ -1207,21 +1208,30 @@ mod tests {
LOCK.get_or_init(|| Mutex::new(())).lock().unwrap()
}
// SAFETY (all env mutation in this module): `set_var`/`remove_var` are
// unsafe as of Rust 2024 because they race with concurrent readers in
// other threads. Every test that reaches these helpers holds
// `process_state_lock()` for the duration, so only one test at a time
// touches the environment and none observes another's edit.
fn with_temp_env_var<F: FnOnce()>(key: &str, value: Option<&str>, f: F) {
let previous = std::env::var_os(key);
match value {
Some(v) => std::env::set_var(key, v),
None => std::env::remove_var(key),
unsafe {
match value {
Some(v) => std::env::set_var(key, v),
None => std::env::remove_var(key),
}
}
f();
restore_env_var(key, previous);
}
fn restore_env_var(key: &str, value: Option<OsString>) {
if let Some(value) = value {
std::env::set_var(key, value);
} else {
std::env::remove_var(key);
unsafe {
if let Some(value) = value {
std::env::set_var(key, value);
} else {
std::env::remove_var(key);
}
}
}
@@ -1266,7 +1276,10 @@ mod tests {
.collect();
for key in KEYS {
std::env::remove_var(key);
// SAFETY: as above — the caller holds `process_state_lock()`.
unsafe {
std::env::remove_var(key);
}
}
f();
@@ -1283,6 +1296,14 @@ mod tests {
assert_eq!(parse_duration("1h"), std::time::Duration::from_secs(3600));
assert_eq!(parse_duration("30"), std::time::Duration::from_secs(30));
assert_eq!(parse_duration(""), std::time::Duration::from_secs(60));
assert_eq!(
parse_duration("307445734561825861m"),
std::time::Duration::from_secs(60)
);
assert_eq!(
parse_duration("5124095576030432h"),
std::time::Duration::from_secs(60)
);
}
#[test]
@@ -1402,12 +1423,18 @@ mod tests {
#[test]
fn test_resolve_config_defaults_dir_to_platform_temp_dir() {
// resolve_config reads HOME/USERPROFILE and the WEED_* set, so it has to
// hold the same lock the mutation helpers take — a concurrent set_var
// during this read is exactly what makes those calls unsafe.
let _guard = process_state_lock();
let cfg = resolve_config(Cli::parse_from(["bin"]));
assert_eq!(cfg.folders, vec![default_volume_dir()]);
}
#[test]
fn test_resolve_config_index_accepts_redb_and_leveldb_aliases() {
// As above: resolve_config reads the environment.
let _guard = process_state_lock();
let pairs = [
("memory", NeedleMapKind::InMemory),
("redb", NeedleMapKind::Redb),
+1
View File
@@ -1,5 +1,6 @@
pub mod config;
pub mod images;
pub mod malloc_tuning;
pub mod metrics;
pub mod remote_storage;
pub mod security;
+18
View File
@@ -39,6 +39,11 @@ const GRPC_MAX_HEADER_LIST_SIZE: u32 = 8 * 1024 * 1024;
const GRPC_MAX_CONCURRENT_STREAMS: u32 = 1000;
fn main() {
// Before anything allocates: stop glibc from training its mmap threshold
// upward on our large EC buffers and turning them into heap it never
// returns. See seaweed_volume::malloc_tuning for the measurements.
let malloc_tuning = seaweed_volume::malloc_tuning::pin_mmap_threshold();
install_default_crypto_provider();
// Initialize tracing
@@ -65,6 +70,19 @@ fn main() {
"SeaweedFS Volume Server (Rust) v{}",
seaweed_volume::version::full_version()
);
match malloc_tuning {
seaweed_volume::malloc_tuning::MallocTuning::Pinned(bytes) => {
info!("pinned glibc M_MMAP_THRESHOLD to {} bytes", bytes)
}
seaweed_volume::malloc_tuning::MallocTuning::DeferredToEnv => info!(
"an allocator mmap-threshold override ({}) is set; leaving glibc's mmap threshold to the environment",
seaweed_volume::malloc_tuning::MMAP_THRESHOLD_ENV
),
seaweed_volume::malloc_tuning::MallocTuning::Failed => {
warn!("mallopt(M_MMAP_THRESHOLD) failed; large freed buffers may stay resident")
}
seaweed_volume::malloc_tuning::MallocTuning::NotApplicable => {}
}
// Register Prometheus metrics
metrics::register_metrics();
+467
View File
@@ -0,0 +1,467 @@
//! Keep glibc from silently converting large short-lived buffers into heap the
//! process never gives back.
//!
//! glibc serves an allocation with `mmap` when it is at least
//! `M_MMAP_THRESHOLD` (128 KiB by default), and `munmap`s it on free, so the
//! pages go straight back to the OS. That threshold is **adaptive**: whenever a
//! block that came from `mmap` is freed, glibc raises the threshold to that
//! block's size — up to 32 MiB — on the theory that a workload repeatedly
//! allocating buffers of that size is better served from the heap.
//!
//! For a volume server that theory is wrong in a specific, expensive way. EC
//! reconstruction and needle reassembly allocate large, short-lived buffers.
//! The first few are mmap'd and freed, which trains the threshold upward; every
//! later buffer of that size is then carved out of the heap instead. Heap
//! memory is only returned to the OS from the top of the arena, so those pages
//! stay resident as anonymous memory for the life of the process. They are
//! still *reusable* — this is not a leak, and a repeat workload does not grow
//! the footprint further — but under a hard cgroup `MemoryMax` they are
//! indistinguishable from a leak, because anonymous pages cannot be reclaimed
//! under pressure the way page cache can. The retained footprint eats exactly
//! the headroom that a burst of maintenance work needs, and the process is
//! OOM-killed while most of its resident memory is free-but-unreturned.
//!
//! Measured on a 17-node cluster (EC 10+4, `--index=redb`), one node, two
//! identical `ec.scrub -mode full` rounds over 10912 EC files each, comparing
//! the same unit restarted with and without a pinned threshold:
//!
//! | | baseline | round 1 | round 2 | 60s idle |
//! |---|---|---|---|---|
//! | default (adaptive) | 10 MB | 84 MB | 88 MB | **88 MB** |
//! | pinned threshold | 10 MB | 13 MB | 14 MB | **14 MB** |
//!
//! 78 MB retained versus 4 MB for identical work. On that cluster's heavier
//! mixed scrub workloads the same effect reached ~600 MB of retained anonymous
//! memory per volume server, against a 3 GiB cap.
//!
//! Calling `mallopt(M_MMAP_THRESHOLD, ...)` sets the threshold *and* disables
//! the dynamic adjustment, which is the documented behaviour of setting it
//! explicitly. We pin it to glibc's own default rather than inventing a value:
//! the goal is to stop the adaptation, not to second-guess the default.
/// glibc's own default `M_MMAP_THRESHOLD`. Pinning to this value changes
/// nothing about which allocations use `mmap` on a freshly started process; it
/// only prevents the threshold from drifting upward later.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
const DEFAULT_MMAP_THRESHOLD: libc::c_int = 128 * 1024;
/// Legacy environment variable glibc reads for the same setting. If an operator
/// has set it, honour their value and do not override it.
pub const MMAP_THRESHOLD_ENV: &str = "MALLOC_MMAP_THRESHOLD_";
/// Modern glibc tunables environment variable. Operators may set the threshold
/// via `GLIBC_TUNABLES=glibc.malloc.mmap_threshold=...` instead of the legacy
/// variable; that override is honoured too.
pub const GLIBC_TUNABLES_ENV: &str = "GLIBC_TUNABLES";
/// The tunable name within `GLIBC_TUNABLES` that maps to `M_MMAP_THRESHOLD`.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
const MMAP_THRESHOLD_TUNABLE: &str = "glibc.malloc.mmap_threshold";
/// Outcome of the tuning attempt, so the caller can log it and tests can assert
/// on it without inspecting global allocator state.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum MallocTuning {
/// Threshold pinned to `DEFAULT_MMAP_THRESHOLD`; dynamic adjustment is off.
Pinned(i32),
/// An allocator override (`MALLOC_MMAP_THRESHOLD_` or
/// `GLIBC_TUNABLES=glibc.malloc.mmap_threshold=...`) was set, so the
/// operator's value wins.
DeferredToEnv,
/// `mallopt` reported failure. Not fatal — the server runs, it just keeps
/// glibc's adaptive behaviour.
Failed,
/// Not glibc, so there is no adaptive threshold to pin.
NotApplicable,
}
/// Pin glibc's mmap threshold unless the operator has set an allocator override.
/// Safe to call more than once; call it before serving traffic, since the point
/// is to prevent the threshold from being trained upward by early allocations.
pub fn pin_mmap_threshold() -> MallocTuning {
pin_mmap_threshold_inner()
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn pin_mmap_threshold_inner() -> MallocTuning {
if operator_mmap_threshold_override_active() {
return MallocTuning::DeferredToEnv;
}
// SAFETY: `mallopt` is a libc entry point that takes two ints and mutates
// only allocator-internal tunables. It has no preconditions and no effect
// on memory this process already owns.
let rc = unsafe { libc::mallopt(libc::M_MMAP_THRESHOLD, DEFAULT_MMAP_THRESHOLD) };
if rc == 1 {
MallocTuning::Pinned(DEFAULT_MMAP_THRESHOLD)
} else {
MallocTuning::Failed
}
}
#[cfg(not(all(target_os = "linux", target_env = "gnu")))]
fn pin_mmap_threshold_inner() -> MallocTuning {
MallocTuning::NotApplicable
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn operator_mmap_threshold_override_active() -> bool {
// MALLOC_MMAP_THRESHOLD_: glibc calls atoi(value) then mallopt, which
// always sets the threshold and disables dynamic adjustment — even for
// empty, negative, or non-numeric values (atoi returns 0). So any presence
// of the variable means the operator's override is in effect.
std::env::var_os(MMAP_THRESHOLD_ENV).is_some()
|| usable_glibc_tunable_threshold(std::env::var_os(GLIBC_TUNABLES_ENV))
}
/// Look for `glibc.malloc.mmap_threshold=<value>` among the colon-separated
/// tunables in `GLIBC_TUNABLES`. glibc's `parse_tunables_string` (elf/dl-tunables.c)
/// rejects the **entire** string (returns -1) if it reaches `\0` before finding
/// `=` in a name (last entry has no `=`), or if any entry's value contains a
/// duplicate `=`. When `parse_tunables_string` returns -1, `parse_tunables`
/// prints a warning and returns immediately without applying ANY tunable —
/// including ones already parsed into the tunables array. We match that by
/// returning `false` for the entire string on any of those conditions.
///
/// glibc parses tunable values with `_dl_strtoul`, which accepts decimal,
/// `0x` hex, `0` octal, an optional sign (negatives wrap to `unsigned long`),
/// and requires the entire value to be consumed; we match that with
/// `dl_strtoul_consumes_all`.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn usable_glibc_tunable_threshold(tunables: Option<std::ffi::OsString>) -> bool {
let s = match tunables.and_then(|v| v.into_string().ok()) {
Some(s) => s,
None => return false,
};
if s.is_empty() {
return false;
}
// Parse the string character-by-character, matching glibc's
// parse_tunables_string logic exactly. Using split(':') would lose the
// distinction between an entry terminated by ':' (skip) and one terminated
// by '\0' with no '=' (reject entire string).
let bytes = s.as_bytes();
let mut pos = 0;
let mut found_threshold = false;
loop {
// Find where the name ends ('=', ':', or end of string).
let name_start = pos;
while pos < bytes.len() && bytes[pos] != b'=' && bytes[pos] != b':' {
pos += 1;
}
// End of string before '=' → glibc returns -1 (reject entire string).
if pos >= bytes.len() {
return false;
}
// ':' before '=' → glibc skips this entry and continues.
if bytes[pos] == b':' {
pos += 1;
continue;
}
// Skip the '='.
let name_end = pos;
pos += 1;
// Find where the value ends ('=', ':', or end of string).
let val_start = pos;
while pos < bytes.len() && bytes[pos] != b'=' && bytes[pos] != b':' {
pos += 1;
}
// '=' in value → glibc returns -1 (reject entire string).
if pos < bytes.len() && bytes[pos] == b'=' {
return false;
}
let key = &s[name_start..name_end];
let val = &s[val_start..pos];
if key == MMAP_THRESHOLD_TUNABLE && dl_strtoul_consumes_all(val) {
found_threshold = true;
}
// End of string → done.
if pos >= bytes.len() {
break;
}
// Skip the ':'.
pos += 1;
}
found_threshold
}
/// Replicate glibc's `_dl_strtoul` (elf/dl-misc.c) just enough to determine
/// whether it would consume the entire string — which is what
/// `tunable_parse_num` checks (`endptr == strval + len`). Returns `true` if
/// glibc would accept the value and apply it.
///
/// `_dl_strtoul` skips leading spaces/tabs, accepts an optional `+`/`-` sign,
/// and parses `0x`-prefixed hex, `0`-prefixed octal, or plain decimal. A
/// negative result wraps to `unsigned long` (`-1` → `SIZE_MAX`). If no digit is
/// found after the sign, the end pointer stays at the current position — which
/// still counts as "consumed" when the string is empty or whitespace-only
/// (value 0). On overflow, `_dl_strtoul` stops at the overflowing digit (endptr
/// does not reach the end), so `tunable_parse_num` rejects the value.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn dl_strtoul_consumes_all(s: &str) -> bool {
let bytes = s.as_bytes();
let mut pos = 0;
// Skip leading whitespace (spaces and tabs, matching _dl_strtoul).
while pos < bytes.len() && (bytes[pos] == b' ' || bytes[pos] == b'\t') {
pos += 1;
}
// Optional sign.
if pos < bytes.len() && (bytes[pos] == b'-' || bytes[pos] == b'+') {
pos += 1;
}
// Must have at least one digit (0-9) to start parsing, unless we're already
// at the end (empty / whitespace-only / sign-only → value 0, consumed).
if pos >= bytes.len() {
return true;
}
if bytes[pos] < b'0' || bytes[pos] > b'9' {
return false;
}
// Determine base: 0x → hex, 0 → octal, else decimal. _dl_strtoul unconditionally
// advances past "0x"/"0X" when the first char is '0' and the next is 'x'/'X',
// even if no hex digit follows — in that case the digit loop breaks immediately,
// endptr reaches the end, and the value is 0.
let base: u32 = if bytes[pos] == b'0'
&& pos + 1 < bytes.len()
&& (bytes[pos + 1] == b'x' || bytes[pos + 1] == b'X')
{
pos += 2; // skip "0x"
16
} else if bytes[pos] == b'0' {
8
} else {
10
};
// Parse digits with overflow detection, matching _dl_strtoul's cutoff/cutlim
// logic. On overflow, _dl_strtoul sets endptr to the overflowing digit and
// returns UINT64_MAX — so the value is NOT fully consumed and
// tunable_parse_num rejects it.
let mut result: u64 = 0;
let cutoff = u64::MAX / base as u64;
let cutlim = u64::MAX % base as u64;
while pos < bytes.len() {
let b = bytes[pos];
let digval: u32 = match digit_value(b, base) {
Some(v) => v,
None => break,
};
if result > cutoff || (result == cutoff && digval as u64 > cutlim) {
// Overflow: _dl_strtoul stops here, endptr points at this digit.
return false;
}
result *= base as u64;
result += digval as u64;
pos += 1;
}
// The entire string must be consumed (matching tunable_parse_num's check).
pos == bytes.len()
}
/// Returns the numeric value of a digit byte in the given base, or `None` if
/// the byte is not a valid digit in that base.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn digit_value(b: u8, base: u32) -> Option<u32> {
if (b'0'..=b'0' + (base - 1).min(9) as u8).contains(&b) {
return Some((b - b'0') as u32);
}
if base == 16 {
if (b'a'..=b'f').contains(&b) {
return Some((b - b'a' + 10) as u32);
}
if (b'A'..=b'F').contains(&b) {
return Some((b - b'A' + 10) as u32);
}
}
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn env_override_constants_match_glibc_names() {
// Verified against the real accessor rather than a copy of the name, so
// renaming the constant cannot silently break the override contract.
assert_eq!(MMAP_THRESHOLD_ENV, "MALLOC_MMAP_THRESHOLD_");
assert_eq!(GLIBC_TUNABLES_ENV, "GLIBC_TUNABLES");
}
#[test]
fn calling_twice_is_stable() {
// Startup paths get re-entered in tests and in `weed mini`; the second
// call must not report a different outcome from the first.
let first = pin_mmap_threshold();
let second = pin_mmap_threshold();
assert_eq!(first, second);
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn pins_threshold_on_glibc_when_no_override_is_set() {
// The env override is not set in the test process, so this exercises the
// mallopt path. If an override happens to be present, defer to it.
if operator_mmap_threshold_override_active() {
assert_eq!(pin_mmap_threshold(), MallocTuning::DeferredToEnv);
return;
}
assert_eq!(
pin_mmap_threshold(),
MallocTuning::Pinned(DEFAULT_MMAP_THRESHOLD),
"mallopt(M_MMAP_THRESHOLD) should succeed on glibc"
);
}
#[cfg(not(all(target_os = "linux", target_env = "gnu")))]
#[test]
fn is_a_noop_off_glibc() {
// No glibc adaptive threshold exists off glibc, so there is nothing to
// pin regardless of any environment variables that happen to be set.
assert_eq!(pin_mmap_threshold(), MallocTuning::NotApplicable);
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn dl_strtoul_consumes_all_matches_glibc_parser() {
// Decimal — any non-empty decimal integer is accepted, including
// negative (wraps to unsigned) and zero.
assert!(dl_strtoul_consumes_all("131072"));
assert!(dl_strtoul_consumes_all("0"));
assert!(dl_strtoul_consumes_all("-1"));
assert!(dl_strtoul_consumes_all("-131072"));
// Values above i64::MAX are valid for glibc's unsigned parser.
assert!(dl_strtoul_consumes_all("9223372036854775808"));
// Hex with 0x prefix.
assert!(dl_strtoul_consumes_all("0x20000"));
assert!(dl_strtoul_consumes_all("0X20000"));
assert!(dl_strtoul_consumes_all("0x0"));
// Octal with leading 0.
assert!(dl_strtoul_consumes_all("010"));
// Leading whitespace (spaces and tabs) is skipped.
assert!(dl_strtoul_consumes_all(" 131072"));
assert!(dl_strtoul_consumes_all("\t0x20000"));
// Empty and whitespace-only strings are accepted (value 0).
assert!(dl_strtoul_consumes_all(""));
assert!(dl_strtoul_consumes_all(" "));
assert!(dl_strtoul_consumes_all("\t"));
// Sign-only strings are accepted: _dl_strtoul skips the sign, finds no
// digit, sets endptr to the position after the sign (== end of string),
// and returns 0. tunable_parse_num sees endptr == strval + len → true.
assert!(dl_strtoul_consumes_all("-"));
assert!(dl_strtoul_consumes_all("+"));
// Trailing garbage is rejected — _dl_strtoul stops at the first
// non-digit and tunable_parse_num requires the entire string consumed.
assert!(!dl_strtoul_consumes_all("131072abc"));
// In hex mode, a-f are digits, so "0x20000abc" is a valid hex number.
// Use a non-hex character like 'g' to test trailing garbage in hex.
assert!(!dl_strtoul_consumes_all("0x20000g"));
assert!(!dl_strtoul_consumes_all("128K"));
// Non-numeric strings are rejected.
assert!(!dl_strtoul_consumes_all("abc"));
// "0x" with no hex digits: _dl_strtoul advances past "0x", the digit loop
// breaks immediately (no hex digit), endptr reaches the end, value is 0.
// tunable_parse_num accepts it.
assert!(dl_strtoul_consumes_all("0x"));
assert!(dl_strtoul_consumes_all("0X"));
// Overflow: _dl_strtoul stops at the overflowing digit (endptr points
// there, not at the end), so tunable_parse_num rejects the value.
assert!(!dl_strtoul_consumes_all("18446744073709551616")); // u64::MAX + 1
assert!(!dl_strtoul_consumes_all("99999999999999999999")); // 20 nines
assert!(!dl_strtoul_consumes_all("0x10000000000000000")); // 2^64
// u64::MAX itself is accepted: the last digit (5) equals cutlim (=5),
// so the overflow check (digval > cutlim) is false.
assert!(dl_strtoul_consumes_all("18446744073709551615")); // u64::MAX
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn usable_glibc_tunable_threshold_detects_mmap_threshold() {
// Decimal, hex, octal, negative, and zero values are all accepted by
// glibc's _dl_strtoul and cause the threshold to be pinned.
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=131072".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=0x20000".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=0".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=-1".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=9223372036854775808".into()
)));
// u64::MAX is accepted by _dl_strtoul (last digit == cutlim, no overflow).
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=18446744073709551615".into()
)));
// Appears alongside other tunables.
assert!(usable_glibc_tunable_threshold(Some(
"glibc.cpu.x=1:glibc.malloc.mmap_threshold=131072".into()
)));
// Leading ':' is accepted — glibc skips the empty entry and continues.
assert!(usable_glibc_tunable_threshold(Some(
":glibc.malloc.mmap_threshold=131072".into()
)));
// Empty value is accepted by _dl_strtoul (value 0).
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=".into()
)));
// Non-numeric values are rejected by _dl_strtoul.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=abc".into()
)));
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=128K".into()
)));
// A malformed sibling entry (duplicate '=') makes glibc reject the
// entire string, so we must not accept the threshold entry either.
// This applies regardless of whether the threshold is before or after
// the malformed entry — parse_tunables_string returns -1, and
// parse_tunables discards all tunables without applying any.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.check=2=2:glibc.malloc.mmap_threshold=131072".into()
)));
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:glibc.malloc.check=2=2".into()
)));
// A trailing entry with no '=' makes glibc reject the entire string
// (parse_tunables_string hits '\0' before '=' and returns -1).
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:glibc.cpu.x".into()
)));
// A trailing ':' makes glibc reject the entire string (the empty entry
// after ':' hits '\0' before '=' and returns -1).
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:".into()
)));
// Unrelated tunables do not count.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.cpu.x=1".into()
)));
assert!(!usable_glibc_tunable_threshold(None));
}
}
+233 -116
View File
@@ -3,10 +3,10 @@
//! Mirrors the Go SeaweedFS volume server metrics.
use prometheus::{
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounterVec, IntGauge, IntGaugeVec,
Opts, Registry, TextEncoder,
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge,
IntGaugeVec, Opts, Registry, TextEncoder,
};
use std::sync::Once;
use std::sync::{LazyLock, Once};
use crate::version;
@@ -16,203 +16,320 @@ pub struct PushGatewayConfig {
pub interval_seconds: u32,
}
lazy_static::lazy_static! {
pub static ref REGISTRY: Registry = Registry::new();
pub static REGISTRY: LazyLock<Registry> = LazyLock::new(Registry::new);
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static ref REQUEST_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_request_total", "Volume server requests"),
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static REQUEST_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_request_total",
"Volume server requests",
),
&["type", "code"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Request duration histogram with label `type` (HTTP method).
pub static ref REQUEST_DURATION: HistogramVec = HistogramVec::new(
/// Request duration histogram with label `type` (HTTP method).
pub static REQUEST_DURATION: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_request_seconds",
"Volume server request duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
/// Handler-level operation counter with label `type`.
pub static ref HANDLER_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_handler_total", "Volume server handler counters"),
/// Handler-level operation counter with label `type`.
pub static HANDLER_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_handler_total",
"Volume server handler counters",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
/// Vacuuming compact counter with label `success` (true/false).
pub static ref VACUUMING_COMPACT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_compact_count", "Counter of volume vacuuming Compact counter"),
/// Vacuuming compact counter with label `success` (true/false).
pub static VACUUMING_COMPACT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_compact_count",
"Counter of volume vacuuming Compact counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming commit counter with label `success` (true/false).
pub static ref VACUUMING_COMMIT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_commit_count", "Counter of volume vacuuming commit counter"),
/// Vacuuming commit counter with label `success` (true/false).
pub static VACUUMING_COMMIT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_commit_count",
"Counter of volume vacuuming commit counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static ref VACUUMING_HISTOGRAM: HistogramVec = HistogramVec::new(
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static VACUUMING_HISTOGRAM: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_vacuuming_seconds",
"Volume vacuuming duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
/// Volumes per collection and type (volume/ec_shards).
pub static ref VOLUME_GAUGE: GaugeVec = GaugeVec::new(
/// Volumes per collection and type (volume/ec_shards).
pub static VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_volumes", "Number of volumes"),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Read-only volumes per collection and type.
pub static ref READ_ONLY_VOLUME_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_read_only_volumes", "Number of read-only volumes."),
/// Read-only volumes per collection and type.
pub static READ_ONLY_VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_read_only_volumes",
"Number of read-only volumes.",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Maximum number of volumes this server can hold.
pub static ref MAX_VOLUMES: IntGauge = IntGauge::new(
/// Maximum number of volumes this server can hold.
pub static MAX_VOLUMES: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_max_volumes",
"Maximum number of volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static ref DISK_SIZE_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_total_disk_size", "Actual disk size used by volumes"),
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static DISK_SIZE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_total_disk_size",
"Actual disk size used by volumes",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
/// Disk resource usage per directory and type (all/used/free/avail).
pub static ref RESOURCE_GAUGE: GaugeVec = GaugeVec::new(
/// Disk resource usage per directory and type (all/used/free/avail).
pub static RESOURCE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_resource", "Server resource usage"),
&["name", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
/// In-flight requests per HTTP method.
pub static ref INFLIGHT_REQUESTS_GAUGE: IntGaugeVec = IntGaugeVec::new(
Opts::new("SeaweedFS_volumeServer_in_flight_requests", "Current number of in-flight requests being handled by volume server."),
/// In-flight requests per HTTP method.
pub static INFLIGHT_REQUESTS_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_in_flight_requests",
"Current number of in-flight requests being handled by volume server.",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent download limit in bytes.
pub static ref CONCURRENT_DOWNLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent download limit in bytes.
pub static CONCURRENT_DOWNLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_download_limit",
"Limit for total concurrent download size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent upload limit in bytes.
pub static ref CONCURRENT_UPLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent upload limit in bytes.
pub static CONCURRENT_UPLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_upload_limit",
"Limit for total concurrent upload size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight download bytes.
pub static ref INFLIGHT_DOWNLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight download bytes.
pub static INFLIGHT_DOWNLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_download_size",
"In flight total download size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight upload bytes.
pub static ref INFLIGHT_UPLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight upload bytes.
pub static INFLIGHT_UPLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_upload_size",
"In flight total upload size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static ref UPLOAD_ERROR_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created");
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static UPLOAD_ERROR_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created")
});
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static ref SCRUB_LAST_TIME_SECONDS: GaugeVec = GaugeVec::new(
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static SCRUB_LAST_TIME_SECONDS: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_last_time_seconds",
"Last scrub execution time, as seconds since UNIX epoch.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_VOLUME_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static SCRUB_VOLUME_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_volume_failures",
"Counter of overall volumes with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_SHARD_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static SCRUB_SHARD_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_shard_failures",
"Counter of overall EC shards with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Legacy aliases for backward compat with existing code ----
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static STORAGE_IO_ERROR_COUNTER: LazyLock<IntCounter> = LazyLock::new(|| {
IntCounter::new(
"SeaweedFS_volumeServer_storage_io_error_total",
"Counter of storage read/write EIO errors on volumes and EC shards.",
)
.expect("metric can be created")
});
/// Total number of volumes on this server (flat gauge).
pub static ref VOLUMES_TOTAL: IntGauge = IntGauge::new(
"volume_server_volumes_total",
"Total number of volumes",
).expect("metric can be created");
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static IO_QUARANTINE_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_io_quarantine",
"Number of volumes or EC shards quarantined due to storage IO errors.",
),
&["kind"],
)
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static ref DISK_SIZE_BYTES: IntGaugeVec = IntGaugeVec::new(
// ---- Legacy aliases for backward compat with existing code ----
/// Total number of volumes on this server (flat gauge).
pub static VOLUMES_TOTAL: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new("volume_server_volumes_total", "Total number of volumes")
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static DISK_SIZE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_size_bytes", "Disk size in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Disk free bytes per directory.
pub static ref DISK_FREE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk free bytes per directory.
pub static DISK_FREE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_free_bytes", "Disk free space in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current number of in-flight requests (flat gauge).
pub static ref INFLIGHT_REQUESTS: IntGauge = IntGauge::new(
/// Current number of in-flight requests (flat gauge).
pub static INFLIGHT_REQUESTS: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_inflight_requests",
"Current number of in-flight requests",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Total number of files stored across all volumes.
pub static ref VOLUME_FILE_COUNT: IntGauge = IntGauge::new(
/// Total number of files stored across all volumes.
pub static VOLUME_FILE_COUNT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_volume_file_count",
"Total number of files stored across all volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Build info (Go: BuildInfo) ----
// ---- Build info (Go: BuildInfo) ----
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static ref BUILD_INFO: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created");
}
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static BUILD_INFO: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created")
});
/// Generate exponential bucket boundaries for histograms.
fn exponential_buckets(start: f64, factor: f64, count: usize) -> Vec<f64> {
@@ -283,6 +400,8 @@ pub fn register_metrics() {
Box::new(SCRUB_LAST_TIME_SECONDS.clone()),
Box::new(SCRUB_VOLUME_FAILURES.clone()),
Box::new(SCRUB_SHARD_FAILURES.clone()),
Box::new(STORAGE_IO_ERROR_COUNTER.clone()),
Box::new(IO_QUARANTINE_GAUGE.clone()),
// Legacy metrics
Box::new(VOLUMES_TOTAL.clone()),
Box::new(DISK_SIZE_BYTES.clone()),
@@ -358,10 +477,8 @@ fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
type_value = Some(label.get_value().to_string());
}
}
if matches_collection {
if let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
if matches_collection && let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
}
}
@@ -173,10 +173,10 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if let IpAddr::V6(v6) = ip
&& let Some(v4) = embedded_transition_ipv4(v6)
{
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
Ok(())
}
@@ -214,9 +214,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
// Authority is everything up to the first '/', '?', or '#'.
let after = &trimmed[scheme_end + 3..];
let authority_end = after
.find(|c| c == '/' || c == '?' || c == '#')
.unwrap_or(after.len());
let authority_end = after.find(['/', '?', '#']).unwrap_or(after.len());
let authority = &after[..authority_end];
// Strip optional userinfo ("user:pass@").
+71 -30
View File
@@ -81,13 +81,17 @@ impl S3TierBackend {
/// Returns (s3_key, file_size) on success.
/// The progress callback receives (bytes_uploaded, percentage).
/// Uses 64MB part size and 5 concurrent uploads (matches Go s3manager).
/// `progress_fn` returning `Err` aborts the upload, mirroring Go's
/// `fn(progressed, percentage) error` in `s3_upload.go`, where the error
/// surfaces out of `ReadAt` and fails the transfer. It is what lets a
/// caller that has hung up stop the work it is no longer waiting for.
pub async fn upload_file<F>(
&self,
file_path: &str,
progress_fn: F,
) -> Result<(String, u64), String>
where
F: FnMut(i64, f32) + Send + Sync + 'static,
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
let key = uuid::Uuid::new_v4().to_string();
@@ -184,8 +188,10 @@ impl S3TierBackend {
let e_tag = upload_part_resp.e_tag().unwrap_or_default().to_string();
// Report progress
{
// Report progress. The lock is released before the result is
// propagated so an aborting callback cannot poison the mutex
// for the other parts still in flight.
let progress_result = {
let mut guard = progress.lock().unwrap();
guard.0 += size as u64;
let uploaded = guard.0;
@@ -194,8 +200,9 @@ impl S3TierBackend {
} else {
100.0
};
(guard.1)(uploaded as i64, pct);
}
(guard.1)(uploaded as i64, pct)
};
progress_result?;
Ok::<_, String>(
CompletedPart::builder()
@@ -206,29 +213,58 @@ impl S3TierBackend {
}));
}
// Collect results, preserving part order
let mut completed_parts = Vec::with_capacity(handles.len());
for handle in handles {
let part = handle
let finish = async {
// Collect results, preserving part order
let mut completed_parts = Vec::with_capacity(handles.len());
for handle in handles {
let part = handle
.await
.map_err(|e| format!("upload task panicked: {}", e))??;
completed_parts.push(part);
}
// Complete multipart upload
let completed_upload = CompletedMultipartUpload::builder()
.set_parts(Some(completed_parts))
.build();
self.client
.complete_multipart_upload()
.bucket(&self.bucket)
.key(&key)
.upload_id(&upload_id)
.multipart_upload(completed_upload)
.send()
.await
.map_err(|e| format!("upload task panicked: {}", e))??;
completed_parts.push(part);
.map_err(|e| format!("failed to complete multipart upload: {}", e))?;
Ok::<(), String>(())
}
.await;
// Complete multipart upload
let completed_upload = CompletedMultipartUpload::builder()
.set_parts(Some(completed_parts))
.build();
self.client
.complete_multipart_upload()
.bucket(&self.bucket)
.key(&key)
.upload_id(&upload_id)
.multipart_upload(completed_upload)
.send()
.await
.map_err(|e| format!("failed to complete multipart upload: {}", e))?;
if let Err(e) = finish {
// An abandoned multipart upload does not appear in an ordinary
// object listing but still accrues storage charges until a
// lifecycle rule reaps it. Now that a departing caller aborts the
// transfer this is a routine path, not a rare one.
if let Err(abort_err) = self
.client
.abort_multipart_upload()
.bucket(&self.bucket)
.key(&key)
.upload_id(&upload_id)
.send()
.await
{
tracing::warn!(
"failed to abort multipart upload {} for key {}: {}",
upload_id,
key,
abort_err
);
}
return Err(e);
}
Ok((key, file_size))
}
@@ -238,6 +274,8 @@ impl S3TierBackend {
///
/// Returns the file size on success.
/// Uses 64MB part size and 5 concurrent downloads (matches Go s3manager).
/// `progress_fn` returning `Err` aborts the download, mirroring Go's
/// `fn(progressed, percentage) error` in `s3_download.go`.
pub async fn download_file<F>(
&self,
dest_path: &str,
@@ -245,7 +283,7 @@ impl S3TierBackend {
progress_fn: F,
) -> Result<u64, String>
where
F: FnMut(i64, f32) + Send + Sync + 'static,
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
// Get file size first
let head_resp = self
@@ -340,8 +378,10 @@ impl S3TierBackend {
.await
.map_err(|e| format!("failed to write to {}: {}", dp, e))?;
// Report progress
{
// Report progress. The lock is released before the result is
// propagated so an aborting callback cannot poison the mutex
// for the other parts still in flight.
let progress_result = {
let mut guard = progress.lock().unwrap();
guard.0 += bytes.len() as u64;
let downloaded = guard.0;
@@ -350,8 +390,9 @@ impl S3TierBackend {
} else {
100.0
};
(guard.1)(downloaded as i64, pct);
}
(guard.1)(downloaded as i64, pct)
};
progress_result?;
Ok::<_, String>(())
}));
+4 -4
View File
@@ -297,10 +297,10 @@ impl Guard {
/// Extract host from "host:port" or "[::1]:port" format.
fn extract_host(addr: &str) -> String {
// Handle IPv6 with brackets
if addr.starts_with('[') {
if let Some(end) = addr.find(']') {
return addr[1..end].to_string();
}
if addr.starts_with('[')
&& let Some(end) = addr.find(']')
{
return addr[1..end].to_string();
}
// Handle host:port
if let Some(pos) = addr.rfind(':') {
File diff suppressed because it is too large Load Diff
+308 -307
View File
@@ -193,10 +193,7 @@ impl http_body::Body for StreamingBody {
}
Ok(Err(e)) => return std::task::Poll::Ready(Some(Err(e))),
Err(e) => {
return std::task::Poll::Ready(Some(Err(std::io::Error::new(
std::io::ErrorKind::Other,
e,
))))
return std::task::Poll::Ready(Some(Err(std::io::Error::other(e))));
}
}
}
@@ -322,10 +319,9 @@ fn parse_url_path(path: &str) -> Option<(VolumeId, NeedleId, Cookie)> {
// Try "vid,fid" or "vid/fid" or "vid/fid/filename" formats
let (vid_str, fid_part) = if let Some(pos) = path.find(',') {
(&path[..pos], &path[pos + 1..])
} else if let Some(pos) = path.find('/') {
(&path[..pos], &path[pos + 1..])
} else {
return None;
let pos = path.find('/')?;
(&path[..pos], &path[pos + 1..])
};
// For fid part, strip extension from the fid (not from filename)
@@ -356,6 +352,10 @@ fn parse_url_path(path: &str) -> Option<(VolumeId, NeedleId, Cookie)> {
#[derive(Clone, Debug, Deserialize)]
struct VolumeLocation {
url: String,
#[serde(rename = "readOnly", default)]
read_only: bool,
#[serde(rename = "readOnlyCanDelete", default)]
read_only_can_delete: bool,
// Master often omits publicUrl when it matches url (Go json omitempty).
#[serde(rename = "publicUrl", default)]
public_url: String,
@@ -405,10 +405,10 @@ async fn lookup_volume(
.json()
.await
.map_err(|e| format!("lookup parse failed: {}", e))?;
if let Some(err) = result.error {
if !err.is_empty() {
return Err(err);
}
if let Some(err) = result.error
&& !err.is_empty()
{
return Err(err);
}
Ok(result.locations.unwrap_or_default())
}
@@ -546,27 +546,39 @@ async fn do_replicated_request(
.await
.map_err(|e| format!("lookup volume failed: {}", e))?;
// Mirror Go's GetWritableRemoteReplications: reject when the master reports fewer replicas than
// the copy count. lookup_volume is uncached, so recovery is immediate once the replica re-registers.
let copy_count = {
let store = state.store.read().unwrap();
store.find_volume(VolumeId(vid)).map_or(1, |(_, v)| {
v.super_block.replica_placement.get_copy_count()
})
store
.find_volume(VolumeId(vid))
.map_or(1, |(_, v)| v.super_block.replica_placement.get_copy_count())
};
if locations.len() < copy_count as usize {
let allow_delete = method == axum::http::Method::DELETE;
let eligible_locations: Vec<_> = locations
.into_iter()
.filter(|loc| {
(!loc.read_only && allow_delete)
|| (!loc.read_only && !allow_delete)
|| (allow_delete && loc.read_only_can_delete)
})
.collect();
if eligible_locations.len() < copy_count as usize {
return Err(format!(
"replicating operations [{}] is less than volume {} replication copy count [{}]",
locations.len(),
eligible_locations.len(),
vid,
copy_count
));
}
let self_http = to_http_address(&state.self_url);
let remote_locations: Vec<_> = locations
let remote_locations: Vec<_> = eligible_locations
.into_iter()
.filter(|loc| {
if (!allow_delete && loc.read_only)
|| (allow_delete && loc.read_only && !loc.read_only_can_delete)
{
return false;
}
to_http_address(&loc.url) != self_http
&& to_http_address(loc.public_or_url()) != self_http
})
@@ -672,7 +684,8 @@ fn build_proxy_request_info(
raw_fid
};
(trimmed[..pos].to_string(), fid.to_string())
} else if let Some(pos) = trimmed.find('/') {
} else {
let pos = trimmed.find('/')?;
let after = &trimmed[pos + 1..];
let fid_part = if let Some(slash) = after.find('/') {
&after[..slash]
@@ -680,8 +693,6 @@ fn build_proxy_request_info(
after
};
(trimmed[..pos].to_string(), fid_part.to_string())
} else {
return None;
};
Some(ProxyRequestInfo {
@@ -821,12 +832,12 @@ fn redirect_request(info: &ProxyRequestInfo, target: &VolumeLocation, scheme: &s
let mut query_params = Vec::new();
if !info.original_query.is_empty() {
for param in info.original_query.split('&') {
if let Some((key, value)) = param.split_once('=') {
if key == "collection" {
query_params.push(format!("collection={}", value));
}
// Intentionally drop readDeleted and other params (Go parity)
if let Some((key, value)) = param.split_once('=')
&& key == "collection"
{
query_params.push(format!("collection={}", value));
}
// Intentionally drop readDeleted and other params (Go parity)
}
}
query_params.push("proxied=true".to_string());
@@ -836,7 +847,7 @@ fn redirect_request(info: &ProxyRequestInfo, target: &VolumeLocation, scheme: &s
let target_http = to_http_address(&target.url);
let raw_target = format!(
"{}/{},{}?{}",
target_http, &info.vid_str, &info.fid_str, query
target_http, info.vid_str, info.fid_str, query
);
let location = match normalize_outgoing_http_url(scheme, &raw_target) {
Ok(url) => url,
@@ -938,12 +949,12 @@ async fn get_or_head_handler_inner(
// so invalid paths with JWT enabled return 401, not 400.
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) =
state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, false)
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, false)
.is_err()
{
let body = serde_json::json!({"error": "wrong jwt"});
return Response::builder()
@@ -999,16 +1010,15 @@ async fn get_or_head_handler_inner(
let should_try_replica =
!query_string.contains("proxied=true") && !state.master_url.is_empty() && {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, vol)| {
store.find_volume(vid).is_some_and(|(_, vol)| {
vol.super_block.replica_placement.get_copy_count() > 1
})
};
if should_try_replica {
if let Some(info) =
if should_try_replica
&& let Some(info) =
build_proxy_request_info(&path, request.headers(), &query_string)
{
return proxy_or_redirect_to_target(&state, info, vid, true).await;
}
{
return proxy_or_redirect_to_target(&state, info, vid, true).await;
}
// Blocking wait loop (Go's waitForDownloadSlot)
@@ -1047,8 +1057,8 @@ async fn get_or_head_handler_inner(
let has_range = headers.contains_key(header::RANGE);
let ext = extract_extension_from_path(&path);
// Go checks resize and crop extensions separately: resize supports .webp, crop does not.
let has_resize_ops =
is_image_resize_ext(&ext) && (query.width.unwrap_or(0) > 0 || query.height.unwrap_or(0) > 0);
let has_resize_ops = is_image_resize_ext(&ext)
&& (query.width.unwrap_or(0) > 0 || query.height.unwrap_or(0) > 0);
// Go's shouldCropImages (L410) requires x2 > x1 && y2 > y1 (x1/y1 default 0).
// Only disable streaming when a real crop will actually happen.
let has_crop_ops = is_image_crop_ext(&ext) && {
@@ -1077,10 +1087,8 @@ async fn get_or_head_handler_inner(
// serves both the "all shards local" fast case and the
// "some intervals need peer fetch + reconstruct" general
// case without paying for the local interval reads twice.
match crate::server::store_ec::read_ec_shard_needle_distributed(
&state, vid, needle_id,
)
.await
match crate::server::store_ec::read_ec_shard_needle_distributed(&state, vid, needle_id)
.await
{
Ok(Some(ec_needle)) => {
n = ec_needle;
@@ -1101,10 +1109,7 @@ async fn get_or_head_handler_inner(
if e.kind() == std::io::ErrorKind::NotFound {
return StatusCode::NOT_FOUND.into_response();
}
return (
StatusCode::INTERNAL_SERVER_ERROR,
format!("ec read: {}", e),
)
return (StatusCode::INTERNAL_SERVER_ERROR, format!("ec read: {}", e))
.into_response();
}
}
@@ -1220,57 +1225,53 @@ async fn get_or_head_handler_inner(
// Build Last-Modified header (RFC 1123 format) — must be done before conditional checks
let last_modified_str = if n.last_modified > 0 {
use chrono::{TimeZone, Utc};
if let Some(dt) = Utc.timestamp_opt(n.last_modified as i64, 0).single() {
Some(dt.format("%a, %d %b %Y %H:%M:%S GMT").to_string())
} else {
None
}
Utc.timestamp_opt(n.last_modified as i64, 0)
.single()
.map(|dt| dt.format("%a, %d %b %Y %H:%M:%S GMT").to_string())
} else {
None
};
// Check If-Modified-Since FIRST (Go checks this before If-None-Match)
if n.last_modified > 0 {
if let Some(ims_header) = headers.get(header::IF_MODIFIED_SINCE) {
if let Ok(ims_str) = ims_header.to_str() {
// Parse HTTP date format: "Mon, 02 Jan 2006 15:04:05 GMT"
if let Ok(ims_time) =
chrono::NaiveDateTime::parse_from_str(ims_str, "%a, %d %b %Y %H:%M:%S GMT")
{
if (n.last_modified as i64) <= ims_time.and_utc().timestamp() {
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
}
if n.last_modified > 0
&& let Some(ims_header) = headers.get(header::IF_MODIFIED_SINCE)
&& let Ok(ims_str) = ims_header.to_str()
{
// Parse HTTP date format: "Mon, 02 Jan 2006 15:04:05 GMT"
if let Ok(ims_time) =
chrono::NaiveDateTime::parse_from_str(ims_str, "%a, %d %b %Y %H:%M:%S GMT")
&& (n.last_modified as i64) <= ims_time.and_utc().timestamp()
{
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
}
// Check If-None-Match SECOND
if let Some(if_none_match) = headers.get(header::IF_NONE_MATCH) {
if let Ok(inm) = if_none_match.to_str() {
if inm == etag {
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
if let Some(if_none_match) = headers.get(header::IF_NONE_MATCH)
&& let Ok(inm) = if_none_match.to_str()
&& inm == etag
{
let mut resp = StatusCode::NOT_MODIFIED.into_response();
if let Some(ref lm) = last_modified_str {
resp.headers_mut()
.insert(header::LAST_MODIFIED, lm.parse().unwrap());
}
// Go sets ETag AFTER the 304 return paths (L235), so 304 does NOT include ETag
return resp;
}
// Chunk manifest expansion (needs full data) — after conditional checks, before response
// Pass ETag so chunk manifest responses include it (matches Go: ETag is set on the
// response writer before tryHandleChunkedFile runs).
if n.is_chunk_manifest() && !bypass_cm {
if let Some(resp) = try_expand_chunk_manifest(
if n.is_chunk_manifest()
&& !bypass_cm
&& let Some(resp) = try_expand_chunk_manifest(
&state,
&n,
&headers,
@@ -1281,27 +1282,26 @@ async fn get_or_head_handler_inner(
&last_modified_str,
)
.await
{
return resp;
}
// If manifest expansion fails (invalid JSON etc.), fall through to raw data
{
return resp;
}
// If manifest expansion fails (invalid JSON etc.), fall through to raw data
let mut response_headers = HeaderMap::new();
response_headers.insert(header::ETAG, etag.parse().unwrap());
// H1: Emit pairs as response headers
if n.has_pairs() && !n.pairs.is_empty() {
if let Ok(pair_map) =
if n.has_pairs()
&& !n.pairs.is_empty()
&& let Ok(pair_map) =
serde_json::from_slice::<std::collections::HashMap<String, String>>(&n.pairs)
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
}
}
@@ -1311,10 +1311,10 @@ async fn get_or_head_handler_inner(
let mut ext = ext;
if n.name_size > 0 && filename.is_empty() {
filename = String::from_utf8_lossy(&n.name).to_string();
if ext.is_empty() {
if let Some(dot_pos) = filename.rfind('.') {
ext = filename[dot_pos..].to_lowercase();
}
if ext.is_empty()
&& let Some(dot_pos) = filename.rfind('.')
{
ext = filename[dot_pos..].to_lowercase();
}
}
@@ -1418,80 +1418,72 @@ async fn get_or_head_handler_inner(
}
// ---- Streaming path: large uncompressed files ----
if can_stream {
if let Some(info) = stream_info {
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
if can_stream && let Some(info) = stream_info {
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
let tracked_bytes = info.data_size as i64;
let tracking_state = if download_guard.is_some() {
let new_val = state
.inflight_download_bytes
.fetch_add(tracked_bytes, Ordering::Relaxed)
+ tracked_bytes;
metrics::INFLIGHT_DOWNLOAD_SIZE.set(new_val);
Some(state.clone())
} else {
let tracked_bytes = info.data_size as i64;
let tracking_state = if download_guard.is_some() {
let new_val = state
.inflight_download_bytes
.fetch_add(tracked_bytes, Ordering::Relaxed)
+ tracked_bytes;
metrics::INFLIGHT_DOWNLOAD_SIZE.set(new_val);
Some(state.clone())
} else {
None
};
let streaming = StreamingBody {
source: info.source,
data_offset: info.data_file_offset,
data_size: info.data_size,
pos: 0,
chunk_size: streaming_chunk_size(state.read_buffer_size_bytes, info.data_size as usize),
_held_read_lease: if state.has_slow_read {
None
};
} else {
Some(info.data_file_access_control.read_lock())
},
data_file_access_control: info.data_file_access_control,
hold_read_lock_for_stream: !state.has_slow_read,
pending: None,
state: tracking_state,
tracked_bytes,
server_state: state.clone(),
volume_id: info.volume_id,
needle_id: info.needle_id,
compaction_revision: info.compaction_revision,
};
let streaming = StreamingBody {
source: info.source,
data_offset: info.data_file_offset,
data_size: info.data_size,
pos: 0,
chunk_size: streaming_chunk_size(
state.read_buffer_size_bytes,
info.data_size as usize,
),
_held_read_lease: if state.has_slow_read {
None
} else {
Some(info.data_file_access_control.read_lock())
},
data_file_access_control: info.data_file_access_control,
hold_read_lock_for_stream: !state.has_slow_read,
pending: None,
state: tracking_state,
tracked_bytes,
server_state: state.clone(),
volume_id: info.volume_id,
needle_id: info.needle_id,
compaction_revision: info.compaction_revision,
};
let body = Body::new(streaming);
let mut resp = Response::new(body);
*resp.status_mut() = StatusCode::OK;
*resp.headers_mut() = response_headers;
return resp;
}
let body = Body::new(streaming);
let mut resp = Response::new(body);
*resp.status_mut() = StatusCode::OK;
*resp.headers_mut() = response_headers;
return resp;
}
if can_handle_head_from_meta {
if let Some(info) = stream_info {
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
return (StatusCode::OK, response_headers).into_response();
}
if can_handle_head_from_meta && let Some(info) = stream_info {
response_headers.insert(
header::CONTENT_LENGTH,
info.data_size.to_string().parse().unwrap(),
);
return (StatusCode::OK, response_headers).into_response();
}
if can_handle_range_from_source {
if let (Some(range_header), Some(info)) = (headers.get(header::RANGE), stream_info) {
if let Ok(range_str) = range_header.to_str() {
return handle_range_request_from_source(
range_str,
info,
response_headers,
track_download.then(|| state.clone()),
);
}
}
if can_handle_range_from_source
&& let (Some(range_header), Some(info)) = (headers.get(header::RANGE), stream_info)
&& let Ok(range_str) = range_header.to_str()
{
return handle_range_request_from_source(
range_str,
info,
response_headers,
track_download.then(|| state.clone()),
);
}
// ---- Buffered path: small files, compressed, images, range requests ----
@@ -1561,15 +1553,15 @@ async fn get_or_head_handler_inner(
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
// Check Range header
if let Some(range_header) = headers.get(header::RANGE) {
if let Ok(range_str) = range_header.to_str() {
return handle_range_request(
range_str,
&data,
response_headers,
track_download.then(|| state.clone()),
);
}
if let Some(range_header) = headers.get(header::RANGE)
&& let Ok(range_str) = range_header.to_str()
{
return handle_range_request(
range_str,
&data,
response_headers,
track_download.then(|| state.clone()),
);
}
if method == Method::HEAD {
@@ -1995,7 +1987,7 @@ fn extract_extension_from_path(path: &str) -> String {
if let Some(dot_pos) = filename.rfind('.') {
return filename[dot_pos..].to_lowercase();
}
} else if parts.len() >= 1 {
} else if !parts.is_empty() {
// 2-segment path: /vid,fid.ext or /vid/fid.ext
// Go's parseURLPath extracts ext from the full path for all formats
let last = parts[parts.len() - 1];
@@ -2118,7 +2110,7 @@ pub async fn post_handler(
// Go's r.ParseForm() returns 400 on malformed query strings
return json_error_with_query(
StatusCode::BAD_REQUEST,
&format!("form parse error: {}", e),
format!("form parse error: {}", e),
Some(&query),
);
}
@@ -2134,11 +2126,12 @@ pub async fn post_handler(
// JWT check for writes
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) = state
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, true)
.is_err()
{
return json_error_with_query(StatusCode::UNAUTHORIZED, "wrong jwt", Some(&query));
}
@@ -2243,7 +2236,10 @@ pub async fn post_handler(
// With a limit configured, an error here means the body exceeded it
// before we buffered the whole thing; report it like the size check.
let msg = if state.file_size_limit_bytes > 0 {
format!("file over the limited {} bytes", state.file_size_limit_bytes)
format!(
"file over the limited {} bytes",
state.file_size_limit_bytes
)
} else {
format!("read body: {}", e)
};
@@ -2265,11 +2261,8 @@ pub async fn post_handler(
.split(';')
.find_map(|part| {
let part = part.trim();
if let Some(val) = part.strip_prefix("boundary=") {
Some(val.trim_matches('"').to_string())
} else {
None
}
part.strip_prefix("boundary=")
.map(|val| val.trim_matches('"').to_string())
})
.unwrap_or_default();
@@ -2411,17 +2404,17 @@ pub async fn post_handler(
} else {
None
};
if let (Some(ref expected_md5), Some(ref actual_md5)) = (&content_md5, &original_content_md5) {
if expected_md5 != actual_md5 {
return json_error_with_query(
StatusCode::BAD_REQUEST,
format!(
"Content-MD5 did not match md5 of file data expected [{}] received [{}] size {}",
expected_md5, actual_md5, original_data_size
),
Some(&query),
);
}
if let (Some(expected_md5), Some(actual_md5)) = (&content_md5, &original_content_md5)
&& expected_md5 != actual_md5
{
return json_error_with_query(
StatusCode::BAD_REQUEST,
format!(
"Content-MD5 did not match md5 of file data expected [{}] received [{}] size {}",
expected_md5, actual_md5, original_data_size
),
Some(&query),
);
}
let now = std::time::SystemTime::now()
@@ -2563,7 +2556,7 @@ pub async fn post_handler(
cookie,
data_size: final_data.len() as u32,
data: final_data,
last_modified: last_modified,
last_modified,
..Needle::default()
};
n.set_has_last_modified_date();
@@ -2581,22 +2574,21 @@ pub async fn post_handler(
}
// Set TTL on needle
if let Some(ref t) = ttl {
if !t.is_empty() {
n.ttl = Some(*t);
n.set_has_ttl();
}
if let Some(ref t) = ttl
&& !t.is_empty()
{
n.ttl = Some(*t);
n.set_has_ttl();
}
// Set pairs on needle
if !pair_map.is_empty() {
if let Ok(pairs_json) = serde_json::to_vec(&pair_map) {
if pairs_json.len() < 65536 {
n.pairs_size = pairs_json.len() as u16;
n.pairs = pairs_json;
n.set_has_pairs();
}
}
if !pair_map.is_empty()
&& let Ok(pairs_json) = serde_json::to_vec(&pair_map)
&& pairs_json.len() < 65536
{
n.pairs_size = pairs_json.len() as u16;
n.pairs = pairs_json;
n.set_has_pairs();
}
// Set filename on needle (matches Go: if len(pu.FileName) < 256)
@@ -2625,9 +2617,9 @@ pub async fn post_handler(
if !is_replicate && write_result.is_ok() && !state.master_url.is_empty() {
let needs_replication = {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, v)| {
v.super_block.replica_placement.get_copy_count() > 1
})
store
.find_volume(vid)
.is_some_and(|(_, v)| v.super_block.replica_placement.get_copy_count() > 1)
};
if needs_replication {
let state_clone = state.clone();
@@ -2650,7 +2642,7 @@ pub async fn post_handler(
let replication_result = replication
.await
.map_err(|e| format!("replication task failed: {}", e))
.and_then(|result| result);
.flatten();
if let Err(e) = replication_result {
tracing::error!("replicated write failed: {}", e);
return json_error_with_query(
@@ -2744,11 +2736,12 @@ pub async fn delete_handler(
// JWT check for writes (deletes use write key)
let file_id = extract_file_id(&path);
let token = extract_jwt(&headers, request.uri());
if let Err(_) = state
if state
.guard
.read()
.unwrap()
.check_jwt_for_file(token.as_deref(), &file_id, true)
.is_err()
{
return json_error_with_query(StatusCode::UNAUTHORIZED, "wrong jwt", Some(&del_query));
}
@@ -2779,14 +2772,14 @@ pub async fn delete_handler(
let count = ec_needle.data_size as i64;
// Step 3: Journal the delete
let mut store = state.store.write().unwrap();
if let Some(ecv) = store.find_ec_volume_mut(vid) {
if let Err(e) = ecv.journal_delete(needle_id) {
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Deletion Failed: {}", e),
Some(&del_query),
);
}
if let Some(ecv) = store.find_ec_volume_mut(vid)
&& let Err(e) = ecv.journal_delete(needle_id)
{
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("Deletion Failed: {}", e),
Some(&del_query),
);
}
let result = DeleteResult { size: count };
return json_response_with_params(
@@ -2932,12 +2925,12 @@ pub async fn delete_handler(
if !is_replicate && delete_result.is_ok() && !state.master_url.is_empty() {
let needs_replication = {
let store = state.store.read().unwrap();
store.find_volume(vid).map_or(false, |(_, v)| {
v.super_block.replica_placement.get_copy_count() > 1
})
store
.find_volume(vid)
.is_some_and(|(_, v)| v.super_block.replica_placement.get_copy_count() > 1)
};
if needs_replication {
if let Err(e) = do_replicated_request(
if needs_replication
&& let Err(e) = do_replicated_request(
&state,
vid.0,
Method::DELETE,
@@ -2947,14 +2940,13 @@ pub async fn delete_handler(
None,
)
.await
{
tracing::error!("replicated delete failed: {}", e);
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("replication failed: {}", e),
Some(&del_query),
);
}
{
tracing::error!("replicated delete failed: {}", e);
return json_error_with_query(
StatusCode::INTERNAL_SERVER_ERROR,
format!("replication failed: {}", e),
Some(&del_query),
);
}
}
@@ -3113,6 +3105,11 @@ pub async fn healthz_handler(State(state): State<Arc<VolumeServerState>>) -> Res
if !state.is_heartbeating.load(Ordering::Relaxed) {
return StatusCode::SERVICE_UNAVAILABLE.into_response();
}
// A server with quarantined local replicas has faulty storage media;
// report degraded so a load balancer can drain it.
if state.store.read().unwrap().has_io_quarantine() {
return StatusCode::SERVICE_UNAVAILABLE.into_response();
}
StatusCode::OK.into_response()
}
@@ -3215,7 +3212,6 @@ pub async fn ui_handler(State(state): State<Arc<VolumeServerState>>) -> Response
// ============================================================================
#[derive(Deserialize)]
#[allow(dead_code)]
struct ChunkManifest {
#[serde(default)]
name: String,
@@ -3235,6 +3231,7 @@ struct ChunkInfo {
}
/// Try to expand a chunk manifest needle. Returns None if manifest can't be parsed.
#[expect(clippy::too_many_arguments)]
async fn try_expand_chunk_manifest(
state: &Arc<VolumeServerState>,
n: &Needle,
@@ -3363,53 +3360,53 @@ async fn try_expand_chunk_manifest(
response_headers.insert(header::ACCEPT_RANGES, "bytes".parse().unwrap());
// Last-Modified — Go sets this on the response writer before tryHandleChunkedFile
if let Some(ref lm) = last_modified_str {
if let Ok(hval) = lm.parse() {
response_headers.insert(header::LAST_MODIFIED, hval);
}
if let Some(lm) = last_modified_str
&& let Ok(hval) = lm.parse()
{
response_headers.insert(header::LAST_MODIFIED, hval);
}
// Pairs — Go sets needle pairs on the response writer before tryHandleChunkedFile
if n.has_pairs() && !n.pairs.is_empty() {
if let Ok(pair_map) =
if n.has_pairs()
&& !n.pairs.is_empty()
&& let Ok(pair_map) =
serde_json::from_slice::<std::collections::HashMap<String, String>>(&n.pairs)
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
{
for (k, v) in &pair_map {
if let (Ok(hname), Ok(hval)) = (
axum::http::HeaderName::from_bytes(k.as_bytes()),
axum::http::HeaderValue::from_str(v),
) {
response_headers.insert(hname, hval);
}
}
}
// S3 response passthrough headers — Go sets these via AdjustPassthroughHeaders
if let Some(ref cc) = query.response_cache_control {
if let Ok(hval) = cc.parse() {
response_headers.insert(header::CACHE_CONTROL, hval);
}
if let Some(ref cc) = query.response_cache_control
&& let Ok(hval) = cc.parse()
{
response_headers.insert(header::CACHE_CONTROL, hval);
}
if let Some(ref ce) = query.response_content_encoding {
if let Ok(hval) = ce.parse() {
response_headers.insert(header::CONTENT_ENCODING, hval);
}
if let Some(ref ce) = query.response_content_encoding
&& let Ok(hval) = ce.parse()
{
response_headers.insert(header::CONTENT_ENCODING, hval);
}
if let Some(ref exp) = query.response_expires {
if let Ok(hval) = exp.parse() {
response_headers.insert(header::EXPIRES, hval);
}
if let Some(ref exp) = query.response_expires
&& let Ok(hval) = exp.parse()
{
response_headers.insert(header::EXPIRES, hval);
}
if let Some(ref cl) = query.response_content_language {
if let Ok(hval) = cl.parse() {
response_headers.insert("Content-Language", hval);
}
if let Some(ref cl) = query.response_content_language
&& let Ok(hval) = cl.parse()
{
response_headers.insert("Content-Language", hval);
}
if let Some(ref cd) = query.response_content_disposition {
if let Ok(hval) = cd.parse() {
response_headers.insert(header::CONTENT_DISPOSITION, hval);
}
if let Some(ref cd) = query.response_content_disposition
&& let Ok(hval) = cd.parse()
{
response_headers.insert(header::CONTENT_DISPOSITION, hval);
}
// Content-Disposition
@@ -3440,7 +3437,6 @@ async fn try_expand_chunk_manifest(
} else {
String::new()
};
let mut result = result;
if is_image_crop_ext(&cm_ext) {
result = maybe_crop_image(&result, &cm_ext, query);
}
@@ -3464,10 +3460,7 @@ async fn try_expand_chunk_manifest(
/// (reconstruct-on-read from surviving shards), or a peer resolved via the
/// master. Mirrors Go's ChunkedFileReader, which looks every chunk up through
/// the master instead of assuming a local regular needle.
async fn read_chunk_needle(
state: &Arc<VolumeServerState>,
fid: &str,
) -> Result<Vec<u8>, String> {
async fn read_chunk_needle(state: &Arc<VolumeServerState>, fid: &str) -> Result<Vec<u8>, String> {
let (vid, nid, cookie) =
parse_url_path(fid).ok_or_else(|| format!("invalid chunk fid: {}", fid))?;
@@ -3719,33 +3712,33 @@ fn extract_jwt(headers: &HeaderMap, uri: &axum::http::Uri) -> Option<String> {
// 1. Check ?jwt= query parameter
if let Some(query) = uri.query() {
for pair in query.split('&') {
if let Some(value) = pair.strip_prefix("jwt=") {
if !value.is_empty() {
return Some(value.to_string());
}
if let Some(value) = pair.strip_prefix("jwt=")
&& !value.is_empty()
{
return Some(value.to_string());
}
}
}
// 2. Check Authorization: Bearer <token> (case-insensitive prefix)
if let Some(auth) = headers.get(header::AUTHORIZATION) {
if let Ok(auth_str) = auth.to_str() {
if auth_str.len() > 7 && auth_str[..7].eq_ignore_ascii_case("bearer ") {
return Some(auth_str[7..].to_string());
}
}
if let Some(auth) = headers.get(header::AUTHORIZATION)
&& let Ok(auth_str) = auth.to_str()
&& auth_str.len() > 7
&& auth_str[..7].eq_ignore_ascii_case("bearer ")
{
return Some(auth_str[7..].to_string());
}
// 3. Check Cookie
if let Some(cookie_header) = headers.get(header::COOKIE) {
if let Ok(cookie_str) = cookie_header.to_str() {
for cookie in cookie_str.split(';') {
let cookie = cookie.trim();
if let Some(value) = cookie.strip_prefix("AT=") {
if !value.is_empty() {
return Some(value.to_string());
}
}
if let Some(cookie_header) = headers.get(header::COOKIE)
&& let Ok(cookie_str) = cookie_header.to_str()
{
for cookie in cookie_str.split(';') {
let cookie = cookie.trim();
if let Some(value) = cookie.strip_prefix("AT=")
&& !value.is_empty()
{
return Some(value.to_string());
}
}
}
@@ -4178,6 +4171,8 @@ mod tests {
url: "volume.internal:8080".to_string(),
public_url: "volume.public:8080".to_string(),
grpc_port: 18080,
read_only: false,
read_only_can_delete: false,
};
let response = redirect_request(&info, &target, "https");
@@ -4204,6 +4199,8 @@ mod tests {
url: "volume.internal:8080.18080".to_string(),
public_url: "volume.public:8080.18080".to_string(),
grpc_port: 18080,
read_only: false,
read_only_can_delete: false,
};
let response = redirect_request(&info, &target, "http");
@@ -4250,15 +4247,19 @@ mod tests {
let app = Router::new().route(
"/dir/lookup",
get(|axum::extract::Query(params): axum::extract::Query<std::collections::HashMap<String, String>>| async move {
assert_eq!(params.get("volumeId").map(String::as_str), Some("31"));
axum::Json(serde_json::json!({
"volumeOrFileId": "31",
"locations": [
{"url": "10.0.0.2:5301", "publicUrl": "10.0.0.2:5301", "grpcPort": 5311}
]
}))
}),
get(
|axum::extract::Query(params): axum::extract::Query<
std::collections::HashMap<String, String>,
>| async move {
assert_eq!(params.get("volumeId").map(String::as_str), Some("31"));
axum::Json(serde_json::json!({
"volumeOrFileId": "31",
"locations": [
{"url": "10.0.0.2:5301", "publicUrl": "10.0.0.2:5301", "grpcPort": 5311}
]
}))
},
),
);
let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await.unwrap();
+138 -49
View File
@@ -19,11 +19,12 @@ use crate::pb::master_pb::seaweed_client::SeaweedClient;
use crate::pb::volume_server_pb;
use crate::remote_storage::s3_tier::{S3TierBackend, S3TierConfig};
use crate::storage::store::Store;
use crate::storage::types::{NeedleId, VolumeId};
use crate::storage::volume_report::VolumeReportKey;
use crate::storage::volume_report_hash::report_hash;
use crate::storage::types::NeedleId;
const DUPLICATE_UUID_RETRY_MESSAGE: &str = "duplicate UUIDs detected, retrying connection";
const VOLUME_IO_ERROR_TOLERANCE: i32 = 3;
const MAX_DUPLICATE_UUID_RETRIES: u32 = 3;
/// Configuration for the heartbeat client.
@@ -92,8 +93,15 @@ pub async fn run_heartbeat_with_state(
SleepDuplicate(Duration),
SleepPulse,
}
let action = match do_heartbeat(&config, &state, &grpc_addr, &target_addr, pulse, &mut shutdown_rx)
.await
let action = match do_heartbeat(
&config,
&state,
&grpc_addr,
&target_addr,
pulse,
&mut shutdown_rx,
)
.await
{
Ok(Some(leader)) => {
info!("Master leader changed to {}", leader);
@@ -181,10 +189,10 @@ pub async fn run_heartbeat_with_state(
pub fn to_grpc_address(master_addr: &str) -> String {
if let Some((host, port_str)) = master_addr.rsplit_once(':') {
// "host:port.grpcPort" — the part after the last '.' is the gRPC port.
if let Some((_, grpc_port)) = port_str.rsplit_once('.') {
if grpc_port.parse::<u16>().is_ok() {
return format!("{}:{}", host, grpc_port);
}
if let Some((_, grpc_port)) = port_str.rsplit_once('.')
&& grpc_port.parse::<u16>().is_ok()
{
return format!("{}:{}", host, grpc_port);
}
if let Ok(port) = port_str.parse::<u16>() {
let grpc_port = port + 10000;
@@ -308,6 +316,10 @@ fn collect_ec_shard_delta_messages(
for (disk_id, loc) in store.locations.iter().enumerate() {
for (_, ec_vol) in loc.ec_volumes() {
let (_, _, quarantined) = ec_vol.get_io_error_state();
if quarantined {
continue;
}
for shard in ec_vol.shards.iter().flatten() {
messages.insert(
(
@@ -418,8 +430,7 @@ async fn do_heartbeat(
// form so Ping admission can recognise it once a leader change moves us
// off the seed list. Mirrors Go's vs.setCurrentMaster(masterAddress).
{
let normalised =
super::volume_server::to_http_address(current_master).into_owned();
let normalised = super::volume_server::to_http_address(current_master).into_owned();
let mut guard = state.current_master_url.write().await;
*guard = normalised;
}
@@ -538,7 +549,12 @@ async fn do_heartbeat(
let mut del_vols = Vec::new();
for (id, vol) in &current_volumes {
if !last_volumes.contains_key(id) {
if let Some(previous) = last_volumes.get(id) {
if previous != vol {
del_vols.push(previous.to_short_message(*id));
new_vols.push(vol.to_short_message(*id));
}
} else {
new_vols.push(vol.to_short_message(*id));
}
}
@@ -730,7 +746,7 @@ fn parse_bool_property(value: Option<&String>) -> bool {
/// information message the heartbeat carries. A server holding millions of
/// volumes cannot keep a whole message for each just to notice one leave; the
/// Go report state keeps the same fields for the same reason.
#[derive(Clone)]
#[derive(Clone, PartialEq)]
struct VolumeIdentity {
collection: String,
disk_type: String,
@@ -738,6 +754,8 @@ struct VolumeIdentity {
replica_placement: u32,
ttl: u32,
disk_id: u32,
read_only: bool,
read_only_can_delete: bool,
}
impl VolumeIdentity {
@@ -749,6 +767,8 @@ impl VolumeIdentity {
replica_placement: v.replica_placement,
ttl: v.ttl,
disk_id: v.disk_id,
read_only: v.read_only,
read_only_can_delete: v.read_only_can_delete,
}
}
@@ -761,6 +781,8 @@ impl VolumeIdentity {
ttl: self.ttl,
disk_type: self.disk_type.clone(),
disk_id: self.disk_id,
read_only: self.read_only,
read_only_can_delete: self.read_only_can_delete,
}
}
}
@@ -778,7 +800,10 @@ fn volume_identities(
fn collect_heartbeat_with_snapshot(
config: &HeartbeatConfig,
state: &Arc<VolumeServerState>,
) -> (master_pb::Heartbeat, Vec<master_pb::VolumeInformationMessage>) {
) -> (
master_pb::Heartbeat,
Vec<master_pb::VolumeInformationMessage>,
) {
let mut store = state.store.write().unwrap();
let (ec_shards, deleted_ec_shards) = store.delete_expired_ec_volumes();
build_heartbeat_with_ec_status(
@@ -856,7 +881,10 @@ fn build_heartbeat_with_ec_status(
deleted_ec_shards: Vec<master_pb::VolumeEcShardInformationMessage>,
has_no_ec_shards: bool,
commit_report: bool,
) -> (master_pb::Heartbeat, Vec<master_pb::VolumeInformationMessage>) {
) -> (
master_pb::Heartbeat,
Vec<master_pb::VolumeInformationMessage>,
) {
const MAX_TTL_VOLUME_REMOVAL_DELAY: u32 = 10;
#[derive(Default)]
@@ -872,6 +900,7 @@ fn build_heartbeat_with_ec_status(
// master can tell whether applying what it was sent leaves it current.
// Volumes skipped below -- quarantined, phantom, expired -- are in neither.
let mut volume_digest: u64 = 0;
let mut quarantined_volumes: u32 = 0;
let (send_full_list, report_generation, report_pass) = store.volume_report.begin();
let mut changed_volumes = Vec::new();
let mut max_file_key = NeedleId(0);
@@ -893,10 +922,9 @@ fn build_heartbeat_with_ec_status(
let mut effective_max_count = loc.max_volume_count.load(Ordering::Relaxed);
if loc.is_disk_space_low.load(Ordering::Relaxed) {
let used_slots = loc.volumes_len() as i32
+ ((loc.ec_shard_count()
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
+ loc
.ec_shard_count()
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
effective_max_count = used_slots;
}
@@ -914,6 +942,7 @@ fn build_heartbeat_with_ec_status(
loc.disk_free_bytes.load(Ordering::Relaxed);
let mut delete_vids = Vec::new();
let mut quarantine_vids: Vec<VolumeId> = Vec::new();
for (_, vol) in loc.iter_volumes() {
let cur_max = vol.max_file_key();
if cur_max > max_file_key {
@@ -923,9 +952,18 @@ fn build_heartbeat_with_ec_status(
let volume_size = vol.dat_file_size().unwrap_or(0);
let mut should_delete_volume = false;
if vol.last_io_error().is_some() {
delete_vids.push(vol.id);
should_delete_volume = true;
let (_, io_count, io_quarantined) = vol.get_io_error_state();
if io_quarantined || io_count >= VOLUME_IO_ERROR_TOLERANCE {
if !io_quarantined {
vol.mark_io_quarantined();
warn!(
"Volume {} quarantined after {} consecutive IO errors",
vol.id.0, io_count
);
}
quarantined_volumes += 1;
quarantine_vids.push(vol.id);
continue;
} else if !vol.is_expired(volume_size, volume_size_limit) {
// Detect phantom volumes: the .dat was unlinked from disk but is still
// held open as a deleted FD, so the volume keeps serving and heartbeating
@@ -939,7 +977,9 @@ fn build_heartbeat_with_ec_status(
.duration_since(UNIX_EPOCH)
.unwrap_or(Duration::ZERO)
.as_nanos() as i64;
if now_ns - vol.last_disk_check_ns.load(Ordering::Relaxed) > DISK_CHECK_INTERVAL_NS {
if now_ns - vol.last_disk_check_ns.load(Ordering::Relaxed)
> DISK_CHECK_INTERVAL_NS
{
if !Path::new(&vol.file_name(".dat")).exists() {
warn!("Volume {}: data file {} missing (held open as deleted FD) - not reporting to master", vol.id.0, vol.file_name(".dat"));
continue;
@@ -957,11 +997,15 @@ fn build_heartbeat_with_ec_status(
delete_count: vol.deleted_count() as u64,
deleted_byte_count: vol.deleted_size(),
read_only: vol.is_read_only(),
read_only_can_delete: vol.is_no_write_can_delete(),
replica_placement: vol.super_block.replica_placement.to_byte() as u32,
version: vol.super_block.version.0 as u32,
ttl: vol.super_block.ttl.to_u32(),
compact_revision: vol.super_block.compaction_revision as u32,
modified_at_second: vol.last_modified_ts() as i64,
// The .dat mtime, as Go reports: the shell's quiet-period
// gates read this as "last touched", which a delete has to
// count towards even though the TTL clock ignores it.
modified_at_second: vol.dat_file_mod_time() as i64,
disk_type: loc.disk_type.to_string(),
disk_id: disk_id as u32,
remote_storage_name,
@@ -1012,12 +1056,17 @@ fn build_heartbeat_with_ec_status(
}
}
}
}
for vid in delete_vids {
let _ = loc.delete_volume(vid, false, false);
}
for vid in quarantine_vids {
if let Some(vol) = loc.find_volume_mut(vid) {
vol.set_no_write_or_delete(true);
}
}
}
// Update disk size and read-only gauges
@@ -1074,6 +1123,23 @@ fn build_heartbeat_with_ec_status(
};
let (location_uuids, disk_tags) = collect_location_metadata(store, &disk_max_by_id);
let mut quarantined_ec_shards: u32 = 0;
for loc in &store.locations {
for (_, ec_vol) in loc.ec_volumes() {
let (_, _, quarantined) = ec_vol.get_io_error_state();
if quarantined {
quarantined_ec_shards += ec_vol.shard_count() as u32;
}
}
}
crate::metrics::IO_QUARANTINE_GAUGE
.with_label_values(&["volume"])
.set(quarantined_volumes as i64);
crate::metrics::IO_QUARANTINE_GAUGE
.with_label_values(&["ec_shard"])
.set(quarantined_ec_shards as i64);
let heartbeat = master_pb::Heartbeat {
id: store.id.clone(),
ip: config.ip.clone(),
@@ -1109,6 +1175,10 @@ fn collect_live_ec_shards(
for (disk_id, loc) in store.locations.iter().enumerate() {
for (_, ec_vol) in loc.ec_volumes() {
let (_, _, quarantined) = ec_vol.get_io_error_state();
if quarantined {
continue;
}
for message in ec_vol.to_volume_ec_shard_information_messages(disk_id as u32) {
if update_metrics {
let total_size: u64 = message
@@ -1143,7 +1213,10 @@ fn collect_live_ec_shards(
}
/// Collect EC shard information into a Heartbeat message.
fn collect_ec_heartbeat(config: &HeartbeatConfig, state: &Arc<VolumeServerState>) -> master_pb::Heartbeat {
fn collect_ec_heartbeat(
config: &HeartbeatConfig,
state: &Arc<VolumeServerState>,
) -> master_pb::Heartbeat {
let store = state.store.read().unwrap();
let ec_shards = collect_live_ec_shards(&store, true);
@@ -1166,10 +1239,10 @@ mod tests {
use crate::config::MinFreeSpace;
use crate::config::ReadMode;
use crate::metrics::{
DISK_SIZE_GAUGE, DISK_SIZE_LABEL_DELETED_BYTES, DISK_SIZE_LABEL_EC,
DISK_SIZE_LABEL_NORMAL, READ_ONLY_LABEL_IS_DISK_SPACE_LOW,
READ_ONLY_LABEL_IS_READ_ONLY, READ_ONLY_LABEL_NO_WRITE_CAN_DELETE,
READ_ONLY_LABEL_NO_WRITE_OR_DELETE, READ_ONLY_VOLUME_GAUGE,
DISK_SIZE_GAUGE, DISK_SIZE_LABEL_DELETED_BYTES, DISK_SIZE_LABEL_EC, DISK_SIZE_LABEL_NORMAL,
READ_ONLY_LABEL_IS_DISK_SPACE_LOW, READ_ONLY_LABEL_IS_READ_ONLY,
READ_ONLY_LABEL_NO_WRITE_CAN_DELETE, READ_ONLY_LABEL_NO_WRITE_OR_DELETE,
READ_ONLY_VOLUME_GAUGE,
};
use crate::remote_storage::s3_tier::S3TierRegistry;
use crate::security::{Guard, SigningKey};
@@ -1254,7 +1327,10 @@ mod tests {
fn test_to_grpc_address_explicit_grpc_port() {
// host:port.grpcPort form — gRPC port is what's after the dot.
assert_eq!(to_grpc_address("10.85.183.6:5300.6300"), "10.85.183.6:6300");
assert_eq!(to_grpc_address("master.local:9333.19333"), "master.local:19333");
assert_eq!(
to_grpc_address("master.local:9333.19333"),
"master.local:19333"
);
}
#[test]
@@ -1308,7 +1384,10 @@ mod tests {
heartbeat.disk_tags[0].tags,
vec!["fast".to_string(), "ssd".to_string()]
);
assert_eq!(heartbeat.disk_tags[0].r#type, DiskType::HardDrive.to_string());
assert_eq!(
heartbeat.disk_tags[0].r#type,
DiskType::HardDrive.to_string()
);
assert_eq!(heartbeat.disk_tags[0].max_volume_count, 3);
}
@@ -1347,7 +1426,10 @@ mod tests {
let heartbeat = build_heartbeat(&test_config(), &mut store);
assert_eq!(heartbeat.disk_tags[0].max_volume_count, 1);
assert_eq!(heartbeat.max_volume_counts[&DiskType::HardDrive.to_string()], 1);
assert_eq!(
heartbeat.max_volume_counts[&DiskType::HardDrive.to_string()],
1
);
}
#[test]
@@ -1565,10 +1647,9 @@ mod tests {
let heartbeat = build_heartbeat(&test_config(), &mut store);
assert_eq!(heartbeat.volumes.len(), 2);
let expected = heartbeat
.volumes
.iter()
.fold(0u64, |acc, m| acc ^ crate::storage::volume_report_hash::report_hash(m));
let expected = heartbeat.volumes.iter().fold(0u64, |acc, m| {
acc ^ crate::storage::volume_report_hash::report_hash(m)
});
assert_eq!(heartbeat.volume_digest, Some(expected));
assert_ne!(heartbeat.volume_digest, Some(0));
}
@@ -1663,10 +1744,9 @@ mod tests {
.iter()
.flat_map(|family| family.get_metric().to_vec())
.filter(|metric| {
metric
.get_label()
.iter()
.any(|label| label.get_name() == "collection" && label.get_value() == collection)
metric.get_label().iter().any(|label| {
label.get_name() == "collection" && label.get_value() == collection
})
})
.count()
}
@@ -1800,7 +1880,9 @@ mod tests {
assert_eq!(heartbeat.ec_shards[0].disk_id, 0);
assert_eq!(
heartbeat.ec_shards[0].disk_type,
state.store.read().unwrap().locations[0].disk_type.to_string()
state.store.read().unwrap().locations[0]
.disk_type
.to_string()
);
assert_eq!(heartbeat.ec_shards[0].ec_index_bits, 1);
assert_eq!(heartbeat.ec_shards[0].shard_sizes, vec![8]);
@@ -1936,8 +2018,13 @@ mod tests {
let heartbeat = build_heartbeat(&test_config(), &mut store);
// A sustained IO error quarantines the volume: it stays mounted
// (so healthz can observe the quarantine state) but is not
// advertised to the master.
assert!(heartbeat.volumes.is_empty());
assert!(!store.has_volume(VolumeId(51)));
assert!(store.has_volume(VolumeId(51)));
let (_, volume) = store.find_volume_mut(VolumeId(51)).unwrap();
assert!(volume.is_no_write_or_delete());
}
#[test]
@@ -1968,12 +2055,15 @@ mod tests {
)
.unwrap();
let (_, volume) = store.find_volume_mut(VolumeId(71)).unwrap();
volume.volume_info.files.push(crate::storage::volume::PbRemoteFile {
backend_type: "s3".to_string(),
backend_id: "archive".to_string(),
key: "volumes/71.dat".to_string(),
..Default::default()
});
volume
.volume_info
.files
.push(crate::storage::volume::PbRemoteFile {
backend_type: "s3".to_string(),
backend_id: "archive".to_string(),
key: "volumes/71.dat".to_string(),
..Default::default()
});
volume.refresh_remote_write_mode().unwrap();
let heartbeat = build_heartbeat(&test_config(), &mut store);
@@ -2137,8 +2227,7 @@ mod tests {
.mount_ec_shards(VolumeId(81), "ec_delta_case", &[0], "")
.unwrap();
let current = collect_ec_shard_delta_messages(&store);
let (new_ec_shards, deleted_ec_shards) =
diff_ec_shard_delta_messages(&previous, &current);
let (new_ec_shards, deleted_ec_shards) = diff_ec_shard_delta_messages(&previous, &current);
assert_eq!(new_ec_shards.len(), 1);
assert!(deleted_ec_shards.is_empty());
+323 -180
View File
@@ -134,8 +134,13 @@ pub async fn read_ec_shard_needle_distributed(
Ok(fresh) => {
// A complete reply merges into the cache; an incomplete one
// (< data_shards) is left unwritten — keep the prior cache.
match write_back_shard_locations(state, vid, fresh, snapshot.data_shards as usize)
{
match write_back_shard_locations(
state,
vid,
fresh,
snapshot.data_shards as usize,
snapshot.encode_ts_ns,
) {
Some(merged) => shard_locations = merged,
// An incomplete reply leaves the cache unwritten and its refresh
// time unadvanced, so the mark this refresh consumed goes back.
@@ -168,36 +173,37 @@ pub async fn read_ec_shard_needle_distributed(
let parity_shards = snapshot.parity_shards as usize;
let encode_ts_ns = snapshot.encode_ts_ns;
let intervals = std::mem::take(&mut snapshot.intervals);
let fetched: Vec<io::Result<(Vec<u8>, bool)>> = stream::iter(intervals.into_iter().map(|res| {
let shard_locations = &shard_locations;
async move {
match res {
IntervalResult::Local(buf) => Ok((buf, false)),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
} => {
fetch_one_interval(
state,
vid,
needle_id,
let fetched: Vec<io::Result<(Vec<u8>, bool)>> =
stream::iter(intervals.into_iter().map(|res| {
let shard_locations = &shard_locations;
async move {
match res {
IntervalResult::Local(buf) => Ok((buf, false)),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
encode_ts_ns,
)
.await
} => {
fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
encode_ts_ns,
)
.await
}
}
}
}
}))
.buffered(INTERVAL_READ_CONCURRENCY)
.collect()
.await;
}))
.buffered(INTERVAL_READ_CONCURRENCY)
.collect()
.await;
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(fetched.len());
for res in fetched {
@@ -230,8 +236,10 @@ pub async fn read_ec_shard_needle_distributed(
));
}
let mut n = Needle::default();
n.id = needle_id;
let mut n = Needle {
id: needle_id,
..Needle::default()
};
n.read_bytes(
&bytes,
snapshot.offset.to_actual_offset(),
@@ -256,22 +264,32 @@ pub async fn read_ec_shard_needle_distributed(
pub async fn scrub_ec_volume_distributed(
state: &Arc<VolumeServerState>,
vid: VolumeId,
expected_encode_ts_ns: i64,
force_deleted_needles_check: bool,
recover_unreadable: bool,
) -> (i64, Vec<crate::pb::volume_server_pb::EcShardInfo>, Vec<String>) {
// Phase A — under the Store read lock, run the index scrub and grab the
) -> (
i64,
Vec<crate::pb::volume_server_pb::EcShardInfo>,
Vec<String>,
) {
// Phase A — under the Store read lock, snapshot the index scrub and grab the
// paths/scalars + shard-location staleness; release the lock before any await.
let (
ecx_path,
collection,
seed_errs,
index_plan,
ecx_walk,
encode_ts_ns,
cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
) = {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
// Resolve the runtime matching the anchor's encode generation, not the
// first-match find_ec_volume — otherwise the needle walk can scan an
// older run while the parity half scans the newest.
let ecv = match find_ec_volume_for_scrub(&store, vid, expected_encode_ts_ns) {
Some(v) => v,
None => {
return (
@@ -282,7 +300,26 @@ pub async fn scrub_ec_volume_distributed(
}
};
// full scan means verifying the index as well
let (_, errs) = ecv.scrub_index();
let index_plan = ecv.scrub_index_plan();
// A SECOND .ecx descriptor, opened under the guard for the needle walk
// below. The index plan's handle is consumed by its own structural walk,
// and both seek, so a `dup` would race the cursor. Reopening by PATH
// after the guard is dropped would let a teardown that legitimately
// unlinks or replaces the .ecx (the heartbeat's
// delete_expired_ec_volumes, volume_ec_shards_delete) surface an
// intentional removal as a scrub error or mix index generations within
// one scrub — the same race the vanished-volume policy exists to hide.
// The descriptor outlives the name, the same way the checksum plan's
// shard handles do.
let ecx_walk = fs::File::open(ecv.ecx_file_name());
// Encode-run identity of the volume this scrub started against. The
// per-needle `scrub_snapshot_under_lock` re-resolves the volume by id
// under a fresh guard, so a teardown-and-remount of the same vid between
// two rows would otherwise apply the captured .ecx's offsets to a
// replacement volume's shards. Bind the walk to this generation: if the
// mounted volume's encode_ts_ns no longer matches, abort like a
// mid-scan unmount rather than mixing generations.
let encode_ts_ns = ecv.encode_ts_ns;
// Bind to locals so the inner RwLock/Mutex guards drop before the block ends.
let cached_locations = ecv.shard_locations.read().unwrap().clone();
let cache_refreshed_at = *ecv.shard_locations_refresh_time.lock().unwrap();
@@ -291,13 +328,47 @@ pub async fn scrub_ec_volume_distributed(
(
ecv.ecx_file_name(),
ecv.collection.clone(),
errs,
index_plan,
ecx_walk,
encode_ts_ns,
cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
)
};
// Lock released: walk the index now, before anything else appends to errs,
// so the seeded errors keep their position in the reported details.
//
// `index_plan.run()` reads the whole .ecx synchronously, so run it in the
// blocking pool rather than on this async worker — a large index scan would
// otherwise block unrelated RPC work handled on the same executor. Same
// treatment as the CHECKSUM/LOCAL plans in the gRPC handler.
let (_, seed_errs) = match tokio::task::spawn_blocking(move || index_plan.run()).await {
Ok(v) => v,
Err(e) => {
// A panic is evidence about the volume and counts as broken; a
// cancellation is not — spawn_blocking only reports it when the
// runtime is going down, the volume was never scanned, and the
// caller (FULL/READS) would put a false corruption into
// broken_volume_ids if it reached the errs path. Match the
// record_scrub_join_failure distinction used by the handler arms.
if e.is_panic() {
return (
0,
Vec::new(),
vec![format!(
"EC volume {} index scrub task panicked: {}",
vid.0, e
)],
);
}
// Cancellation: the runtime is shutting down, so this response is
// unlikely to reach anyone. Return clean rather than inventing a
// corruption for a volume that was never scanned.
return (0, Vec::new(), Vec::new());
}
};
let mut errs = seed_errs;
// Refresh the shard-location cache once up front (mirrors Go's
@@ -314,7 +385,9 @@ pub async fn scrub_ec_volume_distributed(
) {
match cached_lookup_ec_shard_locations(state, vid).await {
Ok(fresh) => {
if write_back_shard_locations(state, vid, fresh, data_shards).is_none() {
if write_back_shard_locations(state, vid, fresh, data_shards, expected_encode_ts_ns)
.is_none()
{
mark_shard_locations_stale(state, vid);
return (
0,
@@ -341,53 +414,88 @@ pub async fn scrub_ec_volume_distributed(
// walk, so per-needle snapshots no longer clone it.
let locations: HashMap<ShardId, Vec<String>> = {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
let ecv = match find_ec_volume_for_scrub(&store, vid, expected_encode_ts_ns) {
Some(v) => v,
None => {
return (
0,
Vec::new(),
vec![format!("EC volume id {} not found", vid.0)],
)
);
}
};
let map = ecv.shard_locations.read().unwrap().clone();
map
ecv.shard_locations.read().unwrap().clone()
};
// Walk the .ecx (private fd, no lock) for the row count + live (id, offset, size).
let mut count: i64 = 0;
let mut needles: Vec<(NeedleId, Offset, Size)> = Vec::new();
match fs::File::open(&ecx_path) {
Ok(mut f) => {
if let Err(e) = crate::storage::idx::walk_index_file(&mut f, 0, |id, offset, size| {
count += 1;
// Skip ALL deleted entries: -1 tombstones (runtime delete folded
// into .ecx) and -originalSize entries (a needle deleted on the
// regular volume before EC encode). get_actual_size uses the raw
// signed size, so a negative would yield empty intervals
// (false-positive) or an under-16-byte buffer (parse panic).
if !size.is_deleted() {
needles.push((id, offset, size));
// Walk the .ecx (private fd captured under the lock, no lock held) for the
// row count + live (id, offset, size). Reading through the captured
// descriptor — not a pathname reopen — keeps a concurrent teardown from
// surfacing an intentional removal as a scrub error or mixing index
// generations, the same invariant the vanished-volume policy enforces.
//
// `walk_index_file` reads the full .ecx synchronously, so run it in the
// blocking pool rather than on this async worker — same reason as
// `index_plan.run()` above.
let (count, needles, walk_errs) = match tokio::task::spawn_blocking(
move || -> (i64, Vec<(NeedleId, Offset, Size)>, Vec<String>) {
let mut count: i64 = 0;
let mut needles: Vec<(NeedleId, Offset, Size)> = Vec::new();
let mut walk_errs: Vec<String> = Vec::new();
match ecx_walk {
Ok(mut f) => {
if let Err(e) =
crate::storage::idx::walk_index_file(&mut f, 0, |id, offset, size| {
count += 1;
// Skip ALL deleted entries: -1 tombstones (runtime delete folded
// into .ecx) and -originalSize entries (a needle deleted on the
// regular volume before EC encode). get_actual_size uses the raw
// signed size, so a negative would yield empty intervals
// (false-positive) or an under-16-byte buffer (parse panic).
if !size.is_deleted() {
needles.push((id, offset, size));
}
Ok(())
})
{
walk_errs.push(format!("walk ECX file {}: {}", ecx_path, e));
}
}
Ok(())
}) {
errs.push(format!("walk ECX file {}: {}", ecx_path, e));
Err(e) => walk_errs.push(format!("open ECX file {}: {}", ecx_path, e)),
}
(count, needles, walk_errs)
},
)
.await
{
Ok(v) => v,
Err(e) => {
// A panic is evidence about the volume and counts as broken; a
// cancellation is not — see the index_plan join above for the
// same reasoning.
if e.is_panic() {
return (
0,
Vec::new(),
vec![format!("EC volume {} ecx walk task panicked: {}", vid.0, e)],
);
}
return (0, Vec::new(), Vec::new());
}
Err(e) => errs.push(format!("open ECX file {}: {}", ecx_path, e)),
}
};
errs.extend(walk_errs);
// reads for EC chunks can hit the same shard repeatedly, so dedupe broken shards
let mut broken_shards: HashMap<ShardId, crate::pb::volume_server_pb::EcShardInfo> = HashMap::new();
let mut broken_shards: HashMap<ShardId, crate::pb::volume_server_pb::EcShardInfo> =
HashMap::new();
for (id, offset, size) in needles {
// Per-needle snapshot under the lock from the RAW .ecx (offset, size) so
// logically-deleted needles are still verified; lock dropped before await.
let snapshot = match scrub_snapshot_under_lock(state, vid, offset, size) {
let snapshot = match scrub_snapshot_under_lock(state, vid, offset, size, encode_ts_ns) {
Ok(s) => s,
// Volume unmounted mid-scan: abort with an error rather than skipping
// every remaining needle, which would report a false-CLEAN result.
// Volume unmounted (or remounted as a different encode run) mid-scan:
// abort with an error rather than skipping every remaining needle,
// which would report a false-CLEAN result.
Err(e) if e.kind() == io::ErrorKind::NotFound => {
errs.push(format!("EC volume {} unmounted during scrub: {}", vid.0, e));
break;
@@ -513,7 +621,11 @@ pub async fn scrub_ec_volume_distributed(
// Mirror Go CmpEcShardInfo: sort by (volume_id, shard_id).
let mut broken: Vec<crate::pb::volume_server_pb::EcShardInfo> =
broken_shards.into_values().collect();
broken.sort_by(|a, b| a.volume_id.cmp(&b.volume_id).then(a.shard_id.cmp(&b.shard_id)));
broken.sort_by(|a, b| {
a.volume_id
.cmp(&b.volume_id)
.then(a.shard_id.cmp(&b.shard_id))
});
(count, broken, errs)
}
@@ -551,9 +663,10 @@ fn scrub_snapshot_under_lock(
vid: VolumeId,
offset: Offset,
size: Size,
expected_encode_ts: i64,
) -> io::Result<ScrubSnapshot> {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
let ecv = match find_ec_volume_for_scrub(&store, vid, expected_encode_ts) {
Some(v) => v,
// Volume unmounted mid-scan: a distinct NotFound so the caller aborts
// with an error rather than silently skipping (which would false-CLEAN).
@@ -564,6 +677,28 @@ fn scrub_snapshot_under_lock(
))
}
};
// The volume was torn down and remounted as a DIFFERENT encode run between
// two rows. The .ecx offsets captured at the start of the walk belong to
// the old generation; applying them to the replacement's shards would
// falsely report corruption. Abort like a mid-scan unmount instead of
// mixing generations within one scrub.
//
// `encode_ts_ns == 0` means the .vif carried no encode-run identity (a
// legacy or pre-feature volume). Two such volumes are NOT the same mount
// by this check alone — 0 == 0 would accept a teardown-and-remount and
// apply the old .ecx's offsets to the replacement's shards. Only treat a
// match as verified when the identity is non-zero; when it is zero, fall
// back to the pre-check behavior (no generation binding) rather than
// aborting a scrub that was already running without the guard.
if expected_encode_ts != 0 && ecv.encode_ts_ns != expected_encode_ts {
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!(
"EC volume {} remounted as a different encode run during scrub (was {}, now {})",
vid.0, expected_encode_ts, ecv.encode_ts_ns
),
));
}
let intervals = ecv.locate_ec_shard_needle_interval(offset.to_actual_offset(), size);
if intervals.is_empty() {
return Err(io::Error::new(
@@ -678,9 +813,9 @@ fn needs_refresh(
let ttl = if stale || shard_count < data_shards {
Duration::from_secs(11)
} else if shard_count == total_shards {
Duration::from_secs(37 * 60)
Duration::from_mins(37)
} else {
Duration::from_secs(7 * 60)
Duration::from_mins(7)
};
age >= ttl
}
@@ -734,22 +869,19 @@ async fn cached_lookup_ec_shard_locations(
}
};
if master.is_empty() {
return Err(io::Error::new(
io::ErrorKind::Other,
"no master configured for ec shard lookup",
));
return Err(io::Error::other("no master configured for ec shard lookup"));
}
let grpc_addr = parse_grpc_address(&master)
.map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let grpc_addr =
parse_grpc_address(&master).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(10))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("master connect: {}", e)))?;
.map_err(|e| io::Error::other(format!("master connect: {}", e)))?;
let mut client = SeaweedClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
@@ -758,7 +890,7 @@ async fn cached_lookup_ec_shard_locations(
let resp = client
.lookup_ec_volume(Request::new(LookupEcVolumeRequest { volume_id: vid.0 }))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("lookup_ec_volume: {}", e)))?;
.map_err(|e| io::Error::other(format!("lookup_ec_volume: {}", e)))?;
let resp = resp.into_inner();
let mut out = HashMap::new();
@@ -789,15 +921,34 @@ fn write_back_shard_locations(
vid: VolumeId,
locations: HashMap<ShardId, Vec<String>>,
data_shards: usize,
expected_encode_ts_ns: i64,
) -> Option<HashMap<ShardId, Vec<String>>> {
if locations.len() < data_shards {
return None;
}
let store = state.store.read().unwrap();
let ecv = store.find_ec_volume(vid)?;
let ecv = find_ec_volume_for_scrub(&store, vid, expected_encode_ts_ns)?;
Some(ecv.merge_shard_locations(locations))
}
/// Resolve the runtime matching the scrub's anchor encode generation, not the
/// first-match `find_ec_volume`. When `expected_encode_ts_ns` is 0 (legacy or
/// pre-feature), falls back to first-match so existing behavior is preserved.
fn find_ec_volume_for_scrub(
store: &crate::storage::store::Store,
vid: VolumeId,
expected_encode_ts_ns: i64,
) -> Option<&crate::storage::erasure_coding::EcVolume> {
if expected_encode_ts_ns != 0 {
store
.find_all_ec_volumes(vid)
.into_iter()
.find(|v| v.encode_ts_ns == expected_encode_ts_ns)
} else {
store.find_ec_volume(vid)
}
}
/// Build a SeaweedFS-style `host:httpPort.grpcPort` address from a
/// master `Location` so the result is what `parse_grpc_address` (and
/// the heartbeat path) already understand.
@@ -806,16 +957,17 @@ fn format_location_as_server_address(loc: &master_pb::Location) -> String {
.url
.trim_start_matches("http://")
.trim_start_matches("https://");
if loc.grpc_port > 0 {
if let Some((host, http_port)) = raw.rsplit_once(':') {
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
if loc.grpc_port > 0
&& let Some((host, http_port)) = raw.rsplit_once(':')
{
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
raw.to_string()
}
/// Try direct peer read; on failure, reconstruct via Reed-Solomon
/// from the other shards. Mirrors `readOneEcShardInterval`'s tail.
#[expect(clippy::too_many_arguments)]
async fn fetch_one_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
@@ -829,35 +981,35 @@ async fn fetch_one_interval(
expected_encode_ts_ns: i64,
) -> io::Result<(Vec<u8>, bool)> {
// Direct peer read against the cached locations for this shard.
if let Some(sources) = shard_locations.get(&shard_id) {
if !sources.is_empty() {
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
if let Some(sources) = shard_locations.get(&shard_id)
&& !sources.is_empty()
{
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
}
}
@@ -879,6 +1031,7 @@ async fn fetch_one_interval(
.await
}
#[expect(clippy::too_many_arguments)]
async fn read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
sources: &[String],
@@ -915,6 +1068,7 @@ async fn read_remote_ec_shard_interval(
}))
}
#[expect(clippy::too_many_arguments)]
async fn do_read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
source: &str,
@@ -928,18 +1082,13 @@ async fn do_read_remote_ec_shard_interval(
let grpc_addr =
parse_grpc_address(source).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("connect to {}: {}", source, e),
)
})?;
.map_err(|e| io::Error::other(format!("connect to {}: {}", source, e)))?;
// TODO(grpc-jwt): clusters with `jwt.signing.key` configured will
// reject peer-to-peer VolumeEcShardRead calls until the Rust
@@ -965,10 +1114,10 @@ async fn do_read_remote_ec_shard_interval(
.volume_ec_shard_read(Request::new(req))
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("volume_ec_shard_read {}.{} from {}: {}", vid.0, shard_id, source, e),
)
io::Error::other(format!(
"volume_ec_shard_read {}.{} from {}: {}",
vid.0, shard_id, source, e
))
})?;
let mut stream = resp.into_inner();
@@ -977,19 +1126,16 @@ async fn do_read_remote_ec_shard_interval(
while let Some(msg) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv: {}", e)))?
.map_err(|e| io::Error::other(format!("recv: {}", e)))?
{
// Validate the served shard's identity client-side, so the guard holds even
// against a pre-upgrade server that ignored the request field (returns 0).
// A mismatch fails the read; the caller recovers from parity.
if expected_encode_ts_ns != 0 && msg.encode_ts_ns != expected_encode_ts_ns {
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
),
));
return Err(io::Error::other(format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
)));
}
if msg.is_deleted {
is_deleted = true;
@@ -1025,6 +1171,7 @@ async fn do_read_remote_ec_shard_interval(
Ok((out, false))
}
#[expect(clippy::too_many_arguments)]
async fn recover_one_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
@@ -1038,12 +1185,8 @@ async fn recover_one_remote_ec_shard_interval(
expected_encode_ts_ns: i64,
) -> io::Result<(Vec<u8>, bool)> {
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon init: {:?}", e),
)
})?;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap. An interval whose
@@ -1053,13 +1196,10 @@ async fn recover_one_remote_ec_shard_interval(
.acquire_many((size * data_shards).min(EC_RECOVER_BUDGET) as u32)
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
))
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
@@ -1072,7 +1212,7 @@ async fn recover_one_remote_ec_shard_interval(
let mut available = 0usize;
{
let store = state.store.read().unwrap();
for sid in 0..total_shards {
for (sid, slot) in bufs.iter_mut().enumerate() {
if available >= data_shards {
break;
}
@@ -1085,13 +1225,21 @@ async fn recover_one_remote_ec_shard_interval(
// lenient only when the caller carries no identity (pre-upgrade).
// Mirrors Go's `readLocalEcShardInterval`.
let owner = match store.find_ec_volume_with_shard(vid, sid as u32) {
Some(ecv) if expected_encode_ts_ns == 0 || ecv.encode_ts_ns == expected_encode_ts_ns => ecv,
Some(ecv)
if expected_encode_ts_ns == 0 || ecv.encode_ts_ns == expected_encode_ts_ns =>
{
ecv
}
_ => continue,
};
if let Some(Some(shard)) = owner.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
if shard
.read_at(&mut buf, shard_offset as u64)
.map(|n| n == size)
.unwrap_or(false)
{
*slot = Some(buf);
available += 1;
}
}
@@ -1175,34 +1323,25 @@ async fn recover_one_remote_ec_shard_interval(
if any_deleted {
return Ok((Vec::new(), true));
}
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
),
));
return Err(io::Error::other(format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
)));
}
rs.reconstruct(&mut bufs).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
))
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok((buf, any_deleted)),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
),
)),
None => Err(io::Error::other(format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
))),
}
}
@@ -1344,12 +1483,12 @@ async fn fetch_ec_index_from_one_peer(
let grpc_addr =
parse_grpc_address(peer).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let channel = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?
.map_err(|e| io::Error::other(e.to_string()))?
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("connect {}: {}", peer, e)))?;
.map_err(|e| io::Error::other(format!("connect {}: {}", peer, e)))?;
let mut client = VolumeServerClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE);
@@ -1370,18 +1509,19 @@ async fn fetch_ec_index_from_one_peer(
let stream = client
.copy_file(copy_req(".ecx", false))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("copy .ecx: {}", e)))?
.map_err(|e| io::Error::other(format!("copy .ecx: {}", e)))?
.into_inner();
drain_copy_stream(stream, ecx_path, false).await?;
let meta = fs::metadata(ecx_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("stat copied .ecx: {}", e)))?;
let meta =
fs::metadata(ecx_path).map_err(|e| io::Error::other(format!("stat copied .ecx: {}", e)))?;
if meta.is_dir() || meta.len() == 0 {
let _ = fs::remove_file(ecx_path);
return Err(io::Error::new(
io::ErrorKind::Other,
format!("peer {} served an unusable .ecx (size {})", peer, meta.len()),
));
return Err(io::Error::other(format!(
"peer {} served an unusable .ecx (size {})",
peer,
meta.len()
)));
}
// .ecj is the source peer's deletion journal (appended); .vif carries EC
@@ -1418,18 +1558,21 @@ async fn drain_copy_stream(
) -> io::Result<()> {
use std::io::Write;
let mut file = if append {
fs::OpenOptions::new().create(true).append(true).open(dest_path)
fs::OpenOptions::new()
.create(true)
.append(true)
.open(dest_path)
} else {
fs::File::create(dest_path)
}
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("create {}: {}", dest_path, e)))?;
.map_err(|e| io::Error::other(format!("create {}: {}", dest_path, e)))?;
while let Some(chunk) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv {}: {}", dest_path, e)))?
.map_err(|e| io::Error::other(format!("recv {}: {}", dest_path, e)))?
{
file.write_all(&chunk.file_content)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("write {}: {}", dest_path, e)))?;
.map_err(|e| io::Error::other(format!("write {}: {}", dest_path, e)))?;
}
Ok(())
}
+6 -7
View File
@@ -423,13 +423,12 @@ fn collect_ui_data(
shard_id: shard.shard_id,
size: shard_size,
});
if created_at == "-" {
if let Ok(metadata) = std::fs::metadata(shard.file_name()) {
if let Ok(modified) = metadata.modified() {
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
if created_at == "-"
&& let Ok(metadata) = std::fs::metadata(shard.file_name())
&& let Ok(modified) = metadata.modified()
{
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
let preferred_size = ec_volume.dat_file_size.max(0) as u64;
+8 -10
View File
@@ -312,16 +312,15 @@ async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Requ
)
}
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -358,16 +357,15 @@ async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Req
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
+35 -30
View File
@@ -131,10 +131,10 @@ impl DiskLocation {
for entry in entries {
let entry = entry?;
let name = entry.file_name().into_string().unwrap_or_default();
if let Some((collection, vid)) = parse_volume_filename(&name) {
if seen.insert((collection.clone(), vid)) {
dat_files.push((collection, vid));
}
if let Some((collection, vid)) = parse_volume_filename(&name)
&& seen.insert((collection.clone(), vid))
{
dat_files.push((collection, vid));
}
}
@@ -327,10 +327,10 @@ impl DiskLocation {
.strip_suffix(".cpc")
.or_else(|| name.strip_suffix(".cpd"))
.or_else(|| name.strip_suffix(".cpx"));
if let Some(stem) = stem {
if let Some(key) = parse_collection_volume_id(stem) {
pending.insert(key);
}
if let Some(stem) = stem
&& let Some(key) = parse_collection_volume_id(stem)
{
pending.insert(key);
}
}
}
@@ -426,11 +426,16 @@ impl DiskLocation {
if shard_count == 0 {
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size) {
if actual < expected {
warn!(volume_id = vid.0, actual, expected, "shards smaller than the .dat's full encode; reclaiming the complete .dat");
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size)
&& actual < expected
{
warn!(
volume_id = vid.0,
actual,
expected,
"shards smaller than the .dat's full encode; reclaiming the complete .dat"
);
return false;
}
true
}
@@ -510,10 +515,10 @@ impl DiskLocation {
pub(crate) fn ec_generation_ts_ns(&self, collection: &str, vid: VolumeId) -> Option<i64> {
for dir in [&self.directory, &self.idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Ok(s) = fs::read_to_string(&vif) {
if let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s) {
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if let Ok(s) = fs::read_to_string(&vif)
&& let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s)
{
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if self.directory == self.idx_directory {
break;
@@ -542,6 +547,7 @@ impl DiskLocation {
}
/// Create a new volume in this location.
#[expect(clippy::too_many_arguments)]
pub fn create_volume(
&mut self,
vid: VolumeId,
@@ -777,18 +783,18 @@ impl DiskLocation {
pub fn has_ecx_file_on_disk(&self, collection: &str, vid: VolumeId) -> bool {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
let idx_path = format!("{}.ecx", idx_base);
if let Ok(meta) = fs::metadata(&idx_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&idx_path)
&& !meta.is_dir()
{
return true;
}
if self.idx_directory != self.directory {
let data_base = volume_file_name(&self.directory, collection, vid);
let data_path = format!("{}.ecx", data_base);
if let Ok(meta) = fs::metadata(&data_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&data_path)
&& !meta.is_dir()
{
return true;
}
}
false
@@ -1107,7 +1113,7 @@ impl DiskLocation {
/// Close all volumes.
pub fn close(&mut self) {
for (_, v) in self.volumes.iter_mut() {
for v in self.volumes.values_mut() {
v.close();
}
self.volumes.clear();
@@ -1184,10 +1190,9 @@ fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &st
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
.and_then(|vi| vi.ec_shard_config)
.map(|c| c.data_shards as usize)
&& ds > 0
{
if ds > 0 {
return ds;
}
return ds;
}
if directory == idx_directory {
break;
@@ -1265,7 +1270,7 @@ fn check_dat_file_exists(path: &str) -> bool {
/// True when a `.vif` references remote-tier files: a remote-only volume
/// that has no local `.dat` but must still load via the remote path,
/// rather than be skipped as a lone EC sidecar.
fn vif_references_remote_file(vif_path: &str) -> bool {
pub(crate) fn vif_references_remote_file(vif_path: &str) -> bool {
fs::read_to_string(vif_path)
.ok()
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
@@ -164,16 +164,16 @@ pub fn remove_bitrot_sidecars(base: &str) -> io::Result<()> {
};
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res {
if first_err.is_none() {
first_err = Some(e);
}
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
record(rm(format!("{}{}", base, BITROT_SIDECAR_EXT).into()));
let path = Path::new(base);
if let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) {
let prefix = format!("{}{}.v", fname.to_string_lossy(), BITROT_SIDECAR_EXT);
let prefix = format!("{}{}.v", fname.display(), BITROT_SIDECAR_EXT);
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries.flatten() {
@@ -203,7 +203,7 @@ pub fn new_encode_uuid() -> Vec<u8> {
/// Reports whether `block_size` is a power of two in [1 MiB, MAX_BITROT_BLOCK_SIZE].
pub fn is_pow2_multiple_of_1mib(block_size: u32) -> bool {
block_size >= (1 << 20) && block_size <= MAX_BITROT_BLOCK_SIZE && block_size.count_ones() == 1
((1 << 20)..=MAX_BITROT_BLOCK_SIZE).contains(&block_size) && block_size.count_ones() == 1
}
/// Returns ceil(covered_size / block_size).
@@ -402,7 +402,7 @@ pub fn validate_manifest(
total
));
}
let mut seen = vec![false; MAX_SHARD_COUNT];
let mut seen = [false; MAX_SHARD_COUNT];
for s in &prot.shards {
if s.shard_id >= total as u32 {
return Err(format!(
@@ -505,7 +505,21 @@ pub fn verify_shard_file_blocks(
entry: &EcShardChecksums,
block_size: i64,
) -> io::Result<Vec<usize>> {
let f = File::open(path)?;
verify_shard_blocks(&File::open(path)?, entry, block_size)
}
/// Same verification against an ALREADY-OPEN shard handle.
///
/// Go's `ChecksumScrub` reads through `shard.ReadAt`, i.e. the handle the
/// EcVolumeShard already holds, so a concurrent teardown that unlinks the shard
/// cannot turn an intentional removal into a scrub read error. A scrub that
/// runs with the store lock released has to read the same way — see
/// `EcChecksumScrubPlan`.
pub fn verify_shard_blocks(
f: &File,
entry: &EcShardChecksums,
block_size: i64,
) -> io::Result<Vec<usize>> {
let file_size = f.metadata()?.len() as i64;
let want = unpack_u32_le(&entry.block_crc32c);
@@ -523,7 +537,7 @@ pub fn verify_shard_file_blocks(
break;
}
let to_read = to_read as usize;
read_full_at(&f, &mut buf[..to_read], offset as u64)?;
read_full_at(f, &mut buf[..to_read], offset as u64)?;
if CRC::new(&buf[..to_read]).0 != *want_crc {
mismatched.push(i);
}
@@ -73,6 +73,7 @@ pub fn find_dat_file_size_with_dirs(
/// must live in `dir`. For the cross-disk reconciled layout where
/// shards are split across multiple data dirs of the same node, use
/// [`write_dat_file_from_shards_with_dirs`] instead.
#[expect(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards(
dir: &str,
collection: &str,
@@ -120,7 +121,7 @@ pub fn write_dat_file_from_shards(
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
@@ -145,7 +146,7 @@ pub fn write_dat_file_from_shards_with_dirs(
)
}
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn write_dat_file(
dat_dir: &str,
collection: &str,
@@ -233,10 +234,10 @@ fn write_dat_file(
// Read large blocks
while encoded_remaining >= large_row_size && remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = large_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -255,10 +256,10 @@ fn write_dat_file(
// Read small blocks
while remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = small_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -324,10 +325,7 @@ pub fn write_idx_file_from_ec_index(
// and treat only NotFound as "no journal": Path::exists would also
// swallow a permission/IO error and silently skip deletions, which
// would resurrect deleted needles as live.
let mut idx_file = std::fs::OpenOptions::new()
.write(true)
.append(true)
.open(&tmp_path)?;
let mut idx_file = std::fs::OpenOptions::new().append(true).open(&tmp_path)?;
match std::fs::read(&ecj_path) {
Ok(ecj_data) => {
let count = ecj_data.len() / NEEDLE_ID_SIZE;
@@ -50,7 +50,7 @@ pub fn write_ec_files(
let dat_size = dat_file.metadata()?.len() as i64;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Create shard files
let total_shards = data_shards + parity_shards;
@@ -162,7 +162,7 @@ pub fn rebuild_ec_files(
}
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
@@ -175,7 +175,7 @@ pub fn rebuild_ec_files(
let mut shard_size = 0;
for (i, shard) in shards.iter_mut().enumerate() {
if !missing_shard_ids.contains(&(i as u32)) {
if let Ok(_) = shard.open() {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
@@ -185,7 +185,7 @@ pub fn rebuild_ec_files(
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
if let Ok(_) = alt.open() {
if alt.open().is_ok() {
let size = alt.file_size();
if size > shard_size {
shard_size = size;
@@ -251,12 +251,8 @@ pub fn rebuild_ec_files(
}
// Reconstruct missing shards
rs.reconstruct(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon reconstruct: {:?}", e),
)
})?;
rs.reconstruct(&mut buffers)
.map_err(|e| io::Error::other(format!("reed-solomon reconstruct: {:?}", e)))?;
// Write recovered data into the missing shards
for i in missing_shard_ids {
@@ -284,40 +280,63 @@ pub fn rebuild_ec_files(
/// FULL walk only reads live data-shard intervals, so on its own it can't catch
/// bitrot in a parity shard or an unwalked region. Move to mode 4 (CHECKSUM) and
/// drop it from mode 2 once the `.ecsum` subsystem lands.
///
/// `dirs` is indexed BY SHARD ID: each entry is the directory holding that
/// shard, or `None` when no disk mounts it. A reconciled volume's shards can be
/// split across disks, so a single directory cannot address them all.
pub fn verify_ec_shards(
dir: &str,
dirs: &[Option<String>],
collection: &str,
volume_id: VolumeId,
data_shards: usize,
parity_shards: usize,
) -> io::Result<(Vec<u32>, Vec<String>)> {
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
let mut shards: Vec<Option<EcVolumeShard>> = (0..total_shards)
.map(|i| {
dirs.get(i)
.and_then(|d| d.as_ref())
.map(|d| EcVolumeShard::new(d, collection, volume_id, i as u8))
})
.collect();
let mut shard_size = 0;
let mut broken_shards = std::collections::HashSet::new();
let mut details = Vec::new();
for (i, shard) in shards.iter_mut().enumerate() {
if let Ok(_) = shard.open() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
for (i, slot) in shards.iter_mut().enumerate() {
match slot.as_mut() {
// Not a match guard: a binding is immutable until the guard ends,
// and `open()` needs `&mut self`.
Some(shard) => {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
}
} else {
broken_shards.insert(i as u32);
details.push(format!("failed to open or missing shard {}", i));
}
}
None => {
broken_shards.insert(i as u32);
details.push(format!("shard {} is not mounted on any disk", i));
}
} else {
broken_shards.insert(i as u32);
details.push(format!("failed to open or missing shard {}", i));
}
}
if shard_size == 0 || broken_shards.len() >= parity_shards {
// Can't do much if we don't know the size or have too many missing
return Ok((broken_shards.into_iter().collect(), details));
// Can't do much if we don't know the size or have too many missing.
// Sort like the normal path below: a `HashSet` iteration order would
// make this return shard ids in an arbitrary order, and enough `None`
// entries in `dirs` now reach this branch for a caller to notice.
let mut broken_vec: Vec<u32> = broken_shards.into_iter().collect();
broken_vec.sort_unstable();
return Ok((broken_vec, details));
}
let block_size = ERASURE_CODING_SMALL_BLOCK_SIZE;
@@ -331,7 +350,17 @@ pub fn verify_ec_shards(
let mut read_failed = false;
for i in 0..total_shards {
if !broken_shards.contains(&(i as u32)) {
if let Err(e) = shards[i].read_at(&mut buffers[i], offset) {
// The `None` arm is defensive and unreachable: the open loop
// put every unmounted slot in `broken_shards`, which this
// branch already skipped. Kept because the `Option` forces
// some handling here, and an error is the only shape that
// cannot quietly feed an unread buffer into the parity
// comparison below. Nothing needs to cover it.
let read = match shards[i].as_mut() {
Some(shard) => shard.read_at(&mut buffers[i], offset),
None => Err(io::Error::new(io::ErrorKind::NotFound, "shard not mounted")),
};
if let Err(e) = read {
broken_shards.insert(i as u32);
details.push(format!("read error shard {}: {}", i, e));
read_failed = true;
@@ -345,27 +374,27 @@ pub fn verify_ec_shards(
if !read_failed {
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
if !is_valid {
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
if let Ok(is_valid) = rs.verify(&slice_ptrs)
&& !is_valid
{
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for i in data_shards..total_shards {
verify_buffers[i].fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for buf in &mut verify_buffers[data_shards..total_shards] {
buf.fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
}
}
@@ -377,7 +406,7 @@ pub fn verify_ec_shards(
}
// Close all shards
for shard in &mut shards {
for shard in shards.iter_mut().flatten() {
shard.close();
}
@@ -457,7 +486,7 @@ pub fn rebuild_ecx_file(
.collect();
for (i, shard) in shards.iter_mut().enumerate() {
if let Err(_) = shard.open() {
if shard.open().is_err() {
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
@@ -474,7 +503,7 @@ pub fn rebuild_ecx_file(
}
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("cannot open data shard for ecx rebuild"),
"cannot open data shard for ecx rebuild".to_string(),
));
}
}
@@ -482,7 +511,7 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
let total_data_size = shard_size * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
@@ -505,7 +534,7 @@ pub fn rebuild_ecx_file(
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
(shard_size - 1).max(0)
};
// Read version from superblock (first byte of logical data)
@@ -607,7 +636,6 @@ pub fn rebuild_ecx_file(
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
@@ -681,7 +709,7 @@ const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
dat_size: i64,
@@ -745,7 +773,7 @@ pub(crate) fn encode_dat_file(
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
@@ -757,7 +785,7 @@ fn encode_data(
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
if !block_size.is_multiple_of(buffer_size) {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
@@ -784,7 +812,7 @@ fn encode_data(
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
#[expect(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
@@ -797,21 +825,15 @@ fn encode_one_batch(
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
for (i, buf) in buffers[..data_shards].iter_mut().enumerate() {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
}
let n = read_at_most(dat_file, buf, read_offset)?;
buf[n..].fill(0);
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
)
})?;
rs.encode(&mut *buffers)
.map_err(|e| io::Error::other(format!("reed-solomon encode: {:?}", e)))?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
@@ -1457,4 +1479,111 @@ mod tests {
"should fail when idx_dir doesn't contain .idx"
);
}
/// Write a real 10+4 encoded volume into `dir`.
///
/// Unlike `make_volume_with_needles` and `encode_sample_volume` this seeds
/// a caller-chosen directory, which is what a split-disk test needs: the
/// shards have to be scattered out of the directory they were encoded into.
fn seed_encoded_volume(dir: &str, vid: VolumeId) {
let mut v = Volume::new(
dir,
dir,
"",
vid,
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for i in 1..=8 {
let data = format!("test data for needle {} with a bit more length", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
write_ec_files(dir, dir, "", vid, 10, 4).unwrap();
}
/// Shards split across two directories must all be found. Passing one dir
/// per shard is what lets a reconciled volume's parity be checked at all.
#[test]
fn test_verify_ec_shards_reads_shards_from_multiple_dirs() {
let tmp = TempDir::new().unwrap();
let src = tmp.path().join("src");
let d0 = tmp.path().join("d0");
let d1 = tmp.path().join("d1");
for d in [&src, &d0, &d1] {
std::fs::create_dir_all(d).unwrap();
}
let src_s = src.to_str().unwrap();
seed_encoded_volume(src_s, VolumeId(1));
// Move shards 0..=6 to d0 and 7..=13 to d1.
let mut dirs: Vec<Option<String>> = Vec::new();
for id in 0..14u8 {
let target = if id < 7 { &d0 } else { &d1 };
std::fs::rename(
format!("{}/1.ec{:02}", src_s, id),
format!("{}/1.ec{:02}", target.to_str().unwrap(), id),
)
.unwrap();
dirs.push(Some(target.to_str().unwrap().to_string()));
}
let (broken, details) = verify_ec_shards(&dirs, "", VolumeId(1), 10, 4).unwrap();
assert!(
broken.is_empty(),
"split-dir shards reported broken: {:?}",
details
);
}
/// A shard no disk holds is a missing shard, not a panic and not a silent
/// pass: it is REPORTED, by id, with a message that distinguishes "no disk
/// holds this shard" from "the disk holds it but it won't open".
///
/// Read the scope literally. This does NOT show that the mounted shards
/// verify clean. `dirs[5] = None` puts shard 5 in `broken_shards` before
/// the block loop starts, so every iteration takes the
/// `else { read_failed = true; }` arm and the Reed-Solomon comparison never
/// runs at all. `broken == vec![5]` therefore holds because the other 13
/// were never verified, not because they verified clean -- a parity check
/// over intact shards is what
/// `test_verify_ec_shards_reads_shards_from_multiple_dirs` and the
/// end-to-end split-disk FULL scrub establish.
#[test]
fn test_verify_ec_shards_treats_a_none_dir_as_missing() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
seed_encoded_volume(dir, VolumeId(1));
let mut dirs: Vec<Option<String>> = (0..14).map(|_| Some(dir.to_string())).collect();
dirs[5] = None;
let (broken, details) = verify_ec_shards(&dirs, "", VolumeId(1), 10, 4).unwrap();
assert_eq!(
broken,
vec![5],
"an unmounted shard must be reported, and only it: {:?}",
details
);
// "no disk holds this shard" and "the disk holds it but it won't open"
// are different operator problems, which is why they carry different
// messages. Asserting only the id would let one masquerade as the other.
assert!(
details.iter().any(|d| d.contains("not mounted")),
"an unmounted shard must be distinguished from an unopenable one, got {:?}",
details
);
}
}
@@ -4,6 +4,7 @@ use std::fs::{self, File, OpenOptions};
use std::io::{self, Write};
use crate::storage::types::*;
use crate::storage::volume_open::open_volume_file;
pub const DATA_SHARDS_COUNT: usize = 10;
pub const PARITY_SHARDS_COUNT: usize = 4;
@@ -50,7 +51,7 @@ impl EcVolumeShard {
/// Open the shard file for reading.
pub fn open(&mut self) -> io::Result<()> {
let path = self.file_name();
let file = File::open(&path)?;
let file = open_volume_file(OpenOptions::new().read(true), &path)?;
self.ecd_file_size = file.metadata()?.len() as i64;
self.ecd_file = Some(file);
Ok(())
@@ -59,12 +60,14 @@ impl EcVolumeShard {
/// Create the shard file for writing.
pub fn create(&mut self) -> io::Result<()> {
let path = self.file_name();
let file = OpenOptions::new()
.read(true)
.write(true)
.create(true)
.truncate(true)
.open(&path)?;
let file = open_volume_file(
OpenOptions::new()
.read(true)
.write(true)
.create(true)
.truncate(true),
&path,
)?;
self.ecd_file = Some(file);
self.ecd_file_size = 0;
Ok(())
@@ -75,7 +78,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
#[cfg(unix)]
{
@@ -99,7 +102,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
file.write_all(data)?;
self.ecd_file_size += data.len() as i64;
Ok(())
@@ -109,6 +112,21 @@ impl EcVolumeShard {
self.ecd_file_size
}
/// A duplicate of the mounted shard handle, for a reader that has to
/// outlive the store guard.
///
/// This is the same descriptor `read_at` serves from, so it carries the
/// `O_NOATIME` from `open_volume_file` and keeps pointing at the shard
/// that was mounted, whatever later happens to the path. `dup` shares the
/// kernel file offset, which is why every read through it must be
/// positional (`read_at`), never seek-based.
pub fn try_clone_file(&self) -> io::Result<File> {
self.ecd_file
.as_ref()
.ok_or_else(|| io::Error::other("shard file not open"))?
.try_clone()
}
/// Protobuf descriptor for this shard. Mirrors Go's ToEcShardInfo.
pub fn to_ec_shard_info(&self) -> crate::pb::volume_server_pb::EcShardInfo {
crate::pb::volume_server_pb::EcShardInfo {
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -6,7 +6,7 @@ use crate::storage::needle::needle::get_actual_size;
use crate::storage::types::*;
use std::io::{self, Read, Seek, SeekFrom};
const ROWS_TO_READ: usize = 1024;
pub(crate) const ROWS_TO_READ: usize = 1024;
/// Walk all entries in an .idx file, calling `f` for each.
/// Mirrors Go's `WalkIndexFile()`.
+2
View File
@@ -9,6 +9,8 @@ pub mod store_ec_reconcile;
pub mod super_block;
pub mod types;
pub mod volume;
pub mod volume_idx_rebuild;
pub mod volume_idx_repair;
pub mod volume_open;
pub mod volume_report;
pub mod volume_report_hash;
+4 -2
View File
@@ -21,7 +21,7 @@ impl CRC {
/// Legacy `.Value()` function — deprecated in Go but needed for backward compat check.
/// Formula: (crc >> 15 | crc << 17) + 0xa282ead8
pub fn legacy_value(&self) -> u32 {
(self.0 >> 15 | self.0 << 17).wrapping_add(0xa282ead8)
self.0.rotate_right(15).wrapping_add(0xa282ead8)
}
}
@@ -67,7 +67,9 @@ mod tests {
fn test_crc_legacy_value() {
let crc = CRC(0x12345678);
let v = crc.legacy_value();
let expected = (0x12345678u32 >> 15 | 0x12345678u32 << 17).wrapping_add(0xa282ead8);
// (0x12345678 >> 15 | 0x12345678 << 17) + 0xa282ead8, worked out by hand so
// the test checks the rotate rather than restating it.
let expected = 0x4f730f40_u32;
assert_eq!(v, expected);
}
}
+1
View File
@@ -1,4 +1,5 @@
pub mod crc;
#[expect(clippy::module_inception, reason = "needle/needle.rs mirrors the Go package layout")]
pub mod needle;
pub mod ttl;
+21 -15
View File
@@ -560,7 +560,7 @@ impl Needle {
// Padding to 8-byte alignment
let padding = padding_length(self.size, version).0 as usize;
buf.extend(std::iter::repeat(0u8).take(padding));
buf.extend(std::iter::repeat_n(0u8, padding));
buf
}
@@ -824,11 +824,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v3() {
let mut n = Needle::default();
n.cookie = Cookie(42);
n.id = NeedleId(100);
n.data = b"hello world".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(42),
id: NeedleId(100),
data: b"hello world".to_vec(),
flags: 0,
..Needle::default()
};
n.set_has_name();
n.name = b"test.txt".to_vec();
n.name_size = 8;
@@ -867,11 +869,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v2() {
let mut n = Needle::default();
n.cookie = Cookie(77);
n.id = NeedleId(200);
n.data = b"data v2".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(77),
id: NeedleId(200),
data: b"data v2".to_vec(),
flags: 0,
..Needle::default()
};
let bytes = n.write_bytes(VERSION_2);
let expected_size = get_actual_size(n.size, VERSION_2);
@@ -886,10 +890,12 @@ mod tests {
#[test]
fn test_read_bytes_meta_only_handles_tombstone_v3() {
let mut tombstone = Needle::default();
tombstone.cookie = Cookie(0x1234abcd);
tombstone.id = NeedleId(300);
tombstone.append_at_ns = 999_999;
let mut tombstone = Needle {
cookie: Cookie(0x1234abcd),
id: NeedleId(300),
append_at_ns: 999_999,
..Needle::default()
};
let bytes = tombstone.write_bytes(VERSION_3);
+50 -17
View File
@@ -81,7 +81,7 @@ impl TTL {
return Ok(TTL::EMPTY);
}
let last_byte = s.as_bytes()[s.len() - 1];
let (num_str, unit_byte) = if last_byte >= b'0' && last_byte <= b'9' {
let (num_str, unit_byte) = if last_byte.is_ascii_digit() {
// All digits — default to minutes (matching Go)
(s, b'm')
} else {
@@ -144,40 +144,73 @@ fn fit_ttl_count(count: u32, unit: u8) -> TTL {
const MINUTE_SECS: u64 = 60;
// First pass: try exact fits from largest to smallest
if seconds % YEAR_SECS == 0 && seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
if seconds.is_multiple_of(YEAR_SECS) && seconds / YEAR_SECS < 256 {
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
if seconds % MONTH_SECS == 0 && seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
if seconds.is_multiple_of(MONTH_SECS) && seconds / MONTH_SECS < 256 {
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds % WEEK_SECS == 0 && seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
if seconds.is_multiple_of(WEEK_SECS) && seconds / WEEK_SECS < 256 {
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds % DAY_SECS == 0 && seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
if seconds.is_multiple_of(DAY_SECS) && seconds / DAY_SECS < 256 {
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds % HOUR_SECS == 0 && seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
if seconds.is_multiple_of(HOUR_SECS) && seconds / HOUR_SECS < 256 {
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
// Minutes: truncating division
if seconds / MINUTE_SECS < 256 {
return TTL { count: (seconds / MINUTE_SECS) as u8, unit: TTL_UNIT_MINUTE };
return TTL {
count: (seconds / MINUTE_SECS) as u8,
unit: TTL_UNIT_MINUTE,
};
}
// Second pass: truncating division from smallest to largest
if seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
if seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
TTL::EMPTY
}
File diff suppressed because it is too large Load Diff
@@ -31,7 +31,7 @@ struct CompactEntry {
}
impl CompactEntry {
fn to_needle_value(&self) -> NeedleValue {
fn to_needle_value(self) -> NeedleValue {
NeedleValue {
offset: Offset::from_bytes(&self.offset),
size: self.size,
@@ -18,6 +18,8 @@ use std::fs::{File, OpenOptions};
use std::io;
use std::sync::{Arc, Mutex, OnceLock};
use crate::storage::volume_open::open_volume_file;
/// Descriptors the pool keeps open. Matches Go's `maxPooledIndexFiles`.
pub const MAX_POOLED_INDEX_FILES: usize = 1024;
@@ -68,7 +70,10 @@ impl IndexFilePool {
// Opened outside the lock: a cold open blocks on disk, and holding a
// process-wide mutex across it would serialize every volume's lookups.
let file = Arc::new(OpenOptions::new().read(true).write(writable).open(path)?);
let file = Arc::new(open_volume_file(
OpenOptions::new().read(true).write(writable),
path,
)?);
Ok(self.insert(key, file))
}
@@ -0,0 +1,161 @@
//! Rebuild a needle map's counters from the whole `.idx` history.
//!
//! The volume's garbage ratio is `deleted_size / content_size`, and both are
//! additive over the life of the volume: an overwritten needle keeps its
//! bytes in `content_size` and adds them to `deleted_size`. A backend whose
//! table only holds the final value per key (redb) cannot recover that from
//! the table, so on load the counters come from the `.idx` file instead.
//!
//! This mirrors Go's `needleMapMetricFromIndexFile`: walk the index newest
//! entry first with a bloom filter of the keys already seen, so the memory
//! cost is a few bits per entry instead of a map of every key. The counting
//! rule reproduces what the live `on_put`/`on_delete` path accumulates:
//!
//! - every live entry is one put: `file_count`, `file_byte_count`;
//! - a live entry with a newer entry for the same key was overwritten or
//! deleted later, so it is also one deletion: `deletion_count`,
//! `deletion_byte_count`;
//! - a tombstone only marks its key as seen.
//!
//! A bloom false positive (0.1%) can only add a spurious deletion, which
//! over-reports garbage slightly; it never hides any.
use std::io::{self, Read, Seek, SeekFrom};
use std::sync::atomic::Ordering;
use xxhash_rust::xxh64::xxh64;
use super::NeedleMapMetric;
use crate::storage::types::*;
/// Entries read per batch while walking backwards (64 KiB of index).
const BATCH_ENTRIES: usize = 4096;
/// Same target false-positive rate as the Go server's filter.
const FALSE_POSITIVE_RATE: f64 = 0.001;
/// Minimal bloom filter over needle ids, double hashing with xxh64.
struct SeenKeys {
bits: Vec<u64>,
bit_count: u64,
hashes: u64,
}
impl SeenKeys {
fn new(expected: u64, false_positive_rate: f64) -> Self {
let n = expected.max(1) as f64;
let ln2 = std::f64::consts::LN_2;
let bit_count = (-(n * false_positive_rate.ln()) / (ln2 * ln2))
.ceil()
.max(64.0) as u64;
let hashes = ((bit_count as f64 / n) * ln2).round().clamp(1.0, 16.0) as u64;
SeenKeys {
bits: vec![0u64; bit_count.div_ceil(64) as usize],
bit_count,
hashes,
}
}
/// Whether `key` was (probably) seen before; marks it seen either way.
fn test_and_add(&mut self, key: u64) -> bool {
let bytes = key.to_le_bytes();
let h1 = xxh64(&bytes, 0);
let h2 = xxh64(&bytes, 0x9E37_79B9_7F4A_7C15) | 1;
let mut seen = true;
for i in 0..self.hashes {
let bit = h1.wrapping_add(i.wrapping_mul(h2)) % self.bit_count;
let word = (bit / 64) as usize;
let mask = 1u64 << (bit % 64);
if self.bits[word] & mask == 0 {
seen = false;
self.bits[word] |= mask;
}
}
seen
}
}
/// Walk `reader` (an `.idx` file) newest entry first and return the counters
/// a live volume would hold after applying the same history. A torn partial
/// entry at the tail is ignored, as `walk_index_file` does. The reader is
/// left positioned at the start of the file.
pub(super) fn metrics_from_idx<R: Read + Seek>(
reader: &mut R,
version: Version,
) -> io::Result<NeedleMapMetric> {
let metric = NeedleMapMetric::default();
let file_size = reader.seek(SeekFrom::End(0))?;
let entry_count = file_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let mut seen = SeenKeys::new(entry_count, FALSE_POSITIVE_RATE);
let mut buf = vec![0u8; NEEDLE_MAP_ENTRY_SIZE * BATCH_ENTRIES];
let mut remaining = entry_count;
while remaining > 0 {
let batch = remaining.min(BATCH_ENTRIES as u64) as usize;
let first_entry = remaining - batch as u64;
let len = batch * NEEDLE_MAP_ENTRY_SIZE;
reader.seek(SeekFrom::Start(first_entry * NEEDLE_MAP_ENTRY_SIZE as u64))?;
reader.read_exact(&mut buf[..len])?;
for i in (0..batch).rev() {
let entry = &buf[i * NEEDLE_MAP_ENTRY_SIZE..(i + 1) * NEEDLE_MAP_ENTRY_SIZE];
let (key, offset, size) = idx_entry_from_bytes(entry);
metric.maybe_set_max_needle_end(offset, size, version);
let superseded = seen.test_and_add(key.into());
if offset.is_zero() || size.is_deleted() {
// Tombstone: reserves no bytes, only marks the key as seen.
continue;
}
metric.maybe_set_max_file_key(key);
metric.file_count.fetch_add(1, Ordering::Relaxed);
metric
.file_byte_count
.fetch_add(size.0 as u64, Ordering::Relaxed);
if superseded && size.0 > 0 {
metric.deletion_count.fetch_add(1, Ordering::Relaxed);
metric
.deletion_byte_count
.fetch_add(size.0 as u64, Ordering::Relaxed);
}
}
remaining = first_entry;
}
reader.seek(SeekFrom::Start(0))?;
Ok(metric)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_seen_keys_reports_repeats_and_not_fresh_keys() {
let mut seen = SeenKeys::new(10_000, FALSE_POSITIVE_RATE);
// Fresh keys may occasionally collide (that is the false-positive
// rate), but only rarely.
let fresh_reported_seen = (0..10_000u64)
.filter(|&key| seen.test_and_add(key))
.count();
assert!(
fresh_reported_seen < 50,
"fresh keys reported seen: {fresh_reported_seen}"
);
// A repeated key is never reported fresh: no false negatives.
for key in 0..10_000u64 {
assert!(seen.test_and_add(key), "repeated key {key} reported fresh");
}
// Over 100k never-inserted keys the false-positive rate stays near
// the 0.1% target; allow a generous margin.
let false_positives = (1_000_000..1_100_000u64)
.filter(|k| {
let bytes = k.to_le_bytes();
let h1 = xxh64(&bytes, 0);
let h2 = xxh64(&bytes, 0x9E37_79B9_7F4A_7C15) | 1;
(0..seen.hashes).all(|i| {
let bit = h1.wrapping_add(i.wrapping_mul(h2)) % seen.bit_count;
seen.bits[(bit / 64) as usize] & (1u64 << (bit % 64)) != 0
})
})
.count();
assert!(false_positives < 500, "false positives: {false_positives}");
}
}
@@ -226,10 +226,7 @@ impl SortedFileNeedleMap {
.fail_sdx_mark
.load(std::sync::atomic::Ordering::Relaxed)
{
return Err(io::Error::new(
io::ErrorKind::Other,
"injected .sdx mark failure",
));
return Err(io::Error::other("injected .sdx mark failure"));
}
let mut buf = [0u8; SIZE_SIZE];
TOMBSTONE_FILE_SIZE.to_bytes(&mut buf);
@@ -309,7 +306,7 @@ impl SortedFileNeedleMap {
let rows = rows_per_read.min(entry_count - done) as usize;
let bytes = &mut block[..rows * NEEDLE_MAP_ENTRY_SIZE];
read_exact_at(&file, bytes, done * NEEDLE_MAP_ENTRY_SIZE as u64)?;
for entry in bytes.chunks_exact(NEEDLE_MAP_ENTRY_SIZE) {
for entry in bytes.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0 {
let (key, offset, size) = idx_entry_from_bytes(entry);
if !size.is_valid() || pending.contains_key(&key) {
continue; // deleted in place, or still awaiting that mark
+552 -56
View File
@@ -369,6 +369,7 @@ impl Store {
}
/// Create a new volume, placing it on the location with the most free space.
#[expect(clippy::too_many_arguments)]
pub fn add_volume(
&mut self,
vid: VolumeId,
@@ -383,10 +384,10 @@ impl Store {
return Err(VolumeError::AlreadyExists);
}
let loc_idx = self.find_free_location(&disk_type).ok_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("no free location for disk type {:?}", disk_type),
))
VolumeError::Io(io::Error::other(format!(
"no free location for disk type {:?}",
disk_type
)))
})?;
self.locations[loc_idx].create_volume(
@@ -427,6 +428,26 @@ impl Store {
false
}
/// Reports whether any local volume or EC shard is currently quarantined
/// due to sustained storage-media EIO. Mirrors Go's Store.HasIoQuarantine.
pub fn has_io_quarantine(&self) -> bool {
for loc in &self.locations {
for (_, vol) in loc.iter_volumes() {
let (_, _, quarantined) = vol.get_io_error_state();
if quarantined {
return true;
}
}
for (_, ec_vol) in loc.ec_volumes() {
let (_, _, quarantined) = ec_vol.get_io_error_state();
if quarantined {
return true;
}
}
}
false
}
/// Mount a volume from an existing .dat file.
pub fn mount_volume(
&mut self,
@@ -439,7 +460,7 @@ impl Store {
}
// Find the location where the .dat file exists
for loc in &mut self.locations {
if &loc.disk_type != &disk_type {
if loc.disk_type != disk_type {
continue;
}
let base = crate::storage::volume::volume_file_name(&loc.directory, collection, vid);
@@ -452,10 +473,10 @@ impl Store {
// Fail the mount so the caller (VolumeCopy) treats it as an error.
let note_path = format!("{}.note", base);
if std::path::Path::new(&note_path).exists() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("volume {} copy incomplete: .note still present", vid),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} copy incomplete: .note still present",
vid
))));
}
return loc.create_volume(
vid,
@@ -476,45 +497,185 @@ impl Store {
/// Mount a volume by id only (Go's MountVolume behavior).
/// Scans all locations for a matching .dat file and loads with its collection prefix.
pub fn mount_volume_by_id(&mut self, vid: VolumeId) -> Result<(), VolumeError> {
/// When a collection hint is given, the expected <collection>_<vid>.vif/.idx
/// path is probed directly before falling back to the directory scan.
pub fn mount_volume_by_id(
&mut self,
vid: VolumeId,
collection: Option<&str>,
) -> Result<(), VolumeError> {
if self.find_volume(vid).is_some() {
return Err(VolumeError::AlreadyExists);
}
if let Some((loc_idx, _base_path, collection)) = self.find_volume_file_base(vid) {
let loc = &mut self.locations[loc_idx];
return loc.create_volume(
vid,
&collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
);
// Remember the last non-NotFound error so a caller gets a useful
// message when every candidate fails to open, instead of a generic
// NotFound. A successful mount returns immediately.
let mut last_err: Option<VolumeError> = None;
if let Some(collection) = collection {
// The hint is only an optimization: a collection carrying a path
// separator could route file creation outside the storage
// directory, so skip the shortcut and let the safe directory scan
// resolve the volume instead. Rejecting the exact ".." name is
// sufficient — a collection like "foo..bar" is a valid name and
// stays inside the directory (the ".." is part of the filename, not
// a parent reference, because volume_file_name joins with "_").
let hint_safe = !collection.is_empty()
&& !collection.contains('/')
&& !collection.contains('\\')
&& collection != "..";
if hint_safe {
for loc in &mut self.locations {
let base = crate::storage::volume::volume_file_name(
&loc.directory,
collection,
vid,
);
// Confirm a collection-named sidecar exists before using the
// hint. A lone .vif/.idx (e.g. an EC sidecar whose .ecx is on
// a sibling disk) must NOT mount here: create_volume would
// write an empty .dat and register a phantom normal volume
// that shadows the real EC volume. Match the guard in
// load_existing_volumes: only mount when a real .dat is
// present, or the .vif points at a remote-tiered file.
let sidecar_present = [".vif", ".idx"].iter().any(|ext| {
std::fs::metadata(format!("{}{}", base, ext))
.map(|m| !m.is_dir())
.unwrap_or(false)
});
if !sidecar_present {
continue;
}
let dat_path = format!("{}.dat", base);
let dat_exists = std::fs::metadata(&dat_path)
.map(|m| !m.is_dir())
.unwrap_or(false);
let idx_base = crate::storage::volume::volume_file_name(
&loc.idx_directory,
collection,
vid,
);
let has_remote = crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", base),
) || crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", idx_base),
);
if dat_exists || has_remote {
// A persisting .note means the copy that produced these
// files never completed; mounting it would expose a
// truncated volume. Skip this candidate and keep
// searching (matches load_existing_volumes).
let note_path = format!("{}.note", base);
if std::path::Path::new(&note_path).exists() {
continue;
}
// An open failure on one candidate must not block a
// valid volume on a later disk — remember the error and
// keep scanning (matches open_volumes / Go mountVolume).
match loc.create_volume(
vid,
collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
) {
Ok(()) => return Ok(()),
Err(e) => {
last_err = Some(e);
continue;
}
}
}
// Lone sidecar: leave it for the directory scan below.
}
}
}
Err(VolumeError::Io(io::Error::new(
// Iterate every matching candidate, not just the first. A lone
// sidecar on an earlier disk must not hide a real .dat on a later
// disk (the split-disk EC layout the guard above protects against).
for (loc_idx, base_path, collection) in self.find_volume_file_bases(vid) {
// The scan matches any volume file (.dat/.vif/.idx). A lone .vif/.idx
// sidecar (e.g. an EC sidecar whose .ecx is on a sibling disk) must
// NOT mount here: create_volume would write an empty .dat and
// register a phantom normal volume that shadows the real EC volume.
// Match the guard in load_existing_volumes: only mount when a real
// .dat is present, or the .vif points at a remote-tiered file.
let dat_exists = std::fs::metadata(format!("{}.dat", base_path))
.map(|m| !m.is_dir())
.unwrap_or(false);
let idx_base = crate::storage::volume::volume_file_name(
&self.locations[loc_idx].idx_directory,
&collection,
vid,
);
let has_remote = crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", base_path),
) || crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", idx_base),
);
if dat_exists || has_remote {
// A persisting .note means the copy that produced these files
// never completed; mounting it would expose a truncated volume.
// Skip this candidate and keep searching (matches
// load_existing_volumes).
let note_path = format!("{}.note", base_path);
if std::path::Path::new(&note_path).exists() {
continue;
}
// An open failure on one candidate must not block a valid
// volume on a later disk — remember the error and keep
// scanning (matches open_volumes / Go mountVolume).
let loc = &mut self.locations[loc_idx];
match loc.create_volume(
vid,
&collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
) {
Ok(()) => return Ok(()),
Err(e) => {
last_err = Some(e);
continue;
}
}
}
}
Err(last_err.unwrap_or_else(|| VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("volume {} not found on disk", vid),
)))
))))
}
fn find_volume_file_base(&self, vid: VolumeId) -> Option<(usize, String, String)> {
self.find_volume_file_bases(vid).into_iter().next()
}
/// Collect every location/collection whose directory holds a volume file
/// (.dat/.vif/.idx) for `vid`, in scan order. Callers that need to skip
/// lone sidecars (mount_volume_by_id) must see all candidates so a sidecar
/// on an earlier disk does not hide a real .dat on a later one.
fn find_volume_file_bases(&self, vid: VolumeId) -> Vec<(usize, String, String)> {
let mut results = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
if let Ok(entries) = std::fs::read_dir(&loc.directory) {
for entry in entries.flatten() {
let name = entry.file_name();
let name = name.to_string_lossy();
if let Some((collection, file_vid)) = parse_volume_filename(&name) {
if file_vid == vid {
let base = strip_volume_suffix(&name)?;
let base_path = format!("{}/{}", loc.directory, base);
return Some((loc_idx, base_path, collection));
}
if let Some((collection, file_vid)) = parse_volume_filename(&name)
&& file_vid == vid
&& let Some(base) = strip_volume_suffix(&name)
{
let base_path = format!("{}/{}", loc.directory, base);
results.push((loc_idx, base_path, collection));
}
}
}
}
None
results
}
/// Configure a volume's replica placement on disk.
@@ -681,10 +842,8 @@ impl Store {
let vol_count = loc.volumes_len() as i32;
let loc_ec_shards = loc.ec_shard_count();
let ec_equivalent = ((loc_ec_shards
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
let ec_equivalent = loc_ec_shards
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
let mut max_count = vol_count + ec_equivalent;
@@ -874,6 +1033,20 @@ impl Store {
dirs
}
/// Every per-disk `EcVolume` this store maps for `vid`, in location order.
/// Immutable twin of [`Self::find_all_ec_volumes_mut`].
///
/// Reconciliation can mount one vid as N runtimes holding disjoint shard
/// subsets, and the first-match `find_ec_volume` hides the siblings. Anything
/// that has to reach the whole volume, rather than any one runtime of it,
/// uses this.
pub fn find_all_ec_volumes(&self, vid: VolumeId) -> Vec<&EcVolume> {
self.locations
.iter()
.filter_map(|loc| loc.find_ec_volume(vid))
.collect()
}
pub fn find_all_ec_volumes_mut(&mut self, vid: VolumeId) -> Vec<&mut EcVolume> {
self.locations
.iter_mut()
@@ -919,10 +1092,10 @@ impl Store {
/// first disk and miss shards that live on a sibling.
pub fn find_ec_shard_location(&self, vid: VolumeId, shard_id: u32) -> Option<usize> {
for (i, loc) in self.locations.iter().enumerate() {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(i);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
return Some(i);
}
}
None
@@ -931,16 +1104,12 @@ impl Store {
/// Like [`Self::find_ec_shard_location`] but returns the EcVolume
/// reference directly. Borrows the store immutably for the
/// EcVolume's lifetime.
pub fn find_ec_volume_with_shard(
&self,
vid: VolumeId,
shard_id: u32,
) -> Option<&EcVolume> {
pub fn find_ec_volume_with_shard(&self, vid: VolumeId, shard_id: u32) -> Option<&EcVolume> {
for loc in &self.locations {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(ecv);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
return Some(ecv);
}
}
None
@@ -969,9 +1138,9 @@ impl Store {
if found_vol.is_none() {
found_vol = Some(ecv);
}
for shard_id in 0..max_shard_count {
if dirs[shard_id].is_none() && ecv.has_shard(shard_id as u8) {
dirs[shard_id] = Some(loc.directory.clone());
for (shard_id, dir) in dirs.iter_mut().enumerate() {
if dir.is_none() && ecv.has_shard(shard_id as u8) {
*dir = Some(loc.directory.clone());
}
}
}
@@ -990,12 +1159,19 @@ impl Store {
for (disk_id, loc) in self.locations.iter_mut().enumerate() {
let mut expired_vids = Vec::new();
let mut io_quarantined_vids = Vec::new();
for (vid, ec_vol) in loc.ec_volumes() {
if ec_vol.is_time_to_destroy() {
expired_vids.push(*vid);
} else {
ec_shards
.extend(ec_vol.to_volume_ec_shard_information_messages(disk_id as u32));
let (_, io_count, quarantined) = ec_vol.get_io_error_state();
if quarantined || io_count >= crate::storage::erasure_coding::ec_volume::IO_ERROR_TOLERANCE
{
io_quarantined_vids.push(*vid);
} else {
ec_shards
.extend(ec_vol.to_volume_ec_shard_information_messages(disk_id as u32));
}
}
}
@@ -1016,6 +1192,20 @@ impl Store {
ec_shards.extend(messages);
}
}
for vid in io_quarantined_vids {
if let Some(ec_vol) = loc.find_ec_volume(vid) {
let (_, io_count, quarantined) = ec_vol.get_io_error_state();
if !quarantined {
ec_vol.mark_io_quarantined();
tracing::warn!(
volume_id = vid.0,
io_count,
"ec volume quarantined after consecutive IO errors"
);
}
}
}
}
(ec_shards, deleted)
@@ -1320,9 +1510,10 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
read_only: bool,
}
if let Ok(legacy) = serde_json::from_str::<LegacyVolumeInfo>(&content) {
let mut vif = VifVolumeInfo::default();
vif.read_only = legacy.read_only;
return Ok(vif);
return Ok(VifVolumeInfo {
read_only: legacy.read_only,
..VifVolumeInfo::default()
});
}
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
@@ -1332,7 +1523,7 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
fn save_vif_volume_info(path: &str, info: &VifVolumeInfo) -> Result<(), VolumeError> {
let content = serde_json::to_string_pretty(info)
.map_err(|e| VolumeError::Io(io::Error::new(io::ErrorKind::Other, e.to_string())))?;
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
std::fs::write(path, content)?;
Ok(())
}
@@ -1463,6 +1654,311 @@ mod tests {
assert_eq!(store.total_volume_count(), 1);
}
#[test]
fn test_mount_volume_by_id_hint_skips_lone_sidecar() {
// A lone .vif/.idx sidecar (e.g. an EC sidecar whose .ecx is on a
// sibling disk) must NOT make the collection hint create a phantom
// empty .dat. The hint is skipped and the directory scan finds nothing.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
let base = volume_file_name(dir, "coll", VolumeId(5));
std::fs::write(format!("{}.vif", base), "{}").unwrap();
std::fs::write(format!("{}.idx", base), b"").unwrap();
// No .dat, and the .vif does not reference a remote file.
let err = store
.mount_volume_by_id(VolumeId(5), Some("coll"))
.unwrap_err();
assert!(matches!(err, VolumeError::Io(ref e)
if e.kind() == std::io::ErrorKind::NotFound));
// No phantom .dat was created.
assert!(!std::path::Path::new(&format!("{}.dat", base)).exists());
assert!(store.find_volume(VolumeId(5)).is_none());
}
#[test]
fn test_mount_volume_by_id_hint_rejects_traversal_collection() {
// A collection carrying a parent reference must not route file
// creation outside the storage directory; the hint is dropped.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
let escaped = format!(
"{}/../evil_5.dat",
dir
);
let err = store
.mount_volume_by_id(VolumeId(5), Some("../evil"))
.unwrap_err();
assert!(matches!(err, VolumeError::Io(ref e)
if e.kind() == std::io::ErrorKind::NotFound));
assert!(!std::path::Path::new(&escaped).exists());
assert!(store.find_volume(VolumeId(5)).is_none());
}
#[test]
fn test_mount_volume_by_id_hint_loads_real_volume() {
// With a real .dat on disk, the collection hint mounts the volume
// directly without scanning the directory.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
store
.add_volume(
VolumeId(7),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
// Write a needle so the volume has real data, then unmount it so the
// .dat stays on disk but is no longer registered.
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"hint me".to_vec(),
data_size: 7,
..Needle::default()
};
store
.write_volume_needle(VolumeId(7), &mut n, false)
.unwrap();
assert!(store.unmount_volume(VolumeId(7)));
store
.mount_volume_by_id(VolumeId(7), Some("coll"))
.unwrap();
assert!(store.find_volume(VolumeId(7)).is_some());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = store.read_volume_needle(VolumeId(7), &mut got).unwrap();
assert_eq!(count, 7);
assert_eq!(got.data, b"hint me");
}
#[test]
fn test_mount_volume_by_id_hint_accepts_double_dot_collection() {
// A valid collection like "foo..bar" contains ".." but is not a parent
// reference — volume_file_name joins with "_" so it stays in the dir.
// The hint must NOT be rejected for it.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
store
.add_volume(
VolumeId(9),
"foo..bar",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"dots".to_vec(),
data_size: 4,
..Needle::default()
};
store.write_volume_needle(VolumeId(9), &mut n, false).unwrap();
assert!(store.unmount_volume(VolumeId(9)));
// The hint is accepted and mounts the volume.
store
.mount_volume_by_id(VolumeId(9), Some("foo..bar"))
.unwrap();
assert!(store.find_volume(VolumeId(9)).is_some());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = store.read_volume_needle(VolumeId(9), &mut got).unwrap();
assert_eq!(count, 4);
assert_eq!(got.data, b"dots");
}
#[test]
fn test_mount_volume_by_id_skips_incomplete_note() {
// A persisting .note means a VolumeCopy was interrupted; the volume
// must not mount as live (would expose truncated data). The candidate
// is skipped and mount returns NotFound.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
store
.add_volume(
VolumeId(11),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"partial".to_vec(),
data_size: 7,
..Needle::default()
};
store.write_volume_needle(VolumeId(11), &mut n, false).unwrap();
assert!(store.unmount_volume(VolumeId(11)));
// Simulate an interrupted copy: drop a .note marker.
let base = volume_file_name(dir, "coll", VolumeId(11));
std::fs::write(format!("{}.note", base), "interrupted").unwrap();
// Hint path: skipped because of .note.
let err = store
.mount_volume_by_id(VolumeId(11), Some("coll"))
.unwrap_err();
assert!(matches!(err, VolumeError::Io(ref e)
if e.kind() == std::io::ErrorKind::NotFound));
assert!(store.find_volume(VolumeId(11)).is_none());
// Fallback path (no hint): also skipped because of .note.
let err = store.mount_volume_by_id(VolumeId(11), None).unwrap_err();
assert!(matches!(err, VolumeError::Io(ref e)
if e.kind() == std::io::ErrorKind::NotFound));
assert!(store.find_volume(VolumeId(11)).is_none());
}
#[test]
fn test_mount_volume_by_id_fallback_skips_sidecar_finds_real_dat() {
// A lone sidecar on disk 0 must not hide a real .dat on disk 1.
// The fallback now iterates all candidates instead of stopping at
// the first (sidecar-only) match.
let tmp1 = TempDir::new().unwrap();
let tmp2 = TempDir::new().unwrap();
let dir0 = tmp1.path().to_str().unwrap();
let dir1 = tmp2.path().to_str().unwrap();
// disk 0: lone .vif sidecar, no .dat, no remote.
let base0 = volume_file_name(dir0, "coll", VolumeId(13));
std::fs::write(format!("{}.vif", base0), "{}").unwrap();
// disk 1: real volume with data.
let mut store = make_test_store(&[dir0, dir1]);
store
.add_volume(
VolumeId(13),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"real".to_vec(),
data_size: 4,
..Needle::default()
};
store.write_volume_needle(VolumeId(13), &mut n, false).unwrap();
assert!(store.unmount_volume(VolumeId(13)));
// No hint: the fallback scan finds the sidecar on disk 0 first (skip,
// no .dat), then the real .dat on disk 1 (mount).
store.mount_volume_by_id(VolumeId(13), None).unwrap();
assert!(store.find_volume(VolumeId(13)).is_some());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = store.read_volume_needle(VolumeId(13), &mut got).unwrap();
assert_eq!(count, 4);
assert_eq!(got.data, b"real");
}
#[cfg(unix)]
#[test]
fn test_mount_volume_by_id_continues_past_open_failure() {
// A create_volume failure on an earlier candidate must not block a
// valid volume on a later disk. The scan remembers the error and
// keeps going (matches open_volumes / Go mountVolume).
use std::os::unix::fs::PermissionsExt;
use std::sync::atomic::Ordering;
let tmp1 = TempDir::new().unwrap();
let tmp2 = TempDir::new().unwrap();
let dir0 = tmp1.path().to_str().unwrap();
let dir1 = tmp2.path().to_str().unwrap();
let mut store = make_test_store(&[dir0, dir1]);
// Force add_volume onto disk 1 by marking disk 0 as low on space.
store.locations[0].is_disk_space_low.store(true, Ordering::Relaxed);
store
.add_volume(
VolumeId(15),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"later".to_vec(),
data_size: 5,
..Needle::default()
};
store.write_volume_needle(VolumeId(15), &mut n, false).unwrap();
assert!(store.unmount_volume(VolumeId(15)));
// Clear the low-space flag so mount_volume_by_id considers disk 0.
store.locations[0].is_disk_space_low.store(false, Ordering::Relaxed);
// disk 0: a .dat that exists but is unreadable (chmod 000). The guard
// sees dat_exists=true (metadata succeeds, not a dir), but
// create_volume -> Volume::new -> load fails opening it.
let base0 = volume_file_name(dir0, "coll", VolumeId(15));
std::fs::write(format!("{}.dat", base0), b"").unwrap();
std::fs::set_permissions(
format!("{}.dat", base0),
std::fs::Permissions::from_mode(0o000),
)
.unwrap();
// The fallback scan hits disk 0 first (create_volume fails on the
// unreadable .dat), then disk 1 (succeeds). Mount succeeds from disk 1.
store.mount_volume_by_id(VolumeId(15), None).unwrap();
assert!(store.find_volume(VolumeId(15)).is_some());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = store.read_volume_needle(VolumeId(15), &mut got).unwrap();
assert_eq!(count, 5);
assert_eq!(got.data, b"later");
// Restore permissions so TempDir cleanup can remove the file.
let _ = std::fs::set_permissions(
format!("{}.dat", base0),
std::fs::Permissions::from_mode(0o644),
);
}
#[test]
fn test_store_read_write_delete() {
let tmp = TempDir::new().unwrap();
+127 -52
View File
@@ -80,6 +80,11 @@ struct EcxOwnerInfo {
idx_dir: String,
}
/// One unit of reconcile work: the disk holding orphan shards, the volume
/// they belong to, the shard files, the `.ecx` owner, and whether the
/// mirror already installed sidecars locally (`use_local_idx`).
type OrphanShardLoad = (usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool);
impl Store {
/// Run cross-disk orphan-shard reconciliation. Should be called
/// after every DiskLocation has finished its per-disk EC scan.
@@ -98,7 +103,7 @@ impl Store {
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
let mut to_load: Vec<OrphanShardLoad> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -293,10 +298,10 @@ impl Store {
// may be sole copies of a distributed volume.
let mut node_wide_bits = ev.shard_bits().0;
for other in &self.locations {
if let Some(other_ev) = other.find_ec_volume(*vid) {
if other_ev.collection == ev.collection {
node_wide_bits |= other_ev.shard_bits().0;
}
if let Some(other_ev) = other.find_ec_volume(*vid)
&& other_ev.collection == ev.collection
{
node_wide_bits |= other_ev.shard_bits().0;
}
}
let node_wide = node_wide_bits.count_ones() as usize;
@@ -499,6 +504,53 @@ impl Store {
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id as u8)
{
continue;
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
#[cfg(test)]
mod tests {
use super::*;
@@ -1184,6 +1236,76 @@ mod tests {
assert!(!std::ptr::eq(ev0, ev1));
}
/// `find_ec_volume` returns only disk 0's runtime, which is what hides
/// sibling-disk shards from every scrub mode. The plural lookup must
/// return one runtime per disk holding the vid, in location order.
#[test]
fn test_find_all_ec_volumes_returns_every_disk() {
let (store, _tmp) = build_split_disk_store(7010);
let vid = VolumeId(7010);
let all = store.find_all_ec_volumes(vid);
assert_eq!(all.len(), 2, "expected one EcVolume per disk holding the vid");
// Disk 0 carries shards 0 and 12; disk 1 carries shard 1.
assert!(all[0].has_shard(0));
assert!(all[0].has_shard(12));
assert!(all[1].has_shard(1));
// The singular lookup sees only the first — the bug being fixed.
let first = store.find_ec_volume(vid).unwrap();
assert!(std::ptr::eq(first, all[0]));
// A vid nobody mounts yields an empty vec, not a panic.
assert!(store.find_all_ec_volumes(VolumeId(9999)).is_empty());
}
/// End-to-end: with the vid mounted on two disks, a scrub driven through
/// the Store must reach BOTH disks' shards. Before the aggregation fix
/// `find_ec_volume` returned disk 0 and disk 1's shard 1 was never read.
#[test]
fn test_scrub_plans_reach_every_disk_through_the_store() {
use crate::storage::erasure_coding::ec_volume::{
merge_ec_runtimes, EcChecksumScrubPlan, EcLocalScrubPlan,
};
let (store, _tmp) = build_split_disk_store(7030);
let vid = VolumeId(7030);
let runtimes = store.find_all_ec_volumes(vid);
assert_eq!(runtimes.len(), 2);
// Reachability is the invariant, so assert on the resolved slots rather
// than on scrub message text: shards 0 and 12 live on disk 0, shard 1 on
// disk 1. The old first-match lookup could never see shard 1.
let merged = merge_ec_runtimes(&runtimes).expect("two runtimes merge");
assert!(merged.slots[0].is_some(), "disk 0's shard 0 unreachable");
assert!(merged.slots[12].is_some(), "disk 0's shard 12 unreachable");
assert!(merged.slots[1].is_some(), "disk 1's shard 1 unreachable — the bug");
assert!(merged.skipped.is_empty(), "same generation: {:?}", merged.skipped);
// Shard 1 is owned by the sibling runtime, not the anchor.
let (owner, _) = merged.slots[1].unwrap();
assert!(std::ptr::eq(owner, runtimes[1]));
// Both plans build over the union rather than over disk 0 alone.
assert!(EcChecksumScrubPlan::for_volumes(&runtimes).is_some());
assert!(EcLocalScrubPlan::for_volumes(&runtimes).is_some());
// ...and `is_some()` is a real question: `for_volumes` has exactly one
// `None` (the vanished-volume case), so without this the two lines above
// would hold for any input at all.
assert!(EcChecksumScrubPlan::for_volumes(&[]).is_none());
assert!(EcLocalScrubPlan::for_volumes(&[]).is_none());
// Regression guard: a single-runtime view still sees only its own disk,
// which is exactly what made aggregation necessary.
let disk0 = merge_ec_runtimes(&[runtimes[0]]).unwrap();
assert!(
disk0.slots.get(1).copied().flatten().is_none(),
"disk 0's runtime must not see the sibling's shard"
);
}
/// `Store::unmount_ec_shards` used to return after the first
/// location with the vid, so a request to unmount a shard that
/// lives on a sibling disk became a silent no-op. After the fix,
@@ -1689,50 +1811,3 @@ mod tests {
assert!(std::path::Path::new(&format!("{}.ecx", ec_base)).exists());
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
continue;
}
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
+4 -24
View File
@@ -155,7 +155,7 @@ impl Size {
return 0;
}
if self.0 < 0 {
return (self.0 * -1) as u32;
return -self.0 as u32;
}
self.0 as u32
}
@@ -284,8 +284,9 @@ impl fmt::Display for Offset {
// DiskType
// ============================================================================
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
#[derive(Debug, Clone, PartialEq, Eq, Hash, Default)]
pub enum DiskType {
#[default]
HardDrive,
Ssd,
Custom(String),
@@ -319,12 +320,6 @@ impl fmt::Display for DiskType {
}
}
impl Default for DiskType {
fn default() -> Self {
DiskType::HardDrive
}
}
// ============================================================================
// VolumeId
// ============================================================================
@@ -397,7 +392,7 @@ impl From<u8> for Version {
///
/// Fields are split into request-side options (set by the caller) and response-side
/// flags (set during the read to communicate status back).
#[derive(Debug, Clone)]
#[derive(Debug, Clone, Default)]
pub struct ReadOption {
// -- request --
/// If true, allow reading needles that have been soft-deleted.
@@ -423,21 +418,6 @@ pub struct ReadOption {
pub read_buffer_size: i32,
}
impl Default for ReadOption {
fn default() -> Self {
ReadOption {
read_deleted: false,
attempt_meta_only: false,
must_meta_only: false,
is_meta_only: false,
volume_revision: 0,
is_out_of_range: false,
has_slow_read: false,
read_buffer_size: 0,
}
}
}
// ============================================================================
// NeedleMapEntry helpers (for .idx file)
// ============================================================================
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,376 @@
//! Rebuild a missing .idx from the .dat it indexes. Mirrors
//! `weed/storage/volume_idx_rebuild.go`.
use std::fs::{self, OpenOptions};
use std::io::{BufWriter, Write};
use std::path::Path;
use crate::storage::idx;
use crate::storage::needle::Needle;
use crate::storage::super_block::SuperBlock;
use crate::storage::types::*;
use crate::storage::volume::{
fsync_dir, needle_disk_end, scan_volume_file, Volume, VolumeError, VolumeFileVisitor,
};
/// Writes one .idx row per .dat record, in .dat append order, which is the
/// shape the volume server's own writes leave behind.
struct VolumeFileScanner4RebuildIdx<W: Write> {
writer: W,
dat_size: i64,
version: Version,
stopped: bool,
}
impl<W: Write> VolumeFileVisitor for VolumeFileScanner4RebuildIdx<W> {
fn visit_super_block(&mut self, _sb: &SuperBlock) -> Result<(), VolumeError> {
Ok(())
}
fn read_needle_body(&self) -> bool {
false
}
fn visit_needle(&mut self, n: &Needle, offset: i64) -> Result<(), VolumeError> {
// A record reaching past the end of .dat is a torn append or a corrupt
// header: nothing beyond it is indexable, and a row pointing past EOF
// would fail every read of that needle. The all-zero header case ends
// the walk upstream.
if self.stopped {
return Ok(());
}
if needle_disk_end(Offset::from_actual_offset(offset), n.size, self.version) > self.dat_size
{
self.stopped = true;
return Ok(());
}
let size = if n.size.is_valid() {
n.size
} else {
TOMBSTONE_FILE_SIZE
};
idx::write_index_entry(
&mut self.writer,
n.id,
Offset::from_actual_offset(offset),
size,
)?;
Ok(())
}
}
impl Volume {
/// Regenerate the volume's .idx from its .dat. The whole index is derivable
/// from the data file, so a volume whose index directory has no .idx -- a
/// --dir.idx pointed at an empty directory, or a lost index -- comes back on
/// its own instead of mounting with every needle invisible. The rows go to a
/// temp file that is renamed in, so an interrupted rebuild leaves no partial
/// index behind.
pub(crate) fn rebuild_idx_file(&self) -> Result<(), VolumeError> {
let idx_path = self.file_name(".idx");
let dat_path = self.file_name(".dat");
let tmp_path = format!("{idx_path}.tmp");
let rebuild = || -> Result<(), VolumeError> {
if let Some(parent) = Path::new(&idx_path).parent() {
fs::create_dir_all(parent)?;
}
let tmp_file = OpenOptions::new()
.write(true)
.create(true)
.truncate(true)
.open(&tmp_path)?;
let mut scanner = VolumeFileScanner4RebuildIdx {
writer: BufWriter::new(&tmp_file),
dat_size: fs::metadata(&dat_path)?.len() as i64,
version: self.version(),
stopped: false,
};
scan_volume_file(&dat_path, &mut scanner)?;
scanner.writer.flush()?;
drop(scanner);
tmp_file.sync_all()?;
fs::rename(&tmp_path, &idx_path)?;
fsync_dir(&idx_path)?;
Ok(())
};
let result = rebuild();
if result.is_err() {
let _ = fs::remove_file(&tmp_path);
}
result
}
}
#[cfg(test)]
mod tests {
use crate::storage::needle::crc::CRC;
use crate::storage::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::*;
use crate::storage::volume::Volume;
use std::fs;
use std::path::Path;
use tempfile::TempDir;
fn needle(id: u64) -> Needle {
let data = format!("payload-{id}").into_bytes();
Needle {
id: NeedleId(id),
cookie: Cookie(0x55),
data_size: data.len() as u32,
checksum: CRC::new(&data),
data,
..Needle::default()
}
}
// Pointing --dir.idx at a directory with no .idx used to mount the volume on
// an empty index; the index is derivable from the .dat, so it must be
// rebuilt in place instead.
#[test]
fn test_load_moved_idx_directory_rebuilds_idx() {
let root = TempDir::new().unwrap();
let data_dir = root.path().join("data");
let old_idx_dir = root.path().join("idxA");
let new_idx_dir = root.path().join("idxB");
for dir in [&data_dir, &old_idx_dir, &new_idx_dir] {
fs::create_dir_all(dir).unwrap();
}
let data = data_dir.to_str().unwrap();
let old_idx = old_idx_dir.to_str().unwrap();
let new_idx = new_idx_dir.to_str().unwrap();
let mut v = Volume::new(
data,
old_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for id in 1..=3 {
v.write_needle(&mut needle(id), true, false).unwrap();
}
v.delete_needle(&mut needle(2)).unwrap();
v.sync_to_disk().unwrap();
let (want_count, want_deleted) = (v.file_count(), v.deleted_count());
drop(v);
let seeded = fs::read(format!("{old_idx}/1.idx")).unwrap();
let reopened = Volume::new(
data,
new_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
assert!(
Path::new(&format!("{new_idx}/1.idx")).exists(),
"idx not rebuilt in the new idx dir"
);
assert_eq!(reopened.file_count(), want_count);
assert_eq!(reopened.deleted_count(), want_deleted);
assert_eq!(
fs::read(format!("{new_idx}/1.idx")).unwrap(),
seeded,
"rebuilt idx differs from the one the server wrote"
);
for id in [1, 3] {
let mut got = needle(id);
got.data.clear();
reopened.read_needle(&mut got).unwrap();
assert_eq!(got.data, format!("payload-{id}").into_bytes());
}
}
// A .dat padded with zeros must not be indexed as needle 0 rows: the walk
// stops where the records do.
#[test]
fn test_rebuild_idx_stops_at_zero_padded_dat_tail() {
let root = TempDir::new().unwrap();
let dir = root.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
v.sync_to_disk().unwrap();
drop(v);
let seeded = fs::read(format!("{dir}/1.idx")).unwrap();
let dat = fs::OpenOptions::new()
.write(true)
.open(format!("{dir}/1.dat"))
.unwrap();
let dat_size = dat.metadata().unwrap().len();
dat.set_len(dat_size + 4096).unwrap();
drop(dat);
fs::remove_file(format!("{dir}/1.idx")).unwrap();
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
drop(reopened);
assert_eq!(
fs::read(format!("{dir}/1.idx")).unwrap(),
seeded,
"the zero-padded tail leaked into the rebuilt idx"
);
}
// A .dat whose last append was torn mid-body must not gain an index row
// that points past the end of the file.
#[test]
fn test_rebuild_idx_skips_truncated_dat_tail() {
let root = TempDir::new().unwrap();
let dir = root.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
v.sync_to_disk().unwrap();
let kept = fs::read(format!("{dir}/1.idx")).unwrap();
v.write_needle(&mut needle(2), true, false).unwrap();
v.sync_to_disk().unwrap();
drop(v);
// Chop the second needle's body, leaving its header intact.
let dat = fs::OpenOptions::new()
.write(true)
.open(format!("{dir}/1.dat"))
.unwrap();
let torn_size = dat.metadata().unwrap().len() - 8;
dat.set_len(torn_size).unwrap();
drop(dat);
fs::remove_file(format!("{dir}/1.idx")).unwrap();
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
drop(reopened);
assert_eq!(
fs::read(format!("{dir}/1.idx")).unwrap(),
kept,
"the torn record leaked into the rebuilt idx"
);
}
// A corrupt header carrying a negative size advances the .dat walk
// backwards, which cycles forever between it and the record before it.
#[test]
fn test_rebuild_idx_stops_at_negative_size_header() {
let root = TempDir::new().unwrap();
let dir = root.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
v.sync_to_disk().unwrap();
let kept = fs::read(format!("{dir}/1.idx")).unwrap();
v.write_needle(&mut needle(2), true, false).unwrap();
v.sync_to_disk().unwrap();
drop(v);
let rows = fs::read(format!("{dir}/1.idx")).unwrap();
let (_, offset, _) = idx_entry_from_bytes(&rows[NEEDLE_MAP_ENTRY_SIZE..]);
let corrupt_at = offset.to_actual_offset();
// Overwrite the second needle's size field with a negative i32.
use std::io::{Seek, SeekFrom, Write};
let mut dat = fs::OpenOptions::new()
.write(true)
.open(format!("{dir}/1.dat"))
.unwrap();
dat.seek(SeekFrom::Start(
corrupt_at as u64 + COOKIE_SIZE as u64 + NEEDLE_ID_SIZE as u64,
))
.unwrap();
dat.write_all(&[0xff, 0xff, 0xf0, 0x00]).unwrap();
drop(dat);
fs::remove_file(format!("{dir}/1.idx")).unwrap();
// Pre-fix the rebuild walked backwards from here and never terminated.
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
drop(reopened);
assert_eq!(
fs::read(format!("{dir}/1.idx")).unwrap(),
kept,
"the corrupt record leaked into the rebuilt idx"
);
}
}
+27
View File
@@ -0,0 +1,27 @@
//! Opening volume data and index files without access-time updates.
use std::fs::{File, OpenOptions};
use std::io;
use std::path::Path;
/// Open a volume data or index file with `O_NOATIME`. Nothing reads these
/// files' atime, but without the flag every needle read dirties the inode —
/// even relatime writes atime on the first read after each write, so an
/// actively written volume pays a metadata write per read/write cycle.
/// Matches Go's `backend.OpenVolumeFile`.
pub fn open_volume_file(opts: &OpenOptions, path: impl AsRef<Path>) -> io::Result<File> {
let path = path.as_ref();
#[cfg(target_os = "linux")]
{
use std::os::unix::fs::OpenOptionsExt;
let mut noatime = opts.clone();
noatime.custom_flags(libc::O_NOATIME);
match noatime.open(path) {
// O_NOATIME is refused unless we own the file or hold CAP_FOWNER.
Err(e) if e.raw_os_error() == Some(libc::EPERM) => opts.open(path),
result => result,
}
}
#[cfg(not(target_os = "linux"))]
opts.open(path)
}
@@ -30,6 +30,9 @@ pub fn report_hash(m: &master_pb::VolumeInformationMessage) -> u64 {
if m.read_only {
buf[56] = 1;
}
if m.read_only_can_delete {
buf[56] |= 2;
}
let mut h = xxh64(&buf, 0);
h = fold(h, xxh64(&(m.modified_at_second as u64).to_le_bytes(), 0));
@@ -70,8 +73,10 @@ mod tests {
let empty = master_pb::VolumeInformationMessage::default();
assert_eq!(report_hash(&empty), 10988706248825469653);
let mut one = master_pb::VolumeInformationMessage::default();
one.id = 1;
let one = master_pb::VolumeInformationMessage {
id: 1,
..Default::default()
};
assert_eq!(report_hash(&one), 2035849960016744285);
let full = master_pb::VolumeInformationMessage {
@@ -82,6 +87,7 @@ mod tests {
delete_count: 2,
deleted_byte_count: 99,
read_only: true,
read_only_can_delete: false,
replica_placement: 10,
version: 3,
ttl: 3 << 8,
+1 -1
View File
@@ -66,7 +66,7 @@ fn parse_go_version_number() -> Option<String> {
}
}
match (major, minor) {
(Some(maj), Some(min)) => Some(format!("{}.{}", maj, format!("{:02}", min))),
(Some(maj), Some(min)) => Some(format!("{}.{:02}", maj, min)),
_ => None,
}
}
+13
View File
@@ -5451,6 +5451,14 @@ dependencies = [
"tracing",
]
[[package]]
name = "seaweed-worker-sort"
version = "0.1.0"
dependencies = [
"anyhow",
"seaweed-worker-core",
]
[[package]]
name = "security-framework"
version = "3.7.0"
@@ -6542,14 +6550,19 @@ dependencies = [
"clap",
"futures",
"lance",
"lance-datafusion",
"lance-index",
"lance-linalg",
"lance-table",
"prometheus",
"reqwest 0.12.28",
"rustls",
"seaweed-worker-core",
"seaweed-worker-sort",
"serde",
"serde_json",
"tokio",
"tonic",
"tracing",
"tracing-subscriber",
]
+16 -3
View File
@@ -2,14 +2,27 @@
#
# `core` is the plugin.proto contract and nothing else; a worker crate beside it
# supplies job handlers and a binary. Adding a worker means adding a member here,
# not touching the protocol.
# not touching the protocol. `sort` is neither: it is the sort specification the
# sorting jobs share, so that two of them cannot drift on what an order means.
[workspace]
resolver = "2"
members = ["crates/core", "crates/lance"]
members = ["crates/core", "crates/lance", "crates/sort"]
[workspace.package]
version = "0.1.0"
edition = "2021"
edition = "2024"
# The edition needs 1.85; the dependency tree needs more. Verified with
# `cargo +1.94.1 check --all-targets` (1.94.0 fails on the AWS SDK that
# lance's `aws` feature pulls in).
rust-version = "1.94.1"
[workspace.lints.clippy]
# Every RPC path returns tonic::Status (176 bytes). Boxing it would change
# every handler signature for no gain, so the large-Err lint is off.
result_large_err = "allow"
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[workspace.dependencies]
anyhow = "1"
+6
View File
@@ -14,6 +14,12 @@ one.
## Building
Requires Rust 1.94.1+ (2024 edition), matching `rust-version` in `Cargo.toml`.
The patch release matters: 1.94.0 does not build. The edition itself only needs
1.85; the higher floor comes from the dependency tree — lance's `aws` feature
pulls in the AWS SDK — so it moves with those crates. CI builds on the latest
stable.
`core` compiles `plugin.proto` with the protoc that protoc-bin-vendored ships,
the way seaweed-volume does, so it needs no system install.
+4
View File
@@ -2,6 +2,7 @@
name = "seaweed-worker-core"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
description = "SeaweedFS plugin.proto worker contract"
[lib]
@@ -25,3 +26,6 @@ tonic-build.workspace = true
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry. The same crate seaweed-volume uses.
protoc-bin-vendored = "3"
[lints]
workspace = true
+6 -1
View File
@@ -4,7 +4,12 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
// version. An explicit PROTOC still wins, for packagers supplying their own
// and for the lance crates, whose own build scripts read the same variable.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
// SAFETY: a build script's main runs single-threaded before anything
// else in this process, so no other thread can be reading the
// environment concurrently.
unsafe {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
}
// Compiled straight out of the Go tree, the way seaweed-volume already reads
+4 -4
View File
@@ -14,10 +14,10 @@ pub fn server_to_grpc_address(server: &str) -> Option<String> {
let (host, port_part) = server.rsplit_once(':')?;
// "port.grpcPort" states the gRPC port outright.
if let Some((_, grpc_port)) = port_part.split_once('.') {
if let Ok(port) = grpc_port.parse::<u16>() {
return Some(join_host_port(host, port));
}
if let Some((_, grpc_port)) = port_part.split_once('.')
&& let Ok(port) = grpc_port.parse::<u16>()
{
return Some(join_host_port(host, port));
}
let port: u16 = port_part.parse().ok()?;
@@ -64,6 +64,17 @@ pub fn number_field(name: &str, label: &str, description: &str, min: i64, max: i
}
}
pub fn text_field(name: &str, label: &str, description: &str, placeholder: &str) -> ConfigField {
ConfigField {
name: name.to_string(),
label: label.to_string(),
description: description.to_string(),
field_type: ConfigFieldType::String as i32,
placeholder: placeholder.to_string(),
..Default::default()
}
}
pub fn bool_field(name: &str, label: &str, description: &str) -> ConfigField {
ConfigField {
name: name.to_string(),
+3
View File
@@ -16,6 +16,9 @@ pub mod stream;
/// Generated plugin.proto types.
pub mod pb {
// prost gives every oneof its own enum; the variant sizes are the
// messages' own, not a choice made here.
#![allow(clippy::large_enum_variant)]
tonic::include_proto!("plugin");
}
+12
View File
@@ -2,6 +2,7 @@
name = "weed-lance-worker"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
description = "SeaweedFS maintenance worker for Lance tables"
[lib]
@@ -16,11 +17,17 @@ path = "src/main.rs"
[dependencies]
seaweed-worker-core = { path = "../core" }
seaweed-worker-sort = { path = "../sort" }
prometheus.workspace = true
# Only the S3 backend: the other object stores lance enables by default are
# build time this worker never spends.
lance = { version = "10", default-features = false, features = ["aws"] }
lance-index = "10"
# For the manifest's own fragment type, which the sort marker summarises.
lance-table = "10"
# For the execution options the sort needs; the scanner's own stream helper
# takes the default, which has spilling off.
lance-datafusion = "10"
arrow-schema = "58"
arrow-cast = "58"
chrono = "0.4"
@@ -29,6 +36,7 @@ anyhow.workspace = true
async-trait.workspace = true
clap = { version = "4", features = ["derive", "env"] }
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls"] }
rustls = "0.23"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
tokio.workspace = true
@@ -37,7 +45,11 @@ tracing-subscriber.workspace = true
[dev-dependencies]
tokio = { workspace = true }
tonic.workspace = true
arrow-array = "58"
arrow-schema = "58"
arrow-cast = "58"
lance-linalg = "10"
[lints]
workspace = true

Some files were not shown because too many files have changed in this diff Show More