Compare commits

...
Author SHA1 Message Date
Chris LuandGitHub 4d0afa3286 operation: fold extra dial options into WithVolumeServerClient (#11392)
WithVolumeServerClientOptions duplicated the existing helper only to
append extra dial options. Trail the extras as a variadic parameter on
WithVolumeServerClient instead, keeping the (dialOption, fn) argument
order so all existing callers keep working. TailVolumeFromSource gets
the same treatment.
2026-09-18 13:26:31 -07:00
Chris LuandGitHub 09f835842c Merge branch 'master' into fix/volume-copy-source-guard 2026-09-18 13:24:53 -07:00
Chris LuandGitHub 37bf1cd91d volume: validate copy/tail source addresses before dialing (#11390)
* pb: stop exiting the process on malformed server addresses

ServerToGrpcAddress and GrpcAddressToServerAddress called glog.Fatalf
when hostAndPort could not parse the port, which os.Exit(255)ed the whole
process. A caller-supplied copy or tail source address reached this path
synchronously in the serving goroutine, so one anonymous VolumeCopy with
a non-numeric port terminated the volume server.

Log the parse error and return the input unchanged instead: the dial or
request that consumes the address then fails as an ordinary error.

* volume: validate copy and tail source addresses before dialing

VolumeCopy, VolumeEcShardsCopy and VolumeTailReceiver dial a
caller-supplied source address (SourceDataNode / SourceVolumeServer)
with no endpoint validation, so an anonymous caller could aim the volume
server at loopback, link-local (cloud metadata) or other unintended
destinations and read dial behavior back as a connectivity oracle.

Apply the same peer-target deny list FetchAndWriteNeedle uses for
replica targets: the source must be a bare host:port whose host is not
loopback, link-local or unspecified; cluster peers stay reachable on
private networks, and -volume.allowUntrustedRemoteEndpoints opts out.
The loopback-using copy tests set the flag to keep exercising the copy
path in process.

* rust volume: validate copy and tail source addresses before dialing

Mirror the Go guard on the Rust volume server: volume_copy,
volume_ec_shards_copy and volume_tail_receiver dial a caller-supplied
source address, so run it through validate_replica_target first (bare
host:port; no loopback, link-local or unspecified hosts; private peers
stay allowed). --volume.allowUntrustedRemoteEndpoints opts out; the test
fixture and the Rust test-cluster launcher set it so loopback sources in
tests keep working.

* volume: pin validated copy/tail source addresses at dial time

validateReplicaTarget resolves the source hostname once, but the gRPC
client resolved it again at connect, leaving a DNS-rebinding window for
hostname sources. The copy and tail source dials now run through the
same guardedDialerPolicy the remote-storage path uses, so every resolved
address is re-checked against the replica deny list (private peers
allowed) immediately before the TCP connect. guardedDialerPolicy also
moves to util.OutboundDialContext so the guarded path keeps the -ip.bind
source binding the default gRPC dialer had.

The Rust volume server mirrors this with connect_guarded, a tonic
connector that resolves, re-checks each address, and connects to the
first passing IP; handlers use it whenever the untrusted-endpoint
opt-out is off. A handler-level test now exercises the enabled
validation branches for all three source-taking RPCs.

* pb: return empty server address for malformed grpc addresses

GrpcAddressToServerAddress used to return the unparseable input on a
hostAndPort failure, so a malformed raft address (e.g. "host:abc")
flowed into admin dashboard master maps unchanged. Return an empty
string instead, skip empty conversions at the two raft-cluster merge
sites, and drop the now-stale comment about the fatal exit the earlier
commit removed.

* test: opt erasure-coding loopback clusters out of the remote endpoint guard

The erasure-coding suites drive VolumeEcShardsCopy / VolumeCopy between
volume servers bound to 127.0.0.1, which the copy/tail source guard now
rejects by default. Pass -volume.allowUntrustedRemoteEndpoints to the
test volume launches, matching what the volume_server framework
harnesses already do.

* admin: only claim fallback master leadership on an empty raft response

A nonempty RaftListClusterServers response whose entries were all
rejected left masterMap empty, so the fallback marked the reachable
current master as leader the same way a genuinely empty (non-raft)
response does. Track whether the successful response returned zero
servers and only promote the fallback master then.
2026-09-18 12:55:47 -07:00
Chris LuandGitHub a6d72bc272 s3api: delete orphaned chunks only when the entry is confirmed absent (#11389)
* s3api: test for chunks deleted under an entry the filer committed

Issue #11387: the filer can report a create failure after inserting the
entry (e.g. a parent-directory creation failing post-insert). The error
arrives in the response rather than as a transport status, so it maps
to a definitive error and putToFiler deletes the chunks of the live
entry.

* s3api: confirmCreateLanded also reports a confirmed-absent entry

The verification a failed create runs can answer both directions: the
entry matching the uploaded chunks proves the write landed, and an
authoritative not-found proves the uploaded chunks are orphaned. Return
both outcomes so the cleanup path can gate on the fact rather than the
error class. An empty upload can never prove a landing, so a zero-chunk
entry match no longer upgrades the outcome.

* s3api: delete orphaned chunks only when the entry is confirmed absent

A failed create no longer skips verification based on the error class: the filer can fail after inserting the entry (issue #11387) and a partially-applied routed transaction can leave it behind too, both surfacing as definitive errors. Every failed create now resolves the entry's fate, and the uploaded chunks are deleted only when the entry is confirmed absent; anything unverifiable keeps them for vacuum.

* s3api: confirm absence on every filer the create could have committed on

A lock-path create fails over across filers, so the entry can live on a replica the routed owner has not caught up to; one not-found does not prove absence. The confirmation now queries the owner, the prior owner, and the failover set, declaring absent only when none of them has the entry.

* s3api: bound the reconciliation lookups confirmCreateLanded runs

The lookups ran on context.Background() under the object write lock, so a connected filer that never replies could stall the write path. One timeout now covers the whole enumeration; an expired budget fails the remaining lookups as uncertain, which keeps the chunks.
2026-09-18 12:30:07 -07:00
Chris Lu 934b9b4daf admin: only claim fallback master leadership on an empty raft response
A nonempty RaftListClusterServers response whose entries were all
rejected left masterMap empty, so the fallback marked the reachable
current master as leader the same way a genuinely empty (non-raft)
response does. Track whether the successful response returned zero
servers and only promote the fallback master then.
2026-09-18 12:22:39 -07:00
Chris Lu bd41ce39f7 test: opt erasure-coding loopback clusters out of the remote endpoint guard
The erasure-coding suites drive VolumeEcShardsCopy / VolumeCopy between
volume servers bound to 127.0.0.1, which the copy/tail source guard now
rejects by default. Pass -volume.allowUntrustedRemoteEndpoints to the
test volume launches, matching what the volume_server framework
harnesses already do.
2026-09-18 11:34:05 -07:00
Chris Lu 0d6024e2e0 pb: return empty server address for malformed grpc addresses
GrpcAddressToServerAddress used to return the unparseable input on a
hostAndPort failure, so a malformed raft address (e.g. "host:abc")
flowed into admin dashboard master maps unchanged. Return an empty
string instead, skip empty conversions at the two raft-cluster merge
sites, and drop the now-stale comment about the fatal exit the earlier
commit removed.
2026-09-18 11:34:05 -07:00
Chris Lu 0ae7874ed9 volume: pin validated copy/tail source addresses at dial time
validateReplicaTarget resolves the source hostname once, but the gRPC
client resolved it again at connect, leaving a DNS-rebinding window for
hostname sources. The copy and tail source dials now run through the
same guardedDialerPolicy the remote-storage path uses, so every resolved
address is re-checked against the replica deny list (private peers
allowed) immediately before the TCP connect. guardedDialerPolicy also
moves to util.OutboundDialContext so the guarded path keeps the -ip.bind
source binding the default gRPC dialer had.

The Rust volume server mirrors this with connect_guarded, a tonic
connector that resolves, re-checks each address, and connects to the
first passing IP; handlers use it whenever the untrusted-endpoint
opt-out is off. A handler-level test now exercises the enabled
validation branches for all three source-taking RPCs.
2026-09-18 11:32:28 -07:00
Chris Lu 81778defa1 rust volume: validate copy and tail source addresses before dialing
Mirror the Go guard on the Rust volume server: volume_copy,
volume_ec_shards_copy and volume_tail_receiver dial a caller-supplied
source address, so run it through validate_replica_target first (bare
host:port; no loopback, link-local or unspecified hosts; private peers
stay allowed). --volume.allowUntrustedRemoteEndpoints opts out; the test
fixture and the Rust test-cluster launcher set it so loopback sources in
tests keep working.
2026-09-18 10:57:18 -07:00
Chris Lu 39a8d3253c volume: validate copy and tail source addresses before dialing
VolumeCopy, VolumeEcShardsCopy and VolumeTailReceiver dial a
caller-supplied source address (SourceDataNode / SourceVolumeServer)
with no endpoint validation, so an anonymous caller could aim the volume
server at loopback, link-local (cloud metadata) or other unintended
destinations and read dial behavior back as a connectivity oracle.

Apply the same peer-target deny list FetchAndWriteNeedle uses for
replica targets: the source must be a bare host:port whose host is not
loopback, link-local or unspecified; cluster peers stay reachable on
private networks, and -volume.allowUntrustedRemoteEndpoints opts out.
The loopback-using copy tests set the flag to keep exercising the copy
path in process.
2026-09-18 10:57:18 -07:00
Chris Lu 47f323bbb3 pb: stop exiting the process on malformed server addresses
ServerToGrpcAddress and GrpcAddressToServerAddress called glog.Fatalf
when hostAndPort could not parse the port, which os.Exit(255)ed the whole
process. A caller-supplied copy or tail source address reached this path
synchronously in the serving goroutine, so one anonymous VolumeCopy with
a non-numeric port terminated the volume server.

Log the parse error and return the input unchanged instead: the dial or
request that consumes the address then fails as an ordinary error.
2026-09-18 10:57:14 -07:00
Chris LuandGitHub c72eda50a8 s3: drop implicit reader cache budget that throttled S3 GETs (#11384)
* fix(filer): leave reader cache unbounded without an explicit budget

NewReaderCache silently installed a 256MiB ReaderCacheBudget when the
caller passed none. Only weed mount opts into a budget; every other
caller (S3 gateway, WebDAV, query engine, mq logstore) inherited the
cap. Under ~90 concurrent S3 GETs of medium objects, prefetch wants
far more than 64 chunk buffers, so reserve() serialized chunk fetches,
clients timed out and retried, and the retry re-downloaded chunks the
cancelled request had already fetched.

A nil budget now means unbounded, restoring the pre-4.47 behavior for
callers that never asked for a memory cap; reserve/complete/release
are nil-safe. The mount path is unchanged and still enforces
-readerCacheSizeMB.

Fixes #11380

* feat(s3): expose -s3.readerCacheSizeMB reader buffer budget

Operators who want the S3 gateway read path memory-bounded can now
opt in: -s3.readerCacheSizeMB on weed filer/server/mini and
-readerCacheSizeMB on standalone weed s3, matching the mount flag.
The default 0 keeps the unbounded pre-4.47 behavior; a positive value
installs a shared ReaderCacheBudget across in-flight and retained
chunk buffers for all S3 GETs.

* fix(filer): validate chunk size before consulting the reader budget

A nil budget returned early and skipped the negative chunkSize check,
letting a corrupted size reach mem.Allocate and panic. Also drop the
command-specific flag prefix from the S3 validation error since
standalone weed s3 exposes the option as -readerCacheSizeMB.

* filer: drop chunk buffers once fully consumed

ReaderCache retained every completed chunk buffer in the downloaders
map until the slot limit evicted it, so buffers lingered after all
readers finished with them.

Track attached readers on each SingleChunkCacher and remove the cacher
when the last reader consumes the buffer to its end. In-flight download
deduplication and the prefetch handoff are unchanged: a buffer always
survives until fully read, partial reads keep it available, and an
attached reader pins a consumed buffer until it detaches. Repeat reads
now go through the chunk cache where enabled, or refetch.

* filer: drop consumed buffers on last detach, rechecked under cache lock

Two review findings on the drop-on-consume change:

- Removal only fired when the detaching reader itself reached the chunk
  end. If the end-reaching reader finished first and the last remaining
  reader did a partial read or cancelled, the consumed buffer and its
  budget reservation lingered until eviction. Track a persistent
  consumed flag instead, so any end-reaching read marks the buffer and
  the last detach drops it.
- remove() checked only map identity, so a reader attaching between the
  reader count hitting zero and removal could attach to a cacher that
  was then deleted underneath it. removeConsumed() re-checks identity,
  readers == 0, and consumed under the ReaderCache lock; a raced attach
  keeps the cacher and its own detach retries the removal.
2026-09-18 01:02:32 -07:00
Chris LuandGitHub 87ee3b63a2 s3: abort completed multipart uploads metadata-only (#11385)
* s3: abort a completed upload's leftover directory metadata-only

A .uploads/<id> directory can outlive the object it completed into when
the commit's metadata-only removal failed or the gateway died in between;
the restored part entries then share chunks with the published object.
AbortMultipartUpload deleted the directory recursively, chunks and all,
so aborting such a leftover destroyed a committed object (#11382).

Run the same check s3.clean.uploads gained in #11375 before deleting:
when the object entry or a version file under <key>.versions carries the
upload id, remove .uploads/<id> metadata-only and answer the abort; when
the lookup cannot decide, refuse with InternalError rather than risk
live chunks.

* s3: apply the completed-upload check to lifecycle MPU abort

lifecycleAbortMPU ran the same destructive recursive delete on
.uploads/<id>. Reuse uploadCompleted so a leftover whose object entry or
version file carries the upload id is removed metadata-only, and an
undecidable lookup retries later instead of freeing live chunks.

* s3: serialize abort's upload-dir delete with the object's commit

The completed check alone leaves a race: abort can read completed=false,
then an in-flight completion publishes the object over the same part
chunks before the recursive delete frees them.

Run the check and delete inside the object write lock, which non-routed
completions hold for their whole finalize. With an owner, send the data
delete as an ObjectTransaction on the object's lock key — a routed
commit then either loses its upload-exists precondition after our delete
or has already stamped the object, which the transaction's
IF_EXTENDED_NOT_EQUAL condition detects and falls back to a
metadata-only remove. lifecycleAbortMPU shares removeUploadDir so both
callers get the same ordering.

* s3: check for an empty object before resolving its write owner

* s3: check completion at the abort's resolved object key

An upload record missing ExtMultipartObjectKey skipped the completed
check entirely even though the request's Key names the object.
2026-09-18 01:01:52 -07:00
Chris LuandGitHub 0ca1c19821 s3api: unify auth error handling across s3tables, iceberg and lance (#11381)
* s3api: fail closed when S3 Tables signature verification fails

* s3api: avoid nil Account dereference in S3 Tables auth log

* iceberg: return auth error instead of falling back to DefaultAllow

* lance: return auth error instead of falling back to DefaultAllow

* s3api: stop trusting client-supplied s3-account-id

The header is set by the server after successful authentication; scrub
inbound values alongside the other internal headers, and apply the same
admin guard to the header fallback branch of getAccountID that the
identity branch already has.

* test: cover table-catalog auth wrappers and principal resolution

* test: configure anonymous identity where catalog clients do not sign

* s3api: scrub s3-account-id after signature verification
2026-09-18 01:01:04 -07:00
Chris LuandGitHub f40687b34e s3: tighten STS session token handling (#11383)
* s3api: test that a session token must not reveal its credential

* sts: derive secret access key with HMAC keyed on the signing key

* s3api: stop accepting STS session tokens as bearer credentials

* security: reject STS session tokens on filer and admin gRPC auth

* test: sign s3/iam framework requests with the session credential

* s3api: exercise the real auth pipeline in the end-to-end harness
2026-09-17 23:50:16 -07:00
Chris LuandGitHub 15520f601f s3: commit multipart upload and remove .uploads atomically; purge completed uploads metadata-only (#11375)
* s3: commit versioned multipart upload in one transaction

CompleteMultipartUpload wrote the version file, flipped the .versions
pointer, then removed .uploads/<id> metadata-only as a best-effort
post-commit step. A filer error or gateway crash in that window left the
upload directory referencing the same chunks as the published object, and
the next s3.clean.uploads run purged it with data -- corrupting a
committed object.

Put the version file, remove the upload directory metadata-only (its
chunks are the object's chunks), and recompute the latest pointer in one
ObjectTransaction under the object's per-path lock on the owner filer.
The mutation order keeps every partial state safe: the chunks stay
referenced at all times, and a published object never coexists with the
upload directory the cleaner would purge.

Unused part entries are freed before the transaction, since the
metadata-only directory delete would otherwise leak their chunks.

* s3: remove upload directory inside the multipart object PUT

The same committed-object/stranded-upload window existed on the
suspended and non-versioned paths: writeMultipartObject committed the
object, then a best-effort rm dropped .uploads/<id>. Ride the
metadata-only removal on the routed PUT itself so the two land in one
transaction; the unrouted mkFile fallback keeps post-commit cleanup.

* shell: purge completed uploads metadata-only in s3.clean.uploads

A leftover .uploads/<id> can outlive a committed object when the
completion's metadata-only delete fails or the gateway dies in between;
its part entries then share chunks with the live object, and a recursive
purge frees them out from under it.

Before purging a stale upload, check whether it completed: the object
entry or any version file under <key>.versions carrying the upload id.
If so, delete with skipChunkDeletion. If the lookup fails, skip the
upload for this run rather than risk live chunks.

* s3: abort multipart completion when unused part cleanup fails

Deleting the upload directory metadata-only erases the only metadata
pointing at part entries whose deletion failed, orphaning their chunks.
Propagate the error so the completion fails while the upload directory
still exists and the request remains retriable.

* s3: require the upload directory to exist at multipart commit

A delete that does not take the object lock (abort, lifecycle,
s3.clean.uploads) can remove .uploads/<id> and its chunks between the
prepare step and the commit transaction. The commit now carries an
IF_EXISTS precondition on the upload directory so the race fails the
request with NoSuchUpload instead of publishing an object over freed
chunks.

* s3: keep the version file when the upload directory is gone

The finalize transaction has no rollback, so a failure at the
latest-pointer recompute leaves the version written and .uploads/<id>
removed. Deleting the version then destroys the only remaining record of
the upload, making a retried CompleteMultipartUpload return NoSuchUpload
while the version's chunks leak. Roll back only while the upload
directory survives; otherwise keep the version, which a retry resolves
through SeaweedFSUploadId and the version reconciler promotes.

* s3: keep manifests when a routed object write partially commits

For non-versioned and suspended completions the object PUT precedes the
upload-directory DELETE, so an error can mean the object entry exists
while the response reports failure. Freeing this attempt's manifest
chunks then destroys the committed object. Keep them when the object
entry survived, and after a failed null-marker finalize which always
follows a committed write.

* s3: skip the keep-version path on precondition failure

A rejected precondition means no mutation ran, so there is no version
file to preserve and this attempt's manifests are orphans the error
cleanup should free.

* s3: keep manifests when the object-existence check itself fails

A transient lookup error previously read as absent, letting the error
cleanup free manifest chunks a committed object still references.

* s3: keep the upload directory when post-commit part cleanup fails

Removing it metadata-only after a failed entry delete erases the only
reference to the leftover chunks. Leave the directory so the entries
keep their chunk references for s3.clean.uploads or manual recovery.

* pb: fix filer list entry counting on 32-bit

int(limit) wraps to -1 on 386 when limit is math.MaxUint32, so the
beyond-limit check discarded every streamed entry. Compare in uint64
instead; the semantics are unchanged on 64-bit platforms.

* shell: resolve trailing-slash object keys in s3.clean.uploads

Completion stores a key ending in / inside the directory it names
(<bucket>/dir/dir), but FullPath+DirAndName on the normalized key
looked one level too high. Deriving dir and name with path.Dir and
path.Base mirrors getEntryNameAndDir so the completed-upload check
finds the entry instead of purging its chunks.

* s3: heal a suspended completion hidden behind a delete marker

Removing .uploads/<id> inside the commit transaction means a failed
finalizeSuspendedNullWrite leaves nothing to retry against: the object
entry is committed but the marker still makes the key read as deleted,
and a retried CompleteMultipartUpload can only report NoSuchUpload.
When the upload directory is gone, check the regular path for an entry
carrying the upload id and re-run the marker finalize, so the retry
both succeeds and repairs the key. Only suspended buckets can hold
this state; anything newer owns the key.

* s3: report store errors when resuming a committed multipart object
2026-09-17 21:09:21 -07:00
Chris LuandGitHub bdc37a1e86 mount/shell: bucket allow-empty-folders toggle, mount keeps explicit false (#11370)
* mount: keep a deliberate bucket allow-empty-folders setting

* shell: s3.bucket.allowEmptyFolders toggles the empty folder cleaner

* shell: guard allow-empty-folders toggle with expected extended attrs

* filer: drop cached empty-folder policy on bucket entry update

* mount: guard allow-empty-folders write with expected extended attrs

* filer: skip caching a stale cleanup policy read across an update

* mount, shell: snapshot the full extended attributes for update preconditions

* filer: fail closed and invalidate on all bucket entry events for cleanup policy

* filer: key the cleanup policy generation by bucket

* filer: requeue cleanup when the bucket policy cannot be loaded

* filer: expire idle cleanup policy generations

* filer: skip requeueing cleanup after the cleaner stops

* filer: bound cleanup retries on repeated policy failures

* filer: cover cleanup requeue on repeated policy failures

* filer: keep cleanup policy generations while reads are in flight

* filer: exercise the cleanup queue lifecycle in the retry-cap test
2026-09-17 20:30:53 -07:00
Chris LuandGitHub 2d2619f0b4 ci: make telemetry deploy work on Oracle Linux 7 (#11377)
* telemetry/server: tidy module dependencies

* ci: make telemetry deploy work on Oracle Linux 7

* ci: install telemetry unit and logrotate on every deploy

* ci: abort telemetry deploy on install failures
2026-09-17 20:00:25 -07:00
Chris LuandGitHub ce1e0dc30a s3api: don't delete chunks when CreateEntry outcome is ambiguous (#11376)
* s3api: map ambiguous filer transport errors to retryable 503

Canceled, DeadlineExceeded and Unavailable can be returned after the
filer applied the write, so the outcome is ambiguous. Reporting them as
a 4xx tells the client not to retry; report ServiceUnavailable instead.

* s3api: verify entry existence before deleting orphaned chunks

A failed CreateEntry can still have landed on the filer when the error
is a transport failure, and entryCreated=false would tombstone chunks a
live entry references, leaving a dangling pointer that survives only
because reads pass readDeleted=true until vacuum reclaims the needle.

Before deleting, look the entry up: if it is stored with the same
chunks, the write succeeded; if the lookup cannot be answered, keep the
chunks for vacuum to reclaim; only a confirmed absence still cleans up.

* s3api: regression tests for ambiguous CreateEntry outcomes

Covers the three post-create-failure cases in putToFiler: the entry
landed despite the error (treat as success, keep chunks), the entry is
confirmed absent (delete orphans), and the outcome is unverifiable
(keep chunks, return error).

* volume: count reads served from deleted needles

A readDeleted read succeeding on a tombstoned needle is the signal that
metadata still points at deleted data. Count it under a
readDeletedNeedle handler label in both the Go and Rust volume servers
so the condition is visible before vacuum turns it into a 404.

* s3api: never delete chunks on an ambiguous create error

Review feedback on the first fix showed verification could still go
wrong in both directions: a stale or lagged lookup could report
not-found for a committed entry, a prefix object stores its chunks on a
directory entry, and filer-side manifestization rewrites the top-level
chunk ids the comparison relied on.

Rework the rule so the outcome classes are asymmetric:

- A transport-level error (anything filerErrorToS3Error maps to a
  retryable 503) is ambiguous and never deletes chunks; the lookup can
  only upgrade the write to success.
- Any other error is a definitive filer refusal and still cleans up.

confirmCreateLanded asks the write owner first, resolves the stored
entry through chunk manifests, requires an exact match of the uploaded
file ids, and on success runs the finalize callback the failed create
skipped (under the object write lock, with the same rmObject undo the
create path uses). Zero-chunk writes stay ambiguous since they cannot
be told apart by chunks.

* s3api: cover definitive refusals and stale entries in put tests

The confirmed-failure case now uses a definitive refusal so it still
exercises orphan cleanup, and a new case keeps chunks when the stored
entry belongs to an older object rather than this PUT.

* volume: count deleted-needle reads once per request

Streamed Go reads ran the deleted check in readNeedle and again in
readNeedleDataInto, and non-streamed Rust reads in stream_info and the
full-read fallback, double-counting one request. Count at the single
entry probe each implementation takes per GET: readNeedle in Go,
read_needle_stream_info in Rust.

* s3api: run recovered-write rollback under the object lock

Two follow-ups from review: ResolveChunkManifest returns traversed
manifest blobs in its manifestChunks output, so requiring it empty
rejected every manifestized landing; and the rmObject undo ran after
the object write lock was released, so a concurrent newer write could
be deleted between finalize failure and rollback. Compare only the
resolved data chunks and keep the undo inside the lock.

* s3api: verify, finalize and roll back recovered creates in one lock

A lookup done before the object write lock let a concurrent PUT replace
the entry between the chunk comparison and the finalize/rollback
section, so a failed afterCreate could rmObject a newer write. Run the
owner lookup, manifest resolution, chunk comparison, afterCreate and
the conditional undo inside a single withObjectWriteLock section.
2026-09-17 19:58:49 -07:00
github-actions[bot] d4e11a471d docs: regenerate star history chart 2026-09-18 01:41:46 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
08d5daf0c1 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc from 1.44.0 to 1.45.0 (#11371)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:16:38 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
8d34433308 build(deps): bump go.opentelemetry.io/otel/exporters/zipkin from 1.36.0 to 1.45.0 (#11373)
build(deps): bump go.opentelemetry.io/otel/exporters/zipkin

Bumps [go.opentelemetry.io/otel/exporters/zipkin](https://github.com/open-telemetry/opentelemetry-go) from 1.36.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.36.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/zipkin
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:10:51 -07:00
Chris LuandGitHub 4fd67001d9 security: require go 1.26.6 and bump vulnerable deps (#11374)
* security: require go 1.26.6 and bump vulnerable deps

A dependency scan of the 4.47 release flagged the bundled toolchain
and modules:

- github.com/golang/go < 1.26.6 (CVE-2026-39821, CVE-2026-56853,
  CVE-2026-56859, CVE-2026-56862, CVE-2026-56864, CVE-2026-56865,
  CVE-2026-33818, CVE-2026-46600): raise the go directive to 1.26.6
  so every built artifact requires the fixed toolchain.
- google.golang.org/grpc (CVE-2026-84445, CVE-2026-84304): move to
  the fixed dev pseudo-version; released tags through v1.85.0-dev
  remain in the affected range.
- github.com/pelletier/go-toml/v2 <= v2.4.2 (unbounded parser
  recursion): v2.4.3.
- alpine libcrypto3/libssl3 < 3.5.8-r0 (CVE-2026-75803,
  CVE-2026-63073, CVE-2026-63075, CVE-2026-63076, CVE-2026-63072,
  CVE-2026-54874, CVE-2026-18798, CVE-2026-14456, CVE-2026-14457):
  the release images already apk-upgrade the final stage; extend the
  same to the telemetry and admin-integration images.

Same bumps applied to the test/kafka, test/sftp,
kafka-client-loadtest, and telemetry/server modules.

* telemetry: send integration test report above the 10 GiB floor

The collect endpoint keeps reports only when TotalDiskBytes >=
proto.MinDiskBytes, but the integration test still sent 1 GiB, so the
server counted the report and skipped storing it. No cluster_id series
was ever created and /metrics lacked seaweedfs_telemetry_volume_servers.
Send just above the floor (via proto.MinDiskBytes so it cannot silently
drift again) so the expected per-cluster metrics are exported.
2026-09-17 15:09:55 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c6a3280595 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace from 1.44.0 to 1.45.0 (#11372)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-17 14:36:15 -07:00
799c495226 rust volume: one positional read helper; never seek a dup'd handle on Windows (#11342)
* rust volume: one positional read helper; never seek a dup'd handle on Windows

Positional read-exact was hand-rolled four times: the complete
cross-platform version in needle_map/sorted_file.rs, a Windows-only half
in volume.rs whose unix half was inlined as a
cfg(unix)/cfg(windows)/compile_error! triple at three call sites, a
byte-identical Windows-only copy in ec_volume.rs, and read_full_at in
ec_bitrot.rs. Three more sites -- EcVolumeShard::read_at,
EcLocalShard::read_at and ec_encoder::read_at_most -- hand-rolled the
short-read-permitted variant with a cfg(not(unix)) arm that
try_clone()s the handle and seeks it.

That last arm is wrong. A duplicated descriptor shares one kernel file
offset with the original, so seek-then-read is two syscalls against
state another thread can move in between: a concurrent reader or an
append repositions the offset and the read returns bytes from somewhere
else entirely. EcLocalShard::read_at documents that it must never seek,
one line above the seek. Windows seek_read carries its own offset in a
single call, so that window does not exist.

All seven now go through storage::io::{read_exact_at, read_at}, whose
module doc records why duplicating a handle is not a way to get a
private file position -- opening the file again is, as
Volume::dat_scan_plan already does. read_at_most keeps its own
fill-until-EOF loop; only the per-iteration positional read changes.

Behaviour on unix is unchanged: every unix arm was already
FileExt::read_exact_at or FileExt::read_at. The one exception is
ec_bitrot::verify_shard_blocks, which now retries on EINTR (std's
read_exact_at does; the loop it replaces did not) and, on unix, reports
the standard "failed to fill whole buffer" text instead of "short read
on shard block". The Windows arm still says "unexpected EOF in
seek_read"; both carry ErrorKind::UnexpectedEof, as before.

NeedleStreamSource::read_exact_at and Volume::read_exact_at_backend keep
their signatures; only their bodies shrink.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: retry Interrupted in Windows read_exact_at

Unix std's FileExt::read_exact_at ignores ErrorKind::Interrupted and
retries, but the Windows seek_read loop propagated it, so the shared
exact-read contract differed by platform. seek_read can surface
ERROR_OPERATION_ABORTED, which std maps to Interrupted.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 11:47:38 -07:00
Chris LuandGitHub 4ec564469a s3tables: hide inaccessible catalog resources (#11365)
* s3tables: hide inaccessible table buckets

* s3tables: hide inaccessible namespaces

* s3tables: hide inaccessible tables

* s3tables: hide inaccessible resources in rename and namespace delete

RenameTable/RenameView denied on the source now report the same
not-found as a missing source, and the destination name conflict is
checked only after destination authorization so a denied caller cannot
distinguish an existing destination namespace or name from a missing
one. DeleteNamespace denials use the same formatted message as a
missing namespace.
2026-09-17 11:43:52 -07:00
Chris LuandGitHub 66f1754896 s3: enforce dedicated Object Lock actions (#11362)
s3: enforce dedicated object lock actions
2026-09-16 20:34:06 -07:00
Chris LuandGitHub 994e1f7d64 admin: replace Font Awesome with MIT-licensed icons (#11364)
admin: replace Font Awesome with MIT icons
2026-09-16 20:29:14 -07:00
github-actions[bot] 74eeac6b66 docs: regenerate star history chart 2026-09-17 00:46:37 +00:00
0eb638f503 fix(ec): BatchDelete cookie fail-closed via locate_data geometry (#11348)
* fix(ec): BatchDelete cookie fail-closed via locate_data geometry

* fix(ec): honor skip_cookie_check, require full cookie header

* fix(ec): retry short cookie header reads, still fail closed on EOF

* chore(ec): trim cookie validation comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-16 16:25:03 -07:00
David ChristopherandGitHub 1a285c1334 filer: join shutdown paths before closing metadata store (#11363)
Serve can return when its listener closes while HTTP requests are still draining. The main path could then close the metadata store before those requests finish.

Make signal, context, and Serve-exit paths join one shutdown sequence. Drain gRPC and HTTP concurrently with 15-second default limits, then close the store. Test both completion orders.
2026-09-16 16:20:39 -07:00
Eliah RusinandGitHub caf3d157e6 fix(ec): encode drops tombstoned needles, last-wins replay (#11347)
* fix(ec): encode drops tombstoned needles, last-wins replay

* fix(ec): drop zero-offset rows in encode, match readNeedleMap
2026-09-16 15:40:22 -07:00
Chris LuandGitHub 3ebc05930d s3: separate Object Lock configuration permission (#11361)
* s3: separate object lock configuration permission

* test: synchronize manifest cancellation setup
2026-09-16 15:27:09 -07:00
Chris LuandGitHub 0c7beec697 server: add filer-specific disableHttp flag (#11360) 2026-09-16 14:18:57 -07:00
David ChristopherandGitHub a859f0a019 filer: preserve accepted metadata log records on shutdown (#11359)
fix: flush metadata log before closing filer store

Serialize sealed-batch handoffs with shutdown, reject late appends, and wait for log-buffer workers before closing the filer metadata store.

Cover queued writes, interval and explicit flushes, late-write rejection, and pending persistence with shutdown tests.
2026-09-16 12:23:59 -07:00
Eliah RusinandGitHub def25ca84d fix(ec): validate ShardId at gRPC boundary, reject >=32 (#11346) 2026-09-16 08:33:22 -07:00
Eliah RusinandGitHub 701e397337 fix(volume): reject negative Size, recover poisoned store lock (#11345) 2026-09-16 08:12:11 -07:00
Eliah RusinandGitHub 4fc9ada2ec ci: run seaweed-volume unit tests on Windows (#11349) 2026-09-16 08:06:21 -07:00
a73ba3adbb rust volume: parse vid/fid paths once; the proxy redirect drops the extension like Go (#11341)
handlers.rs split needle URLs in three places and the three disagreed.
Go does it once, in parseURLPath (weed/server/common.go:218-249), and
dispatches on the slash count: /vid/fid/filename takes the extension off
the filename and leaves the fid whole, /vid/fid takes it off the fid, and
the comma form splits the last segment on its last comma and dot.

Two of the Rust copies got that wrong:

- extract_file_id returned the path unchanged when it found no comma, so
  a JWT fid claim, which Go compares against vid + "," + fid for every URL
  form (volume_server_handlers.go:361-364), could never match a slash-form
  request. With a JWT key configured, every read, write or delete of
  /3/01637037d6 was a 401.
- build_proxy_request_info's slash branch had no extension handling, so a
  redirect for /3/01637037d6.jpg sent the client to /3,01637037d6.jpg.
  Go's proxyReqToTargetServer formats "%s/%s,%s" from the already-stripped
  fid (volume_server_handlers_read.go:128-137) and so emits
  /3,01637037d6. The peer still serves either form, since the comma form
  strips the extension again, so this one is parity rather than breakage.

Replace all three with one parse_needle_path returning vid, fid, ext and
filename borrowed from the path. The fid keeps its _delta suffix, as in
Go: parse_needle_id_cookie applies it and the JWT check strips it. The
leading slash stays optional, so chunk manifest fids still parse.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 01:13:06 -07:00
71f8128d75 shell: fs.verify -pruneEntries deletes entries whose needles are lost (#11338)
* shell: fs.verify -pruneEntries deletes entries whose needles are lost

* shell: harden fs.verify -pruneEntries guards; VolumeNeedleStatus returns NotFound for absent needles

* shell: resolve chunk manifests in fs.verify metadata path; require confirmed deletion before counting prunes

* shell: anchor fs.verify legacy missing-needle error matching

* shell: keep fs.verify metadata scan alive on manifest resolution failures

* shell: classify EC missing needles and keep manifest failures unverified

VolumeNeedleStatus now canonicalizes erasure_coding.NotFoundError to
codes.NotFound, so absent needles in EC volumes reach the prune path
through the same stable contract as regular volumes. The client-side
isNeedleMissingError keeps recognizing the legacy wrapped EC shape
("locate in local ec volume: ... needle not found") for mixed-version
clusters.

A chunk manifest that fails to resolve is now an entry-level
verification failure even when the raw top-level chunks are healthy:
the file is not fully readable without the manifest. Raw chunks are
still verified on a resolution failure so a missing top-level manifest
needle is classified and can be pruned. The per-entry logic is
extracted into resolveAndVerify for testability.

* shell: trim fs.verify prune comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 20:42:18 -07:00
github-actions[bot] 01545fc4ff docs: regenerate star history chart 2026-09-16 00:50:47 +00:00
1f037e48f9 s3: a list marker that sorts before the prefix excludes nothing (#11322)
* s3: a list marker that sorts before the prefix excludes nothing

ListObjects `marker` and ListObjectsV2 `start-after` are a plain key
cutoff: list the keys that sort after it. A marker that sorts before the
prefix and is not under it therefore excludes no key carrying the prefix,
and the listing must equal the one with no marker at all.

normalizePrefixMarker treated every marker that does not start with the
prefix as "something wrong" and the listing came back empty. Clients send
this shape routinely: docker/distribution's S3 storage driver walks
prefix "<root>/<path>/" with start-after "<root>" (its rootdirectory), so
on SeaweedFS a registry walk saw an empty bucket. zot read that as "no
repositories": /v2/_catalog was empty, GC/scrub/retention never saw a
repo, and on restart its storage parse deleted every repository's
metadata as "no longer in storage".

listFilerEntries now lists as if no marker were given when the marker
sorts before the prefix; the response still echoes the marker the client
sent. A marker that sorts after the prefix's subtree is left alone: it may
legitimately sit inside a partial-name prefix's match set, which
normalizePrefixMarker already handles, and otherwise correctly lists
nothing.

Reproduce on 4.44 and 4.47:

  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot/zot/"  # all keys
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot"       # KeyCount 0
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=a"         # KeyCount 0

* s3: keep the prefix's own key excluded by a marker that names it

Fold the before-prefix marker rule into normalizePrefixMarker, which now also
derives prefixEndsOnDelimiter from the effective marker instead of each cursor
rebuilding the expression.

A marker equal to the prefix is no longer trimmed to a subtree cutoff:
start-after "a/b/" with prefix "a/b/" excludes only the "a/b/" key, so the walk
starts inside that directory and its children still list.

Adds a listing-level test that walks the whole path for both start-after shapes
a registry sends, and covers the new normalization cases.

* s3: leading slashes do not hide a marker that names the prefix

* s3: echo the V1 marker the client sent, not the walk's cutoff

* s3: filter only the walk's cutoff from the V1 page, not the echoed marker

* s3: skip the key an exclusive marker names as it streams

---------

Co-authored-by: Zuse <be9c90a8-c104-4be2-b7a4-9f92eb833ac8@forge.local>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 16:50:58 -07:00
563c729e70 rust volume: stream the tail scan and release the store lock (#11275)
* rust volume: add a .dat scan plan that runs without the store lock

DatScanPlan captures a fresh .dat handle, the version, the start offset
and an end bound while the caller holds a store guard, then visits one
record at a time with positional reads that never touch the Volume, the
way Go's ScanVolumeFileFrom feeds a scanner.

The handle pins the inode the offset was resolved against: a vacuum
commit renames .cpd over .dat and destroy unlinks it, and neither
rewrites the pinned bytes. The end bound is read while no writer can
hold store.write(), so the scan never meets a partial append. It is a
fresh open, not try_clone, because on Windows read_exact_at uses
seek_read, which moves a cursor a clone shares with the writer.

A header whose size is negative, or does not fit before the end bound,
ends the pass before the body length is computed or anything is
allocated. In today's scan a negative size reaches needle_body_length
and either overflows the buffer size or walks the scan from a wrong
offset. A size near i32::MAX overflows padding_length's i32 arithmetic,
which panics in debug builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: stream the tail scan with the store lock released

volume_tail_sender read every needle from the start offset to EOF into a
Vec while holding store.read(). volume.merge tails from zero, so that was
the whole volume in memory. And because needle writes and the heartbeat
take store.write() on a lock that prefers writers, the whole node stopped
serving until the scan finished: the failure #11235 fixed for EC scrub.

Each pass now runs on a blocking thread. Under one store guard it
resolves the start offset and captures a DatScanPlan, then drops the
guard and sends each needle as it is read, as Go's
VolumeFileScanner4Tailing does. This replaces the one-guard-across-
search-and-scan rule from the previous commit with a stronger invariant:
the offset, the handle and the end bound come from the same guard, and
the handle pins the inode, so a vacuum commit mid-scan cannot point the
offset into the compacted file.

A scan error now ends the stream with Status::internal instead of a
clean EOF, as Go's `streamFollow: %w` does. Once needles stream, a clean
EOF after a partial pass would let volume.move treat a truncated tail as
complete. A panic in the pass is reported the same way. A receiver that
hangs up is also noticed between skipped needles, not only on a send.

Unchanged: the append_at_ns filter, the header on every 2MB chunk, the
caught-up heartbeat without a scan, and the draining countdown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: fail the tail pass on a short read below the snapshot end

DatScanPlan::scan treated an UnexpectedEof on the header or body read as
the end of the data and returned Ok. Every byte below the captured end
existed when the plan was taken, so a short read there can only mean the
inode was truncated under the plan: an unmount followed by a VolumeCopy
of the same volume id reopens .dat with truncate(true). The pass then
reported Scanned, the next pass found the volume gone, and the stream
ended cleanly after a prefix of the planned records, which volume.move
would take as a complete tail.

Both short-read arms now fail the scan with an I/O error that names the
offset and the snapshot end, so tail_pass reports Status::internal as it
does for every other read failure. The break arms were carried over from
scan_raw_needles_from, where the whole scan ran under the store guard and
nothing could truncate the file.

Found by the Devin and Greptile reviews on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: sum the needle padding in i64 so a corrupt size cannot overflow

padding_length added the header, checksum and timestamp widths to the
needle size in i32. A size read from a corrupt header can sit near
i32::MAX, and that sum then overflows: a panic with overflow checks, a
wrapped padding without. DatScanPlan::scan bounds the size against the
bytes left before computing the body length, but that only keeps such a
size out of the arithmetic while under 2 GiB of the file remains, so on a
large volume the scan could still reach the overflow and, in release,
size a buffer from garbage.

Sum in i64 in both version branches. The result is at most
NEEDLE_PADDING_SIZE, so it still fits Size. The scan comment no longer
claims the bound check prevents the overflow.

Found by the CodeRabbit review on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: propagate dat scan parse failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:40:06 -07:00
Chris LuandGitHub 55367afded Update README with performance details of 'weed' binary
Clarified the performance characteristics of the 'weed' binary, emphasizing O(1) complexity for read and write operations.
2026-09-15 14:24:27 -07:00
a9ecfeef45 helm: roll master pods when master config changes (#11331)
* helm: roll master pods when master config changes

The master loads master.toml once at startup (startAdminScripts reads
master.maintenance.scripts and sleep_minutes via viper with no config
watching), and the master ConfigMap is mounted with subPath, which
kubelet never refreshes in a running pod. So a change to
.Values.master.config today updates the ConfigMap but running masters
keep executing the old configuration until something else restarts
them.

Add a checksum/config annotation on the master pod template, following
the existing checksum/s3config pattern on the filer and s3 pods, so a
master config change triggers a rolling restart of the masters.

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Guard against duplicate keys

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Add checksum to deployment as well

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Always ensure the annotation is set

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Update comments

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Soften stance

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* helm: merge pod annotations before checksums

---------

Signed-off-by: Evans Mungai <mbuevans@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:11:58 -07:00
Chris LuandGitHub beaf96a51d s3: cover object lock retention on version deletes (#11335)
* s3: cover WORM guarded version deletes

* s3: trim version delete comments
2026-09-15 13:31:32 -07:00
Chris LuandGitHub 93d4a6aefd s3: drain request body before error response (#11334)
* s3: drain request body before error response

* s3: keep oversized request bodies drainable
2026-09-15 13:09:07 -07:00
166af06a2b rust: cargo fmt both crates, with a commented-out fmt --check CI step (#11329)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-worker: run cargo fmt

Layout only; no token in the workspace changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: run cargo fmt

Layout only; no token in the crate changes. Every earlier Rust PR here
formatted only the blocks it touched so as not to drown its diff in
this one, and this commit is that debt paid in a single place. rustfmt
needed two passes to settle one block in handlers.rs; the committed
form is the fixed point, so `cargo fmt --check` is clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out cargo fmt --check step to both Rust workflows

Same shape as the commented clippy step from #11312: the check is
written out so that making formatting a gate is a one-line uncomment,
and whether to do that stays a maintainer call.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-15 09:29:22 -07:00
517f60e875 rust-volume: fold the 8–15-argument functions into parameter structs (#11328)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 09:04:16 -07:00
Eliah RusinandGitHub 49a680dd64 rust: tonic 0.14 / prost 0.14, drop the result_large_err allow (#11327) 2026-09-15 04:40:46 -07:00
87332eb60b Cloud/remote storage & tiering: configurable multipart upload/download concurrency (#11319)
* pb: add multipart concurrency fields to RemoteConf and tier move requests

RemoteConf gains upload_concurrency/download_concurrency (0 = client
default); VolumeTierMoveDatToRemote/FromRemote requests gain a
concurrency field (0 = backend default).

* remote storage: honor RemoteConf upload/download concurrency in s3 and azure clients

s3 client: ReadFile passes conf download_concurrency to the downloader,
WriteFile uses upload_concurrency for the uploader; previously
hard-coded 1 upload / 5 download parts. 0 keeps defaults. Same for
azure client.

* storage: plumb concurrency through backend interface and tier upload/download

BackendStorage.CopyFile/DownloadFile take a concurrency hint (<=0 =
backend configured default); s3 backend reads
upload_concurrency/download_concurrency from scaffold config with
parseConcurrency fallback, rclone updated to the new signature. Tier
move gRPC handlers forward the request concurrency to the backend.

* shell: -upload_concurrency/-download_concurrency for remote.configure, -concurrent for volume.tier

remote.configure exposes upload/download concurrency persisted into
RemoteConf; volume.tier move/evict commands forward -concurrent to the
tier move requests. Documented in master-cloud.toml scaffold.

* test: cover concurrency propagation in remote tier integration test

* remote.configure: merge existing config on partial update

Load the stored RemoteConf before saving so a partial update (e.g. only
-upload_concurrency) preserves credentials, endpoints, and type instead
of replacing them with new-config defaults. Only treat a confirmed
ErrNotFound as a new configuration; propagate all other load errors so a
transient filer failure does not overwrite stored settings.

On a type transition, reset backend-specific fields to the destination
type's new-config defaults rather than inheriting the old backend's
empty values. Bound configured concurrency to a sane maximum.

* remote storage: honor configured download concurrency in S3 and Azure

ReadFileWithConcurrency now resolves a zero request override against the
client's configured download_concurrency (new downloadConcurrency()
helpers), so the remote-mount/cache read path honors
RemoteConf.DownloadConcurrency instead of the hard-coded default.

Azure also clamps the resolved value to math.MaxUint16 regardless of
whether the fallback was used, preventing uint16 wraparound when a
configured value exceeds 65535.

* shell: rename -concurrent to -concurrency and validate tier transfer bounds

Rename the -concurrent flag to -concurrency across volume.tier.upload,
volume.tier.download, and volume.tier.compact to match the proto field and
RemoteConf field names. Add validateTierConcurrency to reject values that
would wrap int32 or exceed a 1024 cap before constructing the request.

* server: clamp tier move concurrency in gRPC handlers

Add clampTierConcurrency to both VolumeTierMoveDatToRemote and
VolumeTierMoveDatFromRemote handlers so a direct gRPC caller cannot spawn
an unbounded number of network workers.

* trim verbose comments added with concurrency feature

Remove redundant doc comments on the backend interface, rclone backend,
s3_backend parseConcurrency, and test helpers that restated the obvious.

* remote.configure: apply type defaults before re-parse so explicit flags win

applyTypeDefaults ran after the second flag parse, overwriting explicit
destination flags (e.g. -s3.region=eu-west-1) with new-config defaults.
Move the type-transition default reset before the re-parse so user-supplied
flags override the destination defaults.

* remote.configure: only treat explicit -type as a type transition

The first parse defaults -type to s3, so a concurrency-only update on an
existing non-S3 config captured requestedType=s3 and wrongly triggered a
type transition, resetting the stored backend to S3. Use fs.Visit to
detect whether -type was explicitly supplied; an omitted -type keeps the
stored backend.

---------

Co-authored-by: Jack Meredith <9480542+jackusm@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 22:09:08 -07:00
github-actions[bot] e735c12869 docs: regenerate star history chart 2026-09-15 00:48:17 +00:00
e4ca0d09e7 s3: preserve versions for POST policy uploads (#11316)
* s3: preserve versions for POST policy uploads

Route POST policy uploads through the existing version-aware write helpers
and validate promoted Object Lock headers before writing.

Return the generated version ID when versioning is enabled, return
x-amz-version-id: null when versioning is suspended, and omit the header
when versioning has never been enabled.

* s3: reuse versioning helpers in POST policy handler

Route the POST policy handler through the existing getVersioningState
and isObjectLockEnabled helpers instead of open-coding the object-lock
forces-versioning-enabled rule, matching the PUT path.

Drop the x-amz-version-id: null response header for suspended
versioning; the PUT handler omits it and the S3 PutObject sample
response for suspended buckets does not include it. Trim the moved
fileSize comment.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:36:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
01433e801d build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0 (#11306)
* build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0

Bumps [github.com/redis/go-redis/v9](https://github.com/redis/go-redis) from 9.21.0 to 9.22.0.
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/redis/go-redis/compare/v9.21.0...v9.22.0)

---
updated-dependencies:
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.22.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* test(redis_conf): track go-redis 9.22.0 default read timeout of 5s

go-redis 9.22.0 raised the default ReadTimeout from 3s to 5s (part of the
cross-SDK configuration alignment). Update TestUnsetKeepsGoRedisDefaults to
expect the new default so the bump in #11306 stops failing CI.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:16:35 -07:00
1f61097d4d helm: grant List to the generated read-only S3 identity (#11318)
* helm: grant List to the generated read-only S3 identity

The chart's anvReadOnly identity only carried the Read action, so its
credentials could GetObject and HeadObject but every ListObjects request
was denied: List is a separate action and the identity check is an exact
match. Add List so the read-only credentials can list buckets and objects.
Writes stay denied. Update the README example to match.

Bump the chart to 4.47.1. The chart label is part of the s3 and
all-in-one pod templates, so the upgrade rolls the gateways and they
reload the identity config, which is only read at startup.

Fixes #11317

* helm: roll standalone S3 and all-in-one on s3 config changes

Mirror the filer checksum/s3config pod annotation in the standalone S3
and all-in-one deployments so a changed generated S3 secret triggers a
rollout during a normal helm upgrade without relying on a chart version
bump.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:11:22 -07:00
Chris LuandGitHub 0e82b4e351 s3: populate Initiated timestamp in ListMultipartUploads (#11313)
* s3: populate Initiated timestamp in ListMultipartUploads

ListMultipartUploads returned each upload with only Key and UploadId,
omitting the Initiated timestamp. Clients such as GeeseFS rely on this
field to expire stale uploads and crash on its absence. Set Initiated
from the upload directory entry creation time so repeated listings
preserve the original initiation time.

* test/s3: verify Initiated timestamp in ListMultipartUploads

Add an integration test that initiates a multipart upload, lists it,
and asserts the Initiated field is populated and preserved across
repeated listings rather than reflecting the listing time.
2026-09-14 16:03:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0bd048b76f build(deps): bump google.golang.org/api from 0.296.0 to 0.297.0 (#11307)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.296.0 to 0.297.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.296.0...v0.297.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.297.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 15:09:14 -07:00
Chris LuandGitHub c997e54096 admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback (#11314)
* admin: extract isFlagExplicitlySet helper from applyViperFallback

No behavior change; the inline flag-visit check becomes a reusable
helper so the upcoming bind-address default can share it.

* admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback

PR #11185 made the admin HTTP server default to 127.0.0.1 to stop
exposing the unauthenticated admin API on the network by accident.
That also locked out operators who already authenticated with
-adminPassword: their UI became unreachable from the network after
upgrade unless they added -ip=0.0.0.0 (see #11303).

An authenticated deployment is safe to expose, so auto-upgrade the -ip
default to 0.0.0.0 when -adminPassword or [https.admin] mTLS is
configured. The loopback default stays for the unauthenticated case, so
the unauthenticated API is never exposed on the network. An explicit
-ip is always honored.

The worker gRPC control plane has no password auth (only mTLS), so it
must not follow the HTTP upgrade. Give it a separate bind address that
stays on loopback unless -ip is explicit, so adminPassword no longer
re-exposes the unauthenticated worker stream.

* admin: hint loopback-only bind in startup banner

When the admin server binds to loopback (the default for the
unauthenticated case), print a one-line hint that it is not reachable
from other hosts and how to expose it. This helps operators who, after
the #11185 loopback default, can no longer reach the UI from another
machine quickly see the cause and the fix without reading the docs.

* admin: keep worker gRPC on loopback, decouple from https.admin mTLS

The worker gRPC auto-upgrade to 0.0.0.0 was gated on hasMTLS, which
reads the https.admin (HTTP) mTLS config. The worker gRPC mTLS comes
from grpc.admin + grpc.ca, a separate config, so:

- https.admin mTLS without grpc.admin mTLS widened the worker gRPC to
  0.0.0.0 unauthenticated (re-exposing the control plane), and
- grpc.admin mTLS without https.admin mTLS left the worker gRPC on
  loopback, blocking authenticated remote workers.

Drop the worker gRPC auto-upgrade entirely. The worker gRPC keeps the
raw -ip value (loopback by default), matching the pre-existing
behavior; an operator who wants remote workers sets -ip explicitly.
Only the HTTP admin listener auto-upgrades to 0.0.0.0 when
authenticated.

Addresses review feedback on #11314 from Devin and Greptile.
2026-09-14 14:04:35 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2aa6af033d build(deps): bump github.com/go-sql-driver/mysql from 1.10.0 to 1.10.1 (#11308)
Bumps [github.com/go-sql-driver/mysql](https://github.com/go-sql-driver/mysql) from 1.10.0 to 1.10.1.
- [Release notes](https://github.com/go-sql-driver/mysql/releases)
- [Changelog](https://github.com/go-sql-driver/mysql/blob/master/CHANGELOG.md)
- [Commits](https://github.com/go-sql-driver/mysql/compare/v1.10.0...v1.10.1)

---
updated-dependencies:
- dependency-name: github.com/go-sql-driver/mysql
  dependency-version: 1.10.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 14:04:10 -07:00
Chris LuandGitHub 02749c1192 s3api: configurable trusted-proxy allowlist for aws:SourceIp (#11302) (#11315)
* s3api: add TrustedProxies allowlist helper for aws:SourceIp extraction

Introduces a policy_engine.TrustedProxies type that parses a
comma-separated list of bare IPs and CIDRs (mirroring Guard.UpdateWhiteList)
and extracts the client IP for aws:SourceIp condition evaluation.

When the direct TCP peer is in the allowlist, X-Forwarded-For is walked
right-to-left skipping trusted hops (then X-Real-Ip); otherwise the direct
peer address is returned. This is the building block for restoring
configurable forwarded-header trust removed in b88156f (#11231), as
proposed in #11302.

* s3api: honor trusted-proxy allowlist in bucket/IAM policy engine

Make ExtractConditionValuesFromRequest a method on *PolicyEngine so it
can use the engine TrustedProxies when resolving aws:SourceIp. With no
allowlist configured the behavior is unchanged from b88156f: the direct
TCP peer is used and forwarded headers are ignored. When an allowlist is
configured via SetTrustedProxies, requests from a trusted peer honor
X-Forwarded-For (right-to-left) then X-Real-Ip.

Update the two call sites (auth_credentials.go, s3api_bucket_policy_engine.go)
and the engine tests to the method form, and add a regression test for the
trusted-proxy path.

* s3api: honor trusted-proxy allowlist in IAM role/session policies

Make extractRequestContext and extractSourceIP methods on
*S3IAMIntegration so they can use the integration TrustedProxies when
resolving aws:SourceIp. With no allowlist configured the behavior is
unchanged from b88156f: the direct TCP peer is used and forwarded
headers are ignored. When an allowlist is configured via
SetTrustedProxies, requests from a trusted peer honor X-Forwarded-For
(right-to-left) then X-Real-Ip.

Update the call site in isActionExplicitlyDeniedByIAM to type-assert
the integration and use the method, and add a regression test for the
trusted-proxy path.

* s3api: load [s3.trusted_proxies] from security.toml and wire to engines

Read s3.trusted_proxies.white_list (comma-separated IPs/CIDRs) from
security.toml and propagate the allowlist to the bucket policy engine,
the IAM policy engine (persisted across rebuilds via
IdentityAccessManagement.SetTrustedProxies), and the IAM integration.
Reloaded on SIGHUP alongside the JWT signing keys. Document the new
section in the scaffold security.toml.

Closes #11302.

* s3api: harden TrustedProxies parsing and X-Forwarded-For traversal

Canonicalize bare IP entries (via net.ParseIP + String) so non-canonical
IPv6 allowlist entries such as 2001:0db8::1 match peers rendered as
2001:db8::1, and log+skip unparseable bare entries instead of storing
them inertly.

When walking X-Forwarded-For right-to-left, stop at the first malformed
(non-empty, unparseable) entry instead of skipping it, and only fall
back to the leftmost valid IP when the chain was well-formed. This
prevents a malformed hop from masking a forged IP to its left.

Addresses review feedback on #11315.

* s3api: make TrustedProxies reload race-free via atomic.Pointer

Store the trusted-proxy allowlist behind sync/atomic.Pointer in
PolicyEngine and S3IAMIntegration so SIGHUP reloads (which swap the
allowlist) cannot race with concurrent request handlers reading it.
This mirrors the existing Guard guardState pattern. The
IdentityAccessManagement copy is already protected by iam.m.

Addresses review feedback on #11315.
2026-09-14 13:54:26 -07:00
ac03d3fd78 shell: warn when fs.mergeVolumes source holds only orphan needles (#11310)
* shell: warn when fs.mergeVolumes source holds only orphan needles

fs.mergeVolumes traverses filer entries, so a source volume whose
needles are all orphans — filer entries lost to a crashed write or a
wiped filer store — produces only the plan header and exits 0: no move,
no skip, no error. Operators read that as a successful merge while the
real cleanup (volume.fsck) never runs, and dat>idx volumes keep coming
back read-only after restarts.

Count the source-volume needles seen during traversal and, when a plan
source was never seen but its index still reports needles, print a
warning pointing at volume.fsck. Dry-run warns too.

* shell: make needle counting concurrency-safe and count manifest sub-chunks

TraverseBfs runs its callbacks from five workers, so the plain
needlesSeen map raced between source-heavy merges (fatal concurrent
map writes). All increments now funnel through a mutex-guarded
recordSeen closure.

Manifest sub-chunks that live on planned source volumes are now
recorded too — rewriteManifestChunk visits them (including dry-run
and capacity-skipped ones) but previously never marked their source,
which produced false 'orphan needles' warnings for sources whose
chunks were all reached through manifests.

* shell: extract sourceNeedleCounter so the concurrency test covers the production path

The orphan-warning recording was a closure local to Do, so
TestWarnUnreferencedSources_ConcurrentRecording could only exercise a
test-local copy of it — a regression in the production mutex would pass
the test. Lift the map and mutex into a sourceNeedleCounter type with
record/count methods and use it from Do and the test, so the -race test
now drives the actual recording path. Trim the verbose comments added
with the warning while here.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:48 -07:00
adaf3534fa rust: clippy-clean both crates and adopt the std APIs the 1.91 MSRV allows (#11312)
* rust: apply clippy --fix to both crates

The mechanical part of a clippy sweep: `cargo clippy --all-targets --fix`
on seaweed-volume and the seaweed-worker workspace, hand-reviewed. Both
manifests declare their MSRV (1.91.1 and 1.94.1), so every suggestion
clippy applied is within it: the collapsible_if sites become let chains
(1.88, edition 2024), `% n == 0` becomes is_multiple_of (1.87),
chunks_exact with a constant becomes as_chunks (1.88), repeat().take()
becomes repeat_n (1.82), and io::Error::new(Other, ..) becomes
io::Error::other (1.74). The rest is redundant clones, borrows, casts,
closures and field names.

Nothing here changes behaviour. The three let_and_return sites in
needle_map.rs and store_ec.rs deserve a note: the `let result = ..;
result` shape was a deliberate edition-2021 workaround to drop a redb
guard before the table it borrows. Edition 2024 drops tail-expression
temporaries before locals, which is why clippy now flags it, and the
two comments that described the workaround say so instead.

Manual edits on top of the tool output: the blocks clippy rewrote are
re-indented the way rustfmt lays them out (only those blocks — the
crate is not rustfmt-clean and a whole-crate fmt would bury this diff),
the blank lines let_and_return left behind are removed, and the CRC
legacy_value test compares against a literal worked out from the
original shift formula rather than restating rotate_right.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: clear the clippy warnings --fix cannot apply, and say why the rest stay

Hand fixes for the lints clippy only reports. Behaviour is unchanged
throughout; each rewrite is the one clippy names.

- needless_range_loop (7): index loops over shard vectors become
  iterator loops. Where the old code indexed `v[..n]` the new loop
  iterates `v[..n]` so an undersized vector still panics the same way.
- field_reassign_with_default (6): struct literals with `..Default`.
- redundant_pattern_matching (3): `if let Err(_) = guard.check()` becomes
  `.is_err()`, which also releases the read guard at the end of the
  condition instead of at the end of the block.
- manual_strip (2), manual_checked_ops, format_in_format_args,
  redundant_locals, wrong_self_convention (to_vif takes self by value,
  so it is into_vif; CompactEntry is Copy, so to_needle_value takes self).
- type_complexity (2): `OrphanShardLoad` and `RawNeedleEntry` name two
  tuples that were spelled out inline.
- new_without_default: CompactNeedleMap gets a Default that calls new().
- suspicious_open_options: a test helper spells out `.truncate(false)`,
  which is what `.create(true).write(true)` already did.

What stays, and the attribute that says so:

- too_many_arguments (10): `#[expect]` on each function. Folding 8–15
  parameters into a struct is a design change, not a lint fix.
- await_holding_lock / readonly_write_lock: one test holds the store
  write guard across a sleep on purpose, as a barrier that parks the
  copy task at the mount block. `#[expect(.., reason = ..)]` records it.
- module_inception: needle/needle.rs mirrors the Go package layout.

Two lints become crate-wide policy in `[lints.clippy]`, with the reason
next to each: result_large_err, because every RPC path returns
tonic::Status (176 bytes) and boxing it would change every handler
signature; and needless_update, because `..Default::default()` on a
protobuf message literal is what lets a proto gain a field without
touching every constructor (all 11 sites are pb messages). The worker
workspace gets the same table and its members opt in with
`lints.workspace = true`; its generated plugin.rs also allows
large_enum_variant on prost's oneof enums.

Both crates are now clean under `cargo clippy --all-targets -- -D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: use the std APIs the 1.91 MSRV already pays for

The crate declares rust-version 1.91.1, so a few things the code still
worked around are plain std now. All of them come from the 1.85–1.91
release notes; nothing here needs a newer toolchain than the manifest
already requires.

- std::sync::LazyLock (1.80) replaces the lazy_static! block in
  metrics.rs, and the lazy_static dependency goes. Every use site reads
  the same through Deref, so no caller changes.
- Duration::from_mins / from_hours (1.91) replace `from_secs(v * 60)`
  and `from_secs(v * 3600)` in the option parser and the shard-location
  refresh TTLs. One difference for the parser: an absurd count that
  overflows u64 seconds now panics in release builds too, where the
  multiplication used to wrap.
- Result::flatten (1.89) replaces `.and_then(|r| r)` on the replication
  join handle.
- OsStr::display (1.87) replaces `to_string_lossy()` where the name was
  only being formatted; the output is byte-identical.
- `#[allow]` becomes `#[expect]` (1.81) on the suppressions that are
  meant to be permanent, so a suppression that stops being needed
  becomes a warning rather than lingering. Doing that found four that
  already had: dead_code on ChunkManifest, base_name and last_io_error,
  and too_many_arguments on read_from_data_shards, which is down to
  seven parameters. Those attributes are deleted. The three allows that
  depend on cfg (a unix-only mutation, a linux-only field set, a
  profiling-only parameter) stay as allow, because expect would be
  unfulfilled on the other platforms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out clippy step to both Rust workflows

Both crates are warning-free under `cargo clippy --all-targets
-D warnings` now. Whether that becomes a gate is a policy call, so the
step is present but commented out; uncommenting it is the whole change.
The comment points at the `[lints.clippy]` table where crate-wide
exceptions are recorded, so the gate does not become a reason to
sprinkle allows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: guard parse_duration against overflow panics

Duration::from_mins/from_hours panic when the count overflows u64
seconds. Use checked_mul so an oversized CLI value falls back to the
parser default instead of crashing volume startup.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:29 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
49f20489e4 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.20.1 to 1.20.4 (#11305)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.20.1 to 1.20.4.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.20.1...service/mq/v1.20.4)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.20.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:13:14 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bdec508da9 build(deps): bump golang.org/x/image from 0.45.0 to 0.46.0 (#11304)
Bumps [golang.org/x/image](https://github.com/golang/image) from 0.45.0 to 0.46.0.
- [Commits](https://github.com/golang/image/compare/v0.45.0...v0.46.0)

---
updated-dependencies:
- dependency-name: golang.org/x/image
  dependency-version: 0.46.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:45:16 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fd33c07843 build(deps): bump github/codeql-action from 4.37.9 to 4.38.0 (#11311)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.9 to 4.38.0.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.9...v4.38.0)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:41:36 -07:00
Chris LuandGitHub cf38c01978 admin: bind worker gRPC listener to -ip instead of wildcard (#11300)
* admin: bind worker gRPC listener to -ip instead of wildcard

The worker/plugin gRPC control plane called net.Listen("tcp", ":port")
directly, so it wildcard-bound every interface and ignored the -ip setting.
A cluster bound to loopback still exposed the unauthenticated
WorkerService/PluginControlService streams on 0.0.0.0. Bind through
util.JoinHostPort(bindIp, port) so the listener honors -ip like the
master, filer, and volume gRPC listeners.

* admin: warn when worker gRPC is exposed off loopback without mTLS

The worker gRPC stream has no password auth, so grpc.admin mTLS is the
only effective control once the listener leaves loopback. An operator who
sets -adminPassword and binds -ip=0.0.0.0 authenticates the HTTP API but
still exposes the unauthenticated worker control plane. Log a startup
warning naming the port and the mTLS knobs so the exposure is not silent.

* admin: address review on worker gRPC bind fix

- mini: reserve the admin gRPC port with util.JoinHostPort so an IPv6
  bindIp (e.g. ::1) does not form an invalid unbracketed address and
  lose the reservation.
- worker gRPC: track whether grpc.admin mTLS credentials actually loaded
  rather than only whether they were configured, and gate the
  non-loopback exposure warning on that. A cert/key that fails to load
  now still warns instead of silently suppressing.
2026-09-13 21:48:14 -07:00
Chris LuandGitHub f4bad510c9 test/ec: pin rack in seedAndSpread volume.grow to stop silent no-ops (#11299)
seedAndSpread() calls `volume.grow -dataNode X` without pinning the rack.
The master's grow picks the rack by weighted-random when -rack is unset,
and only one of the three racks holds the requested data node, so an
unpinned grow lands on the wrong rack two times out of three. The
VolumeGrow RPC swallows the "No matching data node" failure for
non-cache collections, so those grows count as success without creating
a volume. The per-server cap (maxGrowsPerServer=4) is then exhausted by
silent no-ops before the volumes ever spread, and seedAndSpread times
out with "volumes never spread across >=2 disks on all 3 nodes".

Pin -dataCenter dc1 and -rack rack{i} alongside -dataNode so every grow
reaches the target node. This removes the timing-sensitive assumption
that made TestECVacuumDuplicateShardClaimAcrossDisks flaky.
2026-09-13 21:05:55 -07:00
Chris LuandGitHub 15d9f6c6fe rust-volume: fix Windows build of find_needle_from_ecx (#11298)
* rust-volume: fix Windows build of find_needle_from_ecx

The .ecx binary-search fallback path used on non-Unix targets
(Seek + Read, both &mut self receivers) requires the ecx_file
binding to be mutable. On Unix the read_exact_at path takes &self,
so the mut would be unused there — gate that warning with
#[cfg_attr(unix, allow(unused_mut))].

Without this the build-rust-volume-windows CI job fails with
E0596 at ec_volume.rs:1033, breaking the weed-volume_windows_amd64
release asset.

* rust-volume: use positional seek_read for .ecx lookups on Windows

The previous fix (making ecx_file mut) compiled but left the Windows
fallback using Seek + Read on the shared .ecx file cursor. Concurrent
find_needle_from_ecx calls could interleave seek/read and read the
wrong index entry, corrupting the binary search (raised by Devin and
Greptile review on the PR).

Switch the Windows path to std::os::windows::fs::FileExt::seek_read,
which is positional (offset passed via OVERLAPPED, cursor untouched)
and takes &self — so the binding no longer needs mut, and concurrent
callers on the cached handle can't interfere. Mirrors the existing
read_exact_at helper in storage::volume.

Add a compile_error fallback for non-unix/non-windows targets to
match the convention in storage::volume.
2026-09-13 20:08:01 -07:00
ea179963c0 filer: clean up manifest resolve error propagation and add webdav tes… (#11297)
filer: clean up manifest resolve error propagation and add webdav test (#78)

Drop GitHub issue references from comments and trim verbose comments.
Replace the viewFromChunksOrErr helper with the existing
NonOverlappingVisibleIntervals + ViewFromVisibleIntervals at the stream
call sites, and add a WebDavFile.Read regression test for the manifest
resolution failure path.

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-13 18:34:09 -07:00
358 changed files with 15557 additions and 8205 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.9
uses: github/codeql-action/init@v4.38.0
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.9
uses: github/codeql-action/autobuild@v4.38.0
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.9
uses: github/codeql-action/analyze@v4.38.0
+1 -1
View File
@@ -405,7 +405,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.0
if: always()
with:
sarif_file: trivy-results.sarif
@@ -456,7 +456,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.0
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
+45 -35
View File
@@ -35,11 +35,48 @@ jobs:
cd telemetry/server
go mod tidy
echo "Building telemetry server..."
GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
CGO_ENABLED=0 GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
cd ../..
ls -la telemetry-server
echo "Build completed successfully"
- name: Generate Service Configuration
if: github.event_name == 'workflow_dispatch' && (inputs.setup || inputs.deploy)
env:
REMOTE_USER: ${{ secrets.TELEMETRY_USER }}
run: |
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/bin/sh -c 'exec /home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353 >>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log 2>>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log'
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
- name: First-time Server Setup
if: github.event_name == 'workflow_dispatch' && inputs.setup
env:
@@ -61,40 +98,6 @@ jobs:
touch ~/seaweedfs-telemetry/logs/telemetry.log ~/seaweedfs-telemetry/logs/telemetry.error.log && \
chmod 644 ~/seaweedfs-telemetry/logs/*.log"
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353
Restart=always
RestartSec=5
StandardOutput=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log
StandardError=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
# Copy configuration files
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
@@ -137,11 +140,18 @@ jobs:
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Copy updated service and logrotate files
scp -i ~/.ssh/deploy_key telemetry.service telemetry_logrotate $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Check if service exists and deploy accordingly
ssh -i ~/.ssh/deploy_key $REMOTE_USER@$REMOTE_HOST "
if systemctl list-unit-files telemetry.service >/dev/null 2>&1; then
echo 'Service exists, performing update...'
set -e
sudo systemctl stop telemetry.service
sudo mv ~/seaweedfs-telemetry/telemetry.service /etc/systemd/system/
sudo mv ~/seaweedfs-telemetry/telemetry_logrotate /etc/logrotate.d/seaweedfs-telemetry
sudo systemctl daemon-reload
mkdir -p ~/seaweedfs-telemetry/bin
mv ~/seaweedfs-telemetry/tmp/telemetry-server ~/seaweedfs-telemetry/bin/
chmod +x ~/seaweedfs-telemetry/bin/telemetry-server
@@ -27,6 +27,29 @@ permissions:
jobs:
changes:
name: Detect changed paths
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
outputs:
rust: ${{ steps.filter.outputs.rust }}
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
fetch-depth: 0
- name: Filter changed paths
id: filter
uses: dorny/paths-filter@v3
with:
filters: |
rust:
- 'seaweed-volume/**'
- '.github/workflows/rust-volume-server-tests.yml'
rust-unit-tests:
name: Rust Unit Tests
runs-on: ubuntu-22.04
@@ -59,6 +82,56 @@ jobs:
- name: Build Rust volume server
run: cd seaweed-volume && cargo build --release
# The crate is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[lints.clippy]` in
# seaweed-volume/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-volume && cargo clippy --all-targets -- -D warnings
# The crate is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-volume && cargo fmt --check
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
- name: Run Rust unit tests (redb experimental cursor)
run: cd seaweed-volume && cargo test --features redb-experimental-cursor --lib storage::needle_map
rust-unit-tests-windows:
name: Rust Unit Tests (Windows)
runs-on: windows-latest
timeout-minutes: 30
needs: [changes]
if: needs.changes.outputs.rust == 'true'
defaults:
run:
shell: bash
steps:
- name: Checkout code
uses: actions/checkout@v7
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# No glibc on Windows: key the cache on the toolchain and OS only.
- name: Fingerprint build toolchain
id: toolchain
run: echo "fingerprint=windows-rustc-$(rustc -V | awk '{print $2}')" >> "$GITHUB_OUTPUT"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-windows-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-windows-${{ steps.toolchain.outputs.fingerprint }}-
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
+11
View File
@@ -73,6 +73,17 @@ jobs:
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The workspace is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[workspace.lints.clippy]`
# in seaweed-worker/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-worker && cargo clippy --workspace --all-targets -- -D warnings
# The workspace is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-worker && cargo fmt --all --check
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
+1 -3
View File
@@ -17,7 +17,7 @@ SeaweedFS is a simple and highly scalable distributed file system. There are two
1. to store billions of files!
2. to serve the files fast!
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently.
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently. Both read and write operations have O(1) complexity and can run at the full speed supported by the underlying hardware.
- [Download Binaries for different platforms](https://github.com/seaweedfs/seaweedfs/releases/latest)
- [Wiki Documentation](https://github.com/seaweedfs/seaweedfs/wiki)
@@ -400,8 +400,6 @@ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
The text of this page is available for modification and reuse under the terms of the Creative Commons Attribution-Sharealike 3.0 Unported License and the GNU Free Documentation License (unversioned, with no invariant sections, front-cover texts, or back-cover texts).
[Back to TOC](#table-of-contents)
# Sponsors #
+5 -2
View File
@@ -1,11 +1,14 @@
FROM alpine:latest
# Install required packages
RUN apk add --no-cache \
RUN apk upgrade --no-cache && \
apk add --no-cache \
ca-certificates \
fuse \
curl \
jq
jq \
libcrypto3 \
libssl3
# Copy our locally built binary
COPY weed-local /usr/bin/weed
+2
View File
@@ -30,3 +30,5 @@ sleep_minutes = 17 # sleep minutes between each script execution
bucket = "volume_bucket" # an existing bucket
endpoint = "http://server2:8333"
storage_class = "STANDARD_IA"
# upload_concurrency = 5 # concurrent multipart part uploads per volume (volume.tier.upload -concurrent overrides)
# download_concurrency = 5 # concurrent multipart part downloads per volume (volume.tier.download -concurrent overrides)
+34 -34
View File
@@ -1,6 +1,6 @@
module github.com/seaweedfs/seaweedfs
go 1.26.0
go 1.26.6
require (
cloud.google.com/go v0.123.0 // indirect
@@ -25,7 +25,7 @@ require (
github.com/facebookgo/subset v0.0.0-20200203212716-c811ad88dec4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-redsync/redsync/v4 v4.17.0
github.com/go-sql-driver/mysql v1.10.0
github.com/go-sql-driver/mysql v1.10.1
github.com/go-zookeeper/zk v1.0.4 // indirect
github.com/golang/protobuf v1.5.4
github.com/golang/snappy v1.0.0
@@ -92,16 +92,16 @@ require (
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.56.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.45.0
golang.org/x/image v0.46.0
golang.org/x/net v0.58.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/sys v0.48.0
golang.org/x/text v0.42.0 // indirect
golang.org/x/tools v0.49.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.296.0
google.golang.org/api v0.297.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7
google.golang.org/protobuf v1.36.12
gopkg.in/inf.v0 v0.9.1 // indirect
modernc.org/b v1.0.0 // indirect
@@ -122,9 +122,9 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2 v1.47.0
github.com/aws/aws-sdk-go-v2/config v1.32.35
github.com/aws/aws-sdk-go-v2/credentials v1.20.1
github.com/aws/aws-sdk-go-v2/credentials v1.20.4
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
@@ -146,7 +146,7 @@ require (
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rclone/rclone v1.75.1
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/redis/go-redis/v9 v9.22.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.7
@@ -160,7 +160,7 @@ require (
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
golang.org/x/sync v0.23.0
golang.org/x/tools/godoc v0.1.0-deprecated
google.golang.org/grpc/security/advancedtls v1.0.0
)
@@ -185,7 +185,7 @@ require (
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
@@ -283,19 +283,19 @@ require (
github.com/xeipuuv/gojsonreference v0.0.0-20180127040603-bd5ef7bd5415 // indirect
github.com/xo/terminfo v0.0.0-20220910002029-abceb7e1c41e // indirect
github.com/zeebo/xxh3 v1.1.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/mock v0.5.2 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/mod v0.38.0 // indirect
golang.org/x/mod v0.41.0 // indirect
gonum.org/v1/gonum v0.17.0 // indirect
)
require (
cel.dev/expr v0.25.2 // indirect
cel.dev/expr v0.25.3 // indirect
cloud.google.com/go/auth v0.23.2 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
@@ -310,7 +310,7 @@ require (
github.com/Azure/go-ntlmssp v0.1.1 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 // indirect
github.com/Files-com/files-sdk-go/v3 v3.3.194 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/resourcemapping v0.57.0 // indirect
github.com/IBM/go-sdk-core/v5 v5.23.1 // indirect
@@ -327,20 +327,20 @@ require (
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0
github.com/aws/smithy-go v1.28.1
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
@@ -363,7 +363,7 @@ require (
github.com/elastic/gosigar v0.14.3 // indirect
github.com/emersion/go-message v0.18.2 // indirect
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.37.0 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be // indirect
github.com/envoyproxy/protoc-gen-validate v1.3.3 // indirect
github.com/fatih/color v1.18.0 // indirect
github.com/felixge/httpsnoop v1.1.0 // indirect
@@ -388,7 +388,7 @@ require (
github.com/gogo/protobuf v1.3.2 // indirect
github.com/golang-jwt/jwt/v4 v4.5.2 // indirect
github.com/google/s2a-go v0.1.9 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.20 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.21 // indirect
github.com/gorilla/schema v1.4.1 // indirect
github.com/gorilla/securecookie v1.1.2 // indirect
github.com/gorilla/sessions v1.4.0
@@ -438,7 +438,7 @@ require (
github.com/oracle/oci-go-sdk/v65 v65.121.0 // indirect
github.com/panjf2000/ants/v2 v2.12.1 // indirect
github.com/patrickmn/go-cache v2.1.0+incompatible // indirect
github.com/pelletier/go-toml/v2 v2.4.1 // indirect
github.com/pelletier/go-toml/v2 v2.4.3 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.29
@@ -489,9 +489,9 @@ require (
go.etcd.io/bbolt v1.5.0 // indirect
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 // indirect
go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.45.0 // indirect
@@ -501,7 +501,7 @@ require (
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
+66 -66
View File
@@ -6,8 +6,8 @@ atomicgo.dev/keyboard v0.2.9 h1:tOsIid3nlPLZ3lwgG8KZMp/SFmr7P0ssEN5JUsm78K8=
atomicgo.dev/keyboard v0.2.9/go.mod h1:BC4w9g00XkxH/f1HXhW2sXmJFOCWbKn9xrOunSFtExQ=
atomicgo.dev/schedule v0.1.0 h1:nTthAbhZS5YZmgYbb2+DH8uQIZcTlIrd4eYr3UQxEjs=
atomicgo.dev/schedule v0.1.0/go.mod h1:xeUa3oAkiuHYh8bKiQBRojqAMq3PXXbJujjb0hw8pEU=
cel.dev/expr v0.25.2 h1:K6j46C81hXtZQfuX60cVWQFBJahKSE2gfRbNuvr5bFs=
cel.dev/expr v0.25.2/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cel.dev/expr v0.25.3 h1:A2jO8jwOugrrovveCWfj0KEZOfqiLgAcwjpHPhzIGw0=
cel.dev/expr v0.25.3/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cloud.google.com/go v0.26.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.34.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.38.0/go.mod h1:990N+gfupTy94rShfmMCWGDn0LpTmnzTp2qbd1dvSRU=
@@ -593,8 +593,8 @@ github.com/FilenCloudDienste/filen-sdk-go v0.0.39 h1:tgV5jYL6dsXop9TpDTIQU6UwJjw
github.com/FilenCloudDienste/filen-sdk-go v0.0.39/go.mod h1:0cBhKXQg49XbKZZfk5TCDa3sVLP+xMxZTWL+7KY0XR0=
github.com/Files-com/files-sdk-go/v3 v3.3.194 h1:dtOFxSTWWRpkmvXa6ycNiw8dVDu1wkgzcXyVV1VafNc=
github.com/Files-com/files-sdk-go/v3 v3.3.194/go.mod h1:rl0WumSN9gSo775DgvQv+wMQ8rlb0ES/1hU5jkMtLXg=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 h1:yzIYdwuro811Z27D3T80Wkd3rqZzb0K43nner7Eh1yE=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0/go.mod h1:pJTkW8hEUIIi3Pf65lPZOnn4Y81yCllX6IWk2jNXdkM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 h1:bN1gA3of5bXtbnLsRPrwfmbbe7A5UWFlcTHseujLnpc=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0/go.mod h1:Yj5vHEz/aAepZGliRJsA6uvHAVAQyEwajq9ORCHPxzM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 h1:jLdiS1vO+XJFyDSWRHBx56r4s/NNtcl5J6KyCcWUX/w=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0/go.mod h1:8lmpHY+1VRoteiOwyrQMDt1YGXOrFKCz+1wJW7n3ODY=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/cloudmock v0.57.0 h1:cSjUzZ7KU8hicTgzaSv9NmSyM9fTVK3y5lsBUl3wOis=
@@ -710,48 +710,48 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2 v1.47.0 h1:0jsHallhJCeaU0Ko48c/3FK1ctOQ7NpzggxriJOQ8MQ=
github.com/aws/aws-sdk-go-v2 v1.47.0/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 h1:LAfOuhAH331fmOjTQpAaOlH+Ftn7RzSDJ2VFwjdMMy4=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18/go.mod h1:4e5xhuXHx1e4U9EthvbPP1r/DIMp5c2823OL8karzcM=
github.com/aws/aws-sdk-go-v2/config v1.32.35 h1:UEzXuET8E42lxBPijuACu/tEK7v5lFPlk0Q+GT5WD9E=
github.com/aws/aws-sdk-go-v2/config v1.32.35/go.mod h1:KaMtJpFa2JlL2BStjjHQVwQpzZEmw+ND/EgVrfFoo2g=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1 h1:Z8GRNEx0u9sDkZOq4PUnN8mjGwbUQGRzMSXpvt3d8xQ=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1/go.mod h1:uBIK00kFo95dnemqfFMTWx0X8YRqsh6ecIoCjjOkZqM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 h1:YIEBqcqRnpi4Pfv0YHImtgi6czGCwKHANC7SwmUAVD0=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1/go.mod h1:imEf0oufgAo8KAkCHhrOdqGEC0YWx1PPBQH82shSxGw=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4 h1:hTvrJJseKbvw32kmiE0G+u/9ZqpqscjDrTigHIXP2qs=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4/go.mod h1:gWp9O1ZBWwpcIrgV+mVHk4gZUurAEDkgypu/OXOlIaw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 h1:AM4hHjww+PSFtt6E+UrBrPlZkWsePCLEt9AjkfQX+yM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0/go.mod h1:3x/yXezeQjpOvBb4jEMxrS8SXvpdvJ5abv6l5c1gWM8=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11 h1:eBXB8KZgzQ8A9QB4iJS4aw/u6+4OY3i2hQXPABeAIOg=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11/go.mod h1:N9+5pG27Fy61GUL5YXVLXDTLmUudMrgwsuDbgBMNLxQ=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 h1:pc138gM1CW+XPc60rEwUlwwuwWFQK16CI1T7v1F9Oec=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1/go.mod h1:1+koxpPIbfBdfzP6vojm5/zTpTQ/micYwlxIiNB3TxI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 h1:K0JsbZQj+1h208Ro1zHeA4l7bMp0NvRffHQ91q8Ol1s=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1/go.mod h1:W3/vL6EtCIatICGy9ab29QhMuae+cOKPWcMxv02CO+Q=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 h1:yhw5KD1phVyP9vijxOUzDfEtJx+bt+L63k+VfuiYFAA=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1/go.mod h1:ZW2e0d7DYlRxlS9hEiMXE47gTdX5KRN4byUiNbUpG+Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 h1:Hp/VgjP0BysR3OgLlR057Vz2LcbbVnoWeJ+3qWiS/fY=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3/go.mod h1:nwGV5qw7F1IZPgxCvA/ph8N2TAuz+BkRG/bXn808qMA=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 h1:MUaM4f+kj1ZIBPZfUS8cxP1GKXXZtHJjAthy93AN7SM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3/go.mod h1:6YmVmEVRI5ZZzRjCSsb9SryKH0hAlMRdgA7kG9aDvBU=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 h1:fuSCw4Z2qfRCztMPO3GXJNSiEp6Wee+WOLwrHHUMy9c=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3/go.mod h1:6SxcHheD1pPR5+kWm1wGvjlL/YqUsh267sAfEmN4K7A=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 h1:uZOinZb+h7lZw8IYzP1z1IuEnueB76/EFkcf/fEW4Ag=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31/go.mod h1:NRtwAM/p5VRt03TlEUs0pH3TeWamWdf4YyJpSrzPYLc=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 h1:RmmWQPREQdk9U+PfqeHW3MqZaBaNK7TpV9W3RY+b+7g=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1/go.mod h1:0A3W4F+68ZnNk5XcNL/e9HFMwnP8RlEicFfy6eOEDyw=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 h1:bON1rJf67TSTDCKg816AAIE4xSTtoo9tl0XRkO72R+I=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3/go.mod h1:c5BBpjJcQXpfeq9iASyVKA3T6vX6B6LEXY4mL/gklDY=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 h1:HLPAVrlLDaN2boN0xJx7MgaQDNEO3Q+c9L6kl/8m47Q=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39/go.mod h1:Pg/dVfsNkm1hsIDK/gMvCKtmyNfNTV12mrgHqVE/6Oo=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3 h1:IKoCZqfWfZzSBi16QFQ+QcbQ3LRQ7QgB1S5tDAyPBQQ=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3/go.mod h1:RBpRcXiM4s2pOInVs32GsBonnje+fiAj4mcrStRmlCA=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 h1:mdMtSVKdQ3+mzBh+l0ogrFYZVQUCg6pJZOirA2ARsYE=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1/go.mod h1:9IqUlsJDbUPcg6cgx3WEzXdjrbWzLDQrak0aaSqlTcI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 h1:ZD5qFpWcaOKdTuhBi431pIDkCgrMkMlMT6jlpSPoIRI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0/go.mod h1:8Nuuf+tR346PjJ3MvZPh9pekbLiLQFWJhzMXfwy7alA=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 h1:B6WFn91tobD6gG4724ONHaqrpKsoETGnv98LHe/yIGM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1/go.mod h1:tWuiVBUtPBr8/rgRiYS8Uf85sHcAN+G7XS3D3CEoUh8=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 h1:6yeYCWFvgbI2TI3K6jr9LtBNhXgJ7g4xqD+DEiaDDmM=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1/go.mod h1:naFe83jSMuYkH+QjQPX8n1MLhBkeCFM5Lsnh5m5wz3c=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1 h1:Sv2xPnRHlThSUtVujYuUBPI/Il8si6UPHXL8DMiB/F0=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1/go.mod h1:mKo/CzaCz8qytGW70NG4vIIGAx1HXTlb5lHNkC5k3lk=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 h1:JGeeBcMlhg1xtOXYpeCaTQBZObtXMPQCUqBcmr65NRA=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0/go.mod h1:XwteswG9EOMRFm73UT0t+MbTwyLxMrEXkU6e+v92Lzo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 h1:obhahQXDEdVEv8y5bTKXR30LVaxYe1kyYM0L7l2Iq+k=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0/go.mod h1:6twZZ/aXHNy1vXUO8koUbp++MYzMASkOgEBdkbJYmO0=
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0 h1:khXV3+K5D3f4e8xtplaRdSFn1bEg3gj5EBHQvbCOZbQ=
github.com/aws/aws-sdk-go-v2/service/sts v1.50.0/go.mod h1:/8JRcdTt//hG0Q4BTmGbuOplT7ABe+5rdtqUHqXvYIM=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
@@ -987,8 +987,8 @@ github.com/envoyproxy/go-control-plane v0.10.3/go.mod h1:fJJn/j26vwOu972OllsvAgJ
github.com/envoyproxy/go-control-plane v0.11.0/go.mod h1:VnHyVMpzcLvCFt9yUz1UnCwHLhwx1WguiVDV7pTG/tI=
github.com/envoyproxy/go-control-plane v0.14.0 h1:hbG2kr4RuFj222B6+7T83thSPqLjwBIfQawTkC++2HA=
github.com/envoyproxy/go-control-plane v0.14.0/go.mod h1:NcS5X47pLl/hfqxU70yPwL9ZMkUlwlKxtAohpi2wBEU=
github.com/envoyproxy/go-control-plane/envoy v1.37.0 h1:u3riX6BoYRfF4Dr7dwSOroNfdSbEPe9Yyl09/B6wBrQ=
github.com/envoyproxy/go-control-plane/envoy v1.37.0/go.mod h1:DReE9MMrmecPy+YvQOAOHNYMALuowAnbjjEMkkWOi6A=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be h1:SWe0x6yfglnxuvOiYgTTnNq7QD/yvqthOh1RBI8Bj8w=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be/go.mod h1:PYEOlng9XcrulfyWpm49jECTPV0LT4q8cO7fLW/xwgk=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0 h1:/G9QYbddjL25KvtKTv3an9lx6VBE2cnb8wp1vEGNYGI=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0/go.mod h1:Wk+tMFAFbCXaJPzVVHnPgRKdUdwW/KdbRt94AzgRee4=
github.com/envoyproxy/protoc-gen-validate v0.1.0/go.mod h1:iSmxcyjqTsJpI2R4NaDN7+kN2VEUnK/pcBlmesArF7c=
@@ -1109,8 +1109,8 @@ github.com/go-redsync/redsync/v4 v4.17.0 h1:FFJ+uxZs44y4Sq10//IFKic9T94AYl+u3Sog
github.com/go-redsync/redsync/v4 v4.17.0/go.mod h1:CKVA6qwT07S/916i+Yd9h1/8YFQhCCpPYTQhvvYytJo=
github.com/go-resty/resty/v2 v2.17.2 h1:FQW5oHYcIlkCNrMD2lloGScxcHJ0gkjshV3qcQAyHQk=
github.com/go-resty/resty/v2 v2.17.2/go.mod h1:kCKZ3wWmwJaNc7S29BRtUhJwy7iqmn+2mLtQrOyQlVA=
github.com/go-sql-driver/mysql v1.10.0 h1:Q+1LV8DkHJvSYAdR83XzuhDaTykuDx0l6fkXxoWCWfw=
github.com/go-sql-driver/mysql v1.10.0/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-stack/stack v1.8.0/go.mod h1:v0f6uXyyMGvRgIKkXu+yp6POWl0qKG85gN/melR3HDY=
github.com/go-task/slim-sprig v0.0.0-20230315185526-52ccab3ef572 h1:tfuBGBXKqDEevZMzYi5KSi8KkcZtzBcTgAUUtapy0OI=
github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI=
@@ -1262,8 +1262,8 @@ github.com/googleapis/enterprise-certificate-proxy v0.1.0/go.mod h1:17drOmN3MwGY
github.com/googleapis/enterprise-certificate-proxy v0.2.0/go.mod h1:8C0jb7/mgJe/9KK8Lm7X9ctZC2t60YyIpYEI16jx0Qg=
github.com/googleapis/enterprise-certificate-proxy v0.2.1/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.2.3/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.3.20 h1:t/xL64VUoN69MuMRQuJETqYGOw4Z9mSRJK9epIEtwFk=
github.com/googleapis/enterprise-certificate-proxy v0.3.20/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/enterprise-certificate-proxy v0.3.21 h1:OFdQ3tnCX/zaQ0Cedur3D3z7kI6HiLX9g3TiAN4/DFU=
github.com/googleapis/enterprise-certificate-proxy v0.3.21/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/gax-go/v2 v2.1.0/go.mod h1:Q3nei7sK6ybPYH7twZdmQpAd1MKb7pfu6SK+H1/DsU0=
@@ -1652,8 +1652,8 @@ github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144T
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
github.com/patrickmn/go-cache v2.1.0+incompatible/go.mod h1:3Qf8kWWT7OJRJbdiICTKqZju1ZixQ/KpMGzzAfe6+WQ=
github.com/pborman/getopt v0.0.0-20170112200414-7148bc3a4c30/go.mod h1:85jBQOZwpVEaDAr341tbn15RS4fCAsIst0qp7i8ex1o=
github.com/pelletier/go-toml/v2 v2.4.1 h1:j5OMOImsH+j2k7GJ5YO+RxfWwohNiH6t5zB/+h3bagc=
github.com/pelletier/go-toml/v2 v2.4.1/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pelletier/go-toml/v2 v2.4.3 h1:GTRvJQutkOSftxIFD5xw9aepkYNuPWmVJpffdDPYVpY=
github.com/pelletier/go-toml/v2 v2.4.3/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 h1:XeOYlK9W1uCmhjJSsY78Mcuh7MVkNjTzmHx1yBzizSU=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14/go.mod h1:jVblp62SafmidSkvWrXyxAme3gaTfEtWwRPGz5cpvHg=
github.com/peterh/liner v1.2.2 h1:aJ4AOodmL+JxOZZEL2u9iJf8omNRpqHc/EbrK+3mAXw=
@@ -1770,8 +1770,8 @@ github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 h1:N/ElC8H3+5X
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475/go.mod h1:bCqnVzQkZxMG4s8nGwiZ5l3QUCyqpo9Y+/ZMZ9VjZe4=
github.com/rdleal/intervalst v1.5.0 h1:SEB9bCFz5IqD1yhfH1Wv8IBnY/JQxDplwkxHjT6hamU=
github.com/rdleal/intervalst v1.5.0/go.mod h1:xO89Z6BC+LQDH+IPQQw/OESt5UADgFD41tYMUINGpxQ=
github.com/redis/go-redis/v9 v9.21.0 h1:FPBE4hhbAke+TLmcY3WkpbDffJEomdqPn3HYiqAtL9E=
github.com/redis/go-redis/v9 v9.21.0/go.mod h1:v/M13XI1PVCDcm01VtPFOADfZtHf8YW3baQf57KlIkA=
github.com/redis/go-redis/v9 v9.22.0 h1:laDvpYXTJtZLloinw1fA5Kqd6HAEH2XKxOkG/PDq2F0=
github.com/redis/go-redis/v9 v9.22.0/go.mod h1:y2g0Wj8rQvuK0ELM+oxSudcLtC09JScs98I/X9gRWY4=
github.com/redis/rueidis v1.0.76 h1:RdDWuvlYBSp+bTrBvaXqJnNEL3VVzsnjo+0psPFgLc4=
github.com/redis/rueidis v1.0.76/go.mod h1:UsfHPSbomB6QAVMk4iiFkzRy0nh9o7scDGa+SitvBY4=
github.com/redis/rueidis/rueidiscompat v1.0.76 h1:7LikbiqCQqCsZXeZ+akgZMnjIV/J0VHih9PIX4gGZC4=
@@ -2103,30 +2103,30 @@ go.opencensus.io v0.24.0 h1:y73uSU6J157QMP2kn2r30vwW1A2W2WFwSCGnAVxeaD0=
go.opencensus.io v0.24.0/go.mod h1:vNK8G9p7aAivkbmorf4v+7Hgx+Zs0yY+0fOtgBfjQKo=
go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64=
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 h1:NmLfL734pJhM0JKaYd2Y28+nY9dPRWYAAbxhRCrKXPw=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0/go.mod h1:tNAsgd8avTGke1+MndXlU5Cru4PQ9Ai/cCNWQv/ZJ/s=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 h1:9jR0ZPRok9ryaOQ2Wx8rg5F7Aon59mxrqbVI60/vlBk=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0/go.mod h1:VSme3o2fvSg5bVg0dRzyHaj4Z5EVhG+g2Fde6LKzmQA=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 h1:2yEATaop1/a1I4psnSLgWVPLWwCzkqWakgJy7xTDVy0=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0/go.mod h1:D7J12YRapIekYyPWgGPlA/23pRmpSEZC5xJC/TTLI9U=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0 h1:MCcYL7J6Vt/X0kjqbMZkekCmwsurbQRbL69vkiye2lk=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0/go.mod h1:3jnStNwSufK+f5ktjL4EPcwtig4rtd81NS70lqHuXl8=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 h1:LMuyCAyfalSjDyjdC65nK6N0zoTT63+E/u95X0JovZI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0/go.mod h1:085m8qbm4hgc8rZWGDEa4vmyyo2c3nPxUslYUKUIU04=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 h1:SUplec5dp06reu1zaXmOXdvqH398taqrDXqUl99jxSc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0/go.mod h1:ho2g4N+ane+swq5I/VBkKWnRDY4kUINH3FuqyZqX/Ug=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 h1:lgh3PiVrRUWMLOVSkQicxzZll5NjF1r+AtsX1XRIHw0=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0/go.mod h1:5Cnhth3m/AgOeTgE3ex12pPmiu/gGtZit03kSzx9X7s=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 h1:hqxVTu/GtBF+vJ8d1fzW7fRxZFvgoDjWcxwwCaFDYpU=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0/go.mod h1:z5fVEF4X5v0ESvlJqBrrFlBVoj5EQuefZpzsu7R+x5Q=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 h1:s0n95ya5tOG03exJ5JySOdJFtwGo4ZQ+KeY7Zro4CLI=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0/go.mod h1:m9wRxtKA2MZ1HcnNC4BKI+9aYe434qRZTCvI7QGUN7Y=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 h1:KN3btaILMTxR4QDHVGAO87lq5ButzK7l+kIfLuxQ1oA=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0/go.mod h1:yNcodmUclM4InyWoOwX/YW4Jri0Gj5FWAlM+NqCrtqY=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/metric/x v0.67.0 h1:PcicCNZFkZ4bXfSooXdo3WN7RBOVOtjVdo1wD358Uns=
@@ -2140,8 +2140,8 @@ go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbi
go.opentelemetry.io/proto/otlp v0.7.0/go.mod h1:PqfVotwruBrMGOCsRd/89rSnXhoiJIqeYNgFYFoEGnI=
go.opentelemetry.io/proto/otlp v0.15.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v0.19.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk=
go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/atomic v1.6.0/go.mod h1:sABNBOSYdrvTF6hTgEIbc7YasKWGhgEQZyfxyTvoXHQ=
go.uber.org/atomic v1.7.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
go.uber.org/atomic v1.9.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
@@ -2225,8 +2225,8 @@ golang.org/x/image v0.0.0-20210607152325-775e3b0c77b9/go.mod h1:023OzeP/+EPmXeap
golang.org/x/image v0.0.0-20210628002857-a66eb6448b8d/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20211028202545-6944b10bf410/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20220302094943-723b81ca9867/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.45.0 h1:FMb1nTbH5H9vF55SriQHgFw5GnNL9Jg6L25BwXKzhB0=
golang.org/x/image v0.45.0/go.mod h1:n62x/7RqlwXDvGsSU4u6IUTUf6KghUZ9Bt7cG/T9Fx4=
golang.org/x/image v0.46.0 h1:b1+oYj0Jbp6K5MDT4i4/eZpYlk3V8SJhhDKh6LBHAyQ=
golang.org/x/image v0.46.0/go.mod h1:3B3W05VGVQyuXucLINLjXKrqISASfi4Xj+iCVkLMwew=
golang.org/x/lint v0.0.0-20181026193005-c67002cb31c3/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
golang.org/x/lint v0.0.0-20190227174305-5b3e6a55c961/go.mod h1:wehouNa3lNwaWXcvxsM5YxQ5yQlVC4a0KAMCusXpPoU=
golang.org/x/lint v0.0.0-20190301231843-5614ed5bae6f/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
@@ -2258,8 +2258,8 @@ golang.org/x/mod v0.8.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.9.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.12.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.13.0/go.mod h1:hTbmBsO62+eylJbnUtE2MGJUyE7QWk4xUqPFrRgJ+7c=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c=
golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o=
golang.org/x/net v0.0.0-20180724234803-3673e40ba225/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180826012351-8a410e7b638d/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
@@ -2374,8 +2374,8 @@ golang.org/x/sync v0.0.0-20220929204114-8fcdb60fdcc0/go.mod h1:RxMgew5VJxzue5/jJ
golang.org/x/sync v0.1.0/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.3.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.4.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.0.0-20180810173357-98c5dad5d1a0/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180830151530-49385e6e1522/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180905080454-ebe1bf3edb33/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
@@ -2477,8 +2477,8 @@ golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.8.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.12.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.13.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210220032956-6a3ed077a48d/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210615171337-6886f2dfbf5b/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
@@ -2511,8 +2511,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI=
golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2589,8 +2589,8 @@ golang.org/x/tools v0.6.0/go.mod h1:Xwgl3UAJ/d3gWutnCtw505GrjyAbvKui8lOU390QaIU=
golang.org/x/tools v0.7.0/go.mod h1:4pg6aUX35JBAogB10C9AtvVL+qowtN4pT3CGSQex14s=
golang.org/x/tools v0.13.0/go.mod h1:HvlwmtVNQAhOuCjW7xxvovg8wbNq7LwfXh/k7wXUl58=
golang.org/x/tools v0.14.0/go.mod h1:uYBEerGOWcJyEORxN+Ek8+TT266gXkNlHdJBwexUsBg=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
golang.org/x/tools/godoc v0.1.0-deprecated h1:o+aZ1BOj6Hsx/GBdJO/s815sqftjSnrZZwyYTHODvtk=
golang.org/x/tools/godoc v0.1.0-deprecated/go.mod h1:qM63CriJ961IHWmnWa9CjZnBndniPt4a3CK0PVB9bIg=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -2668,8 +2668,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.296.0 h1:Nn5EHeKdGx70MFClaV/II0gsWUm6xhEjb0xYLylVvaA=
google.golang.org/api v0.296.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/api v0.297.0 h1:WktxTsnnx0yZNnsR6j0q6hR21RnnK81FHTOPy/ux4OE=
google.golang.org/api v0.297.0/go.mod h1:S4m8x0M6OkQpkOzGk1y9JG2sm4fFQrMh6dxzjCTszhE=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2805,8 +2805,8 @@ google.golang.org/genproto v0.0.0-20230222225845-10f96fb3dbec/go.mod h1:3Dl5ZL0q
google.golang.org/genproto v0.0.0-20230306155012-7f2fa6fef1f4/go.mod h1:NWraEVixdDnqcqQ30jipen1STv2r/n24Wb7twVTGR4s=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d h1:C9v1o0/4quuhOAfmRXA2j+we0PqZIp8traLdeogF3Ms=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d/go.mod h1:Wz2wFJntZFmLGo7pLDXZ3wYk5hyc0Mb+SkHhDDXT+lU=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d h1:QwnJwPte4XXAkhPu26LTDIahnsMSUV0kK8HkxbC+Pc4=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d/go.mod h1:WRrQ7/7N19PypuT0fxLOL5Lq0waoiRri4FbtHDEKrGE=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 h1:lrupDmKL3p5kEX1M92oan027eCKcouzjuPbH6YBK+Rs=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1/go.mod h1:q/3oV3jAi5vwelxsVAprMBC8BcM2zmNe+IjRGd+9/ks=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 h1:cYNAzI2sUwhmCcoj9TxvihSrqsxt6uIkj3rDRhSDmW4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
@@ -2849,8 +2849,8 @@ google.golang.org/grpc v1.51.0/go.mod h1:wgNDFcnuBGmxLKI/qn4T+m5BtEBYXJPvibbUPsA
google.golang.org/grpc v1.52.0/go.mod h1:pu6fVzoFb+NBYNAvQL08ic+lvB2IojljRYuun5vorUY=
google.golang.org/grpc v1.53.0/go.mod h1:OnIrk0ipVdj4N5d9IUoFUx72/VlD7+jUsHwZgwSMQpw=
google.golang.org/grpc v1.55.0/go.mod h1:iYEXKGkEBhg1PjZQvoYEVPTDkHo1/bjTnfwTeGONTY8=
google.golang.org/grpc v1.85.0-dev h1:HxkDyKIIZPpFnroC56tQv5gNuKTmVvi0t7TzOf5zt7g=
google.golang.org/grpc v1.85.0-dev/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7 h1:5+EEM1fC0yjOZID0NUZVrE2+8M/+1TclNrSz/l1xMYs=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7/go.mod h1:Ovl0ECo4xx5r4kn/6d4BPSNB7OIFuu6EAjOzjtVAKaM=
google.golang.org/grpc/cmd/protoc-gen-go-grpc v1.1.0/go.mod h1:6Kw0yEErY5E/yWrBtf03jp27GLLJujG4z/JK95pnjjw=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6 h1:ExN12ndbJ608cboPYflpTny6mXSzPrDLh0iTaVrRrds=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6/go.mod h1:6ytKWczdvnpnO+m+JiG9NjEDzR1FJfsnmJdG7B8QVZ8=
+1 -1
View File
@@ -3,4 +3,4 @@ description: SeaweedFS
name: seaweedfs
appVersion: "4.47"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.47.0
version: 4.47.1
+1 -1
View File
@@ -286,7 +286,7 @@ metadata:
app.kubernetes.io/component: s3
stringData:
# this key must be an inline json config file
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read","List"]}]}'
```
#### Source S3 credentials from an existing Secret
@@ -37,13 +37,17 @@ spec:
{{- with .Values.allInOne.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.allInOne.podAnnotations | default dict) }}
{{- $existingS3ConfigSecret := or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret }}
{{- if $existingS3ConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace $existingS3ConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.allInOne.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.allInOne.restartPolicy }}
{{- if .Values.allInOne.affinity }}
@@ -43,19 +43,15 @@ spec:
{{- with .Values.filer.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.filer.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.filer.podAnnotations | default dict) }}
{{- if .Values.filer.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
checksum/s3config: {{ $configSecret | toYaml | sha256sum }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
checksum/s3config: {{ include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.filer.restartPolicy }}
{{- if .Values.filer.affinity }}
@@ -43,13 +43,10 @@ spec:
{{- with .Values.master.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.master.podAnnotations | default dict) }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.master.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.master.restartPolicy }}
{{- if .Values.master.affinity }}
@@ -35,13 +35,15 @@ spec:
{{- with .Values.s3.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.s3.podAnnotations | default dict) }}
{{- if .Values.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.s3.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.s3.restartPolicy }}
{{- if .Values.s3.affinity }}
@@ -60,7 +60,7 @@ stringData:
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read","List"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
{{ toJson .Values.filer.s3.auditLogConfig | nindent 4 }}
+11
View File
@@ -190,6 +190,8 @@ master:
podLabels: {}
# Annotations to be added to the master pods
# The chart sets checksum/master-config on master pods; other checksum/* keys
# can be used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the master resources
@@ -773,6 +775,8 @@ filer:
podLabels: {}
# Annotations to be added to the filer pods
# The chart sets checksum/s3config on filer pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the filer resource
@@ -1078,6 +1082,8 @@ s3:
podLabels: {}
# Annotations to be added to the s3 pods
# The chart sets checksum/s3config on s3 pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the s3 resources
@@ -1788,6 +1794,8 @@ allInOne:
initContainers: "" # Init containers
sidecars: "" # Sidecar containers
annotations: {} # Annotations for the deployment
# The chart sets checksum/master-config and checksum/s3config on all-in-one
# pods; other checksum/* keys can be used for custom rollouts.
podAnnotations: {} # Annotations for the pods
podLabels: {} # Labels for the pods
@@ -1922,6 +1930,9 @@ certificates:
# Labels to be added to all the created pods
podLabels: {}
# Annotations to be added to all the created pods
# The chart sets checksum/master-config and checksum/s3config on pods whose
# rendered ConfigMaps or Secrets should trigger rollouts. Other checksum/* keys
# can be used for custom rollout annotations.
podAnnotations: {}
networkPolicy:
+1026 -1055
View File
File diff suppressed because it is too large Load Diff

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 53 KiB

+112 -127
View File
@@ -503,7 +503,7 @@ dependencies = [
"rustls-pki-types",
"tokio",
"tokio-rustls",
"tower 0.5.3",
"tower",
"tracing",
]
@@ -628,13 +628,13 @@ dependencies = [
[[package]]
name = "axum"
version = "0.7.9"
version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "edca88bc138befd0323b20752846e6587272d3b03b0343c8ea28a6f819e6e71f"
checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90"
dependencies = [
"async-trait",
"axum-core",
"bytes",
"form_urlencoded",
"futures-util",
"http 1.4.0",
"http-body 1.0.1",
@@ -648,14 +648,13 @@ dependencies = [
"multer",
"percent-encoding",
"pin-project-lite",
"rustversion",
"serde",
"serde_core",
"serde_json",
"serde_path_to_error",
"serde_urlencoded",
"sync_wrapper",
"tokio",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -663,19 +662,17 @@ dependencies = [
[[package]]
name = "axum-core"
version = "0.4.5"
version = "0.5.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09f2bd6146b97ae3359fa0cc6d6b376d9539582c7b4220f041a33ec24c226199"
checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1"
dependencies = [
"async-trait",
"bytes",
"futures-util",
"futures-core",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"mime",
"pin-project-lite",
"rustversion",
"sync_wrapper",
"tower-layer",
"tower-service",
@@ -1654,19 +1651,13 @@ dependencies = [
"futures-core",
"futures-sink",
"http 1.4.0",
"indexmap 2.13.1",
"indexmap",
"slab",
"tokio",
"tokio-util",
"tracing",
]
[[package]]
name = "hashbrown"
version = "0.12.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a9ee70c43aaf417c914396645a0fa852624801b24ebb7ae78fe8272889ac888"
[[package]]
name = "hashbrown"
version = "0.14.5"
@@ -1860,7 +1851,7 @@ dependencies = [
"libc",
"percent-encoding",
"pin-project-lite",
"socket2 0.6.3",
"socket2",
"tokio",
"tower-service",
"tracing",
@@ -2027,16 +2018,6 @@ dependencies = [
"quick-error",
]
[[package]]
name = "indexmap"
version = "1.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bd070e393353796e801d209ad339e89596eb4c8d430d18ede6a1cced8fafbd99"
dependencies = [
"autocfg",
"hashbrown 0.12.3",
]
[[package]]
name = "indexmap"
version = "2.13.1"
@@ -2250,9 +2231,9 @@ dependencies = [
[[package]]
name = "matchit"
version = "0.7.3"
version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0e7465ac9959cc2b1404e8e2367b43684a6d13790fe23056cc8c6c5a6b7bcb94"
checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3"
[[package]]
name = "md-5"
@@ -2619,17 +2600,18 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4c5cc86750666a3ed20bdaf5ca2a0344f9c67674cae0515bec2da16fbaa47db"
dependencies = [
"fixedbitset 0.4.2",
"indexmap 2.13.1",
"indexmap",
]
[[package]]
name = "petgraph"
version = "0.7.1"
version = "0.8.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3672b37090dbd86368a4145bc067582552b29c27377cad4e0a306c97f9bd7772"
checksum = "8701b58ea97060d5e5b155d383a69952a60943f0e6dfe30b04c287beb0b27455"
dependencies = [
"fixedbitset 0.5.7",
"indexmap 2.13.1",
"hashbrown 0.15.5",
"indexmap",
]
[[package]]
@@ -2836,12 +2818,12 @@ dependencies = [
[[package]]
name = "prost"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2796faa41db3ec313a31f7624d9286acf277b52de526150b7e69f3debf891ee5"
checksum = "528ac67416ff8646872a3c02cad9cc4ee5dc9f9540c9b10771855c95cb2e5ae1"
dependencies = [
"bytes",
"prost-derive 0.13.5",
"prost-derive 0.14.4",
]
[[package]]
@@ -2867,19 +2849,20 @@ dependencies = [
[[package]]
name = "prost-build"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042"
dependencies = [
"heck",
"itertools 0.14.0",
"log",
"multimap",
"once_cell",
"petgraph 0.7.1",
"petgraph 0.8.3",
"prettyplease",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"pulldown-cmark",
"pulldown-cmark-to-cmark",
"regex",
"syn",
"tempfile",
@@ -2900,9 +2883,9 @@ dependencies = [
[[package]]
name = "prost-derive"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf"
dependencies = [
"anyhow",
"itertools 0.14.0",
@@ -2922,11 +2905,11 @@ dependencies = [
[[package]]
name = "prost-types"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "52c2c1bf36ddb1a1c396b3601a3cec27c2462e45f07c386894ec3ccf5332bd16"
checksum = "f94967dc7688f3054c7fac87473ffae4cc4c3904800e2d9f5b857246d8963b0a"
dependencies = [
"prost 0.13.5",
"prost 0.14.4",
]
[[package]]
@@ -2993,6 +2976,26 @@ version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.11.0",
"memchr",
"unicase",
]
[[package]]
name = "pulldown-cmark-to-cmark"
version = "22.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab1ad36992cead65f02aa399a373a42730922f1525d988172634fdefdecb8a60"
dependencies = [
"pulldown-cmark",
]
[[package]]
name = "pxfm"
version = "0.1.28"
@@ -3018,7 +3021,7 @@ dependencies = [
"quinn-udp",
"rustc-hash",
"rustls",
"socket2 0.6.3",
"socket2",
"thiserror 2.0.18",
"tokio",
"tracing",
@@ -3056,7 +3059,7 @@ dependencies = [
"cfg_aliases",
"libc",
"once_cell",
"socket2 0.6.3",
"socket2",
"tracing",
"windows-sys 0.60.2",
]
@@ -3262,8 +3265,8 @@ dependencies = [
"tokio",
"tokio-rustls",
"tokio-util",
"tower 0.5.3",
"tower-http 0.6.8",
"tower",
"tower-http",
"tower-service",
"url",
"wasm-bindgen",
@@ -3719,16 +3722,6 @@ version = "1.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1b6b67fb9a61334225b5b790716f609cd58395f895b3fe8b328786812a40bc3b"
[[package]]
name = "socket2"
version = "0.5.10"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e22376abed350d73dd1cd119b57ffccad95b4e585a7cda43e286245ce23c0678"
dependencies = [
"libc",
"windows-sys 0.52.0",
]
[[package]]
name = "socket2"
version = "0.6.3"
@@ -3990,7 +3983,7 @@ dependencies = [
"parking_lot 0.12.5",
"pin-project-lite",
"signal-hook-registry",
"socket2 0.6.3",
"socket2",
"tokio-macros",
"windows-sys 0.61.2",
]
@@ -4077,7 +4070,7 @@ version = "0.22.27"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
dependencies = [
"indexmap 2.13.1",
"indexmap",
"serde",
"serde_spanned",
"toml_datetime",
@@ -4093,11 +4086,10 @@ checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
[[package]]
name = "tonic"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "877c5b330756d856ffcc4553ab34a5684481ade925ecc54bcd1bf02b1d0d4d52"
checksum = "ac2a5518c70fa84342385732db33fb3f44bc4cc748936eb5833d2df34d6445ef"
dependencies = [
"async-stream",
"async-trait",
"axum",
"base64",
@@ -4111,13 +4103,12 @@ dependencies = [
"hyper-util",
"percent-encoding",
"pin-project",
"prost 0.13.5",
"rustls-pemfile",
"socket2 0.5.10",
"socket2",
"sync_wrapper",
"tokio",
"tokio-rustls",
"tokio-stream",
"tower 0.4.13",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -4125,49 +4116,55 @@ dependencies = [
[[package]]
name = "tonic-build"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9557ce109ea773b399c9b9e5dca39294110b74f1f342cb347a80d1fce8c26a11"
checksum = "c68f61875ac5293cf72e6c8cf0158086428c82c37229e98c840878f1706b0322"
dependencies = [
"prettyplease",
"proc-macro2",
"prost-build 0.13.5",
"prost-types 0.13.5",
"quote",
"syn",
]
[[package]]
name = "tonic-reflection"
version = "0.12.3"
name = "tonic-prost"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "878d81f52e7fcfd80026b7fdb6a9b578b3c3653ba987f87f0dce4b64043cba27"
checksum = "50849f68853be452acf590cde0b146665b8d507b3b8af17261df47e02c209ea0"
dependencies = [
"prost 0.13.5",
"prost-types 0.13.5",
"tokio",
"tokio-stream",
"bytes",
"prost 0.14.4",
"tonic",
]
[[package]]
name = "tower"
version = "0.4.13"
name = "tonic-prost-build"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8fa9be0de6cf49e536ce1851f987bd21a43b771b09473c3549a6c853db37c1c"
checksum = "654e5643eff75d7f8c99197ce1440ed19a3474eada74c12bbac488b2cafdae27"
dependencies = [
"futures-core",
"futures-util",
"indexmap 1.9.3",
"pin-project",
"pin-project-lite",
"rand 0.8.7",
"slab",
"prettyplease",
"proc-macro2",
"prost-build 0.14.4",
"prost-types 0.14.4",
"quote",
"syn",
"tempfile",
"tonic-build",
]
[[package]]
name = "tonic-reflection"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "acccd136a4bf19810a1fde9c74edc6129b42a66b44d0c1c8aaa67aeb49a146a7"
dependencies = [
"prost 0.14.4",
"prost-types 0.14.4",
"tokio",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
"tokio-stream",
"tonic",
"tonic-prost",
]
[[package]]
@@ -4178,26 +4175,12 @@ checksum = "ebe5ef63511595f1344e2d5cfa636d973292adc0eec1f0ad45fae9f0851ab1d4"
dependencies = [
"futures-core",
"futures-util",
"indexmap",
"pin-project-lite",
"slab",
"sync_wrapper",
"tokio",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
name = "tower-http"
version = "0.5.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1e9cd434a998747dd2c4276bc96ee2e0c7a2eadf3cae88e52be55a05fa9053f5"
dependencies = [
"bitflags 2.11.0",
"bytes",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"pin-project-lite",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
@@ -4216,9 +4199,10 @@ dependencies = [
"http-body 1.0.1",
"iri-string",
"pin-project-lite",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
@@ -4495,7 +4479,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bb0e353e6a2fbdc176932bbaab493762eb1255a7900fe0fea1a2f96c296cc909"
dependencies = [
"anyhow",
"indexmap 2.13.1",
"indexmap",
"wasm-encoder",
"wasmparser",
]
@@ -4521,7 +4505,7 @@ checksum = "47b807c72e1bac69382b3a6fb3dbe8ea4c0ed87ff5629b8685ae6b9a611028fe"
dependencies = [
"bitflags 2.11.0",
"hashbrown 0.15.5",
"indexmap 2.13.1",
"indexmap",
"semver",
]
@@ -4582,7 +4566,6 @@ dependencies = [
"image",
"jsonwebtoken",
"kamadak-exif",
"lazy_static",
"libc",
"md-5",
"memmap2",
@@ -4591,8 +4574,8 @@ dependencies = [
"parking_lot 0.12.5",
"pprof",
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"protoc-bin-vendored",
"rand 0.10.2",
"redb",
@@ -4613,13 +4596,15 @@ dependencies = [
"tokio-stream",
"toml",
"tonic",
"tonic-build",
"tonic-prost",
"tonic-prost-build",
"tonic-reflection",
"tower 0.4.13",
"tower-http 0.5.2",
"tower",
"tower-http",
"tracing",
"tracing-subscriber",
"uuid",
"windows-sys 0.61.2",
"x509-parser",
"xxhash-rust",
]
@@ -4966,7 +4951,7 @@ checksum = "b7c566e0f4b284dd6561c786d9cb0142da491f46a9fbed79ea69cdad5db17f21"
dependencies = [
"anyhow",
"heck",
"indexmap 2.13.1",
"indexmap",
"prettyplease",
"syn",
"wasm-metadata",
@@ -4997,7 +4982,7 @@ checksum = "9d66ea20e9553b30172b5e831994e35fbde2d165325bec84fc43dbf6f4eb9cb2"
dependencies = [
"anyhow",
"bitflags 2.11.0",
"indexmap 2.13.1",
"indexmap",
"log",
"serde",
"serde_derive",
@@ -5016,7 +5001,7 @@ checksum = "ecc8ac4bc1dc3381b7f59c34f00b67e18f910c2c0f50015669dde7def656a736"
dependencies = [
"anyhow",
"id-arena",
"indexmap 2.13.1",
"indexmap",
"log",
"semver",
"serde",
+19 -9
View File
@@ -23,6 +23,11 @@ default = ["5bytes"]
# Pulls redb's experimental_cursor (and therefore experimental-api-5).
redb-experimental-cursor = ["redb/experimental_cursor"]
[lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[dependencies]
# Async runtime
tokio = { version = "1", features = ["full"] }
@@ -30,25 +35,25 @@ tokio-stream = { version = "0.1", features = ["net"] }
tokio-io-timeout = "1"
# gRPC + protobuf
tonic = { version = "0.12", features = ["tls"] }
tonic-reflection = "0.12"
prost = "0.13"
prost-types = "0.13"
tonic = { version = "0.14", features = ["tls-aws-lc"] }
tonic-prost = "0.14"
tonic-reflection = "0.14"
prost = "0.14"
prost-types = "0.14"
# HTTP server
axum = { version = "0.7", features = ["multipart"] }
axum = { version = "0.8", features = ["multipart"] }
http-body = "1"
hyper = { version = "1", features = ["full"] }
hyper-util = { version = "0.1", features = ["tokio", "service", "server-auto", "http1", "http2"] }
tower = "0.4"
tower-http = { version = "0.5", features = ["cors", "trace"] }
tower = { version = "0.5", features = ["util"] }
tower-http = { version = "0.6", features = ["cors", "trace"] }
# CLI
clap = { version = "4", features = ["derive"] }
# Metrics
prometheus = { version = "0.13", default-features = false, features = ["process"] }
lazy_static = "1"
# JWT
jsonwebtoken = { version = "10", features = ["rust_crypto"] }
@@ -138,11 +143,16 @@ aws-types = "1"
[target.'cfg(unix)'.dependencies]
pprof = { version = "0.15", features = ["prost-codec"] }
# GetDiskFreeSpaceExW for per-path disk capacity on Windows (0.61.2 already
# in the tree via tempfile/mio, so this unifies rather than adding a version).
[target.'cfg(windows)'.dependencies]
windows-sys = { version = "0.61", features = ["Win32_Storage_FileSystem"] }
[dev-dependencies]
tempfile = "3"
[build-dependencies]
tonic-build = "0.12"
tonic-prost-build = "0.14"
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry.
+1 -1
View File
@@ -12,7 +12,7 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
tonic_build::configure()
tonic_prost_build::configure()
.build_server(true)
.build_client(true)
// filer.proto uses proto3 optional, which protoc rejects without this
+83 -74
View File
@@ -371,17 +371,18 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
if arg == "--" {
break;
}
if arg.starts_with("--") {
let key = if let Some(eq) = arg.find('=') {
arg[2..eq].to_string()
if let Some(long) = arg.strip_prefix("--") {
let key = if let Some(eq) = long.find('=') {
long[..eq].to_string()
} else {
arg[2..].to_string()
long.to_string()
};
cli_flags.insert(key);
} else if arg.starts_with('-') && arg.len() > 2 {
} else if arg.len() > 2
&& let Some(without_dash) = arg.strip_prefix('-')
{
// Single-dash long option (already normalized to -- at this point,
// but handle both for safety)
let without_dash = &arg[1..];
let key = if let Some(eq) = without_dash.find('=') {
without_dash[..eq].to_string()
} else {
@@ -401,15 +402,14 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
}
// Split on first `=`, ` `, or `:`
let (name, value) =
if let Some(pos) = trimmed.find(|c: char| c == '=' || c == ' ' || c == ':') {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
let (name, value) = if let Some(pos) = trimmed.find(['=', ' ', ':']) {
(
trimmed[..pos].trim().to_string(),
trimmed[pos + 1..].trim().to_string(),
)
} else {
(trimmed.to_string(), String::new())
};
// Strip leading dashes from name
let name = name.trim_start_matches('-').to_string();
@@ -436,10 +436,8 @@ fn merge_options_file(args: Vec<String>) -> Vec<String> {
/// Extract the options file path from args (looks for --options or -options).
fn find_options_arg(args: &[String]) -> String {
for i in 1..args.len() {
if args[i] == "--options" || args[i] == "-options" {
if i + 1 < args.len() {
return args[i + 1].clone();
}
if (args[i] == "--options" || args[i] == "-options") && i + 1 < args.len() {
return args[i + 1].clone();
}
if let Some(rest) = args[i].strip_prefix("--options=") {
return rest.to_string();
@@ -457,20 +455,22 @@ fn parse_duration(s: &str) -> std::time::Duration {
if s.is_empty() {
return std::time::Duration::from_secs(60);
}
if let Some(secs) = s.strip_suffix('s') {
if let Ok(v) = secs.parse::<u64>() {
return std::time::Duration::from_secs(v);
}
if let Some(secs) = s.strip_suffix('s')
&& let Ok(v) = secs.parse::<u64>()
{
return std::time::Duration::from_secs(v);
}
if let Some(mins) = s.strip_suffix('m') {
if let Ok(v) = mins.parse::<u64>() {
return std::time::Duration::from_secs(v * 60);
}
if let Some(mins) = s.strip_suffix('m')
&& let Ok(v) = mins.parse::<u64>()
&& let Some(seconds) = v.checked_mul(60)
{
return std::time::Duration::from_secs(seconds);
}
if let Some(hours) = s.strip_suffix('h') {
if let Ok(v) = hours.parse::<u64>() {
return std::time::Duration::from_secs(v * 3600);
}
if let Some(hours) = s.strip_suffix('h')
&& let Ok(v) = hours.parse::<u64>()
&& let Some(seconds) = v.checked_mul(3600)
{
return std::time::Duration::from_secs(seconds);
}
// Fallback: try parsing as raw seconds
if let Ok(v) = s.parse::<u64>() {
@@ -503,40 +503,40 @@ fn parse_min_free_spaces(min_free_space: &str, min_free_space_percent: &str) ->
}
// Try parsing human-readable bytes: e.g. "10GiB", "500MiB", "1TiB"
let s_upper = s.to_uppercase();
if let Some(rest) = s_upper.strip_suffix("TIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0 * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("KIB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1024.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("TB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("GB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB") {
if let Ok(v) = rest.trim().parse::<f64>() {
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
if let Some(rest) = s_upper.strip_suffix("MB")
&& let Ok(v) = rest.trim().parse::<f64>()
{
return MinFreeSpace::Bytes((v * 1_000_000.0) as u64);
}
// Default: 1%
MinFreeSpace::Percent(1.0)
@@ -987,7 +987,9 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
},
Section::JwtSigning => match key {
"key" => cfg.jwt_signing_key = value.as_bytes().to_vec(),
"expires_after_seconds" => cfg.jwt_signing_expires = value.parse().unwrap_or(10),
"expires_after_seconds" => {
cfg.jwt_signing_expires = value.parse().unwrap_or(10)
}
_ => {}
},
Section::HttpsClient => match key {
@@ -1028,20 +1030,20 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
"cipher_suites" => cfg.tls_policy.cipher_suites = value.to_string(),
_ => {}
},
Section::Guard => match key {
"white_list" => {
Section::Guard => {
if key == "white_list" {
cfg.guard_white_list = value
.split(',')
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
.collect();
}
_ => {}
},
Section::Access => match key {
"ui" => cfg.access_ui = value.parse().unwrap_or(false),
_ => {}
},
}
Section::Access => {
if key == "ui" {
cfg.access_ui = value.parse().unwrap_or(false)
}
}
Section::None => {}
}
}
@@ -1188,12 +1190,11 @@ fn apply_env_overrides(cfg: &mut SecurityConfig) {
/// Mirrors Go's `util.DetectedHostAddress()`.
fn detect_host_address() -> String {
// Connect to a remote address to determine the local outbound IP
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0") {
if socket.connect("8.8.8.8:80").is_ok() {
if let Ok(addr) = socket.local_addr() {
return addr.ip().to_string();
}
}
if let Ok(socket) = UdpSocket::bind("0.0.0.0:0")
&& socket.connect("8.8.8.8:80").is_ok()
&& let Ok(addr) = socket.local_addr()
{
return addr.ip().to_string();
}
"localhost".to_string()
}
@@ -1297,6 +1298,14 @@ mod tests {
assert_eq!(parse_duration("1h"), std::time::Duration::from_secs(3600));
assert_eq!(parse_duration("30"), std::time::Duration::from_secs(30));
assert_eq!(parse_duration(""), std::time::Duration::from_secs(60));
assert_eq!(
parse_duration("307445734561825861m"),
std::time::Duration::from_secs(60)
);
assert_eq!(
parse_duration("5124095576030432h"),
std::time::Duration::from_secs(60)
);
}
#[test]
+30 -27
View File
@@ -6,8 +6,8 @@ use seaweed_volume::config::{self, VolumeServerConfig};
use seaweed_volume::metrics;
use seaweed_volume::pb::volume_server_pb::volume_server_server::VolumeServerServer;
use seaweed_volume::security::tls::{
build_rustls_server_config, build_rustls_server_config_with_grpc_client_auth,
install_default_crypto_provider, GrpcClientAuthPolicy, TlsPolicy,
GrpcClientAuthPolicy, TlsPolicy, build_rustls_server_config,
build_rustls_server_config_with_grpc_client_auth, install_default_crypto_provider,
};
use seaweed_volume::security::{Guard, SigningKey};
#[cfg(unix)]
@@ -18,7 +18,7 @@ use seaweed_volume::server::grpc_server::VolumeGrpcService;
use seaweed_volume::server::profiling::CpuProfileSession;
use seaweed_volume::server::request_id::GrpcRequestIdLayer;
use seaweed_volume::server::volume_server::{
build_metrics_router, RuntimeMetricsConfig, VolumeServerState,
RuntimeMetricsConfig, VolumeServerState, build_metrics_router,
};
use seaweed_volume::server::write_queue::WriteQueue;
use seaweed_volume::storage::store::Store;
@@ -671,8 +671,7 @@ async fn run(
})
.await
} else {
let incoming =
tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
let incoming = tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
info!("gRPC server listening on {}", grpc_local_addr);
build_grpc_server_builder()
.layer(GrpcRequestIdLayer)
@@ -1058,15 +1057,17 @@ mod tests {
#[test]
fn test_grpc_server_tls_returns_none_when_files_are_missing() {
assert!(build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none()
);
}
#[test]
@@ -1088,19 +1089,21 @@ mod tests {
"-----BEGIN CERTIFICATE-----\nZmFrZQ==\n-----END CERTIFICATE-----\n",
);
assert!(build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none()
);
}
#[test]
+225 -124
View File
@@ -6,7 +6,7 @@ use prometheus::{
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge,
IntGaugeVec, Opts, Registry, TextEncoder,
};
use std::sync::Once;
use std::sync::{LazyLock, Once};
use crate::version;
@@ -16,220 +16,320 @@ pub struct PushGatewayConfig {
pub interval_seconds: u32,
}
lazy_static::lazy_static! {
pub static ref REGISTRY: Registry = Registry::new();
pub static REGISTRY: LazyLock<Registry> = LazyLock::new(Registry::new);
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static ref REQUEST_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_request_total", "Volume server requests"),
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static REQUEST_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_request_total",
"Volume server requests",
),
&["type", "code"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Request duration histogram with label `type` (HTTP method).
pub static ref REQUEST_DURATION: HistogramVec = HistogramVec::new(
/// Request duration histogram with label `type` (HTTP method).
pub static REQUEST_DURATION: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_request_seconds",
"Volume server request duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
/// Handler-level operation counter with label `type`.
pub static ref HANDLER_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_handler_total", "Volume server handler counters"),
/// Handler-level operation counter with label `type`.
pub static HANDLER_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_handler_total",
"Volume server handler counters",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
/// Vacuuming compact counter with label `success` (true/false).
pub static ref VACUUMING_COMPACT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_compact_count", "Counter of volume vacuuming Compact counter"),
/// Vacuuming compact counter with label `success` (true/false).
pub static VACUUMING_COMPACT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_compact_count",
"Counter of volume vacuuming Compact counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming commit counter with label `success` (true/false).
pub static ref VACUUMING_COMMIT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_commit_count", "Counter of volume vacuuming commit counter"),
/// Vacuuming commit counter with label `success` (true/false).
pub static VACUUMING_COMMIT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_commit_count",
"Counter of volume vacuuming commit counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static ref VACUUMING_HISTOGRAM: HistogramVec = HistogramVec::new(
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static VACUUMING_HISTOGRAM: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_vacuuming_seconds",
"Volume vacuuming duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
/// Volumes per collection and type (volume/ec_shards).
pub static ref VOLUME_GAUGE: GaugeVec = GaugeVec::new(
/// Volumes per collection and type (volume/ec_shards).
pub static VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_volumes", "Number of volumes"),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Read-only volumes per collection and type.
pub static ref READ_ONLY_VOLUME_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_read_only_volumes", "Number of read-only volumes."),
/// Read-only volumes per collection and type.
pub static READ_ONLY_VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_read_only_volumes",
"Number of read-only volumes.",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Maximum number of volumes this server can hold.
pub static ref MAX_VOLUMES: IntGauge = IntGauge::new(
/// Maximum number of volumes this server can hold.
pub static MAX_VOLUMES: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_max_volumes",
"Maximum number of volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static ref DISK_SIZE_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_total_disk_size", "Actual disk size used by volumes"),
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static DISK_SIZE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_total_disk_size",
"Actual disk size used by volumes",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
/// Disk resource usage per directory and type (all/used/free/avail).
pub static ref RESOURCE_GAUGE: GaugeVec = GaugeVec::new(
/// Disk resource usage per directory and type (all/used/free/avail).
pub static RESOURCE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_resource", "Server resource usage"),
&["name", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
/// In-flight requests per HTTP method.
pub static ref INFLIGHT_REQUESTS_GAUGE: IntGaugeVec = IntGaugeVec::new(
Opts::new("SeaweedFS_volumeServer_in_flight_requests", "Current number of in-flight requests being handled by volume server."),
/// In-flight requests per HTTP method.
pub static INFLIGHT_REQUESTS_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_in_flight_requests",
"Current number of in-flight requests being handled by volume server.",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent download limit in bytes.
pub static ref CONCURRENT_DOWNLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent download limit in bytes.
pub static CONCURRENT_DOWNLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_download_limit",
"Limit for total concurrent download size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent upload limit in bytes.
pub static ref CONCURRENT_UPLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent upload limit in bytes.
pub static CONCURRENT_UPLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_upload_limit",
"Limit for total concurrent upload size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight download bytes.
pub static ref INFLIGHT_DOWNLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight download bytes.
pub static INFLIGHT_DOWNLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_download_size",
"In flight total download size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight upload bytes.
pub static ref INFLIGHT_UPLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight upload bytes.
pub static INFLIGHT_UPLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_upload_size",
"In flight total upload size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static ref UPLOAD_ERROR_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created");
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static UPLOAD_ERROR_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created")
});
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static ref SCRUB_LAST_TIME_SECONDS: GaugeVec = GaugeVec::new(
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static SCRUB_LAST_TIME_SECONDS: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_last_time_seconds",
"Last scrub execution time, as seconds since UNIX epoch.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_VOLUME_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static SCRUB_VOLUME_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_volume_failures",
"Counter of overall volumes with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_SHARD_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static SCRUB_SHARD_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_shard_failures",
"Counter of overall EC shards with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static ref STORAGE_IO_ERROR_COUNTER: IntCounter = IntCounter::new(
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static STORAGE_IO_ERROR_COUNTER: LazyLock<IntCounter> = LazyLock::new(|| {
IntCounter::new(
"SeaweedFS_volumeServer_storage_io_error_total",
"Counter of storage read/write EIO errors on volumes and EC shards.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static ref IO_QUARANTINE_GAUGE: IntGaugeVec = IntGaugeVec::new(
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static IO_QUARANTINE_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_io_quarantine",
"Number of volumes or EC shards quarantined due to storage IO errors.",
),
&["kind"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Legacy aliases for backward compat with existing code ----
// ---- Legacy aliases for backward compat with existing code ----
/// Total number of volumes on this server (flat gauge).
pub static ref VOLUMES_TOTAL: IntGauge = IntGauge::new(
"volume_server_volumes_total",
"Total number of volumes",
).expect("metric can be created");
/// Total number of volumes on this server (flat gauge).
pub static VOLUMES_TOTAL: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new("volume_server_volumes_total", "Total number of volumes")
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static ref DISK_SIZE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk size in bytes per directory.
pub static DISK_SIZE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_size_bytes", "Disk size in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Disk free bytes per directory.
pub static ref DISK_FREE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk free bytes per directory.
pub static DISK_FREE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_free_bytes", "Disk free space in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current number of in-flight requests (flat gauge).
pub static ref INFLIGHT_REQUESTS: IntGauge = IntGauge::new(
/// Current number of in-flight requests (flat gauge).
pub static INFLIGHT_REQUESTS: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_inflight_requests",
"Current number of in-flight requests",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Total number of files stored across all volumes.
pub static ref VOLUME_FILE_COUNT: IntGauge = IntGauge::new(
/// Total number of files stored across all volumes.
pub static VOLUME_FILE_COUNT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_volume_file_count",
"Total number of files stored across all volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Build info (Go: BuildInfo) ----
// ---- Build info (Go: BuildInfo) ----
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static ref BUILD_INFO: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created");
}
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static BUILD_INFO: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created")
});
/// Generate exponential bucket boundaries for histograms.
fn exponential_buckets(start: f64, factor: f64, count: usize) -> Vec<f64> {
@@ -249,6 +349,7 @@ pub const DOWNLOAD_LIMIT_COND: &str = "downloadLimitCondition";
pub const UPLOAD_LIMIT_COND: &str = "uploadLimitCondition";
pub const READ_PROXY_REQ: &str = "readProxyRequest";
pub const READ_REDIRECT_REQ: &str = "readRedirectRequest";
pub const READ_DELETED_NEEDLE: &str = "readDeletedNeedle";
pub const EMPTY_READ_PROXY_LOC: &str = "emptyReadProxyLocaction";
pub const FAILED_READ_PROXY_REQ: &str = "failedReadProxyRequest";
@@ -377,10 +478,8 @@ fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
type_value = Some(label.get_value().to_string());
}
}
if matches_collection {
if let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
if matches_collection && let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
}
}
@@ -427,7 +526,7 @@ pub async fn push_metrics_once(
#[cfg(test)]
mod tests {
use super::*;
use axum::{routing::put, Router};
use axum::{Router, routing::put};
use std::sync::{Arc, Mutex};
#[test]
@@ -499,7 +598,9 @@ mod tests {
register_metrics();
VOLUME_GAUGE.with_label_values(&["pics", "volume"]).set(2.0);
VOLUME_GAUGE.with_label_values(&["pics", "ec_shards"]).set(3.0);
VOLUME_GAUGE
.with_label_values(&["pics", "ec_shards"])
.set(3.0);
READ_ONLY_VOLUME_GAUGE
.with_label_values(&["pics", "volume"])
.set(1.0);
@@ -119,7 +119,11 @@ pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
/// reachable for callers whose target legitimately sits on an internal network
/// (peer volume servers), while still blocking loopback, link-local (IMDS) and
/// unspecified. Mirrors Go's `checkBlockedIPPolicy`.
pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool) -> Result<(), String> {
pub fn check_blocked_ip_policy(
endpoint: &str,
ip: IpAddr,
allow_private: bool,
) -> Result<(), String> {
// Normalize IPv4-mapped IPv6 (`::ffff:a.b.c.d`) to its IPv4 form so the
// IPv4 deny rules apply. The OS routes these to the embedded IPv4 address,
// so without this `::ffff:127.0.0.1` / `::ffff:169.254.169.254` would slip
@@ -173,10 +177,10 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if let IpAddr::V6(v6) = ip
&& let Some(v4) = embedded_transition_ipv4(v6)
{
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
Ok(())
}
@@ -214,9 +218,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
// Authority is everything up to the first '/', '?', or '#'.
let after = &trimmed[scheme_end + 3..];
let authority_end = after
.find(|c| c == '/' || c == '?' || c == '#')
.unwrap_or(after.len());
let authority_end = after.find(['/', '?', '#']).unwrap_or(after.len());
let authority = &after[..authority_end];
// Strip optional userinfo ("user:pass@").
@@ -233,7 +235,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
return Err(format!(
"remote endpoint {:?} has a malformed IPv6 host",
endpoint
))
));
}
}
} else {
@@ -309,7 +311,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
return Err("replica target is empty".to_string());
}
if trimmed.contains("://") || trimmed.contains(['/', '?', '#', '@', '\\']) {
return Err(format!("replica target {:?} must be a bare host:port", target));
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
// Require an explicit host:port, handling `[IPv6]:port`. A bracketless IPv6
@@ -318,12 +323,22 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let host = if let Some(rest) = trimmed.strip_prefix('[') {
match rest.split_once(']') {
Some((h, port)) if port.starts_with(':') && port.len() > 1 => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
} else {
match trimmed.rsplit_once(':') {
Some((h, port)) if !port.is_empty() && !h.contains(':') => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
};
@@ -342,7 +357,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let addrs = resolve_host(host).await?;
if addrs.is_empty() {
return Err(format!("resolve replica target host {:?}: no addresses", host));
return Err(format!(
"resolve replica target host {:?}: no addresses",
host
));
}
for ip in addrs {
check_blocked_ip_policy(target, ip, true)?;
@@ -350,6 +368,56 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
Ok(())
}
/// Resolve `host`, re-apply the replica deny list (private peers allowed) to
/// every resolved address, and connect to the first one that passes -- the
/// connect-time twin of [`validate_replica_target`], so a hostname whose DNS
/// answer flips to a blocked address after the up-front check is still refused.
/// Mirrors Go's `guardedDialerPolicy` with allowPrivate=true.
pub async fn guarded_tcp_connect(
host: &str,
port: u16,
endpoint: &str,
) -> std::io::Result<tokio::net::TcpStream> {
use std::io::{Error, ErrorKind};
let denied = |e: String| Error::new(ErrorKind::PermissionDenied, e);
if is_blocked_imds_host(&host.to_ascii_lowercase()) {
return Err(denied(format!(
"remote endpoint {:?} targets instance metadata service",
endpoint
)));
}
if let Ok(ip) = host.parse::<IpAddr>() {
check_blocked_ip_policy(endpoint, ip, true).map_err(denied)?;
return tokio::net::TcpStream::connect((ip, port)).await;
}
let lookup = tokio::net::lookup_host((host.to_string(), port));
let addrs = tokio::time::timeout(std::time::Duration::from_secs(2), lookup)
.await
.map_err(|_| {
Error::new(
ErrorKind::TimedOut,
format!("resolve remote endpoint host {:?}: timed out", host),
)
})??;
let mut first_block_err: Option<String> = None;
for addr in addrs {
if let Err(e) = check_blocked_ip_policy(endpoint, addr.ip(), true) {
if first_block_err.is_none() {
first_block_err = Some(e);
}
continue;
}
return tokio::net::TcpStream::connect(addr).await;
}
Err(denied(first_block_err.unwrap_or_else(|| {
format!("resolve remote endpoint host {:?}: no addresses", host)
})))
}
#[cfg(test)]
mod tests {
use super::*;
@@ -380,22 +448,30 @@ mod tests {
#[test]
fn rejects_empty_and_bad_scheme() {
assert!(precheck_endpoint("").unwrap_err().contains("empty"));
assert!(precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https"));
assert!(precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https"));
assert!(
precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https")
);
assert!(
precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https")
);
}
#[test]
fn rejects_imds_hostnames() {
assert!(precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service"));
assert!(precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service"));
assert!(
precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service")
);
assert!(
precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service")
);
}
#[test]
@@ -415,27 +491,41 @@ mod tests {
#[test]
fn check_blocked_ip_matches_resolved_categories() {
// Mirror Go's "host resolves to X" cases at the address level.
assert!(check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local"));
assert!(check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT"));
assert!(check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local")
);
assert!(
check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT")
);
assert!(
check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private")
);
assert!(check_blocked_ip("e", ip("52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("2606:4700:4700::1111")).is_ok());
}
@@ -478,33 +568,45 @@ mod tests {
assert!(check_blocked_ip("e", ip("2001::f7f7:f7f7")).is_ok());
assert!(check_blocked_ip("e", ip("::808:808")).is_ok());
// Bracketed transition literal via the full endpoint path.
assert!(precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata"));
assert!(
precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata")
);
}
#[test]
fn rejects_ipv4_mapped_ipv6() {
// IPv4-mapped IPv6 must be unmapped so the IPv4 rules catch it.
assert!(check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private")
);
// A mapped public address still passes, and genuine IPv6 loopback is
// still caught by the V6 path.
assert!(check_blocked_ip("e", ip("::ffff:52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback"));
assert!(
check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback")
);
// Bracketed mapped literal via the full endpoint path.
assert!(precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback"));
assert!(
precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback")
);
}
#[test]
@@ -514,60 +616,86 @@ mod tests {
assert!(check_blocked_ip_policy("e", ip("192.168.1.5"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("100.64.0.42"), true).is_ok());
// Loopback / IMDS / unspecified stay blocked even when private is allowed.
assert!(check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified"));
assert!(
check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified")
);
}
#[tokio::test]
async fn validate_replica_target_rejects_and_allows() {
// A path plus a trailing ?a= would otherwise swallow ?type=replicate.
assert!(validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("")
.await
.unwrap_err()
.contains("empty"));
assert!(
validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("")
.await
.unwrap_err()
.contains("empty")
);
// Legitimate peer volume servers on private networks pass.
assert!(validate_replica_target("10.0.0.7:8080").await.is_ok());
assert!(validate_replica_target("192.168.1.5:8080").await.is_ok());
+3 -1
View File
@@ -7,7 +7,9 @@ pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::{validate_remote_endpoint, validate_replica_target};
pub use endpoint_guard::{
guarded_tcp_connect, validate_remote_endpoint, validate_replica_target,
};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
+1 -1
View File
@@ -2,9 +2,9 @@
//!
//! Works with AWS S3, MinIO, SeaweedFS S3, and all S3-compatible providers.
use aws_sdk_s3::Client;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::Client;
use super::{RemoteEntry, RemoteStorageClient, RemoteStorageError};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
+1 -1
View File
@@ -7,9 +7,9 @@ use std::collections::HashMap;
use std::future::Future;
use std::sync::{Arc, OnceLock, RwLock};
use aws_sdk_s3::Client;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::types::{CompletedMultipartUpload, CompletedPart};
use aws_sdk_s3::Client;
use tokio::io::{AsyncReadExt, AsyncSeekExt, AsyncWriteExt};
use tokio::sync::Semaphore;
+10 -8
View File
@@ -10,7 +10,7 @@ use std::collections::HashSet;
use std::net::IpAddr;
use std::time::{SystemTime, UNIX_EPOCH};
use jsonwebtoken::{decode, encode, Algorithm, DecodingKey, EncodingKey, Header, Validation};
use jsonwebtoken::{Algorithm, DecodingKey, EncodingKey, Header, Validation, decode, encode};
use serde::{Deserialize, Serialize};
// ============================================================================
@@ -297,10 +297,10 @@ impl Guard {
/// Extract host from "host:port" or "[::1]:port" format.
fn extract_host(addr: &str) -> String {
// Handle IPv6 with brackets
if addr.starts_with('[') {
if let Some(end) = addr.find(']') {
return addr[1..end].to_string();
}
if addr.starts_with('[')
&& let Some(end) = addr.find(']')
{
return addr[1..end].to_string();
}
// Handle host:port
if let Some(pos) = addr.rfind(':') {
@@ -481,9 +481,11 @@ mod tests {
let token = gen_jwt(&key, 3600, "3,01637037d6").unwrap();
// Correct file ID
assert!(guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok());
assert!(
guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok()
);
// Wrong file ID
let err = guard.check_jwt_for_file(Some(&token), "4,deadbeef", true);
+3 -3
View File
@@ -3,12 +3,12 @@ use std::fmt;
use std::sync::Arc;
use rustls::client::danger::HandshakeSignatureValid;
use rustls::crypto::aws_lc_rs;
use rustls::crypto::CryptoProvider;
use rustls::crypto::aws_lc_rs;
use rustls::pki_types::UnixTime;
use rustls::pki_types::{CertificateDer, PrivateKeyDer};
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::server::WebPkiClientVerifier;
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::{
CipherSuite, DigitallySignedStruct, DistinguishedName, RootCertStore, ServerConfig,
SignatureScheme, SupportedCipherSuite, SupportedProtocolVersion,
@@ -376,7 +376,7 @@ fn go_tls_version_for_supported(version: &SupportedProtocolVersion) -> GoTlsVers
#[cfg(test)]
mod tests {
use super::{build_supported_versions, common_name_is_allowed, parse_cipher_suites, TlsPolicy};
use super::{TlsPolicy, build_supported_versions, common_name_is_allowed, parse_cipher_suites};
use rustls::crypto::aws_lc_rs;
use std::collections::HashSet;
+2 -2
View File
@@ -1,9 +1,9 @@
use axum::Router;
use axum::body::Body;
use axum::extract::Query;
use axum::http::{header, StatusCode};
use axum::http::{StatusCode, header};
use axum::response::{IntoResponse, Response};
use axum::routing::{any, get};
use axum::Router;
use pprof::protos::Message;
use serde::Deserialize;
+35 -1
View File
@@ -40,7 +40,9 @@ pub fn load_outgoing_grpc_tls(
(&config.grpc_client_cert_file, &config.grpc_client_key_file)
} else {
if !config.grpc_client_cert_file.is_empty() || !config.grpc_client_key_file.is_empty() {
tracing::warn!("grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key");
tracing::warn!(
"grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key"
);
}
(&config.grpc_cert_file, &config.grpc_key_file)
};
@@ -115,6 +117,38 @@ pub fn build_grpc_endpoint(
Ok(endpoint)
}
/// Connect `endpoint` through a connector that re-validates every resolved
/// address at connect time (Go's `guardedDialerPolicy` mirror), pinning a
/// validated copy/tail source against DNS rebinding. `allow_untrusted`
/// preserves the plain connect for operators that opted out.
pub async fn connect_guarded(
endpoint: Endpoint,
target: &str,
allow_untrusted: bool,
) -> Result<Channel, GrpcClientError> {
if allow_untrusted {
return endpoint
.connect()
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)));
}
let target_owned = target.to_string();
let connector = tower::service_fn(move |uri: Uri| {
let target = target_owned.clone();
async move {
let host = uri.host().unwrap_or_default().to_string();
let port = uri.port_u16().unwrap_or(80);
crate::remote_storage::guarded_tcp_connect(&host, port, &target)
.await
.map(hyper_util::rt::TokioIo::new)
}
});
endpoint
.connect_with_connector(connector)
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)))
}
/// Parse a SeaweedFS server address (`"ip:port.grpcPort"` or
/// `"ip:port"`) into the `host:grpcPort` form `build_grpc_endpoint`
/// expects. With the trailing `.grpcPort` segment, that segment IS
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+66 -69
View File
@@ -5,14 +5,14 @@
use std::collections::HashMap;
use std::path::Path;
use std::sync::atomic::Ordering;
use std::sync::Arc;
use std::sync::atomic::Ordering;
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use tokio::sync::broadcast;
use tracing::{error, info, warn};
use super::grpc_client::{build_grpc_endpoint, GRPC_MAX_MESSAGE_SIZE};
use super::grpc_client::{GRPC_MAX_MESSAGE_SIZE, build_grpc_endpoint};
use super::volume_server::VolumeServerState;
use crate::pb::master_pb;
use crate::pb::master_pb::seaweed_client::SeaweedClient;
@@ -119,7 +119,9 @@ pub async fn run_heartbeat_with_state(
if err_msg.contains(DUPLICATE_UUID_RETRY_MESSAGE) {
if duplicate_retry_count >= MAX_DUPLICATE_UUID_RETRIES {
error!("Shut down Volume Server due to persistent duplicate volume directories after 3 retries");
error!(
"Shut down Volume Server due to persistent duplicate volume directories after 3 retries"
);
error!(
"Please check if another volume server is using the same directory"
);
@@ -189,10 +191,10 @@ pub async fn run_heartbeat_with_state(
pub fn to_grpc_address(master_addr: &str) -> String {
if let Some((host, port_str)) = master_addr.rsplit_once(':') {
// "host:port.grpcPort" — the part after the last '.' is the gRPC port.
if let Some((_, grpc_port)) = port_str.rsplit_once('.') {
if grpc_port.parse::<u16>().is_ok() {
return format!("{}:{}", host, grpc_port);
}
if let Some((_, grpc_port)) = port_str.rsplit_once('.')
&& grpc_port.parse::<u16>().is_ok()
{
return format!("{}:{}", host, grpc_port);
}
if let Ok(port) = port_str.parse::<u16>() {
let grpc_port = port + 10000;
@@ -922,10 +924,9 @@ fn build_heartbeat_with_ec_status(
let mut effective_max_count = loc.max_volume_count.load(Ordering::Relaxed);
if loc.is_disk_space_low.load(Ordering::Relaxed) {
let used_slots = loc.volumes_len() as i32
+ ((loc.ec_shard_count()
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
+ loc
.ec_shard_count()
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
effective_max_count = used_slots;
}
@@ -982,7 +983,11 @@ fn build_heartbeat_with_ec_status(
> DISK_CHECK_INTERVAL_NS
{
if !Path::new(&vol.file_name(".dat")).exists() {
warn!("Volume {}: data file {} missing (held open as deleted FD) - not reporting to master", vol.id.0, vol.file_name(".dat"));
warn!(
"Volume {}: data file {} missing (held open as deleted FD) - not reporting to master",
vol.id.0,
vol.file_name(".dat")
);
continue;
}
vol.last_disk_check_ns.store(now_ns, Ordering::Relaxed);
@@ -1248,9 +1253,10 @@ mod tests {
use crate::remote_storage::s3_tier::S3TierRegistry;
use crate::security::{Guard, SigningKey};
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::{DiskType, Version, VolumeId};
use std::sync::atomic::Ordering;
use crate::storage::types::{DiskType, VolumeId};
use crate::storage::volume::VolumeSpec;
use std::sync::RwLock;
use std::sync::atomic::Ordering;
use std::time::{SystemTime, UNIX_EPOCH};
fn test_config() -> HeartbeatConfig {
@@ -1361,12 +1367,11 @@ mod tests {
store
.add_volume(
VolumeId(7),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
@@ -1411,12 +1416,11 @@ mod tests {
store
.add_volume(
VolumeId(7),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
@@ -1477,12 +1481,11 @@ mod tests {
store
.add_volume(
VolumeId(id),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
}
@@ -1531,12 +1534,11 @@ mod tests {
store
.add_volume(
VolumeId(3),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
@@ -1592,12 +1594,11 @@ mod tests {
store
.add_volume(
VolumeId(3),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
@@ -1635,12 +1636,11 @@ mod tests {
store
.add_volume(
vid,
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
}
@@ -1674,12 +1674,11 @@ mod tests {
store
.add_volume(
VolumeId(17),
"heartbeat_metrics_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "heartbeat_metrics_case",
..Default::default()
},
)
.unwrap();
store.locations[0]
@@ -1777,12 +1776,11 @@ mod tests {
store
.add_volume(
VolumeId(21),
collection,
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection,
..Default::default()
},
)
.unwrap();
{
@@ -1950,12 +1948,13 @@ mod tests {
store
.add_volume(
VolumeId(41),
"expired_volume_case",
None,
Some(crate::storage::needle::ttl::TTL::read("20m").unwrap()),
1024,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "expired_volume_case",
ttl: Some(crate::storage::needle::ttl::TTL::read("20m").unwrap()),
preallocate: 1024,
..Default::default()
},
)
.unwrap();
let dat_path = {
@@ -2006,12 +2005,11 @@ mod tests {
store
.add_volume(
VolumeId(51),
"io_error_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "io_error_case",
..Default::default()
},
)
.unwrap();
let (_, volume) = store.find_volume_mut(VolumeId(51)).unwrap();
@@ -2047,12 +2045,11 @@ mod tests {
store
.add_volume(
VolumeId(71),
"remote_volume_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "remote_volume_case",
..Default::default()
},
)
.unwrap();
let (_, volume) = store.find_volume_mut(VolumeId(71)).unwrap();
+3 -3
View File
@@ -29,11 +29,11 @@ impl<S> Layer<S> for GrpcRequestIdLayer {
impl<S, B> Service<http::Request<B>> for GrpcRequestIdService<S>
where
S: Service<http::Request<B>, Response = http::Response<tonic::body::BoxBody>> + Send + 'static,
S: Service<http::Request<B>, Response = http::Response<tonic::body::Body>> + Send + 'static,
S::Future: Send + 'static,
B: Send + 'static,
{
type Response = http::Response<tonic::body::BoxBody>;
type Response = http::Response<tonic::body::Body>;
type Error = S::Error;
type Future = Pin<Box<dyn Future<Output = Result<Self::Response, Self::Error>> + Send>>;
@@ -57,7 +57,7 @@ where
let future = self.inner.call(request);
Box::pin(async move {
let mut response: http::Response<tonic::body::BoxBody> =
let mut response: http::Response<tonic::body::Body> =
scope_request_id(request_id.clone(), future).await?;
if let Ok(value) = HeaderValue::from_str(&request_id) {
response.headers_mut().insert("x-amz-request-id", value);
+187 -223
View File
@@ -38,15 +38,15 @@ use reed_solomon_erasure::galois_8::ReedSolomon;
use tokio::sync::Semaphore;
use tonic::Request;
use crate::pb::master_pb::{self, seaweed_client::SeaweedClient, LookupEcVolumeRequest};
use crate::pb::master_pb::{self, LookupEcVolumeRequest, seaweed_client::SeaweedClient};
use crate::pb::volume_server_pb::{
volume_server_client::VolumeServerClient, CopyFileRequest, VolumeEcShardReadRequest,
CopyFileRequest, VolumeEcShardReadRequest, volume_server_client::VolumeServerClient,
};
use crate::server::grpc_client::{build_grpc_endpoint, parse_grpc_address, GRPC_MAX_MESSAGE_SIZE};
use crate::server::grpc_client::{GRPC_MAX_MESSAGE_SIZE, build_grpc_endpoint, parse_grpc_address};
use crate::server::request_id::outgoing_request_id_interceptor;
use crate::server::volume_server::{to_http_address, VolumeServerState};
use crate::storage::erasure_coding::ec_shard::ShardId;
use crate::storage::needle::needle::{get_actual_size, Needle, NeedleError};
use crate::server::volume_server::{VolumeServerState, to_http_address};
use crate::storage::erasure_coding::ec_shard::{ShardId, shard_id_try_from};
use crate::storage::needle::needle::{Needle, NeedleError, get_actual_size};
use crate::storage::store_ec_reconcile::EcVolumeMissingIndex;
use crate::storage::types::*;
use crate::storage::volume::volume_file_name;
@@ -186,15 +186,19 @@ pub async fn read_ec_shard_needle_distributed(
} => {
fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
encode_ts_ns,
EcInterval {
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns: encode_ts_ns,
},
EcShardMap {
locations: shard_locations,
data_shards,
parity_shards,
},
)
.await
}
@@ -236,8 +240,10 @@ pub async fn read_ec_shard_needle_distributed(
));
}
let mut n = Needle::default();
n.id = needle_id;
let mut n = Needle {
id: needle_id,
..Needle::default()
};
n.read_bytes(
&bytes,
snapshot.offset.to_actual_offset(),
@@ -294,7 +300,7 @@ pub async fn scrub_ec_volume_distributed(
0,
Vec::new(),
vec![format!("EC volume id {} not found", vid.0)],
)
);
}
};
// full scan means verifying the index as well
@@ -309,7 +315,7 @@ pub async fn scrub_ec_volume_distributed(
// one scrub — the same race the vanished-volume policy exists to hide.
// The descriptor outlives the name, the same way the checksum plan's
// shard handles do.
let ecx_walk = fs::File::open(&ecv.ecx_file_name());
let ecx_walk = fs::File::open(ecv.ecx_file_name());
// Encode-run identity of the volume this scrub started against. The
// per-needle `scrub_snapshot_under_lock` re-resolves the volume by id
// under a fresh guard, so a teardown-and-remount of the same vid between
@@ -419,11 +425,10 @@ pub async fn scrub_ec_volume_distributed(
0,
Vec::new(),
vec![format!("EC volume id {} not found", vid.0)],
)
);
}
};
let map = ecv.shard_locations.read().unwrap().clone();
map
ecv.shard_locations.read().unwrap().clone()
};
// Walk the .ecx (private fd captured under the lock, no lock held) for the
@@ -520,18 +525,15 @@ pub async fn scrub_ec_volume_distributed(
} => {
let sources: &[String] =
locations.get(shard_id).map(Vec::as_slice).unwrap_or(&[]);
match read_remote_ec_shard_interval(
state,
sources,
let iv = EcInterval {
vid,
id,
*shard_id,
*shard_offset,
*ssize,
snapshot.encode_ts_ns,
)
.await
{
needle_id: id,
shard_id: *shard_id,
shard_offset: *shard_offset,
size: *ssize,
expected_encode_ts_ns: snapshot.encode_ts_ns,
};
match read_remote_ec_shard_interval(state, sources, iv).await {
// A deleted shard yields no bytes; zero-fill the interval so
// the assembled needle reaches read_bytes -> SizeMismatch{0}
// -> the delete-state suppression (mirrors Go's pre-zeroed buffer).
@@ -559,15 +561,12 @@ pub async fn scrub_ec_volume_distributed(
}
match recover_one_remote_ec_shard_interval(
state,
vid,
id,
*shard_id,
*shard_offset,
*ssize,
&locations,
data_shards,
total_shards - data_shards,
snapshot.encode_ts_ns,
iv,
EcShardMap {
locations: &locations,
data_shards,
parity_shards: total_shards - data_shards,
},
)
.await
{
@@ -673,7 +672,7 @@ fn scrub_snapshot_under_lock(
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("EC volume {} not found (unmounted mid-scan)", vid.0),
))
));
}
};
// The volume was torn down and remounted as a DIFFERENT encode run between
@@ -812,9 +811,9 @@ fn needs_refresh(
let ttl = if stale || shard_count < data_shards {
Duration::from_secs(11)
} else if shard_count == total_shards {
Duration::from_secs(37 * 60)
Duration::from_mins(37)
} else {
Duration::from_secs(7 * 60)
Duration::from_mins(7)
};
age >= ttl
}
@@ -868,22 +867,19 @@ async fn cached_lookup_ec_shard_locations(
}
};
if master.is_empty() {
return Err(io::Error::new(
io::ErrorKind::Other,
"no master configured for ec shard lookup",
));
return Err(io::Error::other("no master configured for ec shard lookup"));
}
let grpc_addr =
parse_grpc_address(&master).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(10))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("master connect: {}", e)))?;
.map_err(|e| io::Error::other(format!("master connect: {}", e)))?;
let mut client = SeaweedClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
@@ -892,7 +888,7 @@ async fn cached_lookup_ec_shard_locations(
let resp = client
.lookup_ec_volume(Request::new(LookupEcVolumeRequest { volume_id: vid.0 }))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("lookup_ec_volume: {}", e)))?;
.map_err(|e| io::Error::other(format!("lookup_ec_volume: {}", e)))?;
let resp = resp.into_inner();
let mut out = HashMap::new();
@@ -902,7 +898,12 @@ async fn cached_lookup_ec_shard_locations(
.iter()
.map(format_location_as_server_address)
.collect();
out.insert(entry.shard_id as ShardId, addrs);
// Defensive: skip out-of-range shard ids from the master instead of
// truncating (256 would alias 0). Valid replies are unaffected.
let Ok(sid) = shard_id_try_from(entry.shard_id) else {
continue;
};
out.insert(sid, addrs);
}
Ok(out)
}
@@ -936,11 +937,11 @@ fn write_back_shard_locations(
/// Resolve the runtime matching the scrub's anchor encode generation, not the
/// first-match `find_ec_volume`. When `expected_encode_ts_ns` is 0 (legacy or
/// pre-feature), falls back to first-match so existing behavior is preserved.
fn find_ec_volume_for_scrub<'a>(
store: &'a crate::storage::store::Store,
fn find_ec_volume_for_scrub(
store: &crate::storage::store::Store,
vid: VolumeId,
expected_encode_ts_ns: i64,
) -> Option<&'a crate::storage::erasure_coding::EcVolume> {
) -> Option<&crate::storage::erasure_coding::EcVolume> {
if expected_encode_ts_ns != 0 {
store
.find_all_ec_volumes(vid)
@@ -959,103 +960,83 @@ fn format_location_as_server_address(loc: &master_pb::Location) -> String {
.url
.trim_start_matches("http://")
.trim_start_matches("https://");
if loc.grpc_port > 0 {
if let Some((host, http_port)) = raw.rsplit_once(':') {
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
if loc.grpc_port > 0
&& let Some((host, http_port)) = raw.rsplit_once(':')
{
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
raw.to_string()
}
/// One shard-relative byte range of a needle on an EC volume, the unit the
/// peer-read and recovery paths work in. Mirrors the argument list of Go's
/// `readOneEcShardInterval`.
#[derive(Clone, Copy, Debug)]
struct EcInterval {
vid: VolumeId,
needle_id: NeedleId,
/// The shard the bytes live on, or the one to rebuild when recovering.
shard_id: ShardId,
shard_offset: i64,
size: usize,
/// Encode run the caller expects the shard to belong to; 0 accepts any,
/// for peers that predate the identity check.
expected_encode_ts_ns: i64,
}
/// Where the shards of one EC volume can be fetched from, and the volume's
/// Reed-Solomon shape, as recovery needs both together.
#[derive(Clone, Copy)]
struct EcShardMap<'a> {
locations: &'a HashMap<ShardId, Vec<String>>,
data_shards: usize,
parity_shards: usize,
}
/// Try direct peer read; on failure, reconstruct via Reed-Solomon
/// from the other shards. Mirrors `readOneEcShardInterval`'s tail.
async fn fetch_one_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
shard_locations: &HashMap<ShardId, Vec<String>>,
data_shards: usize,
parity_shards: usize,
expected_encode_ts_ns: i64,
iv: EcInterval,
map: EcShardMap<'_>,
) -> io::Result<(Vec<u8>, bool)> {
let EcInterval { vid, shard_id, .. } = iv;
// Direct peer read against the cached locations for this shard.
if let Some(sources) = shard_locations.get(&shard_id) {
if !sources.is_empty() {
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
if let Some(sources) = map.locations.get(&shard_id)
&& !sources.is_empty()
{
match read_remote_ec_shard_interval(state, sources, iv).await {
// A deleted needle short-circuits: don't reconstruct (every shard
// would report deleted), let the caller return "deleted".
Ok((buf, is_deleted)) => return Ok((buf, is_deleted)),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
}
}
// Reconstruct: fan-out reads to every other shard at the same
// (shard_offset, size). Mirrors `recoverOneRemoteEcShardInterval`.
recover_one_remote_ec_shard_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
expected_encode_ts_ns,
)
.await
recover_one_remote_ec_shard_interval(state, iv, map).await
}
async fn read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
sources: &[String],
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
expected_encode_ts_ns: i64,
iv: EcInterval,
) -> io::Result<(Vec<u8>, bool)> {
let mut last_err: Option<io::Error> = None;
for src in sources {
match do_read_remote_ec_shard_interval(
state,
src,
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
)
.await
{
match do_read_remote_ec_shard_interval(state, src, iv).await {
Ok(res) => return Ok(res),
Err(e) => last_err = Some(e),
}
@@ -1063,7 +1044,7 @@ async fn read_remote_ec_shard_interval(
Err(last_err.unwrap_or_else(|| {
io::Error::new(
io::ErrorKind::NotFound,
format!("no source for ec shard {}.{}", vid.0, shard_id),
format!("no source for ec shard {}.{}", iv.vid.0, iv.shard_id),
)
}))
}
@@ -1071,28 +1052,26 @@ async fn read_remote_ec_shard_interval(
async fn do_read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
source: &str,
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
expected_encode_ts_ns: i64,
iv: EcInterval,
) -> io::Result<(Vec<u8>, bool)> {
let EcInterval {
vid,
needle_id,
shard_id,
shard_offset,
size,
expected_encode_ts_ns,
} = iv;
let grpc_addr =
parse_grpc_address(source).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
.map_err(|e| io::Error::other(e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("connect to {}: {}", source, e),
)
})?;
.map_err(|e| io::Error::other(format!("connect to {}: {}", source, e)))?;
// TODO(grpc-jwt): clusters with `jwt.signing.key` configured will
// reject peer-to-peer VolumeEcShardRead calls until the Rust
@@ -1118,13 +1097,10 @@ async fn do_read_remote_ec_shard_interval(
.volume_ec_shard_read(Request::new(req))
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"volume_ec_shard_read {}.{} from {}: {}",
vid.0, shard_id, source, e
),
)
io::Error::other(format!(
"volume_ec_shard_read {}.{} from {}: {}",
vid.0, shard_id, source, e
))
})?;
let mut stream = resp.into_inner();
@@ -1133,19 +1109,16 @@ async fn do_read_remote_ec_shard_interval(
while let Some(msg) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv: {}", e)))?
.map_err(|e| io::Error::other(format!("recv: {}", e)))?
{
// Validate the served shard's identity client-side, so the guard holds even
// against a pre-upgrade server that ignored the request field (returns 0).
// A mismatch fails the read; the caller recovers from parity.
if expected_encode_ts_ns != 0 && msg.encode_ts_ns != expected_encode_ts_ns {
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
),
));
return Err(io::Error::other(format!(
"ec shard {}.{} from {} belongs to a different encode run (want {} got {})",
vid.0, shard_id, source, expected_encode_ts_ns, msg.encode_ts_ns
)));
}
if msg.is_deleted {
is_deleted = true;
@@ -1183,19 +1156,25 @@ async fn do_read_remote_ec_shard_interval(
async fn recover_one_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
shard_id_to_recover: ShardId,
shard_offset: i64,
size: usize,
shard_locations: &HashMap<ShardId, Vec<String>>,
data_shards: usize,
parity_shards: usize,
expected_encode_ts_ns: i64,
iv: EcInterval,
map: EcShardMap<'_>,
) -> io::Result<(Vec<u8>, bool)> {
let EcInterval {
vid,
needle_id,
shard_id: shard_id_to_recover,
shard_offset,
size,
expected_encode_ts_ns,
} = iv;
let EcShardMap {
locations: shard_locations,
data_shards,
parity_shards,
} = map;
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap. An interval whose
@@ -1205,13 +1184,10 @@ async fn recover_one_remote_ec_shard_interval(
.acquire_many((size * data_shards).min(EC_RECOVER_BUDGET) as u32)
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
))
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
@@ -1224,7 +1200,7 @@ async fn recover_one_remote_ec_shard_interval(
let mut available = 0usize;
{
let store = state.store.read().unwrap();
for sid in 0..total_shards {
for (sid, slot) in bufs.iter_mut().enumerate() {
if available >= data_shards {
break;
}
@@ -1236,7 +1212,10 @@ async fn recover_one_remote_ec_shard_interval(
// shard from a different encode run must not be fed to Reed-Solomon;
// lenient only when the caller carries no identity (pre-upgrade).
// Mirrors Go's `readLocalEcShardInterval`.
let owner = match store.find_ec_volume_with_shard(vid, sid as u32) {
let Ok(sid_shard) = ShardId::try_from(sid) else {
continue;
};
let owner = match store.find_ec_volume_with_shard(vid, sid_shard) {
Some(ecv)
if expected_encode_ts_ns == 0 || ecv.encode_ts_ns == expected_encode_ts_ns =>
{
@@ -1251,7 +1230,7 @@ async fn recover_one_remote_ec_shard_interval(
.map(|n| n == size)
.unwrap_or(false)
{
bufs[sid] = Some(buf);
*slot = Some(buf);
available += 1;
}
}
@@ -1284,12 +1263,10 @@ async fn recover_one_remote_ec_shard_interval(
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
sid,
shard_offset,
size,
expected_encode_ts_ns,
EcInterval {
shard_id: sid,
..iv
},
)
.await;
(sid, res)
@@ -1335,34 +1312,25 @@ async fn recover_one_remote_ec_shard_interval(
if any_deleted {
return Ok((Vec::new(), true));
}
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
),
));
return Err(io::Error::other(format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
)));
}
rs.reconstruct(&mut bufs).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
),
)
io::Error::other(format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
))
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok((buf, any_deleted)),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
),
)),
None => Err(io::Error::other(format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
))),
}
}
@@ -1504,12 +1472,12 @@ async fn fetch_ec_index_from_one_peer(
let grpc_addr =
parse_grpc_address(peer).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let channel = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?
.map_err(|e| io::Error::other(e.to_string()))?
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("connect {}: {}", peer, e)))?;
.map_err(|e| io::Error::other(format!("connect {}: {}", peer, e)))?;
let mut client = VolumeServerClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE);
@@ -1530,22 +1498,19 @@ async fn fetch_ec_index_from_one_peer(
let stream = client
.copy_file(copy_req(".ecx", false))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("copy .ecx: {}", e)))?
.map_err(|e| io::Error::other(format!("copy .ecx: {}", e)))?
.into_inner();
drain_copy_stream(stream, ecx_path, false).await?;
let meta = fs::metadata(ecx_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("stat copied .ecx: {}", e)))?;
let meta =
fs::metadata(ecx_path).map_err(|e| io::Error::other(format!("stat copied .ecx: {}", e)))?;
if meta.is_dir() || meta.len() == 0 {
let _ = fs::remove_file(ecx_path);
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"peer {} served an unusable .ecx (size {})",
peer,
meta.len()
),
));
return Err(io::Error::other(format!(
"peer {} served an unusable .ecx (size {})",
peer,
meta.len()
)));
}
// .ecj is the source peer's deletion journal (appended); .vif carries EC
@@ -1589,15 +1554,14 @@ async fn drain_copy_stream(
} else {
fs::File::create(dest_path)
}
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("create {}: {}", dest_path, e)))?;
.map_err(|e| io::Error::other(format!("create {}: {}", dest_path, e)))?;
while let Some(chunk) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv {}: {}", dest_path, e)))?
.map_err(|e| io::Error::other(format!("recv {}: {}", dest_path, e)))?
{
file.write_all(&chunk.file_content).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("write {}: {}", dest_path, e))
})?;
file.write_all(&chunk.file_content)
.map_err(|e| io::Error::other(format!("write {}: {}", dest_path, e)))?;
}
Ok(())
}
+6 -7
View File
@@ -423,13 +423,12 @@ fn collect_ui_data(
shard_id: shard.shard_id,
size: shard_size,
});
if created_at == "-" {
if let Ok(metadata) = std::fs::metadata(shard.file_name()) {
if let Ok(modified) = metadata.modified() {
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
if created_at == "-"
&& let Ok(metadata) = std::fs::metadata(shard.file_name())
&& let Ok(modified) = metadata.modified()
{
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
let preferred_size = ec_volume.dat_file_size.max(0) as u64;
+24 -25
View File
@@ -14,12 +14,12 @@ use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32, Ordering};
use std::sync::{Arc, RwLock};
use axum::{
extract::{connect_info::ConnectInfo, Request, State},
http::{header, HeaderValue, Method, StatusCode},
Router,
extract::{Request, State, connect_info::ConnectInfo},
http::{HeaderValue, Method, StatusCode, header},
middleware::{self, Next},
response::{IntoResponse, Response},
routing::{any, get},
Router,
};
use crate::config::ReadMode;
@@ -200,9 +200,7 @@ pub fn to_http_address(addr: &str) -> std::borrow::Cow<'_, str> {
// rather than being silently rewritten. Mirrors the validation already
// done in `to_grpc_address` for the inverse direction.
if let (Ok(_), Ok(_)) = (http_port.parse::<u16>(), grpc_port.parse::<u16>()) {
return std::borrow::Cow::Owned(
addr[..ports_sep_index + 1 + dot_idx].to_string(),
);
return std::borrow::Cow::Owned(addr[..ports_sep_index + 1 + dot_idx].to_string());
}
}
std::borrow::Cow::Borrowed(addr)
@@ -312,16 +310,15 @@ async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Requ
)
}
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -358,16 +355,15 @@ async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Req
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -432,13 +428,13 @@ pub fn build_admin_router_with_ui(state: Arc<VolumeServerState>, ui_enabled: boo
.route("/healthz", get(handlers::healthz_handler))
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(admin_store_handler))
.route("/:path", any(admin_store_handler))
.route("/:vid/:fid", any(admin_store_handler))
.route("/:vid/:fid/:filename", any(admin_store_handler))
.route("/{path}", any(admin_store_handler))
.route("/{vid}/{fid}", any(admin_store_handler))
.route("/{vid}/{fid}/{filename}", any(admin_store_handler))
.fallback(admin_store_handler);
if ui_enabled {
// Note: /stats/* endpoints are commented out in Go's volume_server.go (L130-134).
@@ -455,13 +451,13 @@ pub fn build_public_router(state: Arc<VolumeServerState>) -> Router {
Router::new()
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(public_store_handler))
.route("/:path", any(public_store_handler))
.route("/:vid/:fid", any(public_store_handler))
.route("/:vid/:fid/:filename", any(public_store_handler))
.route("/{path}", any(public_store_handler))
.route("/{vid}/{fid}", any(public_store_handler))
.route("/{vid}/{fid}/{filename}", any(public_store_handler))
.fallback(public_store_handler)
.layer(middleware::from_fn(common_headers_middleware))
.with_state(state)
@@ -516,7 +512,10 @@ mod tests {
// "host:abc.def"), and silently rewriting it would just hide the bug.
assert_eq!(to_http_address("host:abc.def"), "host:abc.def");
assert_eq!(to_http_address("host:9333.notaport"), "host:9333.notaport");
assert_eq!(to_http_address("host:notaport.19333"), "host:notaport.19333");
assert_eq!(
to_http_address("host:notaport.19333"),
"host:notaport.19333"
);
// Out-of-range ports must not be silently truncated either.
assert_eq!(to_http_address("host:99999.19333"), "host:99999.19333");
}
+1 -1
View File
@@ -178,8 +178,8 @@ mod tests {
use crate::server::volume_server::RuntimeMetricsConfig;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::store::Store;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
use std::sync::RwLock;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
let store = Store::new(NeedleMapKind::InMemory);
let guard = Guard::new(&[], SigningKey(vec![]), 0, SigningKey(vec![]), 0);
+254 -181
View File
@@ -15,15 +15,15 @@ use tracing::warn;
use crate::config::MinFreeSpace;
use crate::storage::erasure_coding::ec_bitrot::remove_bitrot_sidecars;
use crate::storage::erasure_coding::ec_shard::{
EcVolumeShard, DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE, ERASURE_CODING_SMALL_BLOCK_SIZE,
EcVolumeShard, ShardId,
};
use crate::storage::erasure_coding::ec_volume::EcVolume;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::super_block::{ReplicaPlacement, SUPER_BLOCK_SIZE};
use crate::storage::super_block::SUPER_BLOCK_SIZE;
use crate::storage::types::*;
use crate::storage::volume::{
remove_volume_files, volume_file_name, VifVolumeInfo, Volume, VolumeError,
VifVolumeInfo, Volume, VolumeError, VolumeSpec, remove_volume_files, volume_file_name,
};
/// A single disk location managing volumes in one directory.
@@ -131,10 +131,10 @@ impl DiskLocation {
for entry in entries {
let entry = entry?;
let name = entry.file_name().into_string().unwrap_or_default();
if let Some((collection, vid)) = parse_volume_filename(&name) {
if seen.insert((collection.clone(), vid)) {
dat_files.push((collection, vid));
}
if let Some((collection, vid)) = parse_volume_filename(&name)
&& seen.insert((collection.clone(), vid))
{
dat_files.push((collection, vid));
}
}
@@ -205,7 +205,6 @@ impl DiskLocation {
continue;
}
// Load existing data only; never create a phantom `.dat`. A lone
// `.vif`/`.idx` (e.g. an EC sidecar whose `.ecx` is on a sibling
// disk) would otherwise have Volume::new write an 8-byte stub that
@@ -280,30 +279,33 @@ impl DiskLocation {
let opened = Mutex::new(Vec::with_capacity(to_load.len()));
std::thread::scope(|scope| {
for _ in 0..workers {
scope.spawn(|| loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
match Volume::new(
&self.directory,
&self.idx_directory,
collection,
*vid,
needle_map_kind,
None, // replica placement read from superblock
None, // TTL read from superblock
0, // no preallocate on load
Version::current(),
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
scope.spawn(|| {
loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
// Replica placement and TTL are read back from the
// superblock, and a load never preallocates.
match Volume::new(
&self.directory,
&self.idx_directory,
*vid,
needle_map_kind,
&VolumeSpec {
collection,
..Default::default()
},
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
}
}
}
}
@@ -327,10 +329,10 @@ impl DiskLocation {
.strip_suffix(".cpc")
.or_else(|| name.strip_suffix(".cpd"))
.or_else(|| name.strip_suffix(".cpx"));
if let Some(stem) = stem {
if let Some(key) = parse_collection_volume_id(stem) {
pending.insert(key);
}
if let Some(stem) = stem
&& let Some(key) = parse_collection_volume_id(stem)
{
pending.insert(key);
}
}
}
@@ -374,8 +376,10 @@ impl DiskLocation {
let mut expected_shard_size: Option<i64> = None;
let dat_exists = match fs::metadata(&dat_path) {
Ok(meta) if meta.len() > SUPER_BLOCK_SIZE as u64 => {
expected_shard_size =
Some(calculate_expected_shard_size(meta.len() as i64, data_shards));
expected_shard_size = Some(calculate_expected_shard_size(
meta.len() as i64,
data_shards,
));
true
}
Ok(_) => false,
@@ -399,7 +403,13 @@ impl DiskLocation {
if size != prev {
// Inconsistent sizes signal corruption or mixed
// generations; not trusted for deletion -> keep.
warn!(volume_id = vid.0, shard = i, size, expected = prev, "EC shard size mismatch; keeping shards");
warn!(
volume_id = vid.0,
shard = i,
size,
expected = prev,
"EC shard size mismatch; keeping shards"
);
return true;
}
} else {
@@ -426,11 +436,16 @@ impl DiskLocation {
if shard_count == 0 {
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size) {
if actual < expected {
warn!(volume_id = vid.0, actual, expected, "shards smaller than the .dat's full encode; reclaiming the complete .dat");
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size)
&& actual < expected
{
warn!(
volume_id = vid.0,
actual,
expected,
"shards smaller than the .dat's full encode; reclaiming the complete .dat"
);
return false;
}
true
}
@@ -510,10 +525,10 @@ impl DiskLocation {
pub(crate) fn ec_generation_ts_ns(&self, collection: &str, vid: VolumeId) -> Option<i64> {
for dir in [&self.directory, &self.idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Ok(s) = fs::read_to_string(&vif) {
if let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s) {
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if let Ok(s) = fs::read_to_string(&vif)
&& let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s)
{
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if self.directory == self.idx_directory {
break;
@@ -545,27 +560,19 @@ impl DiskLocation {
pub fn create_volume(
&mut self,
vid: VolumeId,
collection: &str,
needle_map_kind: NeedleMapKind,
replica_placement: Option<ReplicaPlacement>,
ttl: Option<crate::storage::needle::ttl::TTL>,
preallocate: u64,
version: Version,
spec: &VolumeSpec<'_>,
) -> Result<(), VolumeError> {
let mut v = Volume::new(
&self.directory,
&self.idx_directory,
collection,
vid,
needle_map_kind,
replica_placement,
ttl,
preallocate,
version,
spec,
)?;
v.location_disk_space_low = self.is_disk_space_low.clone();
crate::metrics::VOLUME_GAUGE
.with_label_values(&[collection, "volume"])
.with_label_values(&[spec.collection, "volume"])
.inc();
self.volumes.insert(vid, v);
Ok(())
@@ -669,8 +676,7 @@ impl DiskLocation {
pub fn free_volume_count(&self) -> i32 {
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let max = self.max_volume_count.load(Ordering::Relaxed);
let free_count = (max as i64 - self.volumes.len() as i64)
* DATA_SHARDS_COUNT as i64
let free_count = (max as i64 - self.volumes.len() as i64) * DATA_SHARDS_COUNT as i64
- self.ec_shard_count() as i64;
let effective_free = free_count / DATA_SHARDS_COUNT as i64;
if effective_free > 0 {
@@ -777,18 +783,18 @@ impl DiskLocation {
pub fn has_ecx_file_on_disk(&self, collection: &str, vid: VolumeId) -> bool {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
let idx_path = format!("{}.ecx", idx_base);
if let Ok(meta) = fs::metadata(&idx_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&idx_path)
&& !meta.is_dir()
{
return true;
}
if self.idx_directory != self.directory {
let data_base = volume_file_name(&self.directory, collection, vid);
let data_path = format!("{}.ecx", data_base);
if let Ok(meta) = fs::metadata(&data_path) {
if !meta.is_dir() {
return true;
}
if let Ok(meta) = fs::metadata(&data_path)
&& !meta.is_dir()
{
return true;
}
}
false
@@ -811,7 +817,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
source_disk_type: &str,
) -> Result<(), VolumeError> {
let idx_dir = self.idx_directory.clone();
@@ -833,7 +839,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
idx_dir: &str,
source_disk_type: &str,
) -> Result<(), VolumeError> {
@@ -845,14 +851,10 @@ impl DiskLocation {
// propagate the error to the caller.
let created = !self.ec_volumes.contains_key(&vid);
if created {
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid)
.map_err(VolumeError::Io)?;
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid).map_err(VolumeError::Io)?;
self.ec_volumes.insert(vid, ec_vol);
}
let ec_vol = self
.ec_volumes
.get_mut(&vid)
.expect("just inserted above");
let ec_vol = self.ec_volumes.get_mut(&vid).expect("just inserted above");
// When the orchestrator supplied a source disk type on the Mount
// RPC, override the EC volume's disk type so heartbeats report
// under the source volume's disk type (#9423). When the caller
@@ -871,10 +873,10 @@ impl DiskLocation {
// keep the existing registration (mirrors Go's AddEcVolumeShard
// added=false) — re-adding would replace a serving fd and bump
// the ec_shards gauge without growing the mounted count.
if ec_vol.has_shard(shard_id as u8) {
if ec_vol.has_shard(shard_id) {
continue;
}
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id as u8);
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id);
shard.disk_type = ec_vol.disk_type.clone();
if let Err(e) = ec_vol.add_shard(shard) {
// The shard was dropped (its descriptors closed) inside the
@@ -902,14 +904,14 @@ impl DiskLocation {
/// caller passes a shard that lives on a sibling disk
/// (cross-disk reconcile makes that the common case for the same
/// `vid` after reconciliation).
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[u32]) {
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[ShardId]) {
if let Some(ec_vol) = self.ec_volumes.get_mut(&vid) {
let collection = ec_vol.collection.clone();
for &shard_id in shard_ids {
if !ec_vol.has_shard(shard_id as u8) {
if !ec_vol.has_shard(shard_id) {
continue;
}
ec_vol.remove_shard(shard_id as u8);
let _ = ec_vol.remove_shard(shard_id);
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&collection, "ec_shards"])
.dec();
@@ -969,7 +971,7 @@ impl DiskLocation {
}
entries.sort();
let mut same_volume_shards: Vec<(String, u32)> = Vec::new(); // (filename, shard_id)
let mut same_volume_shards: Vec<(String, ShardId)> = Vec::new(); // (filename, shard_id)
let mut prev_vid: Option<VolumeId> = None;
let mut prev_collection: String = String::new();
@@ -1034,7 +1036,12 @@ impl DiskLocation {
/// Validate + mount a (collection, vid) group when its `.ecx` is
/// found. Mirrors `handleFoundEcxFile` in
/// `weed/storage/disk_location_ec.go`.
fn handle_found_ecx_file(&mut self, shards: &[(String, u32)], collection: &str, vid: VolumeId) {
fn handle_found_ecx_file(
&mut self,
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) {
let base = volume_file_name(&self.directory, collection, vid);
let dat_path = format!("{}.dat", base);
let dat_exists = check_dat_file_exists(&dat_path);
@@ -1048,7 +1055,7 @@ impl DiskLocation {
return;
}
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if let Err(e) = self.mount_ec_shards(vid, collection, &shard_ids, "") {
// A mount failure (corrupt/locked .ecx, EMFILE, transient I/O) is
// not proof the shards are disposable -- validate_ec_volume already
@@ -1057,8 +1064,7 @@ impl DiskLocation {
// delete on a load error.
warn!(
volume_id = vid.0,
"Failed to load EC shards: {}; keeping files for retry",
e,
"Failed to load EC shards: {}; keeping files for retry", e,
);
self.unmount_ec_shards(vid, &shard_ids);
}
@@ -1071,7 +1077,7 @@ impl DiskLocation {
/// distributed-EC shards waiting for cross-disk reconciliation.
fn check_orphaned_shards(
&self,
shards: &[(String, u32)],
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) -> bool {
@@ -1107,7 +1113,7 @@ impl DiskLocation {
/// Close all volumes.
pub fn close(&mut self) {
for (_, v) in self.volumes.iter_mut() {
for v in self.volumes.values_mut() {
v.close();
}
self.volumes.clear();
@@ -1137,10 +1143,45 @@ pub fn get_disk_stats(path: &str) -> (u64, u64) {
}
(0, 0)
}
#[cfg(not(unix))]
#[cfg(windows)]
{
let _ = path;
(0, 0)
use std::os::windows::ffi::OsStrExt;
// Canonicalize so symlinks, `.`/`..` segments, and relative paths
// resolve to the real location before querying. `\\?\`-prefixed
// extended-length paths and UNC (`\\?\UNC\...`) are passed through
// untouched: GetDiskFreeSpaceExW accepts them as-is.
let canonical = match std::fs::canonicalize(path) {
Ok(p) => p,
Err(_) => return (0, 0),
};
// UTF-16 with trailing NUL for the Win32 wide-string call.
let mut wide: Vec<u16> = canonical.as_os_str().encode_wide().collect();
// UNC directory names must end in a backslash for GetDiskFreeSpaceExW.
if !wide.ends_with(&[0x5C]) {
wide.push(0x5C);
}
wide.push(0);
// SAFETY: `wide` is NUL-terminated; the out-params are valid u64
// writes; the call has no other preconditions.
unsafe {
let mut free_available: u64 = 0;
let mut total: u64 = 0;
let ok = windows_sys::Win32::Storage::FileSystem::GetDiskFreeSpaceExW(
wide.as_ptr(),
&mut free_available,
&mut total,
std::ptr::null_mut(),
);
if ok == 0 {
return (0, 0);
}
return (total, free_available);
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("get_disk_stats is implemented for unix and windows only");
}
}
@@ -1176,7 +1217,12 @@ fn rm_if_present(path: String) -> io::Result<()> {
}
}
fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &str, vid: VolumeId) -> usize {
fn ec_data_shards_from_vif(
directory: &str,
idx_directory: &str,
collection: &str,
vid: VolumeId,
) -> usize {
for dir in [directory, idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Some(ds) = fs::read_to_string(&vif)
@@ -1184,10 +1230,9 @@ fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &st
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
.and_then(|vi| vi.ec_shard_config)
.map(|c| c.data_shards as usize)
&& ds > 0
{
if ds > 0 {
return ds;
}
return ds;
}
if directory == idx_directory {
break;
@@ -1223,7 +1268,7 @@ fn parse_collection_volume_id(base: &str) -> Option<(String, VolumeId)> {
/// `pub(crate)` re-export of [`parse_ec_shard_extension`] for the
/// cross-disk reconcile in `store_ec_reconcile.rs`.
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<ShardId> {
parse_ec_shard_extension(ext)
}
@@ -1237,7 +1282,7 @@ pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
/// shardId > 255` guard. The 3-digit form (`.ec100`–`.ec255`) is
/// retained so the parser can still recognise shards from custom
/// 32+ ratios that fit in a u8 even though OSS only ships 10+4.
fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
fn parse_ec_shard_extension(ext: &str) -> Option<ShardId> {
let rest = ext.strip_prefix(".ec")?;
if rest.len() < 2 || rest.len() > 3 {
return None;
@@ -1246,7 +1291,7 @@ fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
if id > 255 {
return None;
}
Some(id)
ShardId::try_from(id).ok()
}
/// Robust check that a `.dat` with actual data exists. An empty `.dat`
@@ -1308,7 +1353,10 @@ fn remove_empty_ec_dat_stub(volume_name: &str, idx_name: &str, vid: VolumeId) ->
return false;
}
warn!(volume_id = vid.0, "removing leftover empty .dat stub for EC volume");
warn!(
volume_id = vid.0,
"removing leftover empty .dat stub for EC volume"
);
let _ = fs::remove_file(&dat_path);
let _ = fs::remove_file(format!("{}.idx", idx_name));
true
@@ -1331,6 +1379,17 @@ mod tests {
use super::*;
use tempfile::TempDir;
/// get_disk_stats must report real capacity for a real path on every
/// platform (Windows included) — consumers treat total==0 as "unknown"
/// and leave available_space at 0, which breaks volume assignment.
#[test]
fn test_get_disk_stats_reports_capacity_for_real_path() {
let tmp = TempDir::new().unwrap();
let (total, free) = get_disk_stats(tmp.path().to_str().unwrap());
assert!(total > 0, "expected total>0, got {total}");
assert!(free > 0, "expected free>0, got {free}");
}
/// When `-dir.idx` is configured the EC `.vif` may live in the idx
/// directory; the sweep must look there too, not only the data dir.
#[test]
@@ -1353,7 +1412,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&vif).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&vif).unwrap(),
)
.unwrap();
assert!(
remove_empty_ec_dat_stub(&vbase, &ibase, VolumeId(42)),
@@ -1369,16 +1432,30 @@ mod tests {
fn test_validate_ec_volume_partial_dat_next_to_full_shards_keeps() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(70));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let full = calculate_expected_shard_size(30 * 1024 * 1024, ds);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(full as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(full as u64)
.unwrap();
}
// Partial .dat: bigger than a superblock so it is not swept as a stub,
// but smaller than what these shards encode.
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(5 * 1024 * 1024).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(5 * 1024 * 1024)
.unwrap();
assert!(
loc.validate_ec_volume("", VolumeId(70)),
"full-size shards beside a smaller (stale/partial) .dat must be kept",
@@ -1392,15 +1469,29 @@ mod tests {
fn test_validate_ec_volume_interrupted_encode_reclaims() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(71));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let dat_size = 30 * 1024 * 1024i64;
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(dat_size as u64).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(dat_size as u64)
.unwrap();
let partial = calculate_expected_shard_size(dat_size, ds) / 3;
assert!(partial > 0);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(partial as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(partial as u64)
.unwrap();
}
assert!(
!loc.validate_ec_volume("", VolumeId(71)),
@@ -1445,7 +1536,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
// A .vif with no EC config reads as generation 0 (recovered/pre-upgrade live volume).
@@ -1454,12 +1549,20 @@ mod tests {
version: 3,
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&no_cfg).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&no_cfg).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(0));
// idx-dir fallback: only the idx dir holds the .vif.
std::fs::remove_file(format!("{}.vif", dbase)).unwrap();
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
}
@@ -1499,16 +1602,8 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 1);
assert!(loc.find_volume(VolumeId(1)).is_some());
@@ -1532,24 +1627,15 @@ mod tests {
Vec::new(),
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(
VolumeId(2),
"test",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "test",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1592,12 +1678,11 @@ mod tests {
.unwrap();
loc.create_volume(
VolumeId(9),
"good",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "good",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1641,26 +1726,10 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(
VolumeId(2),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(VolumeId(2), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 2);
loc.delete_volume(VolumeId(1), false, false).unwrap();
@@ -1684,32 +1753,29 @@ mod tests {
loc.create_volume(
VolumeId(1),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(2),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(3),
"docs",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "docs",
..Default::default()
},
)
.unwrap();
assert_eq!(loc.volumes_len(), 3);
@@ -1772,7 +1838,8 @@ mod tests {
// mount_ec_shards with source_disk_type="ssd" — simulating the
// VolumeEcShardsMount RPC path.
std::fs::write(format!("{}/pics_7.ec00", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd")
.unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume mounted");
assert_eq!(
@@ -1789,7 +1856,9 @@ mod tests {
std::fs::write(format!("{}/pics_7.ec01", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[1], "").unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume still mounted");
let ec_vol = loc
.find_ec_volume(VolumeId(7))
.expect("ec volume still mounted");
assert_eq!(
ec_vol.disk_type,
DiskType::Ssd,
@@ -1871,7 +1940,8 @@ mod tests {
let gauge = crate::metrics::VOLUME_GAUGE.with_label_values(&["dupmount", "ec_shards"]);
let before = gauge.get();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "").unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.expect("a duplicate mount must succeed as a no-op");
@@ -1947,8 +2017,11 @@ mod tests {
let path = format!("{}/{}_{}.ec{:02}", dir, collection, vid.0, sid);
std::fs::write(&path, b"shard data nonempty").unwrap();
}
std::fs::write(format!("{}/{}_{}.ecx", dir, collection, vid.0), vec![0u8; 20])
.unwrap();
std::fs::write(
format!("{}/{}_{}.ecx", dir, collection, vid.0),
vec![0u8; 20],
)
.unwrap();
std::fs::write(format!("{}/{}_{}.ecj", dir, collection, vid.0), b"").unwrap();
std::fs::write(
format!("{}/{}_{}.vif", dir, collection, vid.0),
@@ -30,6 +30,7 @@ use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums, EcShardConfig,
};
use crate::storage::erasure_coding::ec_shard::MAX_SHARD_COUNT;
use crate::storage::io::read_exact_at;
use crate::storage::needle::crc::CRC;
/// Canonical extension for the checksum sidecar. Generation 0 (legacy/fresh
@@ -164,16 +165,16 @@ pub fn remove_bitrot_sidecars(base: &str) -> io::Result<()> {
};
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res {
if first_err.is_none() {
first_err = Some(e);
}
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
record(rm(format!("{}{}", base, BITROT_SIDECAR_EXT).into()));
let path = Path::new(base);
if let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) {
let prefix = format!("{}{}.v", fname.to_string_lossy(), BITROT_SIDECAR_EXT);
let prefix = format!("{}{}.v", fname.display(), BITROT_SIDECAR_EXT);
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries.flatten() {
@@ -203,7 +204,7 @@ pub fn new_encode_uuid() -> Vec<u8> {
/// Reports whether `block_size` is a power of two in [1 MiB, MAX_BITROT_BLOCK_SIZE].
pub fn is_pow2_multiple_of_1mib(block_size: u32) -> bool {
block_size >= (1 << 20) && block_size <= MAX_BITROT_BLOCK_SIZE && block_size.count_ones() == 1
((1 << 20)..=MAX_BITROT_BLOCK_SIZE).contains(&block_size) && block_size.count_ones() == 1
}
/// Returns ceil(covered_size / block_size).
@@ -402,7 +403,7 @@ pub fn validate_manifest(
total
));
}
let mut seen = vec![false; MAX_SHARD_COUNT];
let mut seen = [false; MAX_SHARD_COUNT];
for s in &prot.shards {
if s.shard_id >= total as u32 {
return Err(format!(
@@ -537,7 +538,7 @@ pub fn verify_shard_blocks(
break;
}
let to_read = to_read as usize;
read_full_at(f, &mut buf[..to_read], offset as u64)?;
read_exact_at(f, &mut buf[..to_read], offset as u64)?;
if CRC::new(&buf[..to_read]).0 != *want_crc {
mismatched.push(i);
}
@@ -546,33 +547,6 @@ pub fn verify_shard_blocks(
Ok(mismatched)
}
/// Reads exactly `buf.len()` bytes from `f` at `offset`, erroring on early EOF.
fn read_full_at(f: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
let mut total = 0usize;
while total < buf.len() {
#[cfg(unix)]
let n = {
use std::os::unix::fs::FileExt;
f.read_at(&mut buf[total..], offset + total as u64)?
};
#[cfg(not(unix))]
let n = {
use std::io::{Read, Seek, SeekFrom};
let mut fc = f.try_clone()?;
fc.seek(SeekFrom::Start(offset + total as u64))?;
fc.read(&mut buf[total..])?
};
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"short read on shard block",
));
}
total += n;
}
Ok(())
}
/// Builds the `EcShardConfig` proto for the given layout. The bitrot sidecar
/// carries its own top-level encode_uuid, so the nested config leaves it empty.
pub fn ec_shard_config(data_shards: u32, parity_shards: u32, block_size: i64) -> EcShardConfig {
@@ -627,7 +601,10 @@ mod tests {
save_bitrot_sidecar(path, &prot).unwrap();
let bytes = std::fs::read(path).unwrap();
let hex: String = bytes.iter().map(|b| format!("{:02x}", b)).collect();
assert_eq!(hex, CANONICAL_HEX, "Rust .ecsum bytes drifted from the Go canonical form");
assert_eq!(
hex, CANONICAL_HEX,
"Rust .ecsum bytes drifted from the Go canonical form"
);
let _ = std::fs::remove_file(path);
}
@@ -662,7 +639,11 @@ mod tests {
format!("{}.ecsum.v1", base),
format!("{}.ecsum.v7", base),
] {
assert!(!std::path::Path::new(&p).exists(), "{} should be removed", p);
assert!(
!std::path::Path::new(&p).exists(),
"{} should be removed",
p
);
}
assert!(std::path::Path::new(&keep_shard).exists());
assert!(std::path::Path::new(&keep_other_vid).exists());
@@ -681,7 +662,9 @@ mod tests {
assert!(!is_pow2_multiple_of_1mib(1 << 19)); // 512 KiB, too small
assert!(!is_pow2_multiple_of_1mib(3 << 20)); // 3 MiB, not pow2
assert!(!is_pow2_multiple_of_1mib(128 * 1024 * 1024)); // pow2 but > MAX_BITROT_BLOCK_SIZE
assert!(!is_pow2_multiple_of_1mib(DEFAULT_BITROT_BLOCK_SIZE as u32 + 1));
assert!(!is_pow2_multiple_of_1mib(
DEFAULT_BITROT_BLOCK_SIZE as u32 + 1
));
}
#[test]
@@ -735,12 +718,7 @@ mod tests {
#[test]
fn test_save_load_roundtrip() {
let tmp = tempfile::TempDir::new().unwrap();
let path = tmp
.path()
.join("vol.ecsum")
.to_str()
.unwrap()
.to_string();
let path = tmp.path().join("vol.ecsum").to_str().unwrap().to_string();
let mut builder = ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64);
builder.write(b"hello world");
@@ -901,8 +879,7 @@ mod tests {
assert_eq!(resolve_status(&notfound, 0, 10, 4), BitrotStatus::Off);
// Integrity failure => Invalid.
let bad: Result<EcBitrotProtection, BitrotLoadError> =
Err(BitrotLoadError::BadMagic(0));
let bad: Result<EcBitrotProtection, BitrotLoadError> = Err(BitrotLoadError::BadMagic(0));
assert_eq!(resolve_status(&bad, 0, 10, 4), BitrotStatus::Invalid);
// Generation mismatch => Off.
@@ -67,72 +67,45 @@ pub fn find_dat_file_size_with_dirs(
Ok(dat_size)
}
/// Reconstruct a .dat file from EC data shards.
///
/// Reads from .ec00-.ec09 and writes a new .dat file. All data shards
/// must live in `dir`. For the cross-disk reconciled layout where
/// shards are split across multiple data dirs of the same node, use
/// [`write_dat_file_from_shards_with_dirs`] instead.
pub fn write_dat_file_from_shards(
dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let dirs: Vec<String> = (0..data_shards).map(|_| dir.to_string()).collect();
write_dat_file_from_shards_with_dirs(
dir,
collection,
volume_id,
dat_file_size,
encoded_dat_file_size,
data_shards,
&dirs,
large_block_size,
small_block_size,
)
}
/// Reconstruct a .dat file from EC data shards, taking the source
/// directory for each shard separately.
///
/// `dat_dir` is where the produced `.dat` is written. `shard_dirs[i]`
/// is the directory holding shard `i`. For the simple "all shards in
/// one dir" case both can be the same value.
/// What it takes to rebuild a volume's .dat from its EC data shards.
///
/// Mirrors Go's `WriteDatFile(baseFileName, datFileSize,
/// encodedDatFileSize, shardFileNames)` shape — Go passes per-shard
/// paths so a reconciled volume with shards split across disks of the
/// same volume server can still be decoded back to a regular .dat
/// (seaweedfs/seaweedfs#9252).
#[derive(Clone, Copy, Debug)]
pub struct DatRebuild<'a> {
/// Where the produced `.dat` is written.
pub dat_dir: &'a str,
pub collection: &'a str,
pub volume_id: VolumeId,
/// The number of bytes to write, i.e. the live data extent from
/// [`find_dat_file_size`].
pub dat_file_size: i64,
/// The .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Zero when the .vif does not
/// record the encode-time size; the layout is then inferred from the
/// shard size.
pub encoded_dat_file_size: i64,
pub data_shards: usize,
/// `shard_dirs[i]` is the directory holding shard `i`. `None` means every
/// data shard sits in `dat_dir`.
pub shard_dirs: Option<&'a [String]>,
/// The volume's shard block layout, e.g. `EcVolume::large_block_size()`
/// / `small_block_size()` from its .vif EC config.
pub large_block_size: usize,
pub small_block_size: usize,
}
/// Reconstruct a .dat file from EC data shards.
///
/// `dat_file_size` is the number of bytes to write, i.e. the live data
/// extent from [`find_dat_file_size`]. `encoded_dat_file_size` is the
/// .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Pass zero when the .vif does
/// not record the encode-time size to infer the layout from the shard
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
write_dat_file(
/// Reads from .ec00-.ec09 and writes a new .dat file, from one directory or
/// from the per-shard directories of a cross-disk reconciled volume.
pub fn write_dat_file_from_shards(spec: &DatRebuild<'_>) -> io::Result<()> {
let DatRebuild {
dat_dir,
collection,
volume_id,
@@ -142,21 +115,15 @@ pub fn write_dat_file_from_shards_with_dirs(
shard_dirs,
large_block_size,
small_block_size,
)
}
#[allow(clippy::too_many_arguments)]
fn write_dat_file(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
} = *spec;
let same_dir: Vec<String>;
let shard_dirs: &[String] = match shard_dirs {
Some(dirs) => dirs,
None => {
same_dir = vec![dat_dir.to_string(); data_shards];
&same_dir
}
};
if data_shards == 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
@@ -233,10 +200,10 @@ fn write_dat_file(
// Read large blocks
while encoded_remaining >= large_row_size && remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = large_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -255,10 +222,10 @@ fn write_dat_file(
// Read small blocks
while remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = small_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -324,10 +291,7 @@ pub fn write_idx_file_from_ec_index(
// and treat only NotFound as "no journal": Path::exists would also
// swallow a permission/IO error and silently skip deletions, which
// would resurrect deleted needles as live.
let mut idx_file = std::fs::OpenOptions::new()
.write(true)
.append(true)
.open(&tmp_path)?;
let mut idx_file = std::fs::OpenOptions::new().append(true).open(&tmp_path)?;
match std::fs::read(&ecj_path) {
Ok(ecj_data) => {
let count = ecj_data.len() / NEEDLE_ID_SIZE;
@@ -372,7 +336,7 @@ mod tests {
use crate::storage::erasure_coding::ec_encoder;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -384,13 +348,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -430,16 +390,17 @@ mod tests {
std::fs::remove_file(format!("{}/1.idx", dir)).unwrap();
// Reconstruct from EC shards
write_dat_file_from_shards(
dir,
"",
VolumeId(1),
original_dat_size as i64,
original_dat_size as i64,
write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(1),
dat_file_size: original_dat_size as i64,
encoded_dat_file_size: original_dat_size as i64,
data_shards,
block_size as usize,
block_size as usize,
)
shard_dirs: None,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
})
.unwrap();
write_idx_file_from_ec_index(dir, "", VolumeId(1)).unwrap();
@@ -459,13 +420,9 @@ mod tests {
let v2 = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -485,29 +442,29 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
// No shard files exist, so de-striping must fail and publish nothing:
// neither the final .dat nor a partial .dat.tmp may remain.
let res = write_dat_file_from_shards(
dir,
"",
VolumeId(7),
100,
100,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
);
let res = write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(7),
dat_file_size: 100,
encoded_dat_file_size: 100,
data_shards: 10,
shard_dirs: None,
large_block_size: ERASURE_CODING_LARGE_BLOCK_SIZE,
small_block_size: ERASURE_CODING_SMALL_BLOCK_SIZE,
});
assert!(res.is_err());
assert!(!std::path::Path::new(&format!("{}/7.dat", dir)).exists());
assert!(!std::path::Path::new(&format!("{}/7.dat.tmp", dir)).exists());
}
// Decoding when .vif does not record the encode-time size: the layout is
// inferred from the shard size, except when that is an exact large-block
// multiple and the live extent reaches the ambiguous region.
#[test]
fn test_write_dat_file_fallback_layout() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -545,11 +502,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -567,7 +526,17 @@ mod tests {
-> io::Result<Vec<u8>> {
let out = format!("{}/{}", dir, sub);
std::fs::create_dir_all(&out).unwrap();
write_dat_file(&out, "", VolumeId(1), live, encoded, 10, shard_dirs, LARGE, SMALL)?;
write_dat_file_from_shards(&DatRebuild {
dat_dir: &out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live,
encoded_dat_file_size: encoded,
data_shards: 10,
shard_dirs: Some(shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})?;
Ok(std::fs::read(format!("{}/1.dat", out)).unwrap())
};
@@ -581,14 +550,20 @@ mod tests {
// each shard exactly one large block, indistinguishable from one large row
let (dir, shard_dirs, _) = encode("ambig1", large_row_size - 1);
let err = decode_to(&dir, "out", large_row_size / 2, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
// two-row equivalent: decoding within the agreed prefix still works
let (dir, shard_dirs, original) = encode("ambig2", 2 * large_row_size - 1);
let decoded = decode_to(&dir, "outa", large_row_size, 0, &shard_dirs).unwrap();
assert_eq!(&original[..large_row_size as usize], &decoded[..]);
let err = decode_to(&dir, "outb", large_row_size + 1, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
}
// Decoding after deletions moved the live extent below the large-block row
@@ -597,7 +572,7 @@ mod tests {
#[test]
fn test_write_dat_file_after_tail_deletion() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -637,11 +612,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -657,17 +634,17 @@ mod tests {
std::fs::create_dir(&out_dir).unwrap();
let out = out_dir.to_str().unwrap();
let decode = |live_size: i64, encoded_size: i64| -> Vec<u8> {
write_dat_file(
out,
"",
VolumeId(1),
live_size,
encoded_size,
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live_size,
encoded_dat_file_size: encoded_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.unwrap();
let path = format!("{}/1.dat", out);
let decoded = std::fs::read(&path).unwrap();
@@ -702,17 +679,19 @@ mod tests {
assert_ne!(&original[..(large_row_size / 2) as usize], &control[..]);
// the live extent can never exceed the encode-time size
assert!(write_dat_file(
out,
"",
VolumeId(1),
dat_size + 1,
dat_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
.is_err());
assert!(
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: dat_size + 1,
encoded_dat_file_size: dat_size,
data_shards,
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.is_err()
);
}
}
@@ -5,16 +5,12 @@
use std::fs::File;
use std::io;
#[cfg(not(unix))]
use std::io::{Read, Seek, SeekFrom};
use reed_solomon_erasure::galois_8::ReedSolomon;
use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums,
};
use crate::pb::volume_server_pb::{ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums};
use crate::storage::erasure_coding::ec_bitrot::{
self, ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
self, DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::idx;
@@ -50,7 +46,7 @@ pub fn write_ec_files(
let dat_size = dat_file.metadata()?.len() as i64;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Create shard files
let total_shards = data_shards + parity_shards;
@@ -77,11 +73,13 @@ pub fn write_ec_files(
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
ENCODE_BUFFER_SIZE,
block_size as usize,
block_size as usize,
EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: ENCODE_BUFFER_SIZE,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
},
)?;
// Close all shards
@@ -162,7 +160,7 @@ pub fn rebuild_ec_files(
}
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
@@ -175,7 +173,7 @@ pub fn rebuild_ec_files(
let mut shard_size = 0;
for (i, shard) in shards.iter_mut().enumerate() {
if !missing_shard_ids.contains(&(i as u32)) {
if let Ok(_) = shard.open() {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
@@ -185,7 +183,7 @@ pub fn rebuild_ec_files(
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
if let Ok(_) = alt.open() {
if alt.open().is_ok() {
let size = alt.file_size();
if size > shard_size {
shard_size = size;
@@ -251,12 +249,8 @@ pub fn rebuild_ec_files(
}
// Reconstruct missing shards
rs.reconstruct(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon reconstruct: {:?}", e),
)
})?;
rs.reconstruct(&mut buffers)
.map_err(|e| io::Error::other(format!("reed-solomon reconstruct: {:?}", e)))?;
// Write recovered data into the missing shards
for i in missing_shard_ids {
@@ -296,7 +290,7 @@ pub fn verify_ec_shards(
parity_shards: usize,
) -> io::Result<(Vec<u32>, Vec<String>)> {
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<Option<EcVolumeShard>> = (0..total_shards)
@@ -378,27 +372,27 @@ pub fn verify_ec_shards(
if !read_failed {
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
if !is_valid {
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
if let Ok(is_valid) = rs.verify(&slice_ptrs)
&& !is_valid
{
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for i in data_shards..total_shards {
verify_buffers[i].fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for buf in &mut verify_buffers[data_shards..total_shards] {
buf.fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
}
}
@@ -431,22 +425,23 @@ pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::R
// Read all idx entries
let mut idx_file = File::open(idx_path)?;
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
let mut last: std::collections::HashMap<NeedleId, (Offset, Size)> =
std::collections::HashMap::new();
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
entries.push((key, offset, size));
last.insert(key, (offset, size));
Ok(())
})?;
// Sort by NeedleId, then by actual offset so later entries come last
entries.sort_by_key(|&(key, offset, _)| (key, offset.to_actual_offset()));
// Remove duplicates (keep last/latest entry for each key).
// dedup_by_key keeps the first in each run, so we reverse first,
// dedup, then reverse back.
entries.reverse();
entries.dedup_by_key(|entry| entry.0);
entries.reverse();
let mut entries: Vec<(NeedleId, Offset, Size)> = last
.into_iter()
.filter_map(|(key, (offset, size))| {
if size.is_deleted() || offset.is_zero() {
None
} else {
Some((key, offset, size))
}
})
.collect();
entries.sort_by_key(|&(key, _o, _s)| key);
// Write sorted entries to .ecx
let mut ecx_file = File::create(ecx_path)?;
@@ -490,7 +485,7 @@ pub fn rebuild_ecx_file(
.collect();
for (i, shard) in shards.iter_mut().enumerate() {
if let Err(_) = shard.open() {
if shard.open().is_err() {
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
@@ -507,7 +502,7 @@ pub fn rebuild_ecx_file(
}
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("cannot open data shard for ecx rebuild"),
"cannot open data shard for ecx rebuild".to_string(),
));
}
}
@@ -515,7 +510,7 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
let total_data_size = shard_size * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
@@ -538,7 +533,7 @@ pub fn rebuild_ecx_file(
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
(shard_size - 1).max(0)
};
// Read version from superblock (first byte of logical data)
@@ -587,7 +582,8 @@ pub fn rebuild_ecx_file(
}
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
let needle_id = NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let needle_id =
NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let size = Size::from_bytes(&header_buf[COOKIE_SIZE + NEEDLE_ID_SIZE..header_size]);
// Validate: stop if we hit zero cookie+id (end of data)
@@ -640,7 +636,6 @@ pub fn rebuild_ecx_file(
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
@@ -707,30 +702,50 @@ fn read_from_data_shards(
/// the uniform block is.
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// Shape of one encode run: the Reed-Solomon split and the block sizes that
/// fix where every byte of the .dat lands in the shards. Mirrors Go's
/// `ECContext`. `buffer_size` must divide both block sizes.
#[derive(Clone, Copy, Debug)]
pub(crate) struct EcEncodeLayout {
pub(crate) data_shards: usize,
pub(crate) parity_shards: usize,
/// Bytes of each shard's block handled per sub-batch; bounds memory at
/// `total_shards * buffer_size` however large the blocks are.
pub(crate) buffer_size: usize,
pub(crate) large_block_size: usize,
pub(crate) small_block_size: usize,
}
/// Encode the .dat file data into shard files.
///
/// Uses a two-phase approach matching Go's ec_encoder.go:
/// 1. Process as many large blocks as possible
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
dat_size: i64,
rs: &ReedSolomon,
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
buffer_size: usize,
large_block_size: usize,
small_block_size: usize,
layout: EcEncodeLayout,
) -> io::Result<()> {
let EcEncodeLayout {
data_shards,
parity_shards,
buffer_size,
large_block_size,
small_block_size,
} = layout;
let total_shards = data_shards + parity_shards;
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
.map(|_| vec![0u8; buffer_size])
.collect();
let mut buffers: Vec<Vec<u8>> = (0..total_shards).map(|_| vec![0u8; buffer_size]).collect();
let mut run = EncodeRun {
dat_file,
rs,
buffers: &mut buffers,
shards,
builders,
data_shards,
};
let mut remaining = dat_size;
let mut offset: u64 = 0;
@@ -739,16 +754,7 @@ pub(crate) fn encode_dat_file(
let large_row_size = large_block_size * data_shards;
while remaining >= large_row_size as i64 {
encode_data(
dat_file,
offset,
large_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, large_block_size)?;
offset += large_row_size as u64;
remaining -= large_row_size as i64;
}
@@ -758,16 +764,7 @@ pub(crate) fn encode_dat_file(
while remaining > 0 {
let to_process = remaining.min(small_row_size as i64);
encode_data(
dat_file,
offset,
small_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, small_block_size)?;
offset += to_process as u64;
remaining -= to_process;
}
@@ -775,102 +772,73 @@ pub(crate) fn encode_dat_file(
Ok(())
}
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
/// Everything one encode run streams through: the source .dat, the codec, a
/// buffer per shard, and the per-shard file and checksum sinks.
struct EncodeRun<'a> {
dat_file: &'a File,
rs: &'a ReedSolomon,
buffers: &'a mut [Vec<u8>],
shards: &'a mut [EcVolumeShard],
builders: &'a mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
encode_one_batch(
dat_file,
row_offset + (b * buffer_size) as u64,
block_size,
rs,
buffers,
shards,
builders,
data_shards,
)?;
}
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
impl EncodeRun<'_> {
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches
/// so arbitrarily large blocks never require block-sized allocations.
/// Mirrors Go's encodeData.
fn encode_row(&mut self, row_offset: u64, block_size: usize) -> io::Result<()> {
let buffer_size = self.buffers[0].len();
if !block_size.is_multiple_of(buffer_size) {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
self.encode_one_batch(row_offset + (b * buffer_size) as u64, block_size)?;
}
Ok(())
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
)
})?;
/// Encode one sub-batch: the same buffer-sized slice of every shard's block
/// in this row. Mirrors Go's encodeDataOneBatch.
fn encode_one_batch(&mut self, offset: u64, block_size: usize) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the
// buffers are reused across batches, so the tail must be cleared
// explicitly.
for (i, buf) in self.buffers[..self.data_shards].iter_mut().enumerate() {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(self.dat_file, buf, read_offset)?;
buf[n..].fill(0);
}
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
for (i, buf) in buffers.iter().enumerate() {
shards[i].write_all(buf)?;
builders[i].write(buf);
// Encode parity shards
self.rs
.encode(&mut *self.buffers)
.map_err(|e| io::Error::other(format!("reed-solomon encode: {:?}", e)))?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk
// length.
for (i, buf) in self.buffers.iter().enumerate() {
self.shards[i].write_all(buf)?;
self.builders[i].write(buf);
}
Ok(())
}
Ok(())
}
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let mut n = 0;
while n < buf.len() {
#[cfg(unix)]
let r = {
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buf[n..], offset + n as u64)?
};
#[cfg(not(unix))]
let r = {
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(offset + n as u64))?;
f.read(&mut buf[n..])?
};
let r = crate::storage::io::read_at(dat_file, &mut buf[n..], offset + n as u64)?;
if r == 0 {
break;
}
@@ -884,7 +852,7 @@ mod tests {
use super::*;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -896,13 +864,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -947,13 +911,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=n {
@@ -1026,13 +986,9 @@ mod tests {
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=20 {
@@ -1064,7 +1020,10 @@ mod tests {
let victim = format!("{}/1.ec03", dir);
let full = std::fs::metadata(&victim).unwrap().len();
assert!(full > 0, "encoded shard should be non-empty");
let f = std::fs::OpenOptions::new().write(true).open(&victim).unwrap();
let f = std::fs::OpenOptions::new()
.write(true)
.open(&victim)
.unwrap();
f.set_len(full / 2).unwrap();
drop(f);
@@ -1218,19 +1177,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_uniform_layout() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1260,7 +1215,10 @@ mod tests {
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
let rebuilt = std::fs::read(&ecx_path).unwrap();
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
assert_eq!(
canonical, rebuilt,
"rebuilt .ecx must match the encode-time .ecx"
);
}
// A truncated data shard must FAIL the .ecx rebuild, not publish the
@@ -1268,19 +1226,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(3),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1378,13 +1332,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1462,13 +1412,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1500,13 +1446,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
vid,
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -1597,4 +1539,86 @@ mod tests {
details
);
}
#[test]
fn test_encode_drops_tombstone_last_wins() {
use crate::storage::idx;
use crate::storage::types::{NeedleId, Offset, Size, TOMBSTONE_FILE_SIZE};
let tmp = tempfile::TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let idx_path = format!("{}/t.idx", dir);
let ecx_path = format!("{}/t.ecx", dir);
let key = NeedleId(12345);
{
let mut f = std::fs::File::create(&idx_path).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
}
super::write_sorted_ecx_from_idx(&idx_path, &ecx_path).unwrap();
let mut found = false;
{
let mut f = std::fs::File::open(&ecx_path).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found = true;
}
Ok(())
})
.unwrap();
}
assert!(!found, "tombstoned key must not appear in .ecx");
let idx2 = format!("{}/t2.idx", dir);
let ecx2 = format!("{}/t2.ecx", dir);
{
let mut f = std::fs::File::create(&idx2).unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(2048), Size(200))
.unwrap();
}
super::write_sorted_ecx_from_idx(&idx2, &ecx2).unwrap();
let mut found2 = false;
{
let mut f = std::fs::File::open(&ecx2).unwrap();
idx::walk_index_file(&mut f, 0, |k, o, s| {
if k == key {
found2 = true;
assert_eq!(o.to_actual_offset(), 2048);
assert_eq!(s, Size(200));
}
Ok(())
})
.unwrap();
}
assert!(found2, "re-created key must appear live");
// Zero offset with non-negative size is also a deletion: Go
// readNeedleMap (`if !offset.IsZero() && !size.IsDeleted() { Set }
// else { Delete }`) and CompactNeedleMap::load_from_idx both treat
// it as deleted. Encode must drop it too, or the .ecx live-map
// mismatches replay.
let idx3 = format!("{}/t3.idx", dir);
let ecx3 = format!("{}/t3.ecx", dir);
{
let mut f = std::fs::File::create(&idx3).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), Size(0)).unwrap();
}
super::write_sorted_ecx_from_idx(&idx3, &ecx3).unwrap();
let mut found3 = false;
{
let mut f = std::fs::File::open(&ecx3).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found3 = true;
}
Ok(())
})
.unwrap();
}
assert!(
!found3,
"zero-offset row must not appear in .ecx even with non-negative size"
);
}
}
@@ -16,6 +16,20 @@ pub const ERASURE_CODING_SMALL_BLOCK_SIZE: usize = 1024 * 1024; // 1MB
pub type ShardId = u8;
/// Validate a wire shard id. `ShardId` is `u8` but only 0..MAX_SHARD_COUNT are valid.
/// Rejects 256 (would truncate to 0 and delete .ec00) and 270 (would alias 14).
pub fn shard_id_try_from(v: u32) -> Result<ShardId, String> {
if v < MAX_SHARD_COUNT as u32 {
Ok(v as ShardId)
} else {
Err(format!(
"invalid shard id {} (max {})",
v,
MAX_SHARD_COUNT - 1
))
}
}
/// A single erasure-coded shard file.
pub struct EcVolumeShard {
pub volume_id: VolumeId,
@@ -78,23 +92,9 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(not(unix))]
{
use std::io::{Read, Seek, SeekFrom};
// File::read_at is unix-only; fall back to seek + read.
// We need a mutable reference for seek/read, so clone the handle.
let mut f = file.try_clone()?;
f.seek(SeekFrom::Start(offset))?;
f.read(buf)
}
crate::storage::io::read_at(file, buf, offset)
}
/// Write data to the shard file (appends).
@@ -102,7 +102,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
file.write_all(data)?;
self.ecd_file_size += data.len() as i64;
Ok(())
@@ -123,7 +123,7 @@ impl EcVolumeShard {
pub fn try_clone_file(&self) -> io::Result<File> {
self.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?
.ok_or_else(|| io::Error::other("shard file not open"))?
.try_clone()
}
@@ -251,4 +251,42 @@ mod tests {
let shard = EcVolumeShard::new("/data", "", VolumeId(7), 13);
assert_eq!(shard.file_name(), "/data/7.ec13");
}
#[test]
fn test_shard_id_try_from_u32_rejects_overflow() {
use super::{MAX_SHARD_COUNT, shard_id_try_from};
assert_eq!(shard_id_try_from(0).unwrap(), 0u8);
assert_eq!(shard_id_try_from(14).unwrap(), 14u8);
assert_eq!(shard_id_try_from(31).unwrap(), 31u8);
assert!(shard_id_try_from(32).is_err());
assert!(shard_id_try_from(256).is_err());
assert!(shard_id_try_from(270).is_err());
assert!(shard_id_try_from(u32::MAX).is_err());
assert_eq!(MAX_SHARD_COUNT, 32);
}
#[test]
fn test_shard_batch_validation_is_atomic_rejects_without_partial_prefix() {
use super::shard_id_try_from;
// The mount/unmount handlers pre-validate the ENTIRE req.shard_ids into
// a Vec<ShardId> BEFORE acquiring the write lock or mutating any EC
// state. This test pins the validation half of that contract at the
// unit level: a batch like [0, 32] must fail as a whole, so by
// construction no validated prefix (e.g. shard 0) is ever applied.
// The handler-level tests below assert the no-state-change half.
let batch = vec![0u32, 32u32];
let validated: Result<Vec<_>, _> =
batch.iter().map(|&sid| shard_id_try_from(sid)).collect();
assert!(
validated.is_err(),
"batch {:?} must be rejected as a whole",
batch
);
// A fully-valid batch still validates cleanly.
let ok: Result<Vec<_>, _> = [0u32, 1u32, 13u32]
.iter()
.map(|&sid| shard_id_try_from(sid))
.collect();
assert_eq!(ok.unwrap(), vec![0u8, 1u8, 13u8]);
}
}
@@ -12,7 +12,8 @@ use std::time::{SystemTime, UNIX_EPOCH};
use crate::pb::master_pb;
use crate::storage::erasure_coding::ec_locate;
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::needle::needle::{get_actual_size, Needle, NeedleError};
use crate::storage::io::read_exact_at;
use crate::storage::needle::needle::{Needle, NeedleError, get_actual_size};
use crate::storage::types::*;
use crate::storage::volume_open::open_volume_file;
@@ -565,32 +566,31 @@ impl EcVolume {
// A sidecar written for THIS generation that contradicts the volume's
// geometry is not "no protection" — it says the layout the volume is
// about to serve reads with is wrong. Fail the mount.
if let Ok(prot) = &loaded {
if prot.generation == generation
&& !ec_bitrot::geometry_matches(
prot,
self.data_shards as usize,
self.parity_shards as usize,
if let Ok(prot) = &loaded
&& prot.generation == generation
&& !ec_bitrot::geometry_matches(
prot,
self.data_shards as usize,
self.parity_shards as usize,
self.block_size,
)
{
let cfg = prot.ec_shard_config.as_ref();
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"ec volume {} generation {}: {} records layout {}+{} block {} but the volume is mounted as {}+{} block {}; refusing to serve one of the two layouts",
self.volume_id.0,
generation,
path,
cfg.map(|c| c.data_shards).unwrap_or(0),
cfg.map(|c| c.parity_shards).unwrap_or(0),
cfg.map(|c| c.block_size).unwrap_or(0),
self.data_shards,
self.parity_shards,
self.block_size,
)
{
let cfg = prot.ec_shard_config.as_ref();
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"ec volume {} generation {}: {} records layout {}+{} block {} but the volume is mounted as {}+{} block {}; refusing to serve one of the two layouts",
self.volume_id.0,
generation,
path,
cfg.map(|c| c.data_shards).unwrap_or(0),
cfg.map(|c| c.parity_shards).unwrap_or(0),
cfg.map(|c| c.block_size).unwrap_or(0),
self.data_shards,
self.parity_shards,
self.block_size,
),
));
}
),
));
}
let status = ec_bitrot::resolve_status(
&loaded,
@@ -698,7 +698,7 @@ impl EcVolume {
let mut set = self
.deleted_needles
.write()
.map_err(|_| io::Error::new(io::ErrorKind::Other, "deleted_needles lock poisoned"))?;
.map_err(|_| io::Error::other("deleted_needles lock poisoned"))?;
let mut off: i64 = 0;
while off + NEEDLE_ID_SIZE as i64 <= self.ecj_file_size {
#[cfg(unix)]
@@ -706,6 +706,17 @@ impl EcVolume {
use std::os::unix::fs::FileExt;
ecj_file.read_exact_at(&mut buf, off as u64)?;
}
#[cfg(windows)]
{
// Positional read so concurrent readers of the shared .ecj
// handle can't interleave seek/read. Mirrors the
// read_exact_at helper at the bottom of this file.
read_exact_at(ecj_file, &mut buf, off as u64)?;
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
set.insert(NeedleId::from_bytes(&buf));
off += NEEDLE_ID_SIZE as i64;
}
@@ -744,7 +755,6 @@ impl EcVolume {
// ---- File names ----
#[allow(dead_code)]
fn base_name(&self) -> String {
crate::storage::volume::volume_file_name(&self.dir, &self.collection, self.volume_id)
}
@@ -826,19 +836,29 @@ impl EcVolume {
/// default to the physical location's disk type.
pub fn set_disk_type(&mut self, d: DiskType) {
self.disk_type = d.clone();
for slot in self.shards.iter_mut() {
if let Some(shard) = slot {
shard.disk_type = d.clone();
}
for shard in self.shards.iter_mut().flatten() {
shard.disk_type = d.clone();
}
}
/// Remove and close a shard.
pub fn remove_shard(&mut self, shard_id: ShardId) {
if let Some(ref mut shard) = self.shards[shard_id as usize] {
pub fn remove_shard(&mut self, shard_id: ShardId) -> io::Result<()> {
let idx = shard_id as usize;
if idx >= self.shards.len() {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"invalid shard id {} (max {})",
shard_id,
self.shards.len().saturating_sub(1)
),
));
}
if let Some(ref mut shard) = self.shards[idx] {
shard.close();
}
self.shards[shard_id as usize] = None;
self.shards[idx] = None;
Ok(())
}
/// Get a ShardBits bitmap of locally available shards.
@@ -860,7 +880,7 @@ impl EcVolume {
/// Reports whether `shard_id` is currently registered to this
/// EcVolume (used by the cross-disk reconcile to skip already-
/// loaded shards).
pub fn has_shard(&self, shard_id: u8) -> bool {
pub fn has_shard(&self, shard_id: ShardId) -> bool {
self.shards
.get(shard_id as usize)
.map(|s| s.is_some())
@@ -986,21 +1006,21 @@ impl EcVolume {
pub fn check_read_write_error(&self, err: Option<&io::Error>) {
use std::sync::atomic::Ordering;
if let Some(e) = err {
if crate::storage::volume::is_storage_io_error(e) {
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
if let Some(e) = err
&& crate::storage::volume::is_storage_io_error(e)
{
self.io_error_count.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
}
self.io_error_count.store(0, Ordering::Relaxed);
if let Ok(mut guard) = self.last_io_error.lock() {
if guard.is_some() {
*guard = None;
}
if let Ok(mut guard) = self.last_io_error.lock()
&& guard.is_some()
{
*guard = None;
}
}
@@ -1033,7 +1053,7 @@ impl EcVolume {
let ecx_file = self
.ecx_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecx file not open"))?;
.ok_or_else(|| io::Error::other("ecx file not open"))?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
if entry_count == 0 {
@@ -1049,25 +1069,12 @@ impl EcVolume {
let mid = lo + (hi - lo) / 2;
let file_offset = (mid * NEEDLE_MAP_ENTRY_SIZE) as u64;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
if let Err(e) = ecx_file.read_exact_at(&mut entry_buf, file_offset) {
self.check_read_write_error(Some(&e));
return Err(e);
}
}
#[cfg(not(unix))]
{
use std::io::{Read, Seek, SeekFrom};
if let Err(e) = ecx_file.seek(SeekFrom::Start(file_offset)) {
self.check_read_write_error(Some(&e));
return Err(e);
}
if let Err(e) = ecx_file.read_exact(&mut entry_buf) {
self.check_read_write_error(Some(&e));
return Err(e);
}
// Positional read so concurrent find_needle_from_ecx calls on the
// shared .ecx handle don't interleave a seek and a read and corrupt
// each other's binary search.
if let Err(e) = read_exact_at(ecx_file, &mut entry_buf, file_offset) {
self.check_read_write_error(Some(&e));
return Err(e);
}
let (key, offset, size) = idx_entry_from_bytes(&entry_buf);
@@ -1250,10 +1257,8 @@ impl EcVolume {
/// Get the size of a single shard (all shards are the same size).
fn shard_file_size(&self) -> i64 {
for shard in &self.shards {
if let Some(s) = shard {
return s.file_size();
}
if let Some(s) = self.shards.iter().flatten().next() {
return s.file_size();
}
0
}
@@ -1279,6 +1284,17 @@ impl EcVolume {
use std::os::unix::fs::FileExt;
ecx_file.read_exact_at(&mut entry_buf, file_offset)?;
}
#[cfg(windows)]
{
// Positional read so concurrent readers of the shared .ecx
// handle can't interleave seek/read. Mirrors the
// read_exact_at helper at the bottom of this file.
read_exact_at(ecx_file, &mut entry_buf, file_offset)?;
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
let (_key, _offset, size) = idx_entry_from_bytes(&entry_buf);
// Match Go's Size.Raw(): tombstone (-1) returns 0, other negatives return abs
if !size.is_tombstone() {
@@ -1360,13 +1376,10 @@ impl EcVolume {
/// the index (ignored by callers) and an error on IO failure.
fn tombstone_ecx_entry(&self, needle_id: NeedleId) -> io::Result<bool> {
let ecx_file = self.ecx_file.as_ref().ok_or_else(|| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec volume {} has no open .ecx file (closed or corrupt)",
self.volume_id.0
),
)
io::Error::other(format!(
"ec volume {} has no open .ecx file (closed or corrupt)",
self.volume_id.0
))
})?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
@@ -1385,6 +1398,17 @@ impl EcVolume {
use std::os::unix::fs::FileExt;
ecx_file.read_exact_at(&mut entry_buf, file_offset)?;
}
#[cfg(windows)]
{
// Positional read so concurrent readers of the shared .ecx
// handle can't interleave seek/read. Mirrors the
// read_exact_at helper at the bottom of this file.
read_exact_at(ecx_file, &mut entry_buf, file_offset)?;
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
let (key, _offset, _old_size) = idx_entry_from_bytes(&entry_buf);
if key == needle_id {
let size_offset = file_offset + NEEDLE_ID_SIZE as u64 + OFFSET_SIZE as u64;
@@ -1395,6 +1419,31 @@ impl EcVolume {
use std::os::unix::fs::FileExt;
ecx_file.write_all_at(&size_buf, size_offset)?;
}
#[cfg(windows)]
{
// Positional write so concurrent readers of the shared
// .ecx handle can't observe a moved cursor. Mirrors the
// read_exact_at helper at the bottom of this file, with
// seek_write in place of seek_read.
use std::os::windows::fs::FileExt;
let mut written = 0;
let mut at = size_offset;
while written < size_buf.len() {
let n = ecx_file.seek_write(&size_buf[written..], at)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::WriteZero,
"seek_write wrote nothing",
));
}
written += n;
at += n as u64;
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
return Ok(true);
} else if key < needle_id {
lo = mid + 1;
@@ -1413,7 +1462,7 @@ impl EcVolume {
/// `deleted_needles` instead). The rebuild is atomic with respect to
/// the journal: if any individual write fails the .ecj file is left
/// in place and the error is propagated so tombstones are not lost.
#[allow(dead_code)]
#[expect(dead_code, reason = "no caller yet; see the doc comment")]
fn rebuild_ecx_from_journal(&mut self) -> io::Result<()> {
let ecj_path = self.ecj_file_name();
if !std::path::Path::new(&ecj_path).exists() {
@@ -1507,7 +1556,7 @@ impl EcVolume {
let ecj_file = self
.ecj_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecj file not open"))?;
.ok_or_else(|| io::Error::other("ecj file not open"))?;
let mut buf = [0u8; NEEDLE_ID_SIZE];
needle_id.to_bytes(&mut buf);
ecj_file.write_all(&buf).and_then(|_| ecj_file.sync_all())
@@ -1525,15 +1574,15 @@ impl EcVolume {
// write_all may have extended the file on disk before
// sync_all failed; truncate back to the known-good size so
// the on-disk journal never drifts past `deleted_needles`.
if let Some(ecj) = self.ecj_file.as_mut() {
if let Err(trunc_err) = ecj.set_len(prev_ecj_size as u64) {
tracing::error!(
volume_id = self.volume_id.0,
needle_id = needle_id.0,
truncate_error = %trunc_err,
"failed to truncate ecj after append failure"
);
}
if let Some(ecj) = self.ecj_file.as_mut()
&& let Err(trunc_err) = ecj.set_len(prev_ecj_size as u64)
{
tracing::error!(
volume_id = self.volume_id.0,
needle_id = needle_id.0,
truncate_error = %trunc_err,
"failed to truncate ecj after append failure"
);
}
Err(e)
}
@@ -1547,7 +1596,7 @@ impl EcVolume {
let ecx_file = self
.ecx_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "ecx file not open"))?;
.ok_or_else(|| io::Error::other("ecx file not open"))?;
let entry_count = self.ecx_file_size as usize / NEEDLE_MAP_ENTRY_SIZE;
if entry_count == 0 {
return Ok(None);
@@ -1563,6 +1612,18 @@ impl EcVolume {
use std::os::unix::fs::FileExt;
ecx_file.read_exact_at(&mut entry_buf, file_offset)?;
}
#[cfg(windows)]
{
// Positional read so concurrent find_needle_from_ecx_raw calls
// on the shared .ecx handle don't interleave seek/read and
// corrupt each other's binary search. Mirrors the
// read_exact_at helper at the bottom of this file.
read_exact_at(ecx_file, &mut entry_buf, file_offset)?;
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
let (key, offset, size) = idx_entry_from_bytes(&entry_buf);
if key == needle_id {
return Ok(Some((offset, size)));
@@ -1585,38 +1646,63 @@ impl EcVolume {
) -> io::Result<()> {
// cookie == 0 indicates SkipCookieCheck was requested
if cookie.0 != 0 {
// Try to read the needle's cookie from the EC shards to validate
// Look up the needle in ecx index to find its offset, then read header from shard
if let Ok(Some((offset, size))) = self.find_needle_from_ecx(needle_id) {
if !size.is_deleted() && !offset.is_zero() {
let actual_offset = offset.to_actual_offset() as u64;
// Determine which shard contains this offset and read the cookie
let shard_size = self
.shards
.iter()
.filter_map(|s| s.as_ref())
.map(|s| s.file_size())
.next()
.unwrap_or(0) as u64;
if shard_size > 0 {
let shard_id = (actual_offset / shard_size) as usize;
let shard_offset = actual_offset % shard_size;
if let Some(Some(shard)) = self.shards.get(shard_id) {
let mut header_buf = [0u8; 4]; // cookie is first 4 bytes of needle
if shard.read_at(&mut header_buf, shard_offset).is_ok() {
let needle_cookie =
crate::storage::types::Cookie(u32::from_be_bytes(header_buf));
if needle_cookie != cookie {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("unexpected cookie {:x}", cookie.0),
));
}
}
}
let (offset, size) = match self.find_needle_from_ecx(needle_id)? {
Some((o, s)) => (o, s),
None => return self.journal_delete(needle_id),
};
if size.is_deleted() || offset.is_zero() {
return self.journal_delete(needle_id);
}
let actual_offset = offset.to_actual_offset();
let intervals = self.locate_ec_shard_needle_interval(actual_offset, size);
if intervals.is_empty() {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("cannot verify cookie for needle {}", needle_id.0),
));
}
let (shard_id, shard_offset) = self.interval_to_shard_id_and_offset(&intervals[0]);
let shard = self
.shards
.get(shard_id as usize)
.and_then(|s| s.as_ref())
.ok_or_else(|| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("cannot verify cookie: shard {} not local", shard_id),
)
})?;
// Retry short reads, but fail closed on EOF.
let mut header_buf = [0u8; 4];
let mut filled = 0usize;
while filled < header_buf.len() {
match shard.read_at(
&mut header_buf[filled..],
shard_offset as u64 + filled as u64,
) {
Ok(0) => break,
Ok(n) => filled += n,
Err(e) => {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("cannot verify cookie: {}", e),
));
}
}
}
if filled != header_buf.len() {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
"cannot verify cookie: incomplete header",
));
}
let needle_cookie = crate::storage::types::Cookie(u32::from_be_bytes(header_buf));
if needle_cookie != cookie {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
format!("unexpected cookie {:x}", cookie.0),
));
}
}
self.journal_delete(needle_id)
}
@@ -1726,20 +1812,19 @@ mod tests {
#[test]
fn test_destroy_removes_bitrot_sidecar() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"ec1c",
VolumeId(2074),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "ec1c",
..Default::default()
},
)
.unwrap();
for i in 1..=3 {
@@ -1802,20 +1887,16 @@ mod tests {
fn test_mount_loads_bitrot_sidecar() {
use crate::storage::erasure_coding::ec_bitrot::BitrotStatus;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=5 {
@@ -1860,20 +1941,16 @@ mod tests {
#[test]
fn test_scrub_plans_are_self_contained_and_match_direct_call() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -1953,20 +2030,16 @@ mod tests {
#[test]
fn test_local_scrub_plan_reports_negative_size_ecx_row() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -2045,20 +2118,16 @@ mod tests {
#[test]
fn test_scrub_plans_survive_files_removed_after_snapshot() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -2144,20 +2213,16 @@ mod tests {
#[test]
fn test_checksum_scrub_clean_and_detects_corruption() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -2289,6 +2354,147 @@ mod tests {
assert_eq!((fc, dc), (2, 2));
}
#[test]
fn test_journal_delete_wrong_cookie() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let needle = NeedleId(7);
let entries = vec![(needle, Offset::from_actual_offset(8), Size(100))];
write_ecx_file(dir, "", VolumeId(1), &entries);
let vif = crate::storage::volume::VifVolumeInfo {
dat_file_size: 14000,
..Default::default()
};
let base = crate::storage::volume::volume_file_name(dir, "", VolumeId(1));
std::fs::write(
format!("{}.vif", base),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
let mut shard9 = EcVolumeShard::new(dir, "", VolumeId(1), 9);
shard9.create().unwrap();
shard9.write_all(&[0xAAu8; 2048]).unwrap();
shard9.close();
let mut vol = EcVolume::new(dir, dir, "", VolumeId(1)).unwrap();
vol.add_shard(EcVolumeShard::new(dir, "", VolumeId(1), 9))
.unwrap();
let (off, size) = vol
.find_needle_from_ecx(needle)
.unwrap()
.expect("fixture needle must be indexed");
let intervals = vol.locate_ec_shard_needle_interval(off.to_actual_offset(), size);
assert!(
!intervals.is_empty(),
"fixture must locate to a shard for the test to be meaningful"
);
let (located, _) = vol.interval_to_shard_id_and_offset(&intervals[0]);
assert_ne!(
located, 9,
"fixture must locate away from mounted shard 9, got {}",
located
);
let res = vol.journal_delete_with_cookie(needle, Cookie(0xDEAD_BEEF));
let err = res.expect_err("wrong cookie must Err, not bypass to journal");
assert!(
err.to_string().contains("cannot verify cookie"),
"fail-closed error must say why, got: {}",
err
);
let deleted = vol.read_deleted_needles().unwrap();
assert!(
!deleted.contains(&needle),
"failed delete must not append to .ecj, got {:?}",
deleted
);
vol.journal_delete_with_cookie(needle, Cookie(0)).unwrap();
let deleted = vol.read_deleted_needles().unwrap();
assert!(
deleted.contains(&needle),
"cookie-0 delete must still journal, got {:?}",
deleted
);
}
#[test]
fn test_journal_delete_incomplete_header() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let needle = NeedleId(7);
let entries = vec![(needle, Offset::from_actual_offset(8), Size(100))];
write_ecx_file(dir, "", VolumeId(1), &entries);
let vif = crate::storage::volume::VifVolumeInfo {
dat_file_size: 14000,
..Default::default()
};
let base = crate::storage::volume::volume_file_name(dir, "", VolumeId(1));
std::fs::write(
format!("{}.vif", base),
serde_json::to_string_pretty(&vif).unwrap(),
)
.unwrap();
let probe = EcVolume::new(dir, dir, "", VolumeId(1)).unwrap();
let (off, size) = probe
.find_needle_from_ecx(needle)
.unwrap()
.expect("fixture needle must be indexed");
let intervals = probe.locate_ec_shard_needle_interval(off.to_actual_offset(), size);
assert!(
!intervals.is_empty(),
"fixture must locate to a shard for the test to be meaningful"
);
let (located, located_offset) = probe.interval_to_shard_id_and_offset(&intervals[0]);
drop(probe);
let mut shard = EcVolumeShard::new(dir, "", VolumeId(1), located);
shard.create().unwrap();
shard.write_all(&[0x00u8; 2]).unwrap();
shard.close();
let mut vol = EcVolume::new(dir, dir, "", VolumeId(1)).unwrap();
vol.add_shard(EcVolumeShard::new(dir, "", VolumeId(1), located))
.unwrap();
let shard_ref = vol.shards[located as usize]
.as_ref()
.expect("located shard must be mounted");
assert!(
(shard_ref.file_size()) < located_offset + 4,
"fixture must truncate the header read (file {} bytes, offset {})",
shard_ref.file_size(),
located_offset
);
let res = vol.journal_delete_with_cookie(needle, Cookie(0x1234));
let err = res.expect_err("short header read must Err, not forge-match");
assert!(
err.to_string().contains("incomplete header"),
"short read must report incomplete header, got: {}",
err
);
let deleted = vol.read_deleted_needles().unwrap();
assert!(
!deleted.contains(&needle),
"failed delete must not append to .ecj, got {:?}",
deleted
);
vol.journal_delete_with_cookie(needle, Cookie(0)).unwrap();
let deleted = vol.read_deleted_needles().unwrap();
assert!(
deleted.contains(&needle),
"cookie-0 delete must still journal, got {:?}",
deleted
);
}
#[test]
fn test_ec_volume_shard_bits() {
let tmp = TempDir::new().unwrap();
@@ -2594,18 +2800,14 @@ mod tests {
/// split-disk mount without needing N directories.
fn split_runtimes(dir: &str, vid: VolumeId, subsets: &[&[u8]]) -> Vec<EcVolume> {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let mut v = Volume::new(
dir,
dir,
"",
vid,
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -3312,7 +3514,7 @@ mod tests {
mod uniform_layout_tests {
use super::*;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::{VifEcShardConfig, VifVolumeInfo, Volume};
use crate::storage::volume::{VifEcShardConfig, VifVolumeInfo, Volume, VolumeSpec};
use tempfile::TempDir;
// Write ~26MB of needles so the uniform block size (3MB) diverges from the
@@ -3329,13 +3531,9 @@ mod uniform_layout_tests {
let mut v = Volume::new(
dir,
dir,
"",
vid,
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
let mut expected: Vec<(NeedleId, Vec<u8>)> = Vec::new();
@@ -3362,7 +3560,7 @@ mod uniform_layout_tests {
// Legacy fixture: two-tier encode plus a .vif without a block
// size, the state every pre-upgrade EC volume is in.
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
let base = crate::storage::volume::volume_file_name(dir, "", vid);
@@ -3388,11 +3586,13 @@ mod uniform_layout_tests {
&rs,
&mut shards,
&mut builders,
10,
4,
256 * 1024,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
crate::storage::erasure_coding::ec_encoder::EcEncodeLayout {
data_shards: 10,
parity_shards: 4,
buffer_size: 256 * 1024,
large_block_size: ERASURE_CODING_LARGE_BLOCK_SIZE,
small_block_size: ERASURE_CODING_SMALL_BLOCK_SIZE,
},
)
.unwrap();
for shard in &mut shards {
@@ -3739,10 +3939,10 @@ pub(crate) fn merge_ec_runtimes<'a>(runtimes: &[&'a EcVolume]) -> Option<MergedE
let mut slots: Vec<Option<(&'a EcVolume, &'a EcVolumeShard)>> = vec![None; width];
for v in &merged {
for (id, slot) in v.shards.iter().enumerate() {
if let Some(shard) = slot.as_ref() {
if slots[id].is_none() {
slots[id] = Some((*v, shard));
}
if let Some(shard) = slot.as_ref()
&& slots[id].is_none()
{
slots[id] = Some((*v, shard));
}
}
}
@@ -4097,18 +4297,7 @@ impl EcLocalShard {
.file
.as_ref()
.map_err(|e| io::Error::new(e.kind(), e.to_string()))?;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(not(unix))]
{
use std::io::{Read, Seek, SeekFrom};
let mut f = file.try_clone()?;
f.seek(SeekFrom::Start(offset))?;
f.read(buf)
}
crate::storage::io::read_at(file, buf, offset)
}
}
@@ -4363,13 +4552,10 @@ impl EcLocalScrubPlan {
if read != want {
// Like Go, returning from the walk callback aborts the scan.
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"expected {} bytes for needle {} on volume {}, got {}",
want, id.0, volume_id.0, read
),
));
return Err(io::Error::other(format!(
"expected {} bytes for needle {} on volume {}, got {}",
want, id.0, volume_id.0, read
)));
}
// Only a fully-local needle can be reassembled and CRC-checked.
@@ -4402,7 +4588,7 @@ impl EcLocalScrubPlan {
.filter_map(|sid| shards.get(*sid as usize).and_then(|s| s.as_ref()))
.map(|s| s.info.clone())
.collect();
broken.sort_by(|a, b| a.shard_id.cmp(&b.shard_id));
broken.sort_by_key(|a| a.shard_id);
(count, broken, errs)
}
@@ -11,7 +11,7 @@ pub mod ec_shard;
pub mod ec_volume;
pub use ec_shard::{
EcVolumeShard, ShardId, DATA_SHARDS_COUNT, MAX_SHARD_COUNT, MIN_TOTAL_DISKS,
PARITY_SHARDS_COUNT, TOTAL_SHARDS_COUNT,
DATA_SHARDS_COUNT, EcVolumeShard, MAX_SHARD_COUNT, MIN_TOTAL_DISKS, PARITY_SHARDS_COUNT,
ShardId, TOTAL_SHARDS_COUNT,
};
pub use ec_volume::EcVolume;
+6 -2
View File
@@ -57,7 +57,7 @@ pub fn check_index_file<R: Read + Seek>(
errs.push(format!("walk index file: {}", e));
}
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3 .0.cmp(&b.3 .0)));
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3.0.cmp(&b.3.0)));
// Offset-0 logical tombstones (remote-tier deletes) occupy no physical extent,
// so they cannot overlap anything — exclude them from the overlap check. They
@@ -213,7 +213,11 @@ mod tests {
let size = data.len() as i64;
let (count, errs) = check_index_file(&mut Cursor::new(data), size, Version(3));
assert_eq!(count, 2, "tombstone row is still counted: {:?}", errs);
assert!(errs.is_empty(), "offset-0 tombstone must not overlap: {:?}", errs);
assert!(
errs.is_empty(),
"offset-0 tombstone must not overlap: {:?}",
errs
);
}
#[test]
+142
View File
@@ -0,0 +1,142 @@
//! Positional file reads.
//!
//! Every read here is "these bytes at this offset", never "the next bytes".
//! The handles are shared — `.dat` and `.idx` descriptors are borrowed from
//! [`file_pool`](super::needle_map::file_pool), a mounted EC shard's handle is
//! duplicated into a scrub plan — so no caller may rely on a file position.
//!
//! On unix that is `pread(2)` through `std::os::unix::fs::FileExt`. On Windows
//! it is `seek_read`, which passes the offset through `OVERLAPPED`, so the read
//! itself is independent of the current cursor.
//!
//! What these helpers replace is `try_clone()` + `seek()` + `read()`. A
//! duplicated handle shares one kernel file offset with the original, so that
//! sequence is two syscalls against state another thread can move in between:
//! the seek positions the offset, a concurrent reader or an append moves it,
//! and the read returns bytes from somewhere else entirely. `seek_read` carries
//! its own offset in a single call, so there is no window.
//!
//! `seek_read` does still advance the cursor as a side effect — Windows updates
//! the file pointer even for an `OVERLAPPED` read — which nothing here relies
//! on. A caller that genuinely needs a private position must open the file
//! again rather than duplicate a handle; see `Volume::dat_scan_plan` in
//! [`storage::volume`](super::volume).
use std::fs::File;
use std::io;
/// Reads exactly `buf.len()` bytes from `file` starting at `offset`.
///
/// Fails with [`io::ErrorKind::UnexpectedEof`] if the file ends first.
pub(crate) fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)?;
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = offset;
while filled < buf.len() {
let n = match file.seek_read(&mut buf[filled..], at) {
Ok(n) => n,
Err(err) if err.kind() == io::ErrorKind::Interrupted => continue,
Err(err) => return Err(err),
};
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
filled += n;
at += n as u64;
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
Ok(())
}
/// Reads up to `buf.len()` bytes from `file` starting at `offset`, returning
/// how many were read.
///
/// A short read — including `0` at or past end of file — is not an error; use
/// [`read_exact_at`] when the whole buffer must be filled.
pub(crate) fn read_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
file.seek_read(buf, offset)
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
}
#[cfg(test)]
mod tests {
use super::{read_at, read_exact_at};
use std::io::{ErrorKind, Write};
fn temp_file(bytes: &[u8]) -> tempfile::NamedTempFile {
let mut f = tempfile::NamedTempFile::new().expect("temp file");
f.write_all(bytes).expect("write");
f.flush().expect("flush");
f
}
#[test]
fn read_exact_at_fills_the_whole_buffer() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 10];
read_exact_at(f.as_file(), &mut buf, 0).expect("read");
assert_eq!(&buf, b"0123456789");
}
#[test]
fn read_exact_at_reads_from_the_offset() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 4];
read_exact_at(f.as_file(), &mut buf, 3).expect("read");
assert_eq!(&buf, b"3456");
// The helper is positional: a second read at a lower offset sees the
// bytes at that offset, not wherever the first read left a cursor.
let mut again = [0u8; 4];
read_exact_at(f.as_file(), &mut again, 1).expect("read");
assert_eq!(&again, b"1234");
}
#[test]
fn read_exact_at_short_file_is_unexpected_eof() {
let f = temp_file(b"0123");
let mut buf = [0u8; 8];
let err = read_exact_at(f.as_file(), &mut buf, 0).expect_err("short file");
assert_eq!(err.kind(), ErrorKind::UnexpectedEof);
}
#[test]
fn read_at_allows_a_short_read_at_eof() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 8];
let n = read_at(f.as_file(), &mut buf, 6).expect("read");
assert_eq!(n, 4);
assert_eq!(&buf[..n], b"6789");
// Entirely past the end is zero bytes, not an error.
let n = read_at(f.as_file(), &mut buf, 10).expect("read");
assert_eq!(n, 0);
}
}
+1
View File
@@ -1,6 +1,7 @@
pub mod disk_location;
pub mod erasure_coding;
pub mod idx;
pub(crate) mod io;
pub mod needle;
pub mod needle_map;
pub mod store;
+4 -2
View File
@@ -21,7 +21,7 @@ impl CRC {
/// Legacy `.Value()` function — deprecated in Go but needed for backward compat check.
/// Formula: (crc >> 15 | crc << 17) + 0xa282ead8
pub fn legacy_value(&self) -> u32 {
(self.0 >> 15 | self.0 << 17).wrapping_add(0xa282ead8)
self.0.rotate_right(15).wrapping_add(0xa282ead8)
}
}
@@ -67,7 +67,9 @@ mod tests {
fn test_crc_legacy_value() {
let crc = CRC(0x12345678);
let v = crc.legacy_value();
let expected = (0x12345678u32 >> 15 | 0x12345678u32 << 17).wrapping_add(0xa282ead8);
// (0x12345678 >> 15 | 0x12345678 << 17) + 0xa282ead8, worked out by hand so
// the test checks the rotate rather than restating it.
let expected = 0x4f730f40_u32;
assert_eq!(v, expected);
}
}
+4
View File
@@ -1,4 +1,8 @@
pub mod crc;
#[expect(
clippy::module_inception,
reason = "needle/needle.rs mirrors the Go package layout"
)]
pub mod needle;
pub mod ttl;
+88 -33
View File
@@ -198,8 +198,8 @@ impl Needle {
/// the data payload from disk at all, matching Go's `ReadNeedleMeta`.
pub fn read_paged_meta(
&mut self,
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
offset: i64,
expected_size: Size,
version: Version,
@@ -560,7 +560,7 @@ impl Needle {
// Padding to 8-byte alignment
let padding = padding_length(self.size, version).0 as usize;
buf.extend(std::iter::repeat(0u8).take(padding));
buf.extend(std::iter::repeat_n(0u8, padding));
buf
}
@@ -581,23 +581,19 @@ impl Needle {
// ============================================================================
/// Compute padding to align needle to NEEDLE_PADDING_SIZE (8 bytes).
///
/// The sum is formed in i64: a size read from a corrupt header can sit near
/// `i32::MAX`, and adding the header, checksum and timestamp widths to it in
/// i32 would overflow (a panic with overflow checks, a wrapped padding
/// without). The result is at most NEEDLE_PADDING_SIZE, so it fits `Size`.
pub fn padding_length(needle_size: Size, version: Version) -> Size {
if version == VERSION_3 {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32
+ needle_size.0
+ NEEDLE_CHECKSUM_SIZE as i32
+ TIMESTAMP_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
let fixed = if version == VERSION_3 {
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE + TIMESTAMP_SIZE
} else {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32 + needle_size.0 + NEEDLE_CHECKSUM_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
}
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE
};
let unpadded = fixed as i64 + needle_size.0 as i64;
Size((NEEDLE_PADDING_SIZE as i64 - unpadded % NEEDLE_PADDING_SIZE as i64) as i32)
}
/// Body length = Size + Checksum + [Timestamp] + Padding.
@@ -619,6 +615,30 @@ pub fn get_actual_size(size: Size, version: Version) -> i64 {
NEEDLE_HEADER_SIZE as i64 + needle_body_length(size, version)
}
/// Validate a wire-supplied needle body size before any `as usize` cast.
/// Rejects negative/deleted sizes and bodies larger than the gRPC max message.
/// Size(0) is allowed: empty/anomalous entries and tombstones read as size 0
/// (actual_size = header+checksum+pad > 0, safe alloc, no wrap).
/// Transport cap only: storage paths must NOT use this cap — see volume.rs
/// guards (a >1GiB stored needle from a high-limit cluster must remain
/// readable/compaction-safe). Keep `get_actual_size` unchanged (it
/// intentionally returns negative for deleted index entries).
pub fn validate_wire_size(size: Size) -> Result<(), String> {
if size.0 < 0 {
return Err(format!("invalid needle size {}", size.0));
}
// Keep in sync with canonical `GRPC_MAX_MESSAGE_SIZE` in server/grpc_client.rs:10
// (duplicated here to avoid a storage->server import and prevent drift).
const WIRE_MAX_NEEDLE_SIZE: i32 = 1 << 30;
if size.0 > WIRE_MAX_NEEDLE_SIZE {
return Err(format!(
"needle size {} exceeds max {}",
size.0, WIRE_MAX_NEEDLE_SIZE
));
}
Ok(())
}
/// Read 5 bytes as a u64 (big-endian, zero-padded high bytes).
fn bytes_to_u64_5(bytes: &[u8]) -> u64 {
assert!(bytes.len() >= 5);
@@ -770,7 +790,9 @@ pub fn parse_needle_id_cookie(s: &str) -> Result<(NeedleId, Cookie), String> {
#[derive(Debug, thiserror::Error)]
pub enum NeedleError {
#[error("size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}")]
#[error(
"size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}"
)]
SizeMismatch {
offset: i64,
id: NeedleId,
@@ -824,11 +846,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v3() {
let mut n = Needle::default();
n.cookie = Cookie(42);
n.id = NeedleId(100);
n.data = b"hello world".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(42),
id: NeedleId(100),
data: b"hello world".to_vec(),
flags: 0,
..Needle::default()
};
n.set_has_name();
n.name = b"test.txt".to_vec();
n.name_size = 8;
@@ -867,11 +891,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v2() {
let mut n = Needle::default();
n.cookie = Cookie(77);
n.id = NeedleId(200);
n.data = b"data v2".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(77),
id: NeedleId(200),
data: b"data v2".to_vec(),
flags: 0,
..Needle::default()
};
let bytes = n.write_bytes(VERSION_2);
let expected_size = get_actual_size(n.size, VERSION_2);
@@ -886,10 +912,12 @@ mod tests {
#[test]
fn test_read_bytes_meta_only_handles_tombstone_v3() {
let mut tombstone = Needle::default();
tombstone.cookie = Cookie(0x1234abcd);
tombstone.id = NeedleId(300);
tombstone.append_at_ns = 999_999;
let mut tombstone = Needle {
cookie: Cookie(0x1234abcd),
id: NeedleId(300),
append_at_ns: 999_999,
..Needle::default()
};
let bytes = tombstone.write_bytes(VERSION_3);
@@ -917,6 +945,21 @@ mod tests {
}
}
#[test]
fn padding_length_does_not_overflow_on_a_corrupt_size() {
// A header read from a corrupt or truncated file can carry any i32
// size. The scanners bound it against the bytes left before sizing a
// buffer, but on a volume with more than 2 GiB left a size near
// i32::MAX passes that bound, so the padding arithmetic itself must
// not overflow. Overflow checks are on in test builds, so an i32 sum
// here would panic rather than wrap.
for version in [VERSION_2, VERSION_3] {
let padding = padding_length(Size(i32::MAX), version).0 as i64;
assert!((1..=NEEDLE_PADDING_SIZE as i64).contains(&padding));
assert_eq!(get_actual_size(Size(i32::MAX), version) % 8, 0);
}
}
#[test]
fn test_file_id_parse() {
let fid = FileId::parse("3,01637037d6").unwrap();
@@ -961,4 +1004,16 @@ mod tests {
assert_eq!(fid.key, NeedleId(0x123));
assert_eq!(fid.cookie, Cookie(0));
}
#[test]
fn test_validate_wire_size_boundaries() {
assert!(validate_wire_size(Size(-100)).is_err());
assert!(validate_wire_size(Size(-1)).is_err());
assert!(validate_wire_size(Size(0)).is_ok());
assert!(validate_wire_size(Size(1024)).is_ok());
assert!(validate_wire_size(Size(1)).is_ok());
assert!(validate_wire_size(Size(1 << 30)).is_ok());
assert!(validate_wire_size(Size((1 << 30) + 1)).is_err());
assert!(validate_wire_size(Size(i32::MAX)).is_err());
}
}
+106 -25
View File
@@ -81,7 +81,7 @@ impl TTL {
return Ok(TTL::EMPTY);
}
let last_byte = s.as_bytes()[s.len() - 1];
let (num_str, unit_byte) = if last_byte >= b'0' && last_byte <= b'9' {
let (num_str, unit_byte) = if last_byte.is_ascii_digit() {
// All digits — default to minutes (matching Go)
(s, b'm')
} else {
@@ -144,40 +144,73 @@ fn fit_ttl_count(count: u32, unit: u8) -> TTL {
const MINUTE_SECS: u64 = 60;
// First pass: try exact fits from largest to smallest
if seconds % YEAR_SECS == 0 && seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
if seconds.is_multiple_of(YEAR_SECS) && seconds / YEAR_SECS < 256 {
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
if seconds % MONTH_SECS == 0 && seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
if seconds.is_multiple_of(MONTH_SECS) && seconds / MONTH_SECS < 256 {
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds % WEEK_SECS == 0 && seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
if seconds.is_multiple_of(WEEK_SECS) && seconds / WEEK_SECS < 256 {
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds % DAY_SECS == 0 && seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
if seconds.is_multiple_of(DAY_SECS) && seconds / DAY_SECS < 256 {
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds % HOUR_SECS == 0 && seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
if seconds.is_multiple_of(HOUR_SECS) && seconds / HOUR_SECS < 256 {
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
// Minutes: truncating division
if seconds / MINUTE_SECS < 256 {
return TTL { count: (seconds / MINUTE_SECS) as u8, unit: TTL_UNIT_MINUTE };
return TTL {
count: (seconds / MINUTE_SECS) as u8,
unit: TTL_UNIT_MINUTE,
};
}
// Second pass: truncating division from smallest to largest
if seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
if seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
TTL::EMPTY
}
@@ -225,7 +258,13 @@ mod tests {
// 24h normalizes to 1d via fitTtlCount
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl.to_seconds(), 86400);
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -271,12 +310,24 @@ mod tests {
fn test_ttl_overflow_normalizes() {
// Go's ReadTTL calls fitTtlCount: 300m = 18000s = 5h (exact fit)
let ttl = TTL::read("300m").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_HOUR
}
);
// 256h = 921600s. Doesn't fit in hours (256 >= 256), doesn't fit exact in days.
// Second pass: 921600/86400 = 10 (truncated) < 256 -> 10d
let ttl = TTL::read("256h").unwrap();
assert_eq!(ttl, TTL { count: 10, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 10,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -284,19 +335,49 @@ mod tests {
// Go's ReadTTL calls fitTtlCount which normalizes to coarsest unit.
// 120m -> 2h, 7d -> 1w, 24h -> 1d.
let ttl = TTL::read("120m").unwrap();
assert_eq!(ttl, TTL { count: 2, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 2,
unit: TTL_UNIT_HOUR
}
);
let ttl = TTL::read("7d").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_WEEK });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_WEEK
}
);
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
// Values that don't simplify stay as-is
let ttl = TTL::read("5d").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_DAY
}
);
let ttl = TTL::read("3m").unwrap();
assert_eq!(ttl, TTL { count: 3, unit: TTL_UNIT_MINUTE });
assert_eq!(
ttl,
TTL {
count: 3,
unit: TTL_UNIT_MINUTE
}
);
}
}
+214 -204
View File
@@ -97,12 +97,13 @@ impl NeedleMapMetric {
self.file_byte_count
.fetch_add(new_size.0 as u64, Ordering::Relaxed);
// Go: if oldSize > 0 && oldSize.IsValid() { LogDeletionCounter(oldSize) }
if let Some(old_val) = old {
if old_val.size.0 > 0 && old_val.size.is_valid() {
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
if let Some(old_val) = old
&& old_val.size.0 > 0
&& old_val.size.is_valid()
{
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
}
@@ -225,6 +226,12 @@ pub struct CompactNeedleMap {
idx_file_offset: u64,
}
impl Default for CompactNeedleMap {
fn default() -> Self {
Self::new()
}
}
impl CompactNeedleMap {
/// Create a new empty in-memory map.
pub fn new() -> Self {
@@ -465,9 +472,9 @@ impl RedbNeedleMap {
/// loses at most the writes since the last checkpoint from redb, and
/// the next load replays them from .idx.
fn begin_write_no_fsync(db: &Database) -> io::Result<redb::WriteTransaction> {
let mut txn = db.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = db
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
let _ = txn.set_durability(Durability::None);
Ok(txn)
}
@@ -501,7 +508,7 @@ impl RedbNeedleMap {
pub fn checkpoint(&mut self, sync_idx: bool) -> io::Result<()> {
let txn = self.begin_checkpoint(sync_idx)?;
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
self.writes_since_checkpoint = 0;
Ok(())
}
@@ -516,17 +523,17 @@ impl RedbNeedleMap {
if sync_idx {
self.sync()?;
}
let mut txn = self.db_or_err()?.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = self
.db_or_err()?
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
txn.set_quick_repair(true);
if self.idx_file.is_some() {
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, self.idx_file_offset).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, self.idx_file_offset)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
Ok(txn)
}
@@ -538,22 +545,20 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.create(db_path)
.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb create error: {}", e))
})?;
.map_err(|e| io::Error::other(format!("redb create error: {}", e)))?;
// Ensure tables exist
let txn = Self::begin_write_no_fsync(&db)?;
{
let _table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let _meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table meta: {}", e))
})?;
let _table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let _meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table meta: {}", e)))?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
Ok(RedbNeedleMap {
db: Some(db),
@@ -572,16 +577,14 @@ impl RedbNeedleMap {
fn save_idx_size_meta(&self, idx_size: u64) -> io::Result<()> {
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
{
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, idx_size).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, idx_size)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
txn.commit().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb commit meta: {}", e))
})?;
txn.commit()
.map_err(|e| io::Error::other(format!("redb commit meta: {}", e)))?;
Ok(())
}
@@ -590,22 +593,26 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `meta`; bind the match so the
// temporary Result is dropped before `meta`.
let result = match meta.get(META_IDX_SIZE) {
// ReadableTable::get guard borrows `meta`; edition 2024 drops the tail
// expression's temporaries before `meta`, so no extra binding is needed.
match meta.get(META_IDX_SIZE) {
Ok(Some(guard)) => Ok(Some(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get meta: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get meta: {}", e))),
}
}
/// Test-only read of META `idx_size` through the live handle. See
/// [`test_support::live_meta_idx_size`] for why durability tests use
/// this instead of copying the open `.rdb`.
#[cfg(test)]
pub(crate) fn live_meta_idx_size(&self) -> Option<u64> {
self.read_idx_size_meta().unwrap()
}
/// Load from an .idx file, reusing an existing .rdb if it is consistent.
@@ -648,7 +655,7 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.open(db_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open: {}", e)))?;
let mut nm = RedbNeedleMap {
db: Some(db),
@@ -663,14 +670,11 @@ impl RedbNeedleMap {
let stored_idx_size = nm
.read_idx_size_meta()?
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "no idx_size in redb meta"))?;
.ok_or_else(|| io::Error::other("no idx_size in redb meta"))?;
if stored_idx_size > idx_size {
// .idx shrank — corrupted or truncated, need full rebuild
return Err(io::Error::new(
io::ErrorKind::Other,
"idx file smaller than stored size",
));
return Err(io::Error::other("idx file smaller than stored size"));
}
// Counters come from the whole .idx history, never from the table,
@@ -683,40 +687,37 @@ impl RedbNeedleMap {
let start_entry = stored_idx_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
idx::walk_index_file(reader, start_entry, |key, offset, size| {
let key_u64: u64 = key.into();
if offset.is_zero() || size.is_deleted() {
// Delete: store a tombstone (negative size, original
// offset) over a live value; already deleted is a no-op.
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64) {
if old.size.is_valid() {
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
)
})?;
}
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64)
&& old.size.is_valid()
{
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
} else {
let packed = pack_needle_value(&NeedleValue { offset, size });
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
Ok(())
})?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
}
@@ -734,10 +735,7 @@ impl RedbNeedleMap {
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
@@ -790,13 +788,13 @@ impl RedbNeedleMap {
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
if !unlinked {
table.retain(|_, _| false).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb retain: {}", e))
})?;
table
.retain(|_, _| false)
.map_err(|e| io::Error::other(format!("redb retain: {}", e)))?;
}
#[cfg(not(feature = "redb-experimental-cursor"))]
@@ -804,30 +802,33 @@ impl RedbNeedleMap {
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
}
#[cfg(feature = "redb-experimental-cursor")]
{
let mut cursor = table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
let mut cursor =
table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
cursor.insert_before(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
cursor
.insert_before(key_u64, packed.as_slice())
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
}
cursor.close().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb cursor close: {}", e))
@@ -835,7 +836,7 @@ impl RedbNeedleMap {
}
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
Ok(())
@@ -901,23 +902,16 @@ impl RedbNeedleMap {
Ok(t) => t,
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb open_table: {}", e),
));
return Err(io::Error::other(format!("redb open_table: {}", e)));
}
};
let result = match table.insert(key_u64, packed.as_slice()) {
match table.insert(key_u64, packed.as_slice()) {
Ok(prev) => prev.and_then(|g| packed_to_needle_value(g.value())),
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
};
result
}
};
match txn.commit() {
Ok(()) => old,
@@ -925,8 +919,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -935,12 +928,9 @@ impl RedbNeedleMap {
// visible and redb refuses further writes. Keep
// the .idx row (do NOT truncate) and reopen from
// .idx to repair redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after put commit error failed: {}",
reopen_err
);
tracing::warn!("redb reopen after put commit error failed: {}", reopen_err);
}
return Err(err);
}
@@ -968,39 +958,32 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `table`; bind the match so the
// temporary Result is dropped before `table`.
let result = match table.get(key_u64) {
// ReadableTable::get guard borrows `table`; edition 2024 drops the tail
// expression's temporaries before `table`, so no extra binding is needed.
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
/// Mark a needle as deleted. Appends tombstone to .idx file, negates size in redb.
pub fn delete(&mut self, key: NeedleId, offset: Offset) -> io::Result<Option<Size>> {
let key_u64: u64 = key.into();
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let old = match table.get(key_u64) {
Ok(Some(guard)) => packed_to_needle_value(guard.value()),
Ok(None) => None,
Err(e) => {
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
));
return Err(io::Error::other(format!("redb get: {}", e)));
}
};
let Some(old) = old.filter(|nv| nv.size.is_valid()) else {
@@ -1021,10 +1004,7 @@ impl RedbNeedleMap {
drop(table);
if let Err(e) = insert_res {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
match txn.commit() {
Ok(()) => {}
@@ -1032,8 +1012,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -1042,7 +1021,7 @@ impl RedbNeedleMap {
// and redb refuses further writes. Keep the .idx row
// (do NOT truncate) and reopen from .idx to repair
// redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after delete commit error failed: {}",
@@ -1105,10 +1084,10 @@ impl RedbNeedleMap {
/// after the orphan, `idx_file_offset` advances past it, and a later
/// checkpoint records an offset that makes the reload skip the orphan.
fn truncate_idx_to_offset(&mut self) {
if let Some(ref mut idx_file) = self.idx_file {
if let Err(e) = idx_file.truncate_to(self.idx_file_offset) {
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
if let Some(ref mut idx_file) = self.idx_file
&& let Err(e) = idx_file.truncate_to(self.idx_file_offset)
{
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
}
@@ -1146,18 +1125,12 @@ impl RedbNeedleMap {
let read_file = std::fs::OpenOptions::new()
.read(true)
.open(&idx_path)
.map_err(|e| {
io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e))
})?;
.map_err(|e| io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e)))?;
let actual_idx_size = read_file.metadata()?.len();
let mut reader = io::BufReader::new(read_file);
let reopened = Self::load_from_idx(
&self.rdb_path,
&mut reader,
self.version,
self.cache_bytes,
)?;
let reopened =
Self::load_from_idx(&self.rdb_path, &mut reader, self.version, self.cache_bytes)?;
// Preserve the append writer and the paths/version/cache; adopt the
// repaired database, metrics, and idx_file_offset from the reload.
@@ -1198,10 +1171,10 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let mut file = std::fs::OpenOptions::new()
.write(true)
@@ -1212,18 +1185,17 @@ impl RedbNeedleMap {
// redb iterates in key order (u64 ascending)
let iter = table
.iter()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb iter: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb iter: {}", e)))?;
for entry in iter {
let (key_guard, val_guard) = entry.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb iter next: {}", e))
})?;
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb iter next: {}", e)))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if let Some(nv) = packed_to_needle_value(bytes) {
if nv.size.is_valid() {
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
if let Some(nv) = packed_to_needle_value(bytes)
&& nv.size.is_valid()
{
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
}
file.sync_all()?;
@@ -1512,22 +1484,30 @@ impl NeedleMap {
pub(crate) mod test_support {
use super::*;
/// The `.idx` size recorded in the durable state of the `.rdb` at
/// `rdb_path`, read from a copy taken while the map may still be open:
/// exactly what a crash would leave behind. `None` when nothing durable
/// has been recorded yet.
pub(crate) fn durable_idx_size(rdb_path: &Path) -> Option<u64> {
let copy = rdb_path.with_extension("crash-copy.rdb");
std::fs::copy(rdb_path, &copy).unwrap();
let db = Database::open(&copy).unwrap();
let txn = db.begin_read().unwrap();
let meta = txn.open_table(META_TABLE).ok()?;
let size = meta.get(META_IDX_SIZE).unwrap().map(|g| g.value());
drop(meta);
drop(txn);
drop(db);
let _ = std::fs::remove_file(&copy);
size
/// The `.idx` size in the map's META table, read through the live
/// handle.
///
/// The load path records the `.idx` size with `Durability::None`, and
/// every `put`/`delete` also commits non-durably, so before the first
/// checkpoint this is the load-time value (`Some(0)` for a fresh map) —
/// NOT the crash-durable `None` a copy of the open `.rdb` would show.
/// A live read is the only portable observation: redb 4.2.0 takes an
/// exclusive whole-file lock, which is advisory on Unix but mandatory
/// on Windows, so copying the open `.rdb` fails there with OS error 33.
///
/// It still pins the property under test: the only *durable* META
/// writer is `checkpoint`, so any value other than the load-time one
/// proves a checkpoint recorded progress — and the post-checkpoint
/// value equals the durable one, because checkpoints commit with
/// `Durability::Immediate`. What is lost vs the old copy: strict crash
/// fidelity — a hard crash pre-checkpoint would leave META absent
/// rather than `Some(0)` (loader-equivalent outcomes: full rebuild vs
/// replay-from-0, both correct). A clean close+reopen cannot recover
/// that distinction either: dropping the `Database` flushes pending
/// non-durable commits, so a reopened handle reads `Some(0)` just like
/// the live one.
pub(crate) fn live_meta_idx_size(nm: &RedbNeedleMap) -> Option<u64> {
nm.live_meta_idx_size()
}
}
@@ -1682,6 +1662,7 @@ mod tests {
.read(true)
.write(true)
.create(true)
.truncate(false)
.open(&idx_path)
.unwrap();
let idx_size = idx_file.metadata().unwrap().len();
@@ -2168,8 +2149,14 @@ mod tests {
// server opens one redb database per volume, so the process-wide
// ceiling is roughly (volumes x budget).
assert_eq!(NeedleMapKind::Redb.redb_cache_bytes(), 4 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbMedium.redb_cache_bytes(), 8 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbLarge.redb_cache_bytes(), 16 * 1024 * 1024);
assert_eq!(
NeedleMapKind::RedbMedium.redb_cache_bytes(),
8 * 1024 * 1024
);
assert_eq!(
NeedleMapKind::RedbLarge.redb_cache_bytes(),
16 * 1024 * 1024
);
}
#[test]
@@ -2198,7 +2185,7 @@ mod tests {
#[test]
fn test_redb_checkpoint_is_explicit_and_due_every_interval() {
use test_support::durable_idx_size;
use test_support::live_meta_idx_size;
// Every non-durable redb commit leaves bookkeeping behind until a
// durable one clears it, so a writable map asks for a checkpoint on
@@ -2208,8 +2195,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, _idx_path) = open_writable_redb(dir.path());
for i in 1..EXPECTED_INTERVAL {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
assert!(!nm.checkpoint_due(), "due after only {i} writes");
}
nm.put(
@@ -2219,20 +2210,29 @@ mod tests {
)
.unwrap();
assert!(nm.checkpoint_due());
assert_eq!(durable_idx_size(&db_path), None, "put() must not commit durably");
// No checkpoint taken yet: META still holds the load-time .idx size.
// put() only commits non-durably, so the live value is unchanged.
assert_eq!(
live_meta_idx_size(&nm),
Some(0),
"put() must not record checkpoint progress"
);
nm.checkpoint(true).unwrap();
assert!(!nm.checkpoint_due());
assert_eq!(
durable_idx_size(&db_path),
live_meta_idx_size(&nm),
Some(EXPECTED_INTERVAL * NEEDLE_MAP_ENTRY_SIZE as u64),
"checkpoint records how much of the .idx the table reflects"
);
// Snapshot the .rdb while the map is still open: what a crash leaves.
// Everything is durable after the checkpoint, so the map is closed
// first and the snapshot sees the same bytes on every platform.
// (Copying while open fails on Windows, where redb's file lock is
// mandatory: what a crash leaves.)
drop(nm);
let crash_copy = dir.path().join("crash.rdb");
std::fs::copy(&db_path, &crash_copy).unwrap();
drop(nm);
let db = Database::open(&crash_copy).unwrap();
let txn = db.begin_read().unwrap();
let table = txn.open_table(NEEDLE_TABLE).unwrap();
@@ -2247,8 +2247,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
nm.close();
drop(nm);
@@ -2272,8 +2276,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
// Drop without close(): redb makes the table durable on drop, but the
// recorded .idx size stays at its load-time value (0), so the reload
@@ -2344,12 +2352,14 @@ mod tests {
reloaded.deleted_count(),
reloaded.deleted_size(),
);
assert_eq!(
after, live,
"close_first={close_first} rebuild={rebuild}"
);
assert_eq!(after, live, "close_first={close_first} rebuild={rebuild}");
assert_eq!(reloaded.get(NeedleId(1)).unwrap().unwrap().size, Size(200));
assert!(reloaded.get(NeedleId(2)).unwrap().map_or(true, |v| v.size.is_deleted()));
assert!(
reloaded
.get(NeedleId(2))
.unwrap()
.is_none_or(|v| v.size.is_deleted())
);
}
}
}
@@ -31,7 +31,7 @@ struct CompactEntry {
}
impl CompactEntry {
fn to_needle_value(&self) -> NeedleValue {
fn to_needle_value(self) -> NeedleValue {
NeedleValue {
offset: Offset::from_bytes(&self.offset),
size: self.size,
@@ -205,6 +205,18 @@ mod tests {
use std::os::unix::fs::FileExt;
borrowed.read_exact_at(&mut buf, 0).unwrap();
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = 0;
while filled < buf.len() {
let n = borrowed.seek_read(&mut buf[filled..], at).unwrap();
assert!(n != 0, "unexpected EOF in seek_read");
filled += n;
at += n as u64;
}
}
assert_eq!(&buf, b"first");
}
@@ -132,9 +132,7 @@ mod tests {
let mut seen = SeenKeys::new(10_000, FALSE_POSITIVE_RATE);
// Fresh keys may occasionally collide (that is the false-positive
// rate), but only rarely.
let fresh_reported_seen = (0..10_000u64)
.filter(|&key| seen.test_and_add(key))
.count();
let fresh_reported_seen = (0..10_000u64).filter(|&key| seen.test_and_add(key)).count();
assert!(
fresh_reported_seen < 50,
"fresh keys reported seen: {fresh_reported_seen}"
@@ -18,6 +18,7 @@ use std::sync::{Mutex, RwLock};
use super::file_pool::pooled_index_files;
use crate::storage::idx;
use crate::storage::io::read_exact_at;
use crate::storage::needle_map::{CompactNeedleMap, NeedleMapMetric, NeedleValue};
use crate::storage::types::*;
@@ -133,9 +134,7 @@ impl SortedFileNeedleMap {
}
let file = pooled_index_files()
.borrow(&self.db_file_name, false)
.map_err(|e| {
io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e))
})?;
.map_err(|e| io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e)))?;
match search_sorted_index(&file, self.db_file_size, key)? {
Some((_, offset, size)) => Ok(Some(NeedleValue { offset, size })),
None => Ok(None),
@@ -226,10 +225,7 @@ impl SortedFileNeedleMap {
.fail_sdx_mark
.load(std::sync::atomic::Ordering::Relaxed)
{
return Err(io::Error::new(
io::ErrorKind::Other,
"injected .sdx mark failure",
));
return Err(io::Error::other("injected .sdx mark failure"));
}
let mut buf = [0u8; SIZE_SIZE];
TOMBSTONE_FILE_SIZE.to_bytes(&mut buf);
@@ -309,7 +305,7 @@ impl SortedFileNeedleMap {
let rows = rows_per_read.min(entry_count - done) as usize;
let bytes = &mut block[..rows * NEEDLE_MAP_ENTRY_SIZE];
read_exact_at(&file, bytes, done * NEEDLE_MAP_ENTRY_SIZE as u64)?;
for entry in bytes.chunks_exact(NEEDLE_MAP_ENTRY_SIZE) {
for entry in bytes.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0 {
let (key, offset, size) = idx_entry_from_bytes(entry);
if !size.is_valid() || pending.contains_key(&key) {
continue; // deleted in place, or still awaiting that mark
@@ -525,32 +521,6 @@ fn search_sorted_index(
Ok(None)
}
fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = offset;
while filled < buf.len() {
let n = file.seek_read(&mut buf[filled..], at)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
filled += n;
at += n as u64;
}
Ok(())
}
}
fn write_at(file: &File, buf: &[u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
@@ -703,10 +673,11 @@ mod tests {
// without a reload — the same contract Go's Get has, where callers
// check size.is_deleted().
assert!(m.get(NeedleId(2)).unwrap().unwrap().size.is_deleted());
assert!(m
.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none());
assert!(
m.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none()
);
assert!(!m.get(NeedleId(1)).unwrap().unwrap().size.is_deleted());
}
@@ -1002,7 +973,8 @@ mod tests {
// The retry is a no-op: no second tombstone, no double counting.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
assert_eq!(m.deleted_count(), deleted_before + 2);
@@ -1032,7 +1004,8 @@ mod tests {
);
// And a retry must not append a second tombstone for it.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
}
+170 -224
View File
@@ -6,19 +6,19 @@
use std::collections::HashSet;
use std::io;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::Mutex;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use crate::config::MinFreeSpace;
use crate::pb::master_pb;
use crate::storage::disk_location::DiskLocation;
use crate::storage::erasure_coding::ec_shard::{EcVolumeShard, MAX_SHARD_COUNT};
use crate::storage::erasure_coding::ec_shard::{EcVolumeShard, MAX_SHARD_COUNT, ShardId};
use crate::storage::erasure_coding::ec_volume::EcVolume;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::super_block::ReplicaPlacement;
use crate::storage::types::*;
use crate::storage::volume::{VifVolumeInfo, VolumeError};
use crate::storage::volume::{VifVolumeInfo, VolumeError, VolumeSpec};
/// Top-level storage manager containing all disk locations and their volumes.
pub struct Store {
@@ -300,7 +300,7 @@ impl Store {
collection: &str,
vid: VolumeId,
data_shard_count: u32,
shard_ids: &[u32],
shard_ids: &[ShardId],
) -> Option<usize> {
const TIER_ANY_DISK: u8 = 1;
const TIER_HDD: u8 = 2;
@@ -359,7 +359,7 @@ impl Store {
/// owner (`volume_ec_shards_copy` refuses such a batch instead of
/// guessing). Mirrors `Store.EcShardOwnerDisks` in
/// `weed/storage/store_ec.go`.
pub fn ec_shard_owner_disks(&self, vid: VolumeId, shard_ids: &[u32]) -> Vec<usize> {
pub fn ec_shard_owner_disks(&self, vid: VolumeId, shard_ids: &[ShardId]) -> Vec<usize> {
self.locations
.iter()
.enumerate()
@@ -372,32 +372,20 @@ impl Store {
pub fn add_volume(
&mut self,
vid: VolumeId,
collection: &str,
replica_placement: Option<ReplicaPlacement>,
ttl: Option<crate::storage::needle::ttl::TTL>,
preallocate: u64,
disk_type: DiskType,
version: Version,
spec: &VolumeSpec<'_>,
) -> Result<(), VolumeError> {
if self.find_volume(vid).is_some() {
return Err(VolumeError::AlreadyExists);
}
let loc_idx = self.find_free_location(&disk_type).ok_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("no free location for disk type {:?}", disk_type),
))
VolumeError::Io(io::Error::other(format!(
"no free location for disk type {:?}",
disk_type
)))
})?;
self.locations[loc_idx].create_volume(
vid,
collection,
self.needle_map_kind,
replica_placement,
ttl,
preallocate,
version,
)
self.locations[loc_idx].create_volume(vid, self.needle_map_kind, spec)
}
/// Delete a volume from any location. When keep_remote_data is true the
@@ -459,7 +447,7 @@ impl Store {
}
// Find the location where the .dat file exists
for loc in &mut self.locations {
if &loc.disk_type != &disk_type {
if loc.disk_type != disk_type {
continue;
}
let base = crate::storage::volume::volume_file_name(&loc.directory, collection, vid);
@@ -472,19 +460,18 @@ impl Store {
// Fail the mount so the caller (VolumeCopy) treats it as an error.
let note_path = format!("{}.note", base);
if std::path::Path::new(&note_path).exists() {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::Other,
format!("volume {} copy incomplete: .note still present", vid),
)));
return Err(VolumeError::Io(io::Error::other(format!(
"volume {} copy incomplete: .note still present",
vid
))));
}
return loc.create_volume(
vid,
collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection,
..Default::default()
},
);
}
}
@@ -524,11 +511,8 @@ impl Store {
&& collection != "..";
if hint_safe {
for loc in &mut self.locations {
let base = crate::storage::volume::volume_file_name(
&loc.directory,
collection,
vid,
);
let base =
crate::storage::volume::volume_file_name(&loc.directory, collection, vid);
// Confirm a collection-named sidecar exists before using the
// hint. A lone .vif/.idx (e.g. an EC sidecar whose .ecx is on
// a sibling disk) must NOT mount here: create_volume would
@@ -572,12 +556,11 @@ impl Store {
// keep scanning (matches open_volumes / Go mountVolume).
match loc.create_volume(
vid,
collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection,
..Default::default()
},
) {
Ok(()) => return Ok(()),
Err(e) => {
@@ -600,7 +583,7 @@ impl Store {
// register a phantom normal volume that shadows the real EC volume.
// Match the guard in load_existing_volumes: only mount when a real
// .dat is present, or the .vif points at a remote-tiered file.
let dat_exists = std::fs::metadata(&format!("{}.dat", base_path))
let dat_exists = std::fs::metadata(format!("{}.dat", base_path))
.map(|m| !m.is_dir())
.unwrap_or(false);
let idx_base = crate::storage::volume::volume_file_name(
@@ -608,9 +591,10 @@ impl Store {
&collection,
vid,
);
let has_remote = crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", base_path),
) || crate::storage::disk_location::vif_references_remote_file(
let has_remote = crate::storage::disk_location::vif_references_remote_file(&format!(
"{}.vif",
base_path
)) || crate::storage::disk_location::vif_references_remote_file(
&format!("{}.vif", idx_base),
);
if dat_exists || has_remote {
@@ -628,12 +612,11 @@ impl Store {
let loc = &mut self.locations[loc_idx];
match loc.create_volume(
vid,
&collection,
self.needle_map_kind,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: &collection,
..Default::default()
},
) {
Ok(()) => return Ok(()),
Err(e) => {
@@ -643,10 +626,12 @@ impl Store {
}
}
}
Err(last_err.unwrap_or_else(|| VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("volume {} not found on disk", vid),
))))
Err(last_err.unwrap_or_else(|| {
VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("volume {} not found on disk", vid),
))
}))
}
fn find_volume_file_base(&self, vid: VolumeId) -> Option<(usize, String, String)> {
@@ -664,13 +649,12 @@ impl Store {
for entry in entries.flatten() {
let name = entry.file_name();
let name = name.to_string_lossy();
if let Some((collection, file_vid)) = parse_volume_filename(&name) {
if file_vid == vid {
if let Some(base) = strip_volume_suffix(&name) {
let base_path = format!("{}/{}", loc.directory, base);
results.push((loc_idx, base_path, collection));
}
}
if let Some((collection, file_vid)) = parse_volume_filename(&name)
&& file_vid == vid
&& let Some(base) = strip_volume_suffix(&name)
{
let base_path = format!("{}/{}", loc.directory, base);
results.push((loc_idx, base_path, collection));
}
}
}
@@ -842,10 +826,8 @@ impl Store {
let vol_count = loc.volumes_len() as i32;
let loc_ec_shards = loc.ec_shard_count();
let ec_equivalent = ((loc_ec_shards
+ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT
- 1)
/ crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
let ec_equivalent = loc_ec_shards
.div_ceil(crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT)
as i32;
let mut max_count = vol_count + ec_equivalent;
@@ -900,7 +882,7 @@ impl Store {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
) -> Result<(), VolumeError> {
// Find the location where the EC files live
let loc_idx = self.find_ec_location(vid, collection).ok_or_else(|| {
@@ -921,12 +903,12 @@ impl Store {
&mut self,
vid: VolumeId,
collection: &str,
shard_id: u32,
shard_id: ShardId,
source_disk_type: &str,
) -> Result<(), VolumeError> {
for loc in &mut self.locations {
// Check if the shard file exists on this location
let shard = EcVolumeShard::new(&loc.directory, collection, vid, shard_id as u8);
let shard = EcVolumeShard::new(&loc.directory, collection, vid, shard_id);
if std::path::Path::new(&shard.file_name()).exists() {
loc.mount_ec_shards(vid, collection, &[shard_id], source_disk_type)?;
return Ok(());
@@ -947,7 +929,7 @@ impl Store {
/// the same store (#9252). DiskLocation::unmount_ec_shards
/// already skips shards that aren't mounted, so this is safe to
/// fan out blindly.
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[u32]) {
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[ShardId]) {
for loc in &mut self.locations {
if loc.has_ec_volume(vid) {
loc.unmount_ec_shards(vid, shard_ids);
@@ -960,7 +942,7 @@ impl Store {
pub fn unmount_ec_shard(
&mut self,
vid: VolumeId,
shard_id: u32,
shard_id: ShardId,
req_encode_ts_ns: i64,
) -> Result<(), VolumeError> {
// Walk all locations rather than stopping at the first with the
@@ -968,7 +950,7 @@ impl Store {
// multiple disks, with the target shard on any of them.
for disk_id in 0..self.locations.len() {
let ec_vol = self.locations[disk_id].find_ec_volume(vid);
let has_shard = ec_vol.is_some_and(|ec_vol| ec_vol.has_shard(shard_id as u8));
let has_shard = ec_vol.is_some_and(|ec_vol| ec_vol.has_shard(shard_id));
if !has_shard {
continue;
}
@@ -988,12 +970,7 @@ impl Store {
);
continue;
}
tracing::info!(
volume_id = vid.0,
shard_id,
disk_id,
"UnmountEcShards"
);
tracing::info!(volume_id = vid.0, shard_id, disk_id, "UnmountEcShards");
self.locations[disk_id].unmount_ec_shards(vid, &[shard_id]);
}
// Go returns nil if shard not found (no error)
@@ -1092,12 +1069,12 @@ impl Store {
/// disks (each holding a disjoint subset of the shards). Without
/// this, callers using `find_ec_volume(vid)` would only see the
/// first disk and miss shards that live on a sibling.
pub fn find_ec_shard_location(&self, vid: VolumeId, shard_id: u32) -> Option<usize> {
pub fn find_ec_shard_location(&self, vid: VolumeId, shard_id: ShardId) -> Option<usize> {
for (i, loc) in self.locations.iter().enumerate() {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(i);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id)
{
return Some(i);
}
}
None
@@ -1106,16 +1083,12 @@ impl Store {
/// Like [`Self::find_ec_shard_location`] but returns the EcVolume
/// reference directly. Borrows the store immutably for the
/// EcVolume's lifetime.
pub fn find_ec_volume_with_shard(
&self,
vid: VolumeId,
shard_id: u32,
) -> Option<&EcVolume> {
pub fn find_ec_volume_with_shard(&self, vid: VolumeId, shard_id: ShardId) -> Option<&EcVolume> {
for loc in &self.locations {
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
return Some(ecv);
}
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id)
{
return Some(ecv);
}
}
None
@@ -1144,9 +1117,12 @@ impl Store {
if found_vol.is_none() {
found_vol = Some(ecv);
}
for shard_id in 0..max_shard_count {
if dirs[shard_id].is_none() && ecv.has_shard(shard_id as u8) {
dirs[shard_id] = Some(loc.directory.clone());
for (shard_id, dir) in dirs.iter_mut().enumerate() {
let Ok(sid) = ShardId::try_from(shard_id) else {
continue;
};
if dir.is_none() && ecv.has_shard(sid) {
*dir = Some(loc.directory.clone());
}
}
}
@@ -1171,7 +1147,8 @@ impl Store {
expired_vids.push(*vid);
} else {
let (_, io_count, quarantined) = ec_vol.get_io_error_state();
if quarantined || io_count >= crate::storage::erasure_coding::ec_volume::IO_ERROR_TOLERANCE
if quarantined
|| io_count >= crate::storage::erasure_coding::ec_volume::IO_ERROR_TOLERANCE
{
io_quarantined_vids.push(*vid);
} else {
@@ -1240,12 +1217,12 @@ impl Store {
}
/// Delete EC shard files from disk.
pub fn delete_ec_shards(&mut self, vid: VolumeId, collection: &str, shard_ids: &[u32]) {
pub fn delete_ec_shards(&mut self, vid: VolumeId, collection: &str, shard_ids: &[ShardId]) {
// Delete shard files from disk, tracking which locations actually held one.
let mut deleted_at = vec![false; self.locations.len()];
for (i, loc) in self.locations.iter().enumerate() {
for &shard_id in shard_ids {
let shard = EcVolumeShard::new(&loc.directory, collection, vid, shard_id as u8);
let shard = EcVolumeShard::new(&loc.directory, collection, vid, shard_id);
if std::fs::remove_file(shard.file_name()).is_ok() {
deleted_at[i] = true;
}
@@ -1281,8 +1258,9 @@ impl Store {
collection,
vid,
);
let _ =
crate::storage::erasure_coding::ec_bitrot::remove_bitrot_sidecars(&idx_base);
let _ = crate::storage::erasure_coding::ec_bitrot::remove_bitrot_sidecars(
&idx_base,
);
}
}
}
@@ -1516,9 +1494,10 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
read_only: bool,
}
if let Ok(legacy) = serde_json::from_str::<LegacyVolumeInfo>(&content) {
let mut vif = VifVolumeInfo::default();
vif.read_only = legacy.read_only;
return Ok(vif);
return Ok(VifVolumeInfo {
read_only: legacy.read_only,
..VifVolumeInfo::default()
});
}
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
@@ -1528,7 +1507,7 @@ fn load_vif_volume_info(path: &str) -> Result<VifVolumeInfo, VolumeError> {
fn save_vif_volume_info(path: &str, info: &VifVolumeInfo) -> Result<(), VolumeError> {
let content = serde_json::to_string_pretty(info)
.map_err(|e| VolumeError::Io(io::Error::new(io::ErrorKind::Other, e.to_string())))?;
.map_err(|e| VolumeError::Io(io::Error::other(e.to_string())))?;
std::fs::write(path, content)?;
Ok(())
}
@@ -1579,13 +1558,13 @@ fn ec_free_shard_count(loc: &DiskLocation, data_shard_count: u32) -> i64 {
/// the in-memory registration the read path and heartbeats use.
///
/// Mirrors `ownedEcShardCount` in `weed/storage/store_ec.go`.
fn owned_ec_shard_count(loc: &DiskLocation, vid: VolumeId, shard_ids: &[u32]) -> usize {
fn owned_ec_shard_count(loc: &DiskLocation, vid: VolumeId, shard_ids: &[ShardId]) -> usize {
let Some(ecv) = loc.find_ec_volume(vid) else {
return 0;
};
shard_ids
.iter()
.filter(|&&shard_id| ecv.has_shard(shard_id as u8))
.filter(|&&shard_id| ecv.has_shard(shard_id))
.count()
}
@@ -1644,15 +1623,7 @@ mod tests {
let mut store = make_test_store(&[dir]);
store
.add_volume(
VolumeId(1),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
assert!(store.has_volume(VolumeId(1)));
assert!(!store.has_volume(VolumeId(2)));
@@ -1691,10 +1662,7 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
let escaped = format!(
"{}/../evil_5.dat",
dir
);
let escaped = format!("{}/../evil_5.dat", dir);
let err = store
.mount_volume_by_id(VolumeId(5), Some("../evil"))
.unwrap_err();
@@ -1714,12 +1682,11 @@ mod tests {
store
.add_volume(
VolumeId(7),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "coll",
..Default::default()
},
)
.unwrap();
// Write a needle so the volume has real data, then unmount it so the
@@ -1736,9 +1703,7 @@ mod tests {
.unwrap();
assert!(store.unmount_volume(VolumeId(7)));
store
.mount_volume_by_id(VolumeId(7), Some("coll"))
.unwrap();
store.mount_volume_by_id(VolumeId(7), Some("coll")).unwrap();
assert!(store.find_volume(VolumeId(7)).is_some());
let mut got = Needle {
@@ -1761,12 +1726,11 @@ mod tests {
store
.add_volume(
VolumeId(9),
"foo..bar",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "foo..bar",
..Default::default()
},
)
.unwrap();
let mut n = Needle {
@@ -1776,7 +1740,9 @@ mod tests {
data_size: 4,
..Needle::default()
};
store.write_volume_needle(VolumeId(9), &mut n, false).unwrap();
store
.write_volume_needle(VolumeId(9), &mut n, false)
.unwrap();
assert!(store.unmount_volume(VolumeId(9)));
// The hint is accepted and mounts the volume.
@@ -1805,12 +1771,11 @@ mod tests {
store
.add_volume(
VolumeId(11),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "coll",
..Default::default()
},
)
.unwrap();
let mut n = Needle {
@@ -1820,7 +1785,9 @@ mod tests {
data_size: 7,
..Needle::default()
};
store.write_volume_needle(VolumeId(11), &mut n, false).unwrap();
store
.write_volume_needle(VolumeId(11), &mut n, false)
.unwrap();
assert!(store.unmount_volume(VolumeId(11)));
// Simulate an interrupted copy: drop a .note marker.
@@ -1861,12 +1828,11 @@ mod tests {
store
.add_volume(
VolumeId(13),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "coll",
..Default::default()
},
)
.unwrap();
let mut n = Needle {
@@ -1876,7 +1842,9 @@ mod tests {
data_size: 4,
..Needle::default()
};
store.write_volume_needle(VolumeId(13), &mut n, false).unwrap();
store
.write_volume_needle(VolumeId(13), &mut n, false)
.unwrap();
assert!(store.unmount_volume(VolumeId(13)));
// No hint: the fallback scan finds the sidecar on disk 0 first (skip,
@@ -1909,16 +1877,17 @@ mod tests {
let mut store = make_test_store(&[dir0, dir1]);
// Force add_volume onto disk 1 by marking disk 0 as low on space.
store.locations[0].is_disk_space_low.store(true, Ordering::Relaxed);
store.locations[0]
.is_disk_space_low
.store(true, Ordering::Relaxed);
store
.add_volume(
VolumeId(15),
"coll",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "coll",
..Default::default()
},
)
.unwrap();
let mut n = Needle {
@@ -1928,10 +1897,14 @@ mod tests {
data_size: 5,
..Needle::default()
};
store.write_volume_needle(VolumeId(15), &mut n, false).unwrap();
store
.write_volume_needle(VolumeId(15), &mut n, false)
.unwrap();
assert!(store.unmount_volume(VolumeId(15)));
// Clear the low-space flag so mount_volume_by_id considers disk 0.
store.locations[0].is_disk_space_low.store(false, Ordering::Relaxed);
store.locations[0]
.is_disk_space_low
.store(false, Ordering::Relaxed);
// disk 0: a .dat that exists but is unreadable (chmod 000). The guard
// sees dat_exists=true (metadata succeeds, not a dir), but
@@ -1970,15 +1943,7 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
store
.add_volume(
VolumeId(1),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
// Write
@@ -2044,15 +2009,7 @@ mod tests {
)
.unwrap();
store
.add_volume(
VolumeId(1),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
let mut n = Needle {
id: NeedleId(1),
@@ -2115,26 +2072,10 @@ mod tests {
// Add volumes — should go to location with fewest volumes
store
.add_volume(
VolumeId(1),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
store
.add_volume(
VolumeId(2),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.add_volume(VolumeId(2), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
assert_eq!(store.total_volume_count(), 2);
@@ -2152,34 +2093,31 @@ mod tests {
store
.add_volume(
VolumeId(1),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
store
.add_volume(
VolumeId(2),
"pics",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
store
.add_volume(
VolumeId(3),
"docs",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "docs",
..Default::default()
},
)
.unwrap();
assert_eq!(store.total_volume_count(), 3);
@@ -2215,23 +2153,21 @@ mod tests {
store
.add_volume(
VolumeId(61),
"preallocate_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "preallocate_case",
..Default::default()
},
)
.unwrap();
store
.add_volume(
VolumeId(62),
"preallocate_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "preallocate_case",
..Default::default()
},
)
.unwrap();
for vid in [VolumeId(61), VolumeId(62)] {
@@ -2332,12 +2268,11 @@ mod tests {
store
.add_volume(
VolumeId(71),
"find_free_location_case",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
collection: "find_free_location_case",
..Default::default()
},
)
.unwrap();
@@ -2562,7 +2497,10 @@ mod tests {
// Mount an EC shard on disk 1 so has_ec_volume returns true.
std::fs::write(
format!("{}/{}_{}.ec00", store.locations[1].directory, collection, vid.0),
format!(
"{}/{}_{}.ec00",
store.locations[1].directory, collection, vid.0
),
b"shard data",
)
.unwrap();
@@ -2575,7 +2513,12 @@ mod tests {
std::fs::write(format!("{}.ecx", base), vec![0u8; 20]).unwrap();
let got = store.find_ec_shard_target_location(collection, vid, 10, &[]);
assert_eq!(got, Some(1), "expected the mounted disk to win; got {:?}", got);
assert_eq!(
got,
Some(1),
"expected the mounted disk to win; got {:?}",
got
);
}
/// Cold-volume case: no mount, no `.ecx` anywhere on this server.
@@ -2629,7 +2572,10 @@ mod tests {
// free shard slots remaining; the old formula would have
// rounded that to 0.
std::fs::write(
format!("{}/{}_{}.ec00", store.locations[1].directory, collection, vid.0),
format!(
"{}/{}_{}.ec00",
store.locations[1].directory, collection, vid.0
),
b"shard data",
)
.unwrap();
@@ -2718,7 +2664,7 @@ mod tests {
// Fill disk 0 past its shard-slot budget so ec_free_shard_count is 0.
let filler = VolumeId(10000);
let filler_base = volume_file_name(&store.locations[0].directory, collection, filler);
let filler_shards: Vec<u32> = (0..10).collect();
let filler_shards: Vec<ShardId> = (0..10).collect();
for shard_id in &filler_shards {
std::fs::write(format!("{}.ec{:02}", filler_base, shard_id), b"x").unwrap();
}
@@ -8,7 +8,7 @@ use std::path::Path;
use tracing::{info, warn};
use crate::storage::disk_location::{parse_collection_volume_id_pub, DiskLocation};
use crate::storage::disk_location::{DiskLocation, parse_collection_volume_id_pub};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -286,9 +286,7 @@ fn collect_shard_disk_volumes(loc: &DiskLocation) -> HashMap<EcKey, Vec<String>>
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
out.entry(EcKey { collection, vid })
.or_default()
.push(name);
out.entry(EcKey { collection, vid }).or_default().push(name);
}
out
}
+129 -89
View File
@@ -21,7 +21,7 @@ use std::fs;
use tracing::{error, info, warn};
use crate::storage::disk_location::{is_ec_shard_extension, parse_collection_volume_id_pub};
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
use crate::storage::erasure_coding::ec_shard::{DATA_SHARDS_COUNT, ShardId};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -80,6 +80,11 @@ struct EcxOwnerInfo {
idx_dir: String,
}
/// One unit of reconcile work: the disk holding orphan shards, the volume
/// they belong to, the shard files, the `.ecx` owner, and whether the
/// mirror already installed sidecars locally (`use_local_idx`).
type OrphanShardLoad = (usize, EcKey, Vec<(String, ShardId)>, EcxOwnerInfo, bool);
impl Store {
/// Run cross-disk orphan-shard reconciliation. Should be called
/// after every DiskLocation has finished its per-disk EC scan.
@@ -98,7 +103,7 @@ impl Store {
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
let mut to_load: Vec<OrphanShardLoad> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -117,9 +122,7 @@ impl Store {
let use_local_idx = std::path::Path::new(&local_ecx).exists()
|| std::path::Path::new(&local_ecx_in_data).exists();
if !use_local_idx
&& owner.location == loc_idx
&& owner.idx_dir == loc.idx_directory
if !use_local_idx && owner.location == loc_idx && owner.idx_dir == loc.idx_directory
{
// Same-disk no-op: load_all_ec_shards already
// tried and logged the failure.
@@ -132,7 +135,7 @@ impl Store {
for (loc_idx, key, shards, owner, use_local_idx) in to_load {
let shard_names: Vec<&str> = shards.iter().map(|(n, _)| n.as_str()).collect();
let loc_dir = self.locations[loc_idx].directory.clone();
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if use_local_idx {
info!(
@@ -293,10 +296,10 @@ impl Store {
// may be sole copies of a distributed volume.
let mut node_wide_bits = ev.shard_bits().0;
for other in &self.locations {
if let Some(other_ev) = other.find_ec_volume(*vid) {
if other_ev.collection == ev.collection {
node_wide_bits |= other_ev.shard_bits().0;
}
if let Some(other_ev) = other.find_ec_volume(*vid)
&& other_ev.collection == ev.collection
{
node_wide_bits |= other_ev.shard_bits().0;
}
}
let node_wide = node_wide_bits.count_ones() as usize;
@@ -474,13 +477,13 @@ impl Store {
/// Unlike `reconcile_ec_shards_across_disks` it needs no sibling disk, so a
/// single-disk store recovers once its index has been fetched from a peer.
fn load_orphan_ec_shards_with_local_index(&mut self) {
let mut work: Vec<(usize, EcKey, Vec<u32>)> = Vec::new();
let mut work: Vec<(usize, EcKey, Vec<ShardId>)> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
for (key, shards) in collect_orphan_ec_shards(loc, loc_idx) {
if !loc.has_ecx_file_on_disk(&key.collection, key.vid) {
continue;
}
let ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
work.push((loc_idx, key, ids));
}
}
@@ -499,6 +502,53 @@ impl Store {
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, ShardId)>> {
let mut orphans: HashMap<EcKey, Vec<(String, ShardId)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id)
{
continue;
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
#[cfg(test)]
mod tests {
use super::*;
@@ -541,7 +591,13 @@ mod tests {
std::fs::write(&p, b"shard data nonempty").unwrap();
}
fn write_index_files(idx_dir: &str, collection: &str, vid: u32, data_shards: u32, parity_shards: u32) {
fn write_index_files(
idx_dir: &str,
collection: &str,
vid: u32,
data_shards: u32,
parity_shards: u32,
) {
// Minimal sealed .ecx (the loader only opens the file; it
// doesn't parse it during placement).
std::fs::write(
@@ -895,15 +951,24 @@ mod tests {
// dir1 owns the .ecx and so already has shard 1 mounted via
// its own load_all_ec_shards.
let ev1 = store.locations[1].find_ec_volume(VolumeId(vid));
assert!(ev1.is_some(), "baseline broken: dir1 should have mounted shard 1");
assert!(
ev1.is_some(),
"baseline broken: dir1 should have mounted shard 1"
);
// dir0's shards must be reconciled across to its own
// ec_volumes map, pointing at dir1's idx dir.
let ev0 = store.locations[0]
.find_ec_volume(VolumeId(vid))
.expect("dir0 should now have an EcVolume after reconcile");
assert!(ev0.has_shard(0), "shard 0 missing from dir0 after reconcile");
assert!(ev0.has_shard(12), "shard 12 missing from dir0 after reconcile");
assert!(
ev0.has_shard(0),
"shard 0 missing from dir0 after reconcile"
);
assert!(
ev0.has_shard(12),
"shard 12 missing from dir0 after reconcile"
);
}
/// PR 9244 review case: idx_directory is configured but the
@@ -1012,7 +1077,13 @@ mod tests {
assert!(store.locations[0].find_ec_volume(VolumeId(vid)).is_none());
// Shard files must still exist on disk for operator recovery.
for sid in [0u8, 12u8] {
let p = format!("{}/{}_{}.ec{:02}", dir0.to_str().unwrap(), collection, vid, sid);
let p = format!(
"{}/{}_{}.ec{:02}",
dir0.to_str().unwrap(),
collection,
vid,
sid
);
assert!(
std::path::Path::new(&p).exists(),
"orphan shard {} was destroyed",
@@ -1077,10 +1148,12 @@ mod tests {
assert!(ev1.has_shard(6), "dir1 shard missing");
// Nothing left to recover.
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
#[test]
@@ -1112,10 +1185,12 @@ mod tests {
.unwrap();
}
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
/// Helper: build a 2-disk store where reconcile produces the
@@ -1193,7 +1268,11 @@ mod tests {
let vid = VolumeId(7010);
let all = store.find_all_ec_volumes(vid);
assert_eq!(all.len(), 2, "expected one EcVolume per disk holding the vid");
assert_eq!(
all.len(),
2,
"expected one EcVolume per disk holding the vid"
);
// Disk 0 carries shards 0 and 12; disk 1 carries shard 1.
assert!(all[0].has_shard(0));
@@ -1214,7 +1293,7 @@ mod tests {
#[test]
fn test_scrub_plans_reach_every_disk_through_the_store() {
use crate::storage::erasure_coding::ec_volume::{
merge_ec_runtimes, EcChecksumScrubPlan, EcLocalScrubPlan,
EcChecksumScrubPlan, EcLocalScrubPlan, merge_ec_runtimes,
};
let (store, _tmp) = build_split_disk_store(7030);
@@ -1229,8 +1308,15 @@ mod tests {
let merged = merge_ec_runtimes(&runtimes).expect("two runtimes merge");
assert!(merged.slots[0].is_some(), "disk 0's shard 0 unreachable");
assert!(merged.slots[12].is_some(), "disk 0's shard 12 unreachable");
assert!(merged.slots[1].is_some(), "disk 1's shard 1 unreachable — the bug");
assert!(merged.skipped.is_empty(), "same generation: {:?}", merged.skipped);
assert!(
merged.slots[1].is_some(),
"disk 1's shard 1 unreachable — the bug"
);
assert!(
merged.skipped.is_empty(),
"same generation: {:?}",
merged.skipped
);
// Shard 1 is owned by the sibling runtime, not the anchor.
let (owner, _) = merged.slots[1].unwrap();
@@ -1337,17 +1423,22 @@ mod tests {
let (_ev, dirs) = store.collect_ec_shard_dirs(vid, max_shards).unwrap();
// Shards 0 and 12 → disk 0's directory.
assert_eq!(dirs[0].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(dirs[12].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(
dirs[0].as_deref(),
Some(store.locations[0].directory.as_str())
);
assert_eq!(
dirs[12].as_deref(),
Some(store.locations[0].directory.as_str())
);
// Shard 1 → disk 1's directory.
assert_eq!(dirs[1].as_deref(), Some(store.locations[1].directory.as_str()));
assert_eq!(
dirs[1].as_deref(),
Some(store.locations[1].directory.as_str())
);
// Unmounted shards → None.
for sid in [2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13] {
assert_eq!(
dirs[sid], None,
"shard {} unexpectedly reported a dir",
sid,
);
assert_eq!(dirs[sid], None, "shard {} unexpectedly reported a dir", sid,);
}
}
@@ -1656,11 +1747,7 @@ mod tests {
vec![0u8; 20],
)
.unwrap();
std::fs::write(
ec_dir.join(format!("{}_{}.ecj", collection, vid)),
b"",
)
.unwrap();
std::fs::write(ec_dir.join(format!("{}_{}.ecj", collection, vid)), b"").unwrap();
let mut store = Store::new(NeedleMapKind::InMemory);
store
@@ -1759,50 +1846,3 @@ mod tests {
assert!(std::path::Path::new(&format!("{}.ecx", ec_base)).exists());
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
continue;
}
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
+4 -24
View File
@@ -155,7 +155,7 @@ impl Size {
return 0;
}
if self.0 < 0 {
return (self.0 * -1) as u32;
return -self.0 as u32;
}
self.0 as u32
}
@@ -284,8 +284,9 @@ impl fmt::Display for Offset {
// DiskType
// ============================================================================
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
#[derive(Debug, Clone, PartialEq, Eq, Hash, Default)]
pub enum DiskType {
#[default]
HardDrive,
Ssd,
Custom(String),
@@ -319,12 +320,6 @@ impl fmt::Display for DiskType {
}
}
impl Default for DiskType {
fn default() -> Self {
DiskType::HardDrive
}
}
// ============================================================================
// VolumeId
// ============================================================================
@@ -397,7 +392,7 @@ impl From<u8> for Version {
///
/// Fields are split into request-side options (set by the caller) and response-side
/// flags (set during the read to communicate status back).
#[derive(Debug, Clone)]
#[derive(Debug, Clone, Default)]
pub struct ReadOption {
// -- request --
/// If true, allow reading needles that have been soft-deleted.
@@ -423,21 +418,6 @@ pub struct ReadOption {
pub read_buffer_size: i32,
}
impl Default for ReadOption {
fn default() -> Self {
ReadOption {
read_deleted: false,
attempt_meta_only: false,
must_meta_only: false,
is_meta_only: false,
volume_revision: 0,
is_out_of_range: false,
has_slow_read: false,
read_buffer_size: 0,
}
}
}
// ============================================================================
// NeedleMapEntry helpers (for .idx file)
// ============================================================================
File diff suppressed because it is too large Load Diff
@@ -10,7 +10,7 @@ use crate::storage::needle::Needle;
use crate::storage::super_block::SuperBlock;
use crate::storage::types::*;
use crate::storage::volume::{
fsync_dir, needle_disk_end, scan_volume_file, Volume, VolumeError, VolumeFileVisitor,
Volume, VolumeError, VolumeFileVisitor, fsync_dir, needle_disk_end, scan_volume_file,
};
/// Writes one .idx row per .dat record, in .dat append order, which is the
@@ -105,11 +105,11 @@ impl Volume {
#[cfg(test)]
mod tests {
use crate::storage::needle::crc::CRC;
use crate::storage::needle::Needle;
use crate::storage::needle::crc::CRC;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::*;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use std::fs;
use std::path::Path;
use tempfile::TempDir;
@@ -145,13 +145,9 @@ mod tests {
let mut v = Volume::new(
data,
old_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for id in 1..=3 {
@@ -167,13 +163,9 @@ mod tests {
let reopened = Volume::new(
data,
new_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -207,13 +199,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -233,13 +221,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
@@ -261,13 +245,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -290,13 +270,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
@@ -318,13 +294,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -356,13 +328,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
+23 -15
View File
@@ -9,10 +9,10 @@ use std::path::Path;
use tracing::info;
use crate::storage::idx;
use crate::storage::needle::needle::needle_body_length;
use crate::storage::needle::Needle;
use crate::storage::needle::needle::needle_body_length;
use crate::storage::types::*;
use crate::storage::volume::{fsync_dir, Volume, VolumeError};
use crate::storage::volume::{Volume, VolumeError, fsync_dir};
/// Needles found in the head of .dat, keyed by id, plus the ids in .dat order.
type DatHeadNeedles = (HashMap<NeedleId, (Offset, Size)>, Vec<NeedleId>);
@@ -215,20 +215,19 @@ mod tests {
use super::*;
use crate::storage::needle::crc::CRC;
use crate::storage::needle_map::NeedleMapKind;
use std::os::unix::fs::{FileExt, PermissionsExt};
use crate::storage::volume::VolumeSpec;
use std::io::{Seek, SeekFrom};
#[cfg(unix)]
use std::os::unix::fs::PermissionsExt;
use tempfile::TempDir;
fn open_volume(dir: &str) -> Volume {
Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap()
}
@@ -256,7 +255,7 @@ mod tests {
/// writes (key, offset 0, tombstone) rows over the front of .idx instead of
/// appending them.
fn clobber_idx_head(idx_path: &str, keys: &[u64]) {
let file = OpenOptions::new().write(true).open(idx_path).unwrap();
let mut file = OpenOptions::new().write(true).open(idx_path).unwrap();
for (i, key) in keys.iter().enumerate() {
let mut row = Vec::new();
idx::write_index_entry(
@@ -266,8 +265,13 @@ mod tests {
TOMBSTONE_FILE_SIZE,
)
.unwrap();
file.write_at(&row, (i * NEEDLE_MAP_ENTRY_SIZE) as u64)
// Positional write without Unix-only `FileExt::write_at`, so this
// helper (and the tests using it) also builds on Windows.
// Single-threaded test helper: no concurrent reader can move the
// offset between seek and write.
file.seek(SeekFrom::Start((i * NEEDLE_MAP_ENTRY_SIZE) as u64))
.unwrap();
file.write_all(&row).unwrap();
}
}
@@ -305,6 +309,7 @@ mod tests {
let size_before = idx_size(&idx_path);
// The rewrite replaces .idx wholesale, so it must not widen the mode.
#[cfg(unix)]
fs::set_permissions(&idx_path, fs::Permissions::from_mode(0o600)).unwrap();
// Deletes against needles 9..12 land on the front of .idx and take the
@@ -324,11 +329,14 @@ mod tests {
let want = size_before + 4 * NEEDLE_MAP_ENTRY_SIZE as u64;
assert_eq!(idx_size(&idx_path), want, "idx size after recovery");
assert_eq!(
fs::metadata(&idx_path).unwrap().permissions().mode() & 0o777,
0o600,
"idx mode after recovery"
);
#[cfg(unix)]
{
assert_eq!(
fs::metadata(&idx_path).unwrap().permissions().mode() & 0o777,
0o600,
"idx mode after recovery"
);
}
// The recovered rows go back in front, so .idx is in .dat append order
// again: the fingerprint is gone and the last row is still the .dat tail.
@@ -73,8 +73,10 @@ mod tests {
let empty = master_pb::VolumeInformationMessage::default();
assert_eq!(report_hash(&empty), 10988706248825469653);
let mut one = master_pb::VolumeInformationMessage::default();
one.id = 1;
let one = master_pb::VolumeInformationMessage {
id: 1,
..Default::default()
};
assert_eq!(report_hash(&one), 2035849960016744285);
let full = master_pb::VolumeInformationMessage {
+1 -1
View File
@@ -66,7 +66,7 @@ fn parse_go_version_number() -> Option<String> {
}
}
match (major, minor) {
(Some(maj), Some(min)) => Some(format!("{}.{}", maj, format!("{:02}", min))),
(Some(maj), Some(min)) => Some(format!("{}.{:02}", maj, min)),
_ => None,
}
}
+65 -16
View File
@@ -20,26 +20,71 @@ use std::collections::{HashMap, HashSet};
fn ungated_handlers() -> HashMap<&'static str, &'static str> {
[
// Cluster-internal: issued volume server -> volume server.
("copy_file", "replica sync and EC task pull whole files from a peer"),
("read_needle_blob", "replica sync, vacuum and EC rebuild read needles from a peer"),
("read_needle_meta", "replica sync compares needle metadata across peers"),
(
"copy_file",
"replica sync and EC task pull whole files from a peer",
),
(
"read_needle_blob",
"replica sync, vacuum and EC rebuild read needles from a peer",
),
(
"read_needle_meta",
"replica sync compares needle metadata across peers",
),
("write_needle_blob", "replica sync repairs a peer's needle"),
("receive_file", "EC shard distribution pushes shards to a peer"),
("read_volume_file_status", "the copy path queries the source volume server"),
("volume_ec_shard_read", "a volume server reads EC shards held by a peer"),
("volume_ec_blob_delete", "EC delete is fanned out to the shard holders"),
("volume_ec_shards_info", "EC verification polls shard holders"),
("volume_ec_shards_mount", "EC shard distribution mounts on the receiving peer"),
("volume_incremental_copy", "volume backup pulls increments from a peer"),
("volume_sync_status", "sync compares volume state across peers"),
("volume_tail_sender", "the tail source streams to the receiving peer"),
("volume_status", "replica sync and the master's vacuum loop poll volume status"),
(
"receive_file",
"EC shard distribution pushes shards to a peer",
),
(
"read_volume_file_status",
"the copy path queries the source volume server",
),
(
"volume_ec_shard_read",
"a volume server reads EC shards held by a peer",
),
(
"volume_ec_blob_delete",
"EC delete is fanned out to the shard holders",
),
(
"volume_ec_shards_info",
"EC verification polls shard holders",
),
(
"volume_ec_shards_mount",
"EC shard distribution mounts on the receiving peer",
),
(
"volume_incremental_copy",
"volume backup pulls increments from a peer",
),
(
"volume_sync_status",
"sync compares volume state across peers",
),
(
"volume_tail_sender",
"the tail source streams to the receiving peer",
),
(
"volume_status",
"replica sync and the master's vacuum loop poll volume status",
),
// Read-only or liveness: no state change.
("ping", "liveness probe"),
("get_state", "read-only volume server state"),
("query", "read-only data query"),
("vacuum_volume_check", "read-only garbage ratio; the vacuum steps that act on it are gated"),
("volume_server_status", "read-only status, the gRPC counterpart of the /status page"),
(
"vacuum_volume_check",
"read-only garbage ratio; the vacuum steps that act on it are gated",
),
(
"volume_server_status",
"read-only status, the gRPC counterpart of the /status page",
),
]
.into_iter()
.collect()
@@ -129,5 +174,9 @@ fn volume_server_admin_auth_coverage() {
}
}
assert!(problems.is_empty(), "admin-auth coverage gaps:\n{}", problems.join("\n"));
assert!(
problems.is_empty(),
"admin-auth coverage gaps:\n{}",
problems.join("\n")
);
}
+12 -15
View File
@@ -12,12 +12,13 @@ use tower::ServiceExt; // for `oneshot`
use seaweed_volume::security::{Guard, SigningKey};
use seaweed_volume::server::volume_server::{
build_admin_router, build_admin_router_with_ui, build_metrics_router, build_public_router,
VolumeServerState,
VolumeServerState, build_admin_router, build_admin_router_with_ui, build_metrics_router,
build_public_router,
};
use seaweed_volume::storage::needle_map::NeedleMapKind;
use seaweed_volume::storage::store::Store;
use seaweed_volume::storage::types::{DiskType, Version, VolumeId};
use seaweed_volume::storage::types::{DiskType, VolumeId};
use seaweed_volume::storage::volume::VolumeSpec;
use tempfile::TempDir;
@@ -73,12 +74,11 @@ fn build_test_state(
store
.add_volume(
VolumeId(1),
"",
replica_placement,
None,
0,
DiskType::HardDrive,
Version::current(),
&VolumeSpec {
replica_placement,
..Default::default()
},
)
.expect("failed to create volume");
@@ -957,9 +957,10 @@ async fn replicate_write_does_not_re_replicate() {
#[tokio::test]
async fn chunk_manifest_expands_chunk_stored_on_ec_volume() {
use seaweed_volume::storage::erasure_coding::ec_encoder::write_ec_files;
use seaweed_volume::storage::erasure_coding::ec_shard::ShardId;
use seaweed_volume::storage::needle::needle::{FileId, Needle};
use seaweed_volume::storage::types::{Cookie, NeedleId};
use seaweed_volume::storage::volume::Volume;
use seaweed_volume::storage::volume::{Volume, VolumeSpec};
let (state, tmp) = test_state();
let dir = tmp.path().to_str().unwrap();
@@ -976,13 +977,9 @@ async fn chunk_manifest_expands_chunk_stored_on_ec_volume() {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
let mut n = Needle {
@@ -1002,7 +999,7 @@ async fn chunk_manifest_expands_chunk_stored_on_ec_volume() {
// after ec.encode retired the regular volume.
{
let mut store = state.store.write().unwrap();
let shard_ids: Vec<u32> = (0..14).collect();
let shard_ids: Vec<ShardId> = (0..14).collect();
store.mount_ec_shards(VolumeId(2), "", &shard_ids).unwrap();
}
+102 -209
View File
@@ -394,28 +394,6 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "async-stream"
version = "0.3.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0b5a71a6f37880a80d1d7f19efd781e4b5de42c88f0722cc13bcb6cc2cfe8476"
dependencies = [
"async-stream-impl",
"futures-core",
"pin-project-lite",
]
[[package]]
name = "async-stream-impl"
version = "0.3.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c7c24de15d275a1ecfd47a380fb4d5ec9bfe0933f309ed5e705b775596a3574d"
dependencies = [
"proc-macro2",
"quote",
"syn 2.0.119",
]
[[package]]
name = "async-trait"
version = "0.1.92"
@@ -701,7 +679,7 @@ dependencies = [
"rustls-pki-types",
"tokio",
"tokio-rustls",
"tower 0.5.3",
"tower",
"tracing",
]
@@ -856,13 +834,13 @@ dependencies = [
[[package]]
name = "axum"
version = "0.7.9"
version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "edca88bc138befd0323b20752846e6587272d3b03b0343c8ea28a6f819e6e71f"
checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90"
dependencies = [
"async-trait",
"axum-core",
"bytes",
"form_urlencoded",
"futures-util",
"http 1.5.0",
"http-body 1.1.0",
@@ -875,14 +853,13 @@ dependencies = [
"mime",
"percent-encoding",
"pin-project-lite",
"rustversion",
"serde",
"serde_core",
"serde_json",
"serde_path_to_error",
"serde_urlencoded",
"sync_wrapper",
"tokio",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -890,19 +867,17 @@ dependencies = [
[[package]]
name = "axum-core"
version = "0.4.5"
version = "0.5.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09f2bd6146b97ae3359fa0cc6d6b376d9539582c7b4220f041a33ec24c226199"
checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1"
dependencies = [
"async-trait",
"bytes",
"futures-util",
"futures-core",
"http 1.5.0",
"http-body 1.1.0",
"http-body-util",
"mime",
"pin-project-lite",
"rustversion",
"sync_wrapper",
"tower-layer",
"tower-service",
@@ -1965,7 +1940,7 @@ dependencies = [
"indexmap 2.14.0",
"itertools",
"parking_lot",
"petgraph 0.8.3",
"petgraph",
"tokio",
]
@@ -2793,7 +2768,7 @@ dependencies = [
"libc",
"percent-encoding",
"pin-project-lite",
"socket2 0.6.5",
"socket2",
"tokio",
"tower-service",
"tracing",
@@ -3249,9 +3224,9 @@ dependencies = [
"object_store",
"permutation",
"pin-project",
"prost 0.14.4",
"prost-build 0.14.4",
"prost-types 0.14.4",
"prost",
"prost-build",
"prost-types",
"rand 0.9.5",
"rayon",
"roaring",
@@ -3357,7 +3332,7 @@ dependencies = [
"num_cpus",
"object_store",
"pin-project",
"prost 0.14.4",
"prost",
"quick_cache",
"rand 0.9.5",
"roaring",
@@ -3398,8 +3373,8 @@ dependencies = [
"lance-datagen",
"log",
"pin-project",
"prost 0.14.4",
"prost-build 0.14.4",
"prost",
"prost-build",
"tokio",
"tracing",
]
@@ -3461,8 +3436,8 @@ dependencies = [
"log",
"lz4",
"num-traits",
"prost 0.14.4",
"prost-build 0.14.4",
"prost",
"prost-build",
"rand 0.9.5",
"strum",
"tokio",
@@ -3497,9 +3472,9 @@ dependencies = [
"log",
"num-traits",
"object_store",
"prost 0.14.4",
"prost-build 0.14.4",
"prost-types 0.14.4",
"prost",
"prost-build",
"prost-types",
"tokio",
"tracing",
]
@@ -3554,9 +3529,9 @@ dependencies = [
"ndarray",
"num-traits",
"object_store",
"prost 0.14.4",
"prost-build 0.14.4",
"prost-types 0.14.4",
"prost",
"prost-build",
"prost-types",
"rand 0.9.5",
"rand_distr",
"rangemap",
@@ -3590,7 +3565,7 @@ dependencies = [
"lance-core",
"lance-io",
"lance-select",
"prost-types 0.14.4",
"prost-types",
"roaring",
"serde",
"serde_json",
@@ -3624,7 +3599,7 @@ dependencies = [
"opendal",
"path_abs",
"pin-project",
"prost 0.14.4",
"prost",
"rand 0.9.5",
"serde",
"tempfile",
@@ -3719,9 +3694,9 @@ dependencies = [
"lance-select",
"log",
"object_store",
"prost 0.14.4",
"prost-build 0.14.4",
"prost-types 0.14.4",
"prost",
"prost-build",
"prost-types",
"rand 0.9.5",
"rangemap",
"roaring",
@@ -3938,9 +3913,9 @@ dependencies = [
[[package]]
name = "matchit"
version = "0.7.3"
version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0e7465ac9959cc2b1404e8e2367b43684a6d13790fe23056cc8c6c5a6b7bcb94"
checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3"
[[package]]
name = "matrixmultiply"
@@ -4411,16 +4386,6 @@ version = "0.4.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "df202b0b0f5b8e389955afd5f27b007b00fb948162953f1db9c70d2c7e3157d7"
[[package]]
name = "petgraph"
version = "0.7.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3672b37090dbd86368a4145bc067582552b29c27377cad4e0a306c97f9bd7772"
dependencies = [
"fixedbitset",
"indexmap 2.14.0",
]
[[package]]
name = "petgraph"
version = "0.8.3"
@@ -4563,16 +4528,6 @@ dependencies = [
"thiserror 1.0.69",
]
[[package]]
name = "prost"
version = "0.13.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2796faa41db3ec313a31f7624d9286acf277b52de526150b7e69f3debf891ee5"
dependencies = [
"bytes",
"prost-derive 0.13.5",
]
[[package]]
name = "prost"
version = "0.14.4"
@@ -4580,27 +4535,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "528ac67416ff8646872a3c02cad9cc4ee5dc9f9540c9b10771855c95cb2e5ae1"
dependencies = [
"bytes",
"prost-derive 0.14.4",
]
[[package]]
name = "prost-build"
version = "0.13.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
dependencies = [
"heck",
"itertools",
"log",
"multimap",
"once_cell",
"petgraph 0.7.1",
"prettyplease",
"prost 0.13.5",
"prost-types 0.13.5",
"regex",
"syn 2.0.119",
"tempfile",
"prost-derive",
]
[[package]]
@@ -4613,28 +4548,17 @@ dependencies = [
"itertools",
"log",
"multimap",
"petgraph 0.8.3",
"petgraph",
"prettyplease",
"prost 0.14.4",
"prost-types 0.14.4",
"prost",
"prost-types",
"pulldown-cmark",
"pulldown-cmark-to-cmark",
"regex",
"syn 2.0.119",
"tempfile",
]
[[package]]
name = "prost-derive"
version = "0.13.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
dependencies = [
"anyhow",
"itertools",
"proc-macro2",
"quote",
"syn 2.0.119",
]
[[package]]
name = "prost-derive"
version = "0.14.4"
@@ -4648,22 +4572,13 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "prost-types"
version = "0.13.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "52c2c1bf36ddb1a1c396b3601a3cec27c2462e45f07c386894ec3ccf5332bd16"
dependencies = [
"prost 0.13.5",
]
[[package]]
name = "prost-types"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f94967dc7688f3054c7fac87473ffae4cc4c3904800e2d9f5b857246d8963b0a"
dependencies = [
"prost 0.14.4",
"prost",
]
[[package]]
@@ -4730,6 +4645,26 @@ version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.13.1",
"memchr",
"unicase",
]
[[package]]
name = "pulldown-cmark-to-cmark"
version = "22.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab1ad36992cead65f02aa399a373a42730922f1525d988172634fdefdecb8a60"
dependencies = [
"pulldown-cmark",
]
[[package]]
name = "quick-xml"
version = "0.39.4"
@@ -4775,7 +4710,7 @@ dependencies = [
"quinn-udp",
"rustc-hash",
"rustls",
"socket2 0.6.5",
"socket2",
"thiserror 2.0.20",
"tokio",
"tracing",
@@ -4814,7 +4749,7 @@ dependencies = [
"cfg_aliases",
"libc",
"once_cell",
"socket2 0.6.5",
"socket2",
"tracing",
"windows-sys 0.61.2",
]
@@ -4846,24 +4781,13 @@ version = "0.7.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dc33ff2d4973d518d823d61aa239014831e521c75da58e3df4840d3f47749d09"
[[package]]
name = "rand"
version = "0.8.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "22f6172bdec972074665ed81ed53b71da00bfc44b65a753cfde883ec4c702a1a"
dependencies = [
"libc",
"rand_chacha 0.3.1",
"rand_core 0.6.4",
]
[[package]]
name = "rand"
version = "0.9.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b9ef1d0d795eb7d84685bca4f72f3649f064e6641543d3a8c415898726a57b41"
dependencies = [
"rand_chacha 0.9.0",
"rand_chacha",
"rand_core 0.9.5",
]
@@ -4878,16 +4802,6 @@ dependencies = [
"rand_core 0.10.1",
]
[[package]]
name = "rand_chacha"
version = "0.3.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e6c10a63a0fa32252be49d21e7709d4d4baf8d231c2dbce1eaa8141b9b127d88"
dependencies = [
"ppv-lite86",
"rand_core 0.6.4",
]
[[package]]
name = "rand_chacha"
version = "0.9.0"
@@ -4898,15 +4812,6 @@ dependencies = [
"rand_core 0.9.5",
]
[[package]]
name = "rand_core"
version = "0.6.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ec0be4795e2f6a28069bec0b5ff3e2ac9bafc99e6a9a7dc3547996c5c816922c"
dependencies = [
"getrandom 0.2.17",
]
[[package]]
name = "rand_core"
version = "0.9.5"
@@ -5160,7 +5065,7 @@ dependencies = [
"tokio",
"tokio-rustls",
"tokio-util",
"tower 0.5.3",
"tower",
"tower-http",
"tower-service",
"url",
@@ -5199,7 +5104,7 @@ dependencies = [
"tokio",
"tokio-rustls",
"tokio-util",
"tower 0.5.3",
"tower",
"tower-http",
"tower-service",
"url",
@@ -5309,15 +5214,6 @@ dependencies = [
"security-framework",
]
[[package]]
name = "rustls-pemfile"
version = "2.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dce314e5fee3f39953d46bb63bb8a46d40c2f8fb7cc5a3b6cab2bde9721d6e50"
dependencies = [
"rustls-pki-types",
]
[[package]]
name = "rustls-pki-types"
version = "1.15.1"
@@ -5441,13 +5337,14 @@ dependencies = [
"async-trait",
"axum",
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"prost",
"prost-types",
"protoc-bin-vendored",
"tokio",
"tokio-stream",
"tonic",
"tonic-build",
"tonic-prost",
"tonic-prost-build",
"tracing",
]
@@ -5723,16 +5620,6 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "socket2"
version = "0.5.10"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e22376abed350d73dd1cd119b57ffccad95b4e585a7cda43e286245ce23c0678"
dependencies = [
"libc",
"windows-sys 0.52.0",
]
[[package]]
name = "socket2"
version = "0.6.5"
@@ -6027,7 +5914,7 @@ dependencies = [
"parking_lot",
"pin-project-lite",
"signal-hook-registry",
"socket2 0.6.5",
"socket2",
"tokio-macros",
"windows-sys 0.61.2",
]
@@ -6081,11 +5968,10 @@ dependencies = [
[[package]]
name = "tonic"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "877c5b330756d856ffcc4553ab34a5684481ade925ecc54bcd1bf02b1d0d4d52"
checksum = "ac2a5518c70fa84342385732db33fb3f44bc4cc748936eb5833d2df34d6445ef"
dependencies = [
"async-stream",
"async-trait",
"axum",
"base64 0.22.1",
@@ -6099,13 +5985,12 @@ dependencies = [
"hyper-util",
"percent-encoding",
"pin-project",
"prost 0.13.5",
"rustls-pemfile",
"socket2 0.5.10",
"socket2",
"sync_wrapper",
"tokio",
"tokio-rustls",
"tokio-stream",
"tower 0.4.13",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -6113,36 +5998,41 @@ dependencies = [
[[package]]
name = "tonic-build"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9557ce109ea773b399c9b9e5dca39294110b74f1f342cb347a80d1fce8c26a11"
checksum = "c68f61875ac5293cf72e6c8cf0158086428c82c37229e98c840878f1706b0322"
dependencies = [
"prettyplease",
"proc-macro2",
"prost-build 0.13.5",
"prost-types 0.13.5",
"quote",
"syn 2.0.119",
]
[[package]]
name = "tower"
version = "0.4.13"
name = "tonic-prost"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8fa9be0de6cf49e536ce1851f987bd21a43b771b09473c3549a6c853db37c1c"
checksum = "50849f68853be452acf590cde0b146665b8d507b3b8af17261df47e02c209ea0"
dependencies = [
"futures-core",
"futures-util",
"indexmap 1.9.3",
"pin-project",
"pin-project-lite",
"rand 0.8.7",
"slab",
"tokio",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
"bytes",
"prost",
"tonic",
]
[[package]]
name = "tonic-prost-build"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "654e5643eff75d7f8c99197ce1440ed19a3474eada74c12bbac488b2cafdae27"
dependencies = [
"prettyplease",
"proc-macro2",
"prost-build",
"prost-types",
"quote",
"syn 2.0.119",
"tempfile",
"tonic-build",
]
[[package]]
@@ -6153,9 +6043,12 @@ checksum = "ebe5ef63511595f1344e2d5cfa636d973292adc0eec1f0ad45fae9f0851ab1d4"
dependencies = [
"futures-core",
"futures-util",
"indexmap 2.14.0",
"pin-project-lite",
"slab",
"sync_wrapper",
"tokio",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
@@ -6178,7 +6071,7 @@ dependencies = [
"pin-project-lite",
"tokio",
"tokio-util",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"url",
+11 -5
View File
@@ -16,17 +16,23 @@ edition = "2024"
# lance's `aws` feature pulls in).
rust-version = "1.94.1"
[workspace.lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[workspace.dependencies]
anyhow = "1"
async-trait = "0.1"
prost = "0.13"
prost-types = "0.13"
prost = "0.14"
prost-types = "0.14"
tokio = { version = "1", features = ["full"] }
tokio-stream = "0.1"
tonic = { version = "0.12", features = ["tls"] }
tonic = { version = "0.14", features = ["tls-aws-lc"] }
tonic-prost = "0.14"
# Already in the tree via tonic; named here so the metrics server can use them.
axum = "0.7"
axum = "0.8"
prometheus = { version = "0.13", default-features = false }
tonic-build = "0.12"
tonic-prost-build = "0.14"
tracing = "0.1"
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
+5 -1
View File
@@ -16,13 +16,17 @@ prost-types.workspace = true
tokio.workspace = true
tokio-stream.workspace = true
tonic.workspace = true
tonic-prost.workspace = true
axum.workspace = true
prometheus.workspace = true
tracing.workspace = true
[build-dependencies]
tonic-build.workspace = true
tonic-prost-build.workspace = true
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry. The same crate seaweed-volume uses.
protoc-bin-vendored = "3"
[lints]
workspace = true
+1 -1
View File
@@ -14,7 +14,7 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
// Compiled straight out of the Go tree, the way seaweed-volume already reads
// filer.proto, so the contract cannot drift from a vendored copy.
tonic_build::configure()
tonic_prost_build::configure()
// The server half is only for tests, which stand up a fake admin.
.build_server(true)
.build_client(true)
+4 -4
View File
@@ -14,10 +14,10 @@ pub fn server_to_grpc_address(server: &str) -> Option<String> {
let (host, port_part) = server.rsplit_once(':')?;
// "port.grpcPort" states the gRPC port outright.
if let Some((_, grpc_port)) = port_part.split_once('.') {
if let Ok(port) = grpc_port.parse::<u16>() {
return Some(join_host_port(host, port));
}
if let Some((_, grpc_port)) = port_part.split_once('.')
&& let Ok(port) = grpc_port.parse::<u16>()
{
return Some(join_host_port(host, port));
}
let port: u16 = port_part.parse().ok()?;
@@ -8,7 +8,7 @@
use std::collections::HashMap;
use crate::pb::{
config_value::Kind, ConfigField, ConfigFieldType, ConfigForm, ConfigSection, ConfigValue,
ConfigField, ConfigFieldType, ConfigForm, ConfigSection, ConfigValue, config_value::Kind,
};
pub fn int_value(value: i64) -> ConfigValue {
+3
View File
@@ -16,6 +16,9 @@ pub mod stream;
/// Generated plugin.proto types.
pub mod pb {
// prost gives every oneof its own enum; the variant sizes are the
// messages' own, not a choice made here.
#![allow(clippy::large_enum_variant)]
tonic::include_proto!("plugin");
}
+24 -15
View File
@@ -270,10 +270,10 @@ impl Metrics {
/// logged rather than fatal: a worker that cannot publish metrics should still
/// do its work.
pub async fn serve(metrics: Metrics, addr: SocketAddr) -> Result<()> {
use axum::Router;
use axum::extract::State;
use axum::http::StatusCode;
use axum::routing::get;
use axum::Router;
let app = Router::new()
.route("/health", get(|| async { StatusCode::OK }))
@@ -355,28 +355,37 @@ mod tests {
metrics.stream_connected();
assert!(metrics.is_ready());
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 1"));
assert!(
metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 1")
);
metrics.stream_ended("closed");
assert!(!metrics.is_ready());
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 0"));
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_stream_events_total{event=\"closed\"} 1"));
assert!(
metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 0")
);
assert!(
metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_stream_events_total{event=\"closed\"} 1")
);
}
#[test]
fn build_info_names_the_worker() {
let text = metrics().gather().expect("gather");
assert!(text
.contains("SeaweedFS_worker_build_info{version=\"0.1.0\",worker_id=\"worker-1\"} 1"));
assert!(
text.contains(
"SeaweedFS_worker_build_info{version=\"0.1.0\",worker_id=\"worker-1\"} 1"
)
);
}
// A format's own numbers land on the same registry, so one endpoint serves
+2 -2
View File
@@ -2,8 +2,8 @@ use anyhow::Result;
use tokio::sync::mpsc;
use crate::pb::{
worker_to_admin_message::Body, ActivityEvent, DetectionComplete, DetectionProposals,
JobCompleted, JobProgressUpdate, WorkerObservations, WorkerToAdminMessage,
ActivityEvent, DetectionComplete, DetectionProposals, JobCompleted, JobProgressUpdate,
WorkerObservations, WorkerToAdminMessage, worker_to_admin_message::Body,
};
/// Replies to one detection request.
+5 -5
View File
@@ -1,8 +1,8 @@
use std::sync::Arc;
use std::time::{Duration, Instant};
use anyhow::{anyhow, Context, Result};
use tokio::sync::{mpsc, Semaphore};
use anyhow::{Context, Result, anyhow};
use tokio::sync::{Semaphore, mpsc};
use tokio_stream::wrappers::UnboundedReceiverStream;
use tonic::transport::{Certificate, Channel, ClientTlsConfig, Identity};
use tracing::{info, warn};
@@ -10,11 +10,11 @@ use tracing::{info, warn};
use crate::config::WorkerOptions;
use crate::metrics::Metrics;
use crate::pb::{
ConfigSchemaResponse, ExecuteJobRequest, JobCompleted, ObjectPreviewResponse, PreviewRow,
RequestObjectPreview, RunDetectionRequest, RunningWork, WorkerHeartbeat, WorkerHello,
admin_to_worker_message::Body as AdminBody,
plugin_control_service_client::PluginControlServiceClient,
worker_to_admin_message::Body as WorkerBody, ConfigSchemaResponse, ExecuteJobRequest,
JobCompleted, ObjectPreviewResponse, PreviewRow, RequestObjectPreview, RunDetectionRequest,
RunningWork, WorkerHeartbeat, WorkerHello,
worker_to_admin_message::Body as WorkerBody,
};
use crate::registry::Registry;
use crate::senders::{MeteredSender, StreamSender};
+3
View File
@@ -50,3 +50,6 @@ arrow-array = "58"
arrow-schema = "58"
arrow-cast = "58"
lance-linalg = "10"
[lints]
workspace = true
@@ -7,4 +7,4 @@
pub mod namespace;
pub use namespace::{parse_id, NamespaceClient, TableDescription};
pub use namespace::{NamespaceClient, TableDescription, parse_id};
+1 -1
View File
@@ -8,8 +8,8 @@
use std::collections::HashMap;
use anyhow::{Context, Result};
use lance::dataset::builder::DatasetBuilder;
use lance::dataset::Dataset;
use lance::dataset::builder::DatasetBuilder;
use crate::catalog::{NamespaceClient, TableDescription};
@@ -1,9 +1,9 @@
use std::collections::HashMap;
use anyhow::{anyhow, Context, Result};
use anyhow::{Context, Result, anyhow};
use async_trait::async_trait;
use chrono::{Duration, Utc};
use lance::dataset::cleanup::{cleanup_old_versions, CleanupPolicy};
use lance::dataset::cleanup::{CleanupPolicy, cleanup_old_versions};
use seaweed_worker_core::config_form::{form, int_or, int_value, number_field};
use seaweed_worker_core::pb::{
ConfigValue, DetectionComplete, DetectionProposals, ExecuteJobRequest, JobCompleted,
@@ -13,7 +13,7 @@ use seaweed_worker_core::pb::{
use seaweed_worker_core::{DetectionSender, ExecutionSender, JobHandler};
use tracing::warn;
use crate::catalog::{parse_id, NamespaceClient};
use crate::catalog::{NamespaceClient, parse_id};
use crate::dataset;
use crate::jobs::{clamp, string_list, table_id};
@@ -1,8 +1,8 @@
use std::collections::HashMap;
use anyhow::{anyhow, Context, Result};
use anyhow::{Context, Result, anyhow};
use async_trait::async_trait;
use lance::dataset::optimize::{compact_files, CompactionOptions};
use lance::dataset::optimize::{CompactionOptions, compact_files};
use seaweed_worker_core::config_form::{form, int_or, int_value, number_field, string_value};
use seaweed_worker_core::pb::{
ConfigValue, DetectionComplete, DetectionProposals, ExecuteJobRequest, JobCompleted,
@@ -12,9 +12,9 @@ use seaweed_worker_core::pb::{
use seaweed_worker_core::{DetectionSender, ExecutionSender, JobHandler};
use tracing::warn;
use crate::catalog::{parse_id, NamespaceClient};
use crate::catalog::{NamespaceClient, parse_id};
use crate::dataset;
use crate::jobs::{clamp, observation, string_list, table_id, FORMAT};
use crate::jobs::{FORMAT, clamp, observation, string_list, table_id};
pub const JOB_TYPE: &str = "lance_compact";
@@ -1,6 +1,6 @@
use std::collections::HashMap;
use anyhow::{anyhow, Context, Result};
use anyhow::{Context, Result, anyhow};
use async_trait::async_trait;
use lance::index::DatasetIndexExt;
use lance_index::optimize::OptimizeOptions;
@@ -13,7 +13,7 @@ use seaweed_worker_core::pb::{
use seaweed_worker_core::{DetectionSender, ExecutionSender, JobHandler};
use tracing::warn;
use crate::catalog::{parse_id, NamespaceClient};
use crate::catalog::{NamespaceClient, parse_id};
use crate::dataset::{self, OpenTable};
use crate::jobs::{clamp, string_list, table_id};
+1 -1
View File
@@ -14,8 +14,8 @@ pub mod sort;
use std::collections::HashMap;
use std::sync::Arc;
use seaweed_worker_core::pb::{config_value::Kind, ConfigValue, ObjectObservation, StringList};
use seaweed_worker_core::JobHandler;
use seaweed_worker_core::pb::{ConfigValue, ObjectObservation, StringList, config_value::Kind};
use crate::catalog::parse_id;
+7 -7
View File
@@ -1,14 +1,14 @@
use std::collections::HashMap;
use std::sync::Arc;
use anyhow::{anyhow, Context, Result};
use anyhow::{Context, Result, anyhow};
use async_trait::async_trait;
use lance::dataset::scanner::ColumnOrdering;
use lance::dataset::transaction::Operation;
use lance::dataset::write::{CommitBuilder, InsertBuilder};
use lance::dataset::{WriteDestination, WriteMode, WriteParams};
use lance::index::DatasetIndexExt;
use lance_datafusion::exec::{execute_plan, LanceExecutionOptions};
use lance_datafusion::exec::{LanceExecutionOptions, execute_plan};
use seaweed_worker_core::config_form::{form, int_or, int_value, string_or, string_value};
use seaweed_worker_core::pb::{
ConfigValue, DetectionComplete, DetectionProposals, ExecuteJobRequest, JobCompleted,
@@ -17,15 +17,15 @@ use seaweed_worker_core::pb::{
};
use seaweed_worker_core::{DetectionSender, ExecutionSender, JobHandler};
use seaweed_worker_sort::{
config_fields, resolve, verdict, FragmentSummary, SortSpec, SortState,
CONFIG_MAX_ROWS_PER_FILE, CONFIG_MEMORY_BUDGET_MB, CONFIG_MIN_UNSORTED_ROWS,
CONFIG_SORT_FIELDS, DECLARED_FIELDS_KEY,
CONFIG_SORT_FIELDS, DECLARED_FIELDS_KEY, FragmentSummary, SortSpec, SortState, config_fields,
resolve, verdict,
};
use tracing::warn;
use crate::catalog::{parse_id, NamespaceClient};
use crate::catalog::{NamespaceClient, parse_id};
use crate::dataset;
use crate::jobs::{clamp, observation, string_list, table_id, FORMAT};
use crate::jobs::{FORMAT, clamp, observation, string_list, table_id};
pub const JOB_TYPE: &str = "lance_sort";
@@ -424,7 +424,7 @@ impl JobHandler for SortHandler {
other => {
return Err(anyhow!(
"a sorted rewrite produced {other} instead of an overwrite"
))
));
}
}
@@ -10,8 +10,8 @@ use std::sync::Mutex;
use anyhow::Result;
use seaweed_worker_core::pb::{
config_value::Kind, ActivityEvent, ConfigValue, DetectionComplete, DetectionProposals,
JobCompleted, JobProgressUpdate, JobProposal, ObjectObservation, WorkerObservations,
ActivityEvent, ConfigValue, DetectionComplete, DetectionProposals, JobCompleted,
JobProgressUpdate, JobProposal, ObjectObservation, WorkerObservations, config_value::Kind,
};
use seaweed_worker_core::{DetectionSender, ExecutionSender};
@@ -8,7 +8,7 @@ use std::collections::HashMap;
use anyhow::Result;
use seaweed_worker_core::pb::{
config_value::Kind, ExecuteJobRequest, JobSpec, RunDetectionRequest,
ExecuteJobRequest, JobSpec, RunDetectionRequest, config_value::Kind,
};
use seaweed_worker_core::{JobHandler, PreviewProvider};
use weed_lance_worker::catalog::NamespaceClient;
@@ -18,7 +18,7 @@ use weed_lance_worker::jobs::indices::OptimizeIndicesHandler;
use weed_lance_worker::preview::LancePreview;
mod common;
use common::{fallback, int_config, namespace_url, Recorder};
use common::{Recorder, fallback, int_config, namespace_url};
/// These tests drive one live gateway and one shared catalog: `list_all_tables`
/// sweeps everything, so a table another test is writing shows up in this test's
@@ -121,10 +121,10 @@ async fn seed_table(
// The index is built after the first batch, so everything appended
// afterwards is a row it does not cover.
if with_index && i == 0 {
use lance::index::vector::VectorIndexParams;
use lance::index::DatasetIndexExt;
use lance_index::vector::{ivf::IvfBuildParams, pq::PQBuildParams};
use lance::index::vector::VectorIndexParams;
use lance_index::IndexType;
use lance_index::vector::{ivf::IvfBuildParams, pq::PQBuildParams};
let mut dataset = dataset;
let params = VectorIndexParams::with_ivf_pq_params(
@@ -12,13 +12,13 @@
use std::collections::HashMap;
use seaweed_worker_core::pb::{ConfigValue, ExecuteJobRequest, JobSpec, RunDetectionRequest};
use seaweed_worker_core::JobHandler;
use seaweed_worker_core::pb::{ConfigValue, ExecuteJobRequest, JobSpec, RunDetectionRequest};
use weed_lance_worker::jobs::cleanup::{self, CleanupVersionsHandler};
use weed_lance_worker::jobs::compact::{CompactHandler, JOB_TYPE as COMPACT_JOB_TYPE};
mod common;
use common::{fallback, int_config, namespace_url, Recorder};
use common::{Recorder, fallback, int_config, namespace_url};
fn table() -> Option<String> {
std::env::var("WEED_LANCE_TABLE")
+5 -5
View File
@@ -5,15 +5,15 @@
//! the commit is the half worth testing.
use anyhow::Result;
use seaweed_worker_core::pb::{
config_value::Kind, ConfigValue, ExecuteJobRequest, JobSpec, RunDetectionRequest,
};
use seaweed_worker_core::JobHandler;
use seaweed_worker_core::pb::{
ConfigValue, ExecuteJobRequest, JobSpec, RunDetectionRequest, config_value::Kind,
};
use weed_lance_worker::catalog::NamespaceClient;
use weed_lance_worker::jobs::sort::{SortHandler, JOB_TYPE};
use weed_lance_worker::jobs::sort::{JOB_TYPE, SortHandler};
mod common;
use common::{fallback, namespace_url, Recorder};
use common::{Recorder, fallback, namespace_url};
/// One live gateway and one shared catalog, and `list_all_tables` sweeps
/// everything, so these tests take a lock the way the compaction ones do.
+3
View File
@@ -11,3 +11,6 @@ name = "seaweed_worker_sort"
[dependencies]
seaweed-worker-core = { path = "../core" }
anyhow.workspace = true
[lints]
workspace = true
+5 -5
View File
@@ -14,7 +14,7 @@
use std::collections::{HashMap, HashSet};
use std::fmt;
use anyhow::{bail, Context, Result};
use anyhow::{Context, Result, bail};
use seaweed_worker_core::config_form::{number_field, text_field};
use seaweed_worker_core::pb::ConfigField;
@@ -152,10 +152,10 @@ fn parse_field(entry: &str) -> Result<SortField> {
/// back: sorting by the worker's default order instead of the one the table
/// asked for would silently rewrite the table the wrong way.
pub fn resolve(declared: Option<&str>, configured: &str) -> Result<Option<SortSpec>> {
if let Some(declared) = declared {
if let Some(spec) = SortSpec::parse(declared).context("read the table's declared order")? {
return Ok(Some(spec));
}
if let Some(declared) = declared
&& let Some(spec) = SortSpec::parse(declared).context("read the table's declared order")?
{
return Ok(Some(spec));
}
SortSpec::parse(configured).context("read the configured sort order")
}
+2 -1
View File
@@ -12,7 +12,8 @@ WORKDIR /app/telemetry/server
RUN CGO_ENABLED=0 GOOS=linux go build -a -installsuffix cgo -ldflags '-extldflags "-static"' -o telemetry-server .
FROM alpine:latest
RUN apk --no-cache add ca-certificates \
RUN apk upgrade --no-cache \
&& apk --no-cache add ca-certificates libcrypto3 libssl3 \
&& addgroup -S appgroup \
&& adduser -S appuser -G appgroup
+4 -4
View File
@@ -1,9 +1,10 @@
module github.com/seaweedfs/seaweedfs/telemetry/server
go 1.26
go 1.26.6
require (
github.com/prometheus/client_golang v1.24.1
github.com/prometheus/client_model v0.6.3
github.com/seaweedfs/seaweedfs v0.0.0-00010101000000-000000000000
google.golang.org/protobuf v1.36.12
)
@@ -12,10 +13,9 @@ require (
github.com/beorn7/perks v1.0.1 // indirect
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 // indirect
github.com/prometheus/client_model v0.6.2 // indirect
github.com/prometheus/common v0.70.1 // indirect
github.com/prometheus/procfs v0.21.1 // indirect
golang.org/x/sys v0.47.0 // indirect
github.com/prometheus/procfs v0.22.0 // indirect
golang.org/x/sys v0.48.0 // indirect
)
replace github.com/seaweedfs/seaweedfs => ../..
+8 -8
View File
@@ -6,8 +6,8 @@ github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/klauspost/compress v1.19.1 h1:VsB4HPswih7mmZ8WleSFQ75c/Ui1M4trX5oAsJnhSlk=
github.com/klauspost/compress v1.19.1/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/compress v1.19.2 h1:hMRETovs/pu/dVWN7zIT1PGG8t509MwT6bO7XSi26R8=
github.com/klauspost/compress v1.19.2/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc=
github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw=
github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA=
@@ -16,20 +16,20 @@ github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRI
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_golang v1.24.1 h1:JnJkREXzWxUdCuPFpIWZiPispT9xVV59uiuyR2bPlnU=
github.com/prometheus/client_golang v1.24.1/go.mod h1:F+oSRECHg4sse5ucfYpYDeIv/hu68Zo0uoHKetWnzcE=
github.com/prometheus/client_model v0.6.2 h1:oBsgwpGs7iVziMvrGhE53c/GrLUsZdHnqNwqPLxwZyk=
github.com/prometheus/client_model v0.6.2/go.mod h1:y3m2F6Gdpfy6Ut/GBsUqTWZqCUvMVzSfMLjcu6wAwpE=
github.com/prometheus/client_model v0.6.3 h1:O0jaTVAYNxTHYInEPFJt5I3+sN8zqBtVMPTB1qyxiEo=
github.com/prometheus/client_model v0.6.3/go.mod h1:gpN5P9S7Rr6Yr92PiQ+Ixvhf6JZEkF1dnxsYL2aPBEM=
github.com/prometheus/common v0.70.1 h1:1HvjP4D5oL3t8RsPlwxA9onvvStjtIHYE5XuuwOi/PY=
github.com/prometheus/common v0.70.1/go.mod h1:VdFUQDMZK3VLkurFUVhia6uys/0suUp86TJz5qbJRhc=
github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+OJI=
github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY=
github.com/prometheus/procfs v0.22.0 h1:6q9+/JL9IKAPbCmBrv9n5O5Ty3NKnciV5X7YGw0oics=
github.com/prometheus/procfs v0.22.0/go.mod h1:CvmFr/GVhIjIvWJZW3tgkODBQMRIf0EyWMQLHCHab58=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
google.golang.org/protobuf v1.36.12 h1:pJOKDDOyeXErUroCihFAd5LQuwXBSpVnKGrj5o/fwxc=
google.golang.org/protobuf v1.36.12/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
+1 -1
View File
@@ -206,7 +206,7 @@ func testTelemetryClient() error {
Version: "3.45",
Os: "linux/amd64",
VolumeServerCount: 3,
TotalDiskBytes: 1073741824, // 1GB
TotalDiskBytes: proto.MinDiskBytes + (1 << 30), // just above the reporting floor
TotalVolumeCount: 50,
FilerCount: 2,
BrokerCount: 1,

Some files were not shown because too many files have changed in this diff Show More