Compare commits

..
240 Commits
Author SHA1 Message Date
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Opus 5.5Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
825c3dff8b volume server: fetch only the chunks a manifest Range needs, like Go (#11546)
* volume server: read GET/HEAD needles off the store lock, and only once

The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stream needle chunks without the store lock

StreamingBody::poll_frame took store.read() and find_volume for every
chunk to compare the volume's compaction revision, dup'd the source
handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the
stream also holds a data-file read lease for its whole life, while a
writer waits for that lease under store.write(): the next chunk's
store.read() then waits for the writer and the writer for the stream.

The per-chunk re-lookup was also wrong. The stream reads a handle opened
at plan time, which pins the .dat inode the offset was resolved against;
a vacuum commit renames a new file over .dat and leaves that inode
untouched. The re-looked-up offset belongs to the new file but was read
from the old inode, so a stream whose needle a vacuum moved ended in a
checksum error. The pinned offset stays valid, so the check, and with
it every store access, is dropped, along with the now unused
re_lookup_needle_data_offset and the revision fields of the read plan.

The source is shared as an Arc instead of dup'd per chunk, and the chunk
buffer is a BytesMut that the blocking read hands back with its result,
so its allocation is reclaimed once the previous frame has been written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: split get_or_head_handler_inner into phases

get_or_head_handler_inner was a ~650-line function. Its middle resolved
the needle and set five mutable flags (stream_info, can_stream,
can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that
three if-let reply paths then re-tested, each re-checking stream_info.

It is now a 126-line orchestrator over named phases: reject_read_jwt,
proxy_missing_volume, wait_for_download_slot, parse_read_request,
read_ec_needle / read_volume_needle, etag_and_last_modified,
not_modified_response, read_response_headers, and the reply phases
stream_response, head_from_meta_response, range_from_source_response,
buffered_payload and buffered_response. The read phases return a
ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource,
Buffered) carries the NeedleStreamInfo only on the variants that use it,
so the reply is one match instead of three flag checks.

Pure refactor: every status code, header and header order, error text,
metric increment, lock and data-file lease scope, spawn_blocking
boundary and side-effect order is unchanged. Phases that can end the
request return ControlFlow<Response, T>. A Range header that is not
visible ASCII still falls through to the buffered path, as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a needle stream once its volume becomes unavailable

Taking the store lock out of StreamingBody also dropped its per-chunk
unavailable_error() check. With -hasSlowRead a writer can take the
data-file lease between chunks, fail its fsync and its truncate, and mark
the volume unavailable; the stream then kept serving the rest of the
needle from its pinned handle.

The volume's io_unavailable reason is now an Arc-shared leaf mutex that
the read plan hands to the stream. Each chunk checks it under its
data-file lease, where the writer marks it, and fails with the same
"volume is unavailable: <reason>" error the old check returned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: read a non-ASCII or empty Range header as Go does

A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed
HeaderValue::to_str at both range gates. For a needle served as stored the
handler had already chosen a meta-only read, so it fell through to the
buffered path with no payload and answered 200 with an empty body; a
compressed or EC needle answered 200 with the full body. Go's parseRange
fails on a byte it can neither trim nor parse and answers 416 "invalid
range", and trims Unicode whitespace such as NBSP into a normal 206. An
empty Range value was also a 200 with an empty body, where Go sends the
whole payload.

Read Range once with from_utf8_lossy, dropping an empty value, and hand
that one value to the read plan and to both range gates. A replaced byte
never parses, so it is a 416; str::trim trims the same Unicode whitespace
as strings.TrimSpace. A range read from the data file now always answers
itself instead of falling through with an empty needle. The buffered path
answers HEAD before it looks at Range, as Go's writeResponseContent does,
so an EC HEAD with a Range is a 200 with the full length.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: apply Range to chunk manifests and forward raw headers when proxying

A GET of a chunk manifest assembled the object and always answered 200 with
the whole body, ignoring Range. Go serves the expanded manifest through
writeResponseContent, which answers HEAD first and then hands Range to
ProcessRangeRequest: 206 for one range, multipart/byteranges for several,
416 for an unsatisfiable or unparsable one. try_expand_chunk_manifest now
returns the assembled body and headers, and the caller answers through
buffered_response, the same helper the buffered needle path uses.

A proxied read forwarded a request header only if HeaderValue::to_str
succeeded, so a Range with an obs-text byte was dropped and the target
answered 200 with the full body. Go copies every header value as is.
Forward the raw HeaderValue for every header.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: fetch only the chunks a manifest Range needs, like Go

A ranged GET of a chunk manifest fetched every chunk, assembled the whole
object and then sliced it, so reading a few bytes of a large object cost a
read of all of it, and a 416 still fetched everything. Go serves a manifest
through ChunkedFileReader, which seeks to each range and reads only the
chunks under it.

For a GET with a Range, try_expand_chunk_manifest now parses the ranges
against the manifest size, fetches only the chunks whose declared window
overlaps one of them (none when the reply carries no body), and answers
through handle_range_request_with, the reader-based core that
handle_range_request now wraps, so 206/416/multipart stay one code path.
The reader replays assembly: chunks clamped as before, later chunks over
earlier ones, zeros in gaps. HEAD, no-Range GETs and GETs that crop or
resize an image still assemble the whole object.

A missing chunk outside the requested ranges no longer turns a ranged GET
into a 500, as in Go; a missing chunk inside them still does, before any
headers are sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: reword a comment codespell flags

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: keep only range-covered bytes of fetched manifest chunks

A ranged GET retained every overlapping chunk's full contents; 1,000
overlapping 8 MiB chunks could pin ~8 GiB for a one-byte response. Clip
each fetched chunk to the bytes the requested ranges can actually read,
preserving the later-chunks-overwrite and zero-fill-gap semantics.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: bucket ranged manifest parts by range

Serving a multipart range scanned every retained part. Bucket the kept
intersections by their range so one range only reads its own parts.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:17:19 +08:00
xopGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>CursorChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a7a590474d kafka: honor notification.kafka.event_types (#11601)
* kafka: honor notification.kafka.event_types

Kafka published every filer event and ignored the filter the webhook notifier already uses.

Co-authored-by: Cursor <cursoragent@cursor.com>

* notification: share event-type classification between queues

Kafka duplicated the webhook's event classification verbatim; move it to
the notification package so the two queues cannot drift. Webhook keeps
its typed eventType wrappers over the shared helpers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 21:15:06 +08:00
zhao-ycGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>zhaoyuchenChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
8c67b75190 image: add an optional public image processing gateway (#11593)
* image: add an optional public image processing gateway

* image: fix representation metadata and processing bounds

* image: restrict passthrough to non-executable media types

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: tighten source media-type validation

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: write passthrough body on the inner response writer

CodeQL still flagged the passthrough write: the content type was set on
the wrapper while the body reached w.ResponseWriter, so the validated
header could not be associated with the write. Set headers and copy the
body on the same inner writer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* image: serve processed output on the inner response writer

* image: reject XML source types and unsafe conditional metadata

---------

Co-authored-by: zhaoyuchen <yc.zhao@yinzon.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 19:49:02 +08:00
7a19961074 volume server: read a non-ASCII or empty Range header as Go does (#11534)
* volume server: read GET/HEAD needles off the store lock, and only once

The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stream needle chunks without the store lock

StreamingBody::poll_frame took store.read() and find_volume for every
chunk to compare the volume's compaction revision, dup'd the source
handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the
stream also holds a data-file read lease for its whole life, while a
writer waits for that lease under store.write(): the next chunk's
store.read() then waits for the writer and the writer for the stream.

The per-chunk re-lookup was also wrong. The stream reads a handle opened
at plan time, which pins the .dat inode the offset was resolved against;
a vacuum commit renames a new file over .dat and leaves that inode
untouched. The re-looked-up offset belongs to the new file but was read
from the old inode, so a stream whose needle a vacuum moved ended in a
checksum error. The pinned offset stays valid, so the check, and with
it every store access, is dropped, along with the now unused
re_lookup_needle_data_offset and the revision fields of the read plan.

The source is shared as an Arc instead of dup'd per chunk, and the chunk
buffer is a BytesMut that the blocking read hands back with its result,
so its allocation is reclaimed once the previous frame has been written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: split get_or_head_handler_inner into phases

get_or_head_handler_inner was a ~650-line function. Its middle resolved
the needle and set five mutable flags (stream_info, can_stream,
can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that
three if-let reply paths then re-tested, each re-checking stream_info.

It is now a 126-line orchestrator over named phases: reject_read_jwt,
proxy_missing_volume, wait_for_download_slot, parse_read_request,
read_ec_needle / read_volume_needle, etag_and_last_modified,
not_modified_response, read_response_headers, and the reply phases
stream_response, head_from_meta_response, range_from_source_response,
buffered_payload and buffered_response. The read phases return a
ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource,
Buffered) carries the NeedleStreamInfo only on the variants that use it,
so the reply is one match instead of three flag checks.

Pure refactor: every status code, header and header order, error text,
metric increment, lock and data-file lease scope, spawn_blocking
boundary and side-effect order is unchanged. Phases that can end the
request return ControlFlow<Response, T>. A Range header that is not
visible ASCII still falls through to the buffered path, as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a needle stream once its volume becomes unavailable

Taking the store lock out of StreamingBody also dropped its per-chunk
unavailable_error() check. With -hasSlowRead a writer can take the
data-file lease between chunks, fail its fsync and its truncate, and mark
the volume unavailable; the stream then kept serving the rest of the
needle from its pinned handle.

The volume's io_unavailable reason is now an Arc-shared leaf mutex that
the read plan hands to the stream. Each chunk checks it under its
data-file lease, where the writer marks it, and fails with the same
"volume is unavailable: <reason>" error the old check returned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: read a non-ASCII or empty Range header as Go does

A Range value with a byte >= 0x80 (obs-text, which hyper accepts) failed
HeaderValue::to_str at both range gates. For a needle served as stored the
handler had already chosen a meta-only read, so it fell through to the
buffered path with no payload and answered 200 with an empty body; a
compressed or EC needle answered 200 with the full body. Go's parseRange
fails on a byte it can neither trim nor parse and answers 416 "invalid
range", and trims Unicode whitespace such as NBSP into a normal 206. An
empty Range value was also a 200 with an empty body, where Go sends the
whole payload.

Read Range once with from_utf8_lossy, dropping an empty value, and hand
that one value to the read plan and to both range gates. A replaced byte
never parses, so it is a 416; str::trim trims the same Unicode whitespace
as strings.TrimSpace. A range read from the data file now always answers
itself instead of falling through with an empty needle. The buffered path
answers HEAD before it looks at Range, as Go's writeResponseContent does,
so an EC HEAD with a Range is a 200 with the full length.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-05 19:47:54 +08:00
16e66b1bad fix(s3): initialize destination ACLs for CopyObject (#11599)
* fix(s3): initialize destination ACLs for CopyObject

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* fix(s3): re-check routed self-copy eligibility on the locked read

routeInPlace was decided on the pre-lock entry, but the PATCH body
re-reads the entry. A concurrent write changing file mode or MIME in
between left the routed PATCH installing new ACL keys while Attributes
kept the stale mode. Evaluate eligibility against the re-read entry and
retry the self-copy under the distributed lock when it no longer
qualifies.

* fix(s3): guard metadata self-copies against concurrent writes

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* ci: raise s3api unit-test timeout to 9m

The suite crossed the 5m binary timeout on the hosted runner (local run
is ~4.3m and still growing). The job-level limit is already 10m.

* ci: allow setup time before the S3 API test suite

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-05 18:22:16 +08:00
Chris LuandGitHub 7e809c9991 admin: stop leaking cancelled maintenance task files in -dataDir/tasks (#11597)
* admin: delete persisted state when scan cancels pending tasks

Each detection cycle cancels every pending task of a type before
re-detecting it, and the cancel path saved the cancelled task back to
disk. Nothing ever removed those files, so -dataDir/tasks gained one
orphaned .pb per candidate volume per scan cycle.

Cancelled is terminal, so drop the file the same way CompleteTask does
for completed/failed tasks. The cancelled entry stays in memory for the
UI until the next purge.

Refs #11595

* admin: delete persisted state when CancelTask cancels a pending task

The manual cancel path only updated memory, leaving the pending .pb on
disk where a restart would resurrect the cancelled task as pending and
the file would linger until then. Delete it like the scan-cycle cancel
path now does.

* admin: count cancelled tasks toward task retention cleanup

CleanupOldTasks and ConfigPersistence.CleanupCompletedTasks only
filtered completed/failed tasks, so cancelled entries were exempt from
retention in both memory and on disk. Treat all terminal states alike;
nil CompletedAt entries also count and sort last, so they are pruned
first.

* admin: run task file retention in the periodic cleanup loop

cleanupCompletedTasks had no callers, so the on-disk retention bound
never ran during uptime. Invoke it from performCleanup alongside the
in-memory CleanupOldTasks sweep.

* admin: guard task state writes against stale saves and failed deletes

saveTaskState runs after mq.mutex is released, so the task may have gone
terminal in between; a delayed pending save could then recreate the file
a cancel just deleted and resurrect the task on restart. Skip saving
non-terminal snapshots once the live task is terminal or gone.

If a cancel file removal fails, fall back to writing the cancelled
snapshot so the file is terminal rather than pending. deleteTaskState now
returns its error, and CancelTask captures task.Status while still
holding the queue lock.

* admin: serialize task file check+write against cancel deletes

The saveTaskState guard still had a check-then-write window: a pending
snapshot could pass the terminal check before a cancel deleted the file,
then write it back after. A persistMu on the queue now covers the
check+save and the cancel paths' delete (with its terminal-state
fallback), so the two cannot interleave for the same task.
2026-10-05 12:02:53 +08:00
Chris LuandGitHub 3c17c5146e S3: fix ListObjectVersions losing keys across page boundaries (#11598)
* s3api: thread filer client through the versioned-listing collector

findVersionsRecursively now binds one SeaweedFilerClient for the whole
recursive walk instead of re-resolving a filer on every list/lookup call,
and the collector's list/getEntry/scanLatestVersionEntry/getObjectVersionList
helpers go through it. No behavior change; this also lets tests drive
collectVersions with a stubbed client.

* s3api: keep collecting versions while pending names can sort into the page

ListObjectVersions walked the filer in directory-entry name order and
stopped as soon as maxKeys+1 items were collected, sorting only that
partial set. Filer names do not match key order: "a.copy.versions" sorts
before "a.versions" while key "a.copy" sorts after "a", so a page
boundary inside the earlier-walked sibling's versions permanently skipped
the later key.

Track the largest key collected (maxKey) and, once the collector is full,
keep walking until entry names pass the ceiling of names that can still
resolve to keys at or below it; the ceiling reaches through the prefix
versions of maxKey. Entries whose subtree can only hold keys above maxKey
are skipped. Versions of an in-bound object are collected in full so its
position in the sorted page is exact.

Fixes seaweedfs#11594

* s3api: resume versioned listings at the earliest covering name prefix

computeStartFrom mapped the key marker straight to an entry name (or cut
it at the first '/'), which skips sibling directories that are a prefix
of the marker below '0' - for marker "d.x" the listing resumed at name
"d.x", skipping directory "d" whose keys "d/*" all sort after it.

Resume at the earliest remainder prefix ending at a byte below '0' ('/',
'.', '-' and friends), so every directory whose subtree can still hold
keys past the marker is revisited; already-returned keys inside are
filtered by the existing marker checks as before.

* s3api: regression test for versioned-listing pagination order

Drive collectVersions against a stubbed filer holding the issue-11594
layout - "a.copy.versions" listing before "a.versions", plus a "d/"
subtree next to "d.x" - and assert that every page size from 1 up
reproduces the unpaginated ordering with no lost or duplicated entries.
Also updates TestComputeStartFrom for the new earliest-prefix resume and
gives testFilerClient a LookupDirectoryEntry stub.

* s3api: inject list/getEntry functions into the version collector

Pinning one SeaweedFilerClient for the whole walk dropped per-call
failover: previously each s3a.list resolved a filer through
WithFilerClient, so a mid-walk filer failure could fall back to a
healthy peer. Inject s3a.list/s3a.getEntry as function fields instead -
production keeps the failover behavior, tests can still stub.

* s3api: keep scanning marker for a later covering prefix

A leading byte below '0' (marker .hidden/file) has no non-empty prefix
at index 0, but a deeper separator still does - resuming at .hidden/file
skipped the .hidden directory and its keys after file. Continue the scan
instead of bailing on the first byte.
2026-10-05 11:58:04 +08:00
483dd4b12e s3api: persist ACLs on PutObject uploads (#11592)
* s3api: persist ACLs on PutObject uploads

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: fix PutObject ACL edge cases found in review

- Only enforce BucketOwnerEnforced when explicitly configured; buckets
  without a stored ownership control keep accepting upload ACLs
- Ignore ACL query parameters on SigV2 requests, which do not sign them
- Mirror signed-query ACL values into headers after authentication so
  grant parsing and resolveFileMode agree on presigned uploads
- Validate only caller-supplied grantees against the account registry;
  default grants now work for accounts outside the local registry
- Reject unknown grantee keys and accept comma-separated grantee lists
  without spaces in ParseCustomAclHeader
- Guard against identities without an account

* s3api: harden upload ACL parsing and authorization

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: evaluate upload ACL grantees individually in policies

A comma-joined grant header or a signed query parameter reached policy
conditions as one value, so a deny on a later grantee did not fire. Split
grant headers into per-grantee values for policy evaluation and share the
grantee pair parser with ParseCustomAclHeader.

* s3api: keep raw grant header values visible to policy conditions

Exact-match conditions written against the signed header value stopped
matching once grantees were split for evaluation. Preserve the original
wire values alongside the per-grantee values so deny policies fire on
either granularity.

* s3api: evaluate upload ACL grants as one canonical list in policies

Conditions on s3:x-amz-grant-* now see a single comma-separated canonical
grant list identical for a single line, repeated header lines, or a signed
query parameter. This keeps StringEquals allows and exact-list or
allowlist (StringNotEquals) denies accurate regardless of wire encoding.

* s3api: preserve upload ACL denies and align policy checks

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

* s3api: retain upload owner grants and literal policy values

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>

---------

Signed-off-by: zhaoyuchen <43179751+zhao-yc@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-05 09:13:19 +08:00
ssshr-66GitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
9d1c24d80d [Filer] Support append to inline small files (#11591)
* fix 11586

* Update filer_server_handlers_write_autochunk.go

* filer: fix inline append races, empty files, and stale ETags

Serialize the append read-modify-write on the entry lock so concurrent
appends merge instead of losing content, keep small appends to empty
files inline, tolerate legacy entries whose metadata size differs from
their content, and set the entry digest so appended inline files keep a
real ETag.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-05 09:12:20 +08:00
github-actions[bot] 76ddde6a6d docs: regenerate star history chart 2026-10-05 00:55:58 +00:00
hajimohammadi.netandGitHub 0d93dec145 helm: mount TLS certificates in bucket hook (#11589) 2026-10-04 20:35:08 +08:00
Chris LuandGitHub d9b69a7f76 s3api: fix PutObjectAcl permission scoping and owner grants (#11587)
PutObjectAcl had four authorization and ownership bugs:

- The handler embedded the resource path into the action
  (WriteAcp:bucket/object), and authRequest/CanDo then scoped it to
  the request's bucket/object again. A bucket-wide WriteAcp:bucket
  grant could never match, so legitimate owners got 403.
- After authRequest succeeded via an IAM or bucket policy, a leftover
  identity.CanDo gate re-checked only the legacy Actions list, denying
  identities authorized purely by policies.
- For canned and default ACLs, ExtractAcl generated the FULL_CONTROL
  grant for the requesting account instead of the object owner. An
  admin setting private/public-read on another account's object left
  the owner metadata intact but reassigned full control to the admin.
- Objects without stored owner metadata (e.g. written via the filer
  outside S3) fell back to treating the requester as the owner, so any
  user with a WriteAcp grant could take them over. Non-admins are now
  denied; admins keep the takeover fallback.

Grantee validation now also accepts the object's stored owner even when
that account has been removed from the registry, so canned/XML ACLs for
retired owners keep working.
2026-10-04 19:41:10 +08:00
Chris LuandGitHub 2b5fdc639f filer: stop isSameChunks from sorting caller-owned chunk slices (#11584)
* filer: stop isSameChunks from sorting caller-owned chunk slices

slices.SortFunc reorders the input in place. filer.remote.sync calls IsSameData on a metadata event's NewEntry inside isMetadataOnlyUpdate and later stamps the filer entry under an IF_ENTRY_EQUAL precondition carrying that same entry. The ETag-sorted chunk list never matches the stored entry, so every stamp of a multi-chunk object fails, synced_mtime_ns stays zero, and dirty objects are re-uploaded forever. Sort clones of the slices instead.

* filer: test IsSameData leaves input chunk order unchanged

Guards the clone-then-sort fix: a regression back to in-place sorting would reorder caller-owned chunk slices and reintroduce the remote-sync IF_ENTRY_EQUAL mismatch.
2026-10-04 17:37:26 +08:00
github-actions[bot] d7a02567e3 docs: regenerate star history chart 2026-10-04 01:25:13 +00:00
39bc9cd0ef s3api: copy the trailer checksum before reading the next trailer line (#11583)
* s3api: copy the trailer checksum before reading the next trailer line

parseChunkChecksum kept the checksum value as a sub-slice of the line
returned by bufio.Reader.ReadSlice, which is only valid until the next
read. When the trailer lines arrive in separate TCP segments, reading
x-amz-trailer-signature refills the buffer and overwrites the saved
value, so a correct upload fails with InvalidDigest ("The Content-Md5
you specified is not valid").

The AWS SDK for Java v2 (>= 2.30) on a Linux JDK sends the trailer that
way; about half of its signed streaming uploads failed.

Fixes #11582

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* s3api: reuse crc32 writer and trim comments in trailer split test

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-04 01:14:15 +08:00
Chris LuandGitHub f4ef37e752 filer.sync: resubscribe the metadata stream when a failure pins the offset (#11581)
* filer sink: keep the gRPC status inside wrapped errors

%v stringifies the status, so a peer teardown reported as Canceled ("the
client connection is closing") reached IsTransientError as plain text and
matched nothing: the sync job failed on the first attempt and pinned the
offset. %w keeps the status reachable, so the retry runs on a fresh
connection once the target is back.

* pb: let a consumer drop the metadata stream to force a resubscribe

A MetadataProcessor job that exhausts its retries pins the processed
watermark so the event replays on the next subscribe — but nothing on the
source stream notices a target-side failure, so the replay waited for an
unrelated reconnect or a restart. The new Resubscribe channel cancels the
stream's context; the Recv loop answers it with ErrResubscribe so the
caller's retry loop resubscribes from GetResumeTsNs and replays the pinned
events in order.

* pb: stop the event retry loop once the stream context is done

RetryUntil ignores context, so a subscriber parked on a failing offset
write would keep retrying past a resubscribe signal until the sink came
back. Stop retrying when the stream is being dropped so the resubscribe
takes effect promptly.

* filer.sync: signal resubscribe when a job failure pins the offset

A job that exhausts its in-job retries leaves the event pinned behind oldestFailedTsNs, replayable only on a reconnect. Closing resubscribeCh on the first recorded failure lets the metadata follower drop the stream so the reconnect replays the pinned events instead of waiting for a process restart (#11572).

* filer.sync: wire the resubscribe signal into the follow options

filer.sync, filer.remote.sync, and the remote gateway bucket sync all run their subscription inside an outer retry loop, so ErrResubscribe resurfaces as a resubscribe from the persisted watermark.

* filer.sync: wait for in-flight jobs before signaling resubscribe

* remote sync: never resume past the saved offset when -timeAgo is set

* filer.sync: drop events that arrive after the drain signals resubscribe

* pb: interrupt the event retry backoff when the stream context ends

* filer.sync: stop admitting once a failure pins, and count jobs per timestamp

A pinned watermark only released once the processor went fully quiet, so a busy stream could starve the resubscribe — the failed event would wait for an unrelated reconnect anyway, the wait this mechanism exists to remove. The processor now latches stopped when a job fails: admission drops new events (they replay from the pinned watermark after the reconnect), a broadcast releases blocked waiters, and the resubscribe signals as soon as the jobs already in flight drain. A redelivery of an event still in the failure ledger may still run so its success shrinks the replay, but nothing starts once the signal has fired, or it would race the replay it asked for.

Dropped events no longer inflate the received counters — an event counts only once admitted, and the replay's own admission counts it.

While here: activeJobs keyed by TsNs collapsed events sharing a timestamp, so one completion could empty the map while a same-ts sibling was still running — letting the drain gate and the watermark outrun it. Jobs are now counted per timestamp, and the drain and lazy heap cleanup go through the counts.
2026-10-04 00:06:55 +08:00
10b0f2b8ad volume server: refuse the rest of a grouped run after a durable index failure (#11576)
* volume server: refuse the rest of a grouped run after a durable index failure

A durable write whose needle-map put fails stops the volume taking writes
(#10825): sent on its own, the next write then fails read only before it
appends. The grouped run from #11543 appends and syncs every entry before
publishing any, then kept publishing the entries after the failed one and
acked them once the shared .idx sync went through. When the failed put
tore its .idx row, the rows appended after it land off alignment, so the
next load parses them as garbage and the acked writes are gone.

Once a durable entry fails to publish, refuse every later entry of the run
with ReadOnly, as the per-needle path does. The entries before it stay
acked; their rows go down with the run's one .idx sync. The refused
records stay on the .dat unindexed, as the failed one does on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: refuse a grouped entry staged as a cookie mismatch too

After a durable entry in a grouped run fails to index, the entries
after it are refused as they would be on their own. On its own an entry
meets check_writable before its cookie check, so one staged as a cookie
mismatch now gets the refusal too, instead of keeping its staging error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: trim a torn .idx row back so the next stays aligned

A failed write_index_entry can leave half a row in the .idx. With the
writer appending at the tail, every row written after it lands off
alignment and the next load parses them as garbage, so a write acked
behind a torn row does not come back. Trim the file back to
idx_file_offset on a failed append, in both needle maps, and cover it
with a test that writes past a torn row and reloads.

* volume server: refuse queued Go writes once a durable index update fails

processBatch kept writing after a failed nm.Put, and the single-write
path checked IsReadOnly only outside the volume lock. A durable write
whose index update fails now marks the volume noWriteOrDelete, and each
queued request is checked before it appends, so the ones after a failed
durable entry are refused the way a lone write is. Deletes get the same
noWriteOrDelete refusal a lone delete gets.

* volume server: refuse appends while a torn .idx row cannot be trimmed

When trimming back a half-written .idx row itself fails, the next append
would land after the torn bytes and every later row would parse off
alignment on load. Latch the map as torn and refuse appends until the
trim succeeds, on both CompactNeedleMap and RedbNeedleMap; the same
latch covers an orphan row that could not be trimmed after a failed
redb commit.

The .idx writer is now opened with write+append access so truncate_to
(set_len) works on Windows, where an append-only handle cannot trim.

* volume server: write .idx rows at idx_file_offset, not via append mode

Rust's OpenOptions on Windows strips FILE_WRITE_DATA whenever append is
set so the handle stays strictly append-only, which makes set_len fail -
the torn-row trim could never succeed there. Open the .idx writer with
plain write access and seek to idx_file_offset before each row, the same
positioned-write model the Go server uses.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 21:04:30 +08:00
Chris LuandGitHub eafe79ebff filer: skip UpdateEntry when inline content is unchanged (#11580)
* filer: skip UpdateEntry when inline content is unchanged

SaveInsideFiler rewrites config files (IAM identities, filer.conf,
remote mappings, policies) unconditionally. Each no-op UpdateEntry is a
metadata event the local meta log persists to /topics/.system/log,
which appends a chunk to a volume. A client that rewrites identical
config on a timer, e.g. the seaweedfs-operator 5-minute resync calling
UpdateUser with unchanged actions, keeps .dat/.idx files growing on an
otherwise idle cluster and prevents HDD spindown
(seaweedfs/seaweedfs#11571).

Skip the UpdateEntry when the stored inline content is byte-identical,
so unchanged writes produce no metadata event and no volume writes.

* filer: test that identical SaveInsideFiler writes skip UpdateEntry

* filer: require stamped Md5 before skipping identical writes

An entry holding identical content but no Md5 (written before hashing,
or by a tool that cleared it) would never get the stamp that
IF_ETAG_MATCH conditional writes key off. Skip only when both the
stored content and its Md5 match, so one write still lands to repair
the stamp.
2026-10-03 21:00:51 +08:00
562afa8ec9 filer: resume metadata subscriber from processed watermark on reconnect (#11574)
* filer: resume metadata subscriber from processed watermark on reconnect

* filer: take the reconnect position from GetResumeTsNs verbatim

The callback is the subscriber's durable resume point; falling back to
StartTsNs when it returns zero can resume from a cursor the log-chunk
reader advanced past still-pending work.

* filer: advance the stream cursor once a retried event recovers

RetryForeverOnError resolves the failure inside handleErr, so returning
without moving StartTsNs replays work the event already did when the
stream reconnects before the next one arrives.

* filer: let filtered-progress markers move the processed watermark

A marker means the source examined everything up to its timestamp and
skipped what did not match the subscription. With a resume callback the
marker now reaches the consumer, and AddSyncJob advances the watermark
to it once every earlier job finished and no failure pins the offset.
Idle filtered stretches no longer rescan on every reconnect, while the
guards keep the watermark behind pending or failed work.

* filer: unpin the watermark once a failed event completes

oldestFailedTsNs was only ever set, so a failure that a replay later
fixed still held the resume offset, and every reconnect re-read the
same backlog. Track outstanding failures in a set and recompute the
pin when the failed event's job finally succeeds.

* filer.remote.gateway: resume bucket sync from the processed watermark

The bucket-sync subscriber runs the same MetadataProcessor queue as
filer.remote.sync; give it the same GetResumeTsNs callback so a
reconnect resumes from durably processed work, not the last seen event.

* util: treat a peer-sent gRPC Canceled as transient

A peer tearing down its end of the transport reports codes.Canceled
("the client connection is closing"), which IsTransientError used to
reject: the sync job then failed on the first try and held the offset
until a restart. Caller's own cancels are still excluded up front by
errors.Is(err, context.Canceled), so only teardown-style statuses take
the new branch.

* fix: preserve filtered progress and distinguish caller cancellation

* filer: bound the failed-event ledger past a persistent outage

A destination rejecting every event grew failedTs by one entry per source
event for the life of the processor. Past maxFailedSyncEvents the set now
collapses to a sticky pin at the smallest failure seen, so the watermark
still replays from the oldest failure while memory stays bounded; a
restart re-derives the exact set.

Also keep a resume-callback consumer's chunk-ref replay filter at the
subscribe-time position instead of option.StartTsNs, so a resubscribe does
not filter out events whose async processing is still pending.

* filer: key the failed-event ledger by event, not just timestamp

A success for one event cleared the pin recorded for a different event
that shared its TsNs, letting the watermark pass an unresolved failure.
The ledger now keys on the event's path identity, so recovery unblocks
only the event that actually failed.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 20:57:50 +08:00
8c1ebbee32 volume server: split get_or_head_handler_inner into phases (#11489)
* volume server: read GET/HEAD needles off the store lock, and only once

The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stream needle chunks without the store lock

StreamingBody::poll_frame took store.read() and find_volume for every
chunk to compare the volume's compaction revision, dup'd the source
handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the
stream also holds a data-file read lease for its whole life, while a
writer waits for that lease under store.write(): the next chunk's
store.read() then waits for the writer and the writer for the stream.

The per-chunk re-lookup was also wrong. The stream reads a handle opened
at plan time, which pins the .dat inode the offset was resolved against;
a vacuum commit renames a new file over .dat and leaves that inode
untouched. The re-looked-up offset belongs to the new file but was read
from the old inode, so a stream whose needle a vacuum moved ended in a
checksum error. The pinned offset stays valid, so the check, and with
it every store access, is dropped, along with the now unused
re_lookup_needle_data_offset and the revision fields of the read plan.

The source is shared as an Arc instead of dup'd per chunk, and the chunk
buffer is a BytesMut that the blocking read hands back with its result,
so its allocation is reclaimed once the previous frame has been written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: split get_or_head_handler_inner into phases

get_or_head_handler_inner was a ~650-line function. Its middle resolved
the needle and set five mutable flags (stream_info, can_stream,
can_handle_head_from_meta, can_handle_range_from_source, bypass_cm) that
three if-let reply paths then re-tested, each re-checking stream_info.

It is now a 126-line orchestrator over named phases: reject_read_jwt,
proxy_missing_volume, wait_for_download_slot, parse_read_request,
read_ec_needle / read_volume_needle, etag_and_last_modified,
not_modified_response, read_response_headers, and the reply phases
stream_response, head_from_meta_response, range_from_source_response,
buffered_payload and buffered_response. The read phases return a
ReadPlan whose ReadStrategy enum (Stream, HeadFromMeta, RangeFromSource,
Buffered) carries the NeedleStreamInfo only on the variants that use it,
so the reply is one match instead of three flag checks.

Pure refactor: every status code, header and header order, error text,
metric increment, lock and data-file lease scope, spawn_blocking
boundary and side-effect order is unchanged. Phases that can end the
request return ControlFlow<Response, T>. A Range header that is not
visible ASCII still falls through to the buffered path, as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a needle stream once its volume becomes unavailable

Taking the store lock out of StreamingBody also dropped its per-chunk
unavailable_error() check. With -hasSlowRead a writer can take the
data-file lease between chunks, fail its fsync and its truncate, and mark
the volume unavailable; the stream then kept serving the rest of the
needle from its pinned handle.

The volume's io_unavailable reason is now an Arc-shared leaf mutex that
the read plan hands to the stream. Each chunk checks it under its
data-file lease, where the writer marks it, and fails with the same
"volume is unavailable: <reason>" error the old check returned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: mirror Go order in the buffered read path

- check HEAD before Range in buffered_response (writeResponseContent
  order); an EC-volume HEAD with a Range header answered 206, Go
  answers 200
- treat the proxied flag as an exact query pair like Go's parsed
  lookup, not a substring
- name the phases after their Go counterparts: check_download_limit
  and read_ec_shard_needle; reuse has_replication()
- drop comments that restate the code or cite Go line numbers

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 17:54:46 +08:00
Oliver RahnerandGitHub c9ade3f9fd helm: compare PVC sizes numerically in the volume resize hook (#11575)
include always returns a string, so gt compared the rendered quantities
(e.g. "1.2884901888e+11" vs "6.442450944e+10") lexically. Growing a volume
from 60Gi to 100Gi/120Gi or 500Gi to 1Ti emitted no kubectl patch: the
StatefulSet was recreated with the new volumeClaimTemplate but the PVC kept
its old size. Shrinks such as 120Gi -> 60Gi emitted a patch instead.

Pipe both values through float64 before comparing.
2026-10-03 15:07:32 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
6c07a5fdd0 s3: keep small ranged GETs on range reads, no whole-chunk downloads (#11577)
* filer: keep a ranged read in random mode through its contiguous tail

A far ReadAt on a fresh ReaderPattern left the sequential counter at -1,
so the next buffer of the same ranged request landed on the frontier and
flipped the verdict straight back to sequential — readChunkSliceAt then
paid a whole-chunk fetch for the remainder of the range. Drop the
counter to -ModeChangeLimit when random mode is entered so the verdict
needs sustained sequential evidence to undo, matching the hysteresis an
established sequential stream already gets.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: pin small ranged GETs to range reads

A ranged GET whose first read lands within SeqTolerance of offset 0 is
judged sequential immediately, and even a far-starting range could flip
back mid-request; either way readChunkSliceAt downloads each covered
chunk in full, multiplying disk reads for small ranged reads (measured
~7x). Pin random mode for ranged requests no larger than SeqTolerance so
all of the request's buffer reads stay range fetches. Larger ranges keep
the dynamic pattern, where whole-chunk fetches amortize.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: fetch only the part of a chunk the view covers

Replaces the PinRandomMode size heuristic with a per-chunk coverage rule.
ViewFromVisibleIntervals already clips chunk views to the request window,
so a view that is not IsFullChunk() is one the request only partially
needs; fetch it as a range regardless of the detected read pattern.

This closes the holes a request-size pin left open: ranges larger than
SeqTolerance no longer revert to whole-chunk downloads once their buffers
look sequential, and ranges that fully cover a chunk keep the shared
whole-chunk path instead of fetching 256KiB slices piecemeal. Prefetch
(MaybeCache) skips clipped views so it cannot amplify a range read either.

PinRandomMode is dropped: no caller needs it once coverage drives the
fetch choice. Range fetches route through fetchChunkDataFn so tests
observe them the same way as whole-chunk downloads.

* filer: keep ciphered chunks on the whole-chunk path

A range fetch cannot save bytes for a ciphered chunk: readEncryptedUrl
always downloads and decrypts the whole blob before slicing. Sending
partial views of ciphered chunks through fetchChunkRange would repeat the
full download per buffer, so they keep the shared whole-chunk path where
one download serves every buffer. Prefetch stays enabled for them for
the same reason.

* filer: keep compressed chunks on the whole-chunk path

Like ciphered chunks, a range request on a compressed chunk makes the
volume server read and decompress the whole needle, so range-per-buffer
would repeat the full backend read for each 256KiB window. Route them
through the shared whole-chunk path via ChunkView.CanRangeFetch.

* filer: fall back to range fetch when a chunk exceeds the reader budget

A ciphered or compressed chunk larger than readerCacheSizeMB can never
be read through the whole-chunk path — the budget rejects the buffer —
so its partial views must still range-fetch or the GET fails outright.

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 15:01:04 +08:00
Eliah RusinGitHubClaude Opus 5.5Chris LuChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
07da302da0 volume server: ec.decode verifies, cleans up and compacts like Go, off the runtime (#11547)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode runs its file I/O off the async runtime

VolumeEcShardsToVolume released the store lock before decoding, but read
the .ecx/.ecj, rebuilt the .dat and wrote the .idx inside the async
handler, parking a runtime worker for the length of a volume-sized copy.
The decode now runs in spawn_blocking on inputs snapshotted under the
store lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode checks the rebuilt .dat is complete

Go stats the decoded .dat before writing the .idx (VerifyDecodedDatFile)
and fails the decode when it is shorter than the extent the EC index
references, since the caller deletes the shards once the call returns.
The Rust handler returned success without that check. The rebuild
already fails on a short shard read, so this guards the published file
itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode drops the decoded volume's bitrot sidecars

Go removes <base>.ecsum and <base>.ecsum.v<N> beside the .dat and beside
the .ecx once the .idx is written, so a stale checksum sidecar cannot
pass for the protection of a later re-encode. The Rust handler left them
in place. Removal is best effort, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode compacts the decoded volume

Go ends VolumeEcShardsToVolume with an offline CompactVolumeFiles, so the
decoded volume holds only live needles. The Rust decode left every needle
deleted through the .ecj in the .dat, tombstoned in the .idx, until a
later vacuum reclaimed it.

Store::compact_volume_files loads the unmounted volume, checks free space
the way the vacuum does (the estimate now lives in one helper), and runs
the vacuum's compact-by-index and commit. As in Go a failed compaction is
logged and the decode still succeeds, so the uncompacted .idx rules stay:
the tests that pin them now make the compaction fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps deletes journaled while the .dat is written

The decode read the .ecj journals once, before rebuilding the .dat, so a
delete that reached the EC volume during the rebuild was left out of the
new .idx and the needle came back live. Each journal's read length is now
kept, and the bytes appended since are read just before the .idx is
written, after waiting out any journal append in flight (appends hold
the store write lock), so every delete acknowledged by then is in the
.idx. A delete after that point is still lost, as in Go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Guard overlapping ec decode requests; serialize journal catch-up

volume_ec_shards_to_volume runs its decode in spawn_blocking, so a
dropped request leaves the job running and a retry would race it on the
temporary and final volume files. Claim the vid in a per-server
in-flight set until the blocking job finishes, and return Unavailable
to an overlapping request. The Go handler has the same exposure and
gets the same guard.

Journal appends hold the store write lock through their
sync-or-truncate, so holding a read lock across the catch-up read
guarantees every record it sees is committed: a rolled-back delete can
no longer leave a tombstone in the decoded index.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Reconcile the swap when offline compaction commit fails

A CommitCompact that fails after the .cpc marker may have renamed .dat
but not .idx. cleanup_compact refuses while the marker exists, so the
mismatched pair survived until a restart reconciled it — and the decode
caller treats the failure as non-fatal. Run reconcileCompactState on
commit failure so a decided swap rolls forward and orphan temps are
removed before the volume can mount.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Release the decode claim on panic

* volume: add ec_decodes_in_flight to the integration-test state literal

* volume server: hold the decode tail's lock through compaction

The catch_up read released before the rebuilt .idx was written and the
volume compacted, so a delete synced to .ecj in that window was durably
journaled yet absent from the published index — resurrecting the needle.
Rust now holds the store read lock from catch_up through compact, and Go
mirrors it by holding the volume's journal lock from the journal-
consuming index write through CompactVolumeFiles.

* volume server: serialize ec decode's tail per volume, not per store

Review follow-ups on the decode path:

- Rust: holding the store read lock from journal catch-up through the
  offline compaction stalled every writer on unrelated volumes for the
  whole rewrite. The new ec_decode_tail set marks the vid only while its
  .idx is published and .cpd/.cpx swapped; the two local .ecj append paths
  (VolumeEcBlobDelete, the distributed delete's local journal) wait on a
  Notify for that span — Go's per-volume ecjFileAccessLock semantics
  without the global stall. VolumeMount and the staged-adopt path are also
  held off while a decode claim is in flight so neither can race the swap.

- Rust: the initial journal read ran unlocked, so bytes a rolled-back
  append later truncated could be folded in as phantom tombstones. The
  first pass stays unlocked (a slow journal must not stall the store) and
  a rescan under the quiescing read lock re-reads only committed content;
  catch_up now rebuilds the id set when a regular journal shrank.

- Go: the decode resolved the compaction DiskLocation through
  FindEcVolume while holding the journal lock, inverting DestroyEcVolume's
  map->journal order into a deadlock. The lookup now happens first, and
  DestroyEcVolume/deleteEcVolumeById/DiskLocation.Close destroy outside
  the map lock.

- Go: RebuildEcxFile unlinks .ecj while the volume's ecjFile handle stays
  open, so later deletes could commit to a detached inode. Both call sites
  now fold under the journal lock and ReopenDeletionJournal repoints the
  handle at the live path, working on the volume's resolved .ecx dir
  (EcIndexBaseFileName) rather than the configured index dir.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fence EC remounts behind the destroy tombstone

DestroyEcVolume, deleteEcVolumeById, and the collection-delete sweep now
remove the EcVolume from ecVolumes before destroying it off-lock, so a
concurrent remount could re-open shard files that the in-flight destroy
then unlinks — registering a detached fd.

Each destroy records a per-vid tombstone channel in a new
ecVolumesDestroying map before dropping the map entry and closes it when
Destroy returns. The tombstone intentionally survives as the vid's
destroy generation: loadEcShardWithIdxDir compares it before and after
opening the shard, so a destroy that both started and finished inside the
open window is still detected. A mismatch drops the just-opened shard
(releasing its fd and mount gauge) and retries after the destroy
completes; a successful mount clears the stale tombstone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: rescan the .ecj under the store lock only after a rollback

The decode's second journal pass ran a full rescan under the store read
lock on every decode, stalling unrelated writers for the length of the
scan. Bump a process-wide epoch whenever a failed append truncates its
uncommitted tail; an unchanged epoch between the unlocked read and the
quiesced pass proves every id folded in was committed, so catch_up()
suffices. catch_up() also treats a journal that was read but has since
disappeared as shrunk to zero, so its earlier ids cannot linger.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: check the decode tail under the store write lock on delete

A blob delete waited for the publishing tail before taking the store
write lock, so a decode that claimed the tail while the delete was
parked behind the decoder's read lock could still see the journal append
land after the rebuilt .idx — an acknowledged delete the mount would
miss. Test tail membership under the write lock instead, retrying after
the wait; journal_delete_local reports WouldBlock for the same recheck
on the distributed path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: claim the vid for mount and staged adoption, per volume

VolumeMount and the staged .copying adoption held the
ec_decodes_in_flight set lock through slow file renames and mounts,
stalling every unrelated volume's decode, mount, and adoption. Take the
per-volume claim instead — the same exclusion against a racing decode
for this vid, released when the call returns.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: fail the decode when a compaction commit marker survives

CompactVolumeFiles' caller logged a compaction error and went on to
delete the EC shards. When the commit marker (.cpc) is still on disk the
.dat/.idx swap was decided but could not be reconciled, so the mounted
pair may be mismatched — report the failure instead so the shards are
kept and the caller can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: gate the parked-delete test on the held write lock

The releaser thread and the spawned delete raced for the store write
lock; on a slow runner the delete could acquire it first and commit
before the tail was ever claimed, failing !delete.is_finished() on the
Windows unit-test job. Spawn the delete only after the thread reports
the lock held.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 14:55:15 +08:00
68df7511f6 filer.remote.sync: do not pin the sync offset on completed work (#11569)
* filer.remote.sync: do not pin the sync offset on completed work

* filer.remote.sync: a superseded rename uploads the current entry; typed NotFound for a stamp on a deleted entry

* filer.remote.sync: a superseded rename keeps the old key when it is the only copy and uploads once

* filer.remote.sync: a rename whose content is now remote-only fails the event instead of completing it

* filer.remote.sync: a remote-only rename copies the old object to the destination before deleting it

* filer.remote.sync: the remote-only rename path follows the filer's current entry and verifies the destination object

* filer.remote.sync: an event that described an entry without data is superseded once the filer wrote to it

* filer.remote.sync: a superseded rename does only the work left to do

uploadCurrentEntry met a remote-only current entry with a fixed error, but a
sync plus remote.uncache in the meantime leaves the destination holding the
stamped object; that state is complete, not lost. The remote-only case now
finishes through completeRemoteOnlyRename, which verifies the destination
against the entry stamp and fails only when neither key holds the content.

A current entry whose stamp covers its content was already uploaded by the
superseding event; skip it instead of writing the same bytes again.

* filer.remote.sync: an inherited stamp does not prove the content synced

The stamp-coverage skip in uploadCurrentEntry read LastLocalSyncTsNs as
proof the current content was uploaded, but a rename carries the source
entry's stamp to the destination: a rewrite hidden by that stamp (the case
the fallback upload exists for) carries a LastLocalSyncTsNs at or after its
mtime and would have been skipped. Drop the check; the remote-only path
verifies content at the destination itself through describes.

---------

Co-authored-by: James Sas <james@medable.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 12:34:52 +08:00
Chris LuandGitHub 793ce06b10 s3api: allow unsigned SSE-C customer key headers on presigned requests (#11578)
* s3api: allow unsigned SSE-C customer key headers on presigned requests

AWS requires only x-amz-server-side-encryption-customer-algorithm to be signed on presigned URLs; the key and key-MD5 headers are supplied at request time. Since #9121 rejected any x-amz-* header outside SignedHeaders, SDK-generated presigned SSE-C requests (e.g. .NET GetPreSignedUrlRequest) fail with SignatureDoesNotMatch. Exempt the customer key and copy-source key headers for presigned requests only.

* s3api: test presigned SSE-C requests carrying unsigned key headers
2026-10-03 12:33:52 +08:00
Chris LuandGitHub 0ca484c354 vacuum: bound master vacuum RPCs with phase deadlines (#11579)
* vacuum: bound the commit RPC with a phase deadline

VacuumVolumeCommit ran on context.Background(), so a volume server that
keeps the call pending would hold the topology-wide vacuum guard
forever and every later sweep would be skipped. Give the call a
deadline scaled like the existing phase waits (one minute per GB of
the volume size limit) so a stalled commit ends as an error instead of
blocking the sweep; the timeout is a var so tests can shrink it.

* vacuum: bound the replica status probe with a phase deadline

The VolumeStatus call on replicas that were not compacted also ran on
context.Background(), so a stalled replica could pin the sweep the
same way a stalled commit can. Give it the same per-phase deadline.

* vacuum: bound the cleanup RPC with a phase deadline

VacuumVolumeCleanup also ran on context.Background(); a stalled
server would keep the sweep worker and the shared vacuum guard
pending forever. Give it the same per-phase deadline.

* vacuum: let the check and compact phase waits cancel their RPCs

The coordinator wait timers fired while the check and compact calls
still ran on context.Background(), so the sweep gave up but the RPC
goroutine stayed until the server answered, and a compact stream kept
writing on the server. Share one deadline context between the wait and
the calls so an expired wait actually cancels them.

* vacuum: test that a stalled volume server releases the vacuum guard

A fake volume server keeps one vacuum-phase RPC pending until the
client context is cancelled. Before the phase deadlines, Vacuum never
returned and vacuumLockCounter stayed held; now each phase cancels on
its deadline and the guard is free for the next request.

* volume: stop compaction at the next needle when the client cancels

The progress callback only noticed a gone client when a 128 MiB report
failed to send, so an aborted VacuumVolumeCompact kept copying for up
to a whole interval while the master had already moved on to cleanup.
Check the stream context on every needle, the same early return the
Rust volume server does with tx.is_closed().

* vacuum: assert the stalled phase RPC is cancelled, not just bypassed

The check and compact coordinator waits already returned on timeout
before the deadlines existed, so a regression that put the calls back
on context.Background() would pass unnoticed. Wait for the fake server
to report that the phase RPC context ended.

* vacuum: give the stalled-RPC test room to reach the handler

The 50ms phase budget starts before goroutine scheduling and the gRPC
dial, so a busy test host could expire it before the fake server saw
the call. Raise the override to 250ms; the test still finishes in
about a second.

* vacuum: describe the phase deadline as scaled, not per-GB

The formula keeps the exact expression the check and compact waits
already used (floor plus one at 1 GiB granularity); it is a backstop,
not a per-GB SLO.
2026-10-03 12:32:42 +08:00
3e679e925e filer: keep generated inodes inside the positive signed 64-bit range (#11567)
AsInode derives inodes from HashStringToLong, which is uniform over int64,
so roughly half of the derived values land above math.MaxInt64 once they are
converted to uint64. The Elasticsearch store indexes Entry.Attr.Inode as a
signed long, so those values are rejected with HTTP 400 and the metadata
entry is never written, which the filer then retries forever.

Fold the sign bit off in one place, util.NormalizeInode, and route both
derivation sites through it: FullPath.AsInode (path plus creation time) and
the hard-link branch in ensureEntryInode (HardLinkId hash). Masking keeps
the other 63 hash bits, so distinct paths still get distinct inodes, and it
applies identically to the FUSE mount, which derives the same value.

Co-authored-by: Yi-111-a <34116709+0-xiaosu@users.noreply.github.com>
2026-10-03 09:13:59 +08:00
52fb9f93ff s3: track filer joins and leaves pushed by the master (#11563)
* fix(s3): track filer joins and leaves pushed by the master

The S3 FilerClient replaced its -filer seed with a master snapshot of
filer IPs at boot and refreshed it only every 5 minutes. A rolling
restart replaces every filer well inside that window, leaving S3
servers with only dead addresses and failing every write until the
next poll.

Apply the master's ClusterNodeUpdate pushes to the filer list as they
arrive, keeping the poll as a backstop. The last filer is never
removed, and a poll snapshot requested before a push was applied is
discarded rather than overwriting newer membership.

* Defer last-filer leaves; bump the generation only on real changes

* fix(s3): cancel deferred filer leaves on rejoin and on discovery

A deferred last-filer leave outlived the filer it was recorded for: a
rejoin at the same address looked like a duplicate add, and a discovery
snapshot left the entry behind. The next join then removed a live
filer until the following poll.

A join now cancels any deferred leave for its address, and an applied
snapshot clears them, since it is the master's current membership.

* Bump the push generation when a rejoin cancels a deferred leave

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-03 09:10:43 +08:00
0ff7794c54 volume: compact an oversized .ecj at mount, safely (Rust + Go) (#11555)
* volume: compact an oversized .ecj at mount, safely (Rust + Go)

Restore the mount-time compaction dropped from #11408, Rust + Go parity.
A journal already bloated by repeated shard copies is folded down to the
id set it encodes.

- Trigger after load when file_records > max(threshold, 4x distinct),
  with a 1 MiB floor so small journals are never rewritten. The set is
  written to .ecj.compact.tmp + fsync, the handle dropped, renamed,
  the directory fsynced and the append handle reopened. A failure before
  the rename keeps the original journal and handle; a failure after it
  fails the mount.
- Go never compacts after a failed journal load; the set would be
  partial and the rewrite would drop the unread records.
- A per-path registry (ecj_registry.rs / ecj_registry.go) counts EcVolume
  holders and out-of-band writers of each .ecj. Compaction runs only
  when this volume is the sole holder and no copy is writing; holders
  and writers wait while one runs. This covers shared -dir.idx journals
  and cross-disk reconcile, where another EcVolume may hold the same
  journal.
- VolumeEcShardsCopy and EC index recovery register as writers around
  their .ecj append and partial-file cleanup.
- Under the reservation, re-check that the file on disk is still the
  inode and size that was loaded.
- Publish errors are classified where they happen; a failed rename plus
  a failed restore reports both errors.
- Compaction runs after the .vif / bitrot checks, so a refused mount
  leaves the journal untouched.
- The tmp is opened like other volume files, removed at mount if a crash
  left it, and listed in every EC index cleanup path.

Failure paths are tested through the real mount via injectable fs steps
(open_with / newEcVolumeWith), plus sibling holders, active copies,
changed-after-load, stale tmp cleanup, refused mounts and the Go
load-error guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: fail the mount when the compacted .ecj's directory cannot be synced

The Rust mount synced the journal's directory after renaming the compacted
file over it through the crate's best-effort fsync_dir, which returns Ok
when the directory cannot be opened. A rename needs only write and search
permission, so on a directory without read permission the replacement was
published, never synced, and the mount went on taking deletes against it.

Sync through a helper that propagates the open error, as Go's
util.FsyncDir already does, so that case fails the mount like any other
post-rename sync failure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: test the no-compaction-after-failed-load rule through the Go mount

The test for it handed compactEcjAfterLoad an artificial error on a volume
that had loaded cleanly, so it would not notice NewEcVolume dropping the
real load error on the way to compaction.

Make the journal read one of the injectable ecjFsOps steps and fail it
inside the real mount, after the first chunk, on a journal whose last
entry is an id the first chunk does not hold. The mount must leave the
file byte for byte as it was; a clean remount then compacts and keeps
that id. The Rust mount fails outright on a load error, so it has no
equivalent path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: register ReceiveFile's .ecj writes with the journal registry

ReceiveFile refuses a mounted EC volume only once, when the info message
arrives, then creates the .ecj and streams chunks into it. A volume that
mounted on that journal mid-stream could find a bloated prefix, pass the
inode-and-size re-check and rename a compacted file over it; the rest of
the stream then went to the unlinked inode and was lost.

Register the path as a writer before the file is created, in both the Go
and Rust handlers, and hold it until the file is closed and any partial
copy removed, as the shard-copy and index-recovery appends already do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: skip .ecj compaction when a writer ran since the journal was loaded

Compaction checked only that no writer was active at the reservation, and
that the file was still the loaded inode at the loaded size. A ReceiveFile
truncates and refills the journal in place, so one that ran during the
mount's load, or after it, and finished before the reservation could leave
different ids at the same length; compaction then wrote the stale set over
them.

Give each path a write generation that every writer bumps as it starts. A
holder records it, and whether a writer was active, when it registers,
which is before it opens and loads the journal. It may compact only if no
writer was active then and the generation has not moved. Same rule in Go
and Rust; the journal read becomes an injectable step in Rust as it is in
Go, so both test the in-place rewrite through the real mount.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: match the ReadOnly(VolumeId) variant in write_volume_needles

#11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-03 09:08:19 +08:00
github-actions[bot] 94edd0a6d4 docs: regenerate star history chart 2026-10-03 00:54:20 +00:00
30069f3e45 iam: manage OIDC providers and roles over the filer IAM gRPC service (#11523)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store

Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.

Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
  IAM config file the default stays memory; with none it is the filer, as
  for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
  which outlives the file and may be shared by S3 servers with different
  files. They are served from memory beneath the store, as OIDC providers
  are: a stored role of the same name takes precedence, and deleting it
  restores the file's. A config-file role cannot be changed or deleted
  through the API (UnmodifiableEntity), and removing one from the file
  removes it at the next start. An in-memory store holds them as records,
  as before. They have no creation time, so CreateDate is omitted rather
  than reporting when this server started. SetRoleStore installs a store
  the same way, so a store set after startup keeps the config-file roles,
  as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
  cached filer store otherwise serves a peer's stale role for up to its 5m
  TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
  used to report any failed lookup as "role not found". CreateRole proceeds
  only on a confirmed absence, so an unreadable store cannot let it write
  over an existing role.

IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
  AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
  are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
  changing or deleting it through the API is refused (UnmodifiableEntity)
  rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
  <name>.json in the filer, so this also keeps a name from leaving the role
  store's directory. At most 10 managed policies per role (AWS's default
  quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
  policy, as it already is for users and groups: roles attach policies by
  name, so a policy created later under the deleted one's name would
  otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
  refused rather than dropped.

Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
  kept across updates; a config-file role gets a stable ID derived from its
  name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
  and AssumeRole carry the role's ID (claim "rid"), and a request under a role
  whose current ID differs is denied. Resolving a session's policies by role
  name let a session outlive its role: once a role was deleted, a role later
  created under the same name — with a different trust policy and different
  policies — revived every unexpired session of the old one with the new
  role's permissions. Sessions issued before this change carry no ID and are
  unaffected until they expire.

Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: manage OIDC providers and roles over the filer IAM gRPC service

The filer's SeaweedIdentityAccessManagement service covers users, access
keys, policies and service accounts, but not the OIDC providers and roles
that STS web-identity federation needs. A controller that already manages
IAM over this service (seaweedfs-operator's S3OIDCProvider) has no
transport for them; its swadmin client returns ErrOIDCNotWired and names
this as the recommended fix.

- PutOIDCProvider / GetOIDCProvider / DeleteOIDCProvider / ListOIDCProviders
  and PutRole / GetRole / DeleteRole / ListRoles.
- They write the filer-backed stores at their default paths, which S3
  servers read when configured with a filer-typed "oidcProviderStore" and
  "roleStore"; the S3 servers' /etc/iam subscription applies changes
  without a restart.
- Put is an upsert, so a controller can reconcile to it. Deleting a
  provider or role that does not exist returns NotFound, as DeleteUser does
  for a user; clients treat that as already deleted. The provider's account
  ID travels in the request, since the filer does not know the STS
  accountId.
- PutRole applies the IAM API's rules: AWS role names, at most 10 managed
  policies.
- An S3 server serves the roles and providers of its own IAM config file
  ahead of the store, so a stored entry with the same name has no effect
  on that server.
- PutRole keeps a replaced role's RoleId and gives a role created anew a
  fresh one, so sessions of a deleted role do not carry over to a later role
  of the same name.
- DeletePolicy returns FailedPrecondition while a role attaches the policy
  (see the IAM API's DeleteConflict in the previous change). DeletePolicy on
  this service still does not check user attachments, which predates this.
- PutOIDCProvider requires an https issuer (http only for a loopback host):
  STS fetches the issuer's signing keys from it, so over plain HTTP anyone
  on the network path could substitute their own.
- The OIDC provider and role RPCs refuse to run on an unauthenticated
  service (FailedPrecondition until jwt.filer_signing.key is set). Users and
  policies keep the service's opt-in auth, but these grant STS access
  outright: otherwise anyone who can reach the port could register an issuer
  they control, create a role trusting it, and exchange a token for S3
  credentials. The filer's unauthenticated notice becomes a warning that says
  so.
- A store that cannot be read is Unavailable, never "not found", so a Put
  never writes over an entry it could not see.
- Validation is shared with the IAM API through PrepareRoleDefinition and
  PrepareOIDCProviderRecord.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: bind every role session to its role, and change roles atomically

Review follow-ups.

Session binding
- The role-ID check ran only when a session carried no policy names, and
  AssumeRole embeds the role's attached policies, so those sessions kept
  their permissions after the role was deleted or recreated. The check
  now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
  instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
  the name alone: a different role put in the file under the same name
  gets a new ID, while an unchanged role keeps its sessions across restarts.

Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
  role is unchanged since the read, and otherwise re-reads and retries. The
  filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
  role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
  Attach/DetachRolePolicy all go through it, so two gateways no longer
  overwrite each other's changes, a change racing a delete no longer
  writes the role back, and of two concurrent creates one gets
  EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
  broken stream instead of returning what arrived, so DeletePolicy's
  attachment check sees every role. ListRoles skips a role deleted between
  listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
  InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
  comment now says so. Role and OIDC provider actions write their own stores.

Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: PutRole changes roles atomically and checks its ARN; https issuers' keys stay on https

Review follow-ups on top of the role-store changes.

- PutRole goes through RoleStore.UpdateRole, so the decision to keep an
  existing role's ID or mint a new one is made against the role as it is
  when written. A PutRole racing a DeleteRole can no longer write the
  deleted role back with its old ID, which would revive its sessions. A
  failed store read or write is Unavailable.
- PutRole refuses a role_arn that does not name the role: STS resolves a
  role by the name in the ARN it is given.
- PutOIDCProvider requires an https issuer, but discovery could still name
  a plain-http jwks_uri, and a key fetch could be redirected to http. For
  an https issuer, a non-https jwks_uri from discovery is refused (the
  issuer's own /.well-known/jwks.json is used instead), and the client
  that fetches discovery and keys refuses any https-to-http redirect. An
  operator-set jwksUri is left as configured.

Each has a test that fails without its guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path

Review follow-ups.

- Authorization evaluates the policies of the role definition the session's
  binding was checked against, instead of reading the role again: a role
  replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
  definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
  and take its ID, duration cap and embedded policies from that same
  definition. A role replaced after the caller's trust check by one that does
  not trust the caller now yields AccessDenied, not a session bound to the
  replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
  write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
  routed and locked like the conditional CreateEntry. DeleteRole decides
  against the role it deletes, so a policy attached meanwhile on another
  server is a DeleteConflict, and a delete never removes a role written
  after its check.
- S3 servers watch the role store's configured basePath, not only
  /etc/iam/roles, so a custom path also drops peers' cached roles on change.

Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: state which roles DeletePolicy's attachment check can see

RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: note that a role store set after startup is not watched for peer changes

S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: DeleteRole deletes only the role it saw; issuer URLs are bare

Review follow-ups.

- The filer IAM service's DeleteRole looked the role up, then deleted by
  name, so a PutRole landing in between had its new definition deleted. It
  now deletes through RoleStore.UpdateRole, conditional on the entry it
  read. If the role was replaced meanwhile, it returns Aborted rather than
  deleting the replacement, and the caller decides again.
- PutOIDCProvider refuses an issuer URL with userinfo, a query or a
  fragment. The provider's ARN comes from host and path alone, while STS
  matches a token's iss claim against the stored URL exactly, so such a
  provider shared the bare issuer's ARN and matched no token. A loopback
  "localhost" is now matched without regard to case.

Both have tests that fail without them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: write OIDC providers atomically over the filer IAM gRPC service

PutOIDCProvider read the record, then stored unconditionally; a racing
DeleteOIDCProvider left the put's stale read merged into the rewritten
record. DeleteOIDCProvider read, then deleted unconditionally; a racing
PutOIDCProvider's newer record could be removed instead. These are the
races the role RPCs closed with UpdateRole.

OIDCProviderStore gains UpdateProvider with the same contract: memory
under its lock, filer as a conditional write (IF_ENTRY_EQUAL /
IF_NOT_EXISTS) or conditional delete retrying a changed entry.
PutOIDCProvider merges the fields the request cannot carry against the
record as it is written; DeleteOIDCProvider aborts rather than delete a
record replaced meanwhile.

isRoleWriteConflict is renamed isEntryWriteConflict — the conditional-
write check is shared by both stores now.

* iam: guard PutRole against a nil credential manager, fix its doc comment

PutRole read attached policies through s.credentialManager without the
nil check its sibling handlers make, so a server built without one
panicked on a PutRole naming a policy. It now fails the call as
FailedPrecondition like the others.

The doc comment also had the store/static precedence backwards: a stored
role shadows a same-named config-file role (as the overlay serves it),
not the other way around.

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-03 08:36:19 +08:00
ihnokimGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
fa77cde7da vacuum: check compaction space against live bytes, not volume size (#11524)
* vacuum: size the compaction space check by live bytes, not volume size

ensureCompactVolumeSpace required the volume's current .dat and .idx size as
free space before compacting. That is the size of the garbage, not of what
compaction writes, so on a disk that filled up until its volumes went
read-only every compaction was refused, including all-garbage volumes that
would compact to a superblock and an empty index. The sweep then retried
every volume each cycle and reclaimed nothing (issue #11516).

Estimate the output from what the needle map already tracks: live content
bytes plus a per-needle framing upper bound behind a superblock, and one
index entry per live needle. The estimate never exceeds the current volume
size and preallocate still wins when larger. Volumes whose deleted sizes are
unknown (.sdx converted back to .idx) keep the whole volume as the estimate.

The disk probe moves behind a package variable so the tests can stand in
for a full disk; the tests build real volumes instead of re-implementing
the formula.

* vacuum: space check reserves the index on top of preallocate, checks a separate index disk

Review follow-ups: preallocate only stands in for the new .dat, so the
rebuilt index is added on top of it; with separate index directories the
data disk is checked for the .cpd and the index disk for the .cpx; and the
estimates carry 1/16 headroom because counters rebuilt from an index file
pass through a Bloom filter with a 0.1% false positive rate. Neither
estimate exceeds the current file.

* vacuum: split the space check by filesystem, not by directory name

Two directories can sit on one filesystem and share its free space, so
the data and index estimates are checked separately only when the index
directory is on another device; otherwise the sum must fit. Unknown is
treated as shared.

* vacuum: ask the index directory for its share even when it looks like the same filesystem

A volume mounted under the data directory's drive letter on Windows has
the same volume name, so the identity check calls it shared. Checking the
index directory for the index estimate as well costs one statfs and
catches a full index mount either way.

* vacuum: identify a Windows volume by its GUID, not its path prefix

A volume can be reached through a drive letter and through a folder it is
mounted on, so filepath.VolumeName says nothing about the free-space pool.
Resolve each directory to its mount point and compare the volume GUIDs;
when that fails the two are treated as shared.

* vacuum: keep the framing and disk_space_low coverage the rebase displaced

* rust volume: split the compaction space check across data and index disks

Mirror the Go check: estimate the new .dat and rebuilt .idx separately —
live content plus per-needle framing capped at the current file, with
preallocate standing in for the data file when larger — and check each
directory against its own filesystem's free space. Two directories on one
filesystem are asked for the sum.

* vacuum: tighten comments on the compaction space check

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-03 08:31:56 +08:00
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Opus 5.5Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
35b090a4df volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go) (#11554)
* volume: merge .ecj as a set union on EC shard copy + index recovery (Rust+Go)

An EC volume's deletion journal is a set of needle ids, but shard copy
and index recovery appended the peer's whole journal, doubling the file
on every ec_balance round trip. Fold the peer's ids in as a union
instead: only ids the local journal lacks are appended.

- The journal is never replaced. A mounted EcVolume merges a peer's ids
  through its live handle under the lock deletes take (Go
  MergeJournal / Rust merge_journal), wherever its journal lives.
- An unmounted journal gets only the missing ids appended while mounts
  are excluded; the delta is read outside the lock and re-read if the
  journal changed.
- The source .ecj streams into memory as an id set: no staging files,
  chunked reads, memory proportional to distinct ids.
- Go and Rust agree that a source journal exists when it sends a
  modified time or any bytes. A missing source stays a no-op.
- Rust runs every merge in spawn_blocking and shares one receive/merge
  path between shard copy and index recovery.

The decode path and the journal format are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: route .ecj merges to the runtime that holds the journal open

Disks sharing one index directory all resolved as the journal's owner, so
the last one won and a sibling's mounted runtime was skipped: the merge
appended behind its open handle and the sibling kept serving the peer's
deleted needles until remount. Callers now name the receiving disk by its
data directory; the merge goes through that disk's runtime, else a
sibling runtime whose journal is the target file.

In Go the unmounted append now holds every disk's EC lock (in location
order) while it rechecks for a mount, so a sibling mounting from this
disk's index during the unlocked read is merged through instead.

In Rust a mount that lands during the read is merged through directly and
its added count returned, rather than discarded and reported as zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: sync merged .ecj records outside the disks' EC locks

The unmounted merge held every disk's EC read lock across its fsync, so a
slow sync on one disk held off mounts on all of them, along with the EC
reads queued behind those mounts. Mounts only need to be excluded while
the records are written: the write now happens under the locks and the
fsync after they are released, since a later mount reads the written
records from the page cache. A failed fsync rolls back only if nothing
has mounted the journal or appended to it since the write.

A merge through a mounted volume now keeps only that volume's disk locked
across its fsync.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: roll back an unsynced .ecj merge through a volume mounted mid-sync

If a volume mounted after the unmounted merge wrote its records but before
the fsync failed, the rollback kept the records because the journal was now
open, leaving ids in the volume's deleted set that may never reach disk; a
retried merge then saw them and synced nothing. The rollback now goes
through that volume the way its own failed journal fsync does: truncate
back and drop the ids from the in-memory set, so a retry appends and syncs
them again. It still keeps the records if the volume journaled since, as
truncating would lose that delete. No fsync runs under the disk locks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: decide .ecj merge rollback from the journal's actual length

Two runtimes can hold one journal (cross-disk mounts). The rollback of an
unsynced merge checked one runtime's cached ecjFileSize, which another
runtime's appends leave stale, so it could truncate a delete that runtime
had already synced. The rollback now holds every holder's journal lock and
truncates only if the file's actual length is still the append's end,
then updates each holder's size and deleted set. Otherwise later records
follow the merged ones, so they stay and are rewritten in place and
synced outside the locks, rather than left possibly not durable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: keep unsynced .ecj merge ids out of mounted deleted sets

When a merge's fsync failed, later records blocked the rollback, and the
rewrite-and-sync failed as well, the merged ids stayed in every mounted
volume's deleted set without being shown durable, so a retried merge saw
them as present and synced nothing. They now leave those sets while the
records stay in the file, matching DeleteNeedleFromEcx, which publishes an
id only after its record syncs. The merge returns the error and a retry
appends and syncs them again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: publish merged .ecj ids to every holder of the journal

Two runtimes can journal into the same file when disks share an index
directory. The merge went through only the first holder, leaving a
sibling's in-memory deleted set without the ids, so it could keep
serving a needle the peer deleted until it remounted. Every holder of
the journal now gets the merged ids, in Go and in the volume server.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Publish merged .ecj ids to the journal actually written

mountedEcJournal prefers the receiving disk's own runtime for the vid,
whose journal may live in its data directory while the copied records
name a sibling's journal in the index directory. Publishing by the
requested ecjPath then marked a holder of a different file deleted on
records that file never persisted, resurrecting the needles on remount.
Publish by the picked runtime's journal path instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:46:36 +08:00
MorezMartinGitHubChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
1445960f8c filer: persist pending chunk deletions across restarts (durable deletion ledger) (#11550)
* fix(filer): persist pending chunk deletions across restarts

The in-memory FileIdDeletionQueue and DeletionRetryQueue lose every
queued-but-unconfirmed deletion when the filer process restarts. Because
deletions only enter the pipeline through that queue, a crash between
enqueue and the volume confirming the delete leaks the chunk permanently:
nothing remembers it. In a multi-filer deployment this was observed as
growing collections of orphaned chunks after filer restarts, and — via
meta-replay from a peer that still had the entry — orphans being
"resurrected" as live references on the recovered filer.

This implements the "periodic snapshot with recovery on startup" option
noted in the existing DeletionRetryQueue TODO, using the store's KV layer
(no new iterator API required across the 15+ store backends):

- queueDeletions() is the single entry point that keeps the hot in-memory
  queue and the durable ledger in sync.
- Only terminal outcomes (success / not-found / permanent) remove an id
  from the ledger; retryable failures keep it, which is the point.
- A timer and Shutdown() snapshot the pending set to a single KV key.
- On startup, reloadDeletionLedger() re-queues recovered ids after a
  grace window so the initial peer meta-aggregation settles first. This
  avoids a new hazard: purging a chunk that a lagging peer is about to
  re-reference as live data (stale replay turns a stale read into a
  dangling read otherwise).
- Volume deletes are idempotent (not-found == success), so re-deleting
  after a crash never double-frees.
- Kill switch via viper: filer.deleteQueue.persist=false opts out entirely
  (reload also refuses to recover so a stale ledger never comes back).
  Tunables: filer.deleteQueue.persistInterval, .recoveryGrace.

Adds unit tests covering snapshot+recover, retry-keeps-entry, disabled
switch, and zero-value Filer safety (run green under -race).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)

* filer: harden the deletion ledger

- Scope the ledger key by filer address so filers sharing one store do
  not overwrite each other's pending sets; ledgers written under the
  old unscoped key are claimed once on startup.
- Serialize snapshots on deletionSnapshotLock so an in-flight timer
  snapshot cannot overwrite a newer shutdown snapshot, and wake the
  snapshotter on every queue/forget so a queued id persists within
  milliseconds instead of a full interval.
- Merge recovered ids into the pending set immediately on reload; only
  the queue push waits out the grace window, so an early snapshot
  rewrites the recovered ids rather than dropping them.
- A failed or unparseable ledger read blocks persistence for the run
  instead of letting snapshots overwrite the unread ledger.
- Split the ledger into part keys when it exceeds one 64KB value so
  stores with a size cap (FoundationDB) do not strand the backlog.
- GetReadyItems reports retry-exhausted ids so they are forgotten in
  the ledger instead of replaying after every restart.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: close the remaining deletion-ledger durability gaps

- A manifest referencing a missing part is corruption: surface a wrapped
  error and block persistence instead of treating the ledger as absent.
- Multipart snapshots write generation-scoped part keys and publish the
  manifest last, so a crash never mixes old and new part contents.
- Orphaned parts are tracked in a persisted .stale sidecar and retried.
- Legacy/index ledgers are republished under the scoped key before the
  old keys are removed.
- A ledger index lets a filer restart under a new address claim the
  ledger its previous incarnation left behind.
- Expired and permanently-failed retry items only forget the ledger
  epoch they recorded, so they cannot erase a re-queued id.
- A failed startup read no longer disables persistence: every snapshot
  retries the reload until the store reads again.

* filer: tighten ledger claiming, index updates, and retry epochs

- touchLedgerIndex verifies its write and retries so a concurrent
  filer's merge cannot silently drop this key from the index.
- Foreign-ledger claims abort on any unreadable source instead of
  leaving it stranded once the new scoped key exists.
- A source that republished during the claim is left in place and its
  newer ids merge into the claimant's pending set.
- AddOrUpdate no longer overwrites the ledger epoch of an in-flight
  retry item, so its expiry or permanent outcome cannot forget a record
  that was re-queued after the attempt began.
- The recovery grace wait exits on shutdown instead of re-queueing
  after the filer has stopped.

* filer: requeue surviving records, persist claim deltas, guard index writes

- A dropped retry item (expired or permanent) whose ledger record was
  re-enqueued now pushes the id back through the hot queue instead of
  leaving it pending with nothing scheduled.
- Ids merged from a claim source that republished mid-claim are
  rewritten under our ledger immediately, so they are durable even if
  the claimant crashes before the next snapshot.
- touchLedgerIndex aborts when the index read fails for a real error;
  only ErrKvNotFound means the index is empty, so a transient failure
  can no longer wipe peer entries with a one-key write.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 22:37:31 +08:00
43abe21ffa Helm: Add Opt-in Read-Only Root Filesystem Support (#11562)
* security context changes

Signed-off-by: Subhadeep Maity <smaity@slb.com>

* root file changes

Signed-off-by: Subhadeep Maity <smaity@slb.com>

* Skip tmp mount when extras provide one; use allInOne context for the bucket hook

* Mount tmp for secondary containers; keep user /tmp on the main container

---------

Signed-off-by: Subhadeep Maity <smaity@slb.com>
Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>
Co-authored-by: Subhadeep Maity <smaity@slb.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-02 22:31:48 +08:00
github-actions[bot] 67034dee12 docs: regenerate star history chart 2026-10-02 00:50:58 +00:00
MorezMartinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
8d97284d0a filer: option to store system metadata logs in their own collection (#11551)
* feat(filer): option to store system metadata logs in their own collection

The filer's internal /topics/.system/log chunks are assigned to the
filer's default collection (-collection). In a multi-filer deployment
that default is often empty, so every restart flap, full-sync, or
event-buffered flush grows the default collection with system chunks that
are indistinguishable from user data in collection.list. This is a large
part of what makes the default collection balloon and confuses orphan
analysis.

This keeps the internal log in a dedicated collection when the operator
asks for one, without changing where user data goes:

- New optional override, filer.options.metaLog.collection (and
  .replication), read in NewFiler so both `weed filer` and
  `weed server -filer` honour it. Default "" => exactly today's
  behaviour (log follows the filer default), fully backward compatible.
- Resolution is a small helper: override first, then the filer default,
  then a storage rule matched on the log path. Kept separate from the
  user write path so the internal log targets itself.
- bucketCollection() is hardened the same way it already protects the
  filer's default collection: a bucket that happens to resolve to the
  redirected meta-log collection must not drop it on delete, because it
  backs internal log volumes.
- Scaffold filer.toml documents the new knobs under [filer.options].

Related to the persisted deletion ledger branch (fix/persist-deletion-queue):
together they cut the two sources of post-flap junk in the default
collection — that PR stops orphaned user-chunk leak on filer crash,
this one stops the internal log from living in default at all. They are
independent: no file overlap, no functional dependency; either can merge
first. They are paired only in the narrative of cleaning up default.

Adds unit tests for the collection/replication resolution chain, the
viper keys, and the bucket-delete guard (run green under -race).

Co-Authored-By: Athena 🏛️ <hermes-agent@local> (custom / Qwen3.8-Flash-Next-ROCmFP4)

* filer: collect bucket chunks when its collection survives the delete

bucketCollection returning "" preserves the collection, but the bucket
path still skipped per-entry chunk collection and could skip listing the
children entirely, so a bucket sharing the meta-log (or any preserved)
collection left its object chunks orphaned with no entry pointing at
them. Only the wholesale drop of a deleted collection skips those now.

Note in filer.toml that the meta-log target should stay stable: chunks
written under an older collection are not migrated.

* filer: exercise the metaLog override wiring through NewFiler

The viper test only echoed back the keys it set, so a wrong key in
NewFiler would still pass. It now asserts the fields NewFiler fills
from those keys.

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: tighten comments around the metaLog collection override

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-02 08:50:24 +08:00
d8f926cf46 volume server: stream needle chunks without the store lock (#11488)
* volume server: read GET/HEAD needles off the store lock, and only once

The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stream needle chunks without the store lock

StreamingBody::poll_frame took store.read() and find_volume for every
chunk to compare the volume's compaction revision, dup'd the source
handle, and allocated a fresh chunk buffer. With -hasSlowRead=false the
stream also holds a data-file read lease for its whole life, while a
writer waits for that lease under store.write(): the next chunk's
store.read() then waits for the writer and the writer for the stream.

The per-chunk re-lookup was also wrong. The stream reads a handle opened
at plan time, which pins the .dat inode the offset was resolved against;
a vacuum commit renames a new file over .dat and leaves that inode
untouched. The re-looked-up offset belongs to the new file but was read
from the old inode, so a stream whose needle a vacuum moved ended in a
checksum error. The pinned offset stays valid, so the check, and with
it every store access, is dropped, along with the now unused
re_lookup_needle_data_offset and the revision fields of the read plan.

The source is shared as an Arc instead of dup'd per chunk, and the chunk
buffer is a BytesMut that the blocking read hands back with its result,
so its allocation is reclaimed once the previous frame has been written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a needle stream once its volume becomes unavailable

Taking the store lock out of StreamingBody also dropped its per-chunk
unavailable_error() check. With -hasSlowRead a writer can take the
data-file lease between chunks, fail its fsync and its truncate, and mark
the volume unavailable; the stream then kept serving the rest of the
needle from its pinned handle.

The volume's io_unavailable reason is now an Arc-shared leaf mutex that
the read plan hands to the stream. Each chunk checks it under its
data-file lease, where the writer marks it, and fails with the same
"volume is unavailable: <reason>" error the old check returned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-01 23:15:38 +08:00
8a9563e53d volume: test that only the disk holding the replaced replica gets the free-slot credit (#11486)
* volume server: split volume_copy into phases and type the delete-after-status gate

volume_copy was one ~400-line handler, and the rule that an existing
local replica is deleted only after the source's ReadVolumeFileStatus
succeeded was held by statement order alone. The keep_remote_data=true
that the pre-copy delete and the failed-copy rollback must share was
kept in sync by a comment pointing from one to the other.

The handler is now a ~60-line orchestrator over connect_to_copy_source,
SourceVolumeStatus::fetch, delete_existing_replica,
plan_copy_destination and a VolumeCopyJob whose run() drives
preallocate_dat, transfer_files, finish_copied_files and
mount_and_reply, with cleanup_failed_copy on error.
delete_existing_replica takes a &SourceVolumeStatus, which only fetch
can construct (private field in a child module), so the delete cannot
be called before the status RPC. Both deletes go through
delete_replica_keep_remote.

Pure refactor: call order, status codes and messages, cancellation
checks, throttling, progress reports and cleanup are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: find space for a VolumeCopy before deleting the replica it replaces

VolumeCopy deleted an existing local replica as soon as the source answered
ReadVolumeFileStatus and only then looked for a location with room for the
copy. With no usable location (disk full, low-disk, wrong disk type) the call
errored after the delete, leaving the node with neither the old replica nor
the new one.

Plan the destination first, as Go does: find_free_location_replacing credits
the location holding the replaced volume with that volume's slot, so a disk at
its volume limit that holds the replica still accepts the copy. Only then
delete the replica and write the .note (still after the delete, as in Go).
delete_existing_replica now takes the planned CopyDestination, so the delete
cannot precede the plan. find_free_location_predicate keeps its behaviour.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: describe the replace-credit test against the current VolumeCopy flow

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-01 23:13:17 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
90f8c4378f s3: restrict admin gRPC to local callers when no signing key (#11530)
* s3: restrict admin gRPC to local callers when no signing key

The S3 gateway's gRPC port (default 0.0.0.0:19000, always on) serves the
IAM cache and internal lifecycle admin services. checkAdminAuth was a
no-op when jwt.filer_signing.key was unset, so any reachable host could
PutIdentity an admin identity and take over the bucket data.

Without a shared key callers cannot be distinguished, so admin RPCs are
now limited to unix-socket, loopback, and the server's own interface
addresses. Remote filer-to-S3 propagation and lifecycle workers must set
jwt.filer_signing.key; the Bearer-token path is unchanged.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: fail closed on nil guard and refresh local addresses per call

Review feedback: a nil filerGuard bypassed all checks — treat it like a
missing key and require a local peer. The own-address set was cached
forever, so interfaces added later were rejected; enumerate per call
instead since admin RPCs are rare. Nil ctx is denied rather than panics.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: read the signing key once and bound interface enumeration

Review feedback: reading SigningKey twice could straddle a SIGHUP reload
— an old nonempty key skipped the local-peer check while the new empty
key verified the token. And enumerating interfaces per no-key call is
wasteful for co-located workers dialing the announced address; cache the
address set for 30s so new interfaces still become usable promptly.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: enumerate interface addresses per no-key admin call

A cached address set keeps trusting an IP after it is removed from the
host and reassigned to another machine — that host would then hold
unauthenticated admin access for the cache TTL. Per-call enumeration only
runs for non-loopback TCP peers on the no-key path, which is low-volume
admin traffic, so the freshness is worth the syscall.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 23:12:23 +08:00
eeec9ec09a volume server: refuse to compact a volume tiered to remote storage (#11545)
* volume server: refuse a tier move while compacting, and a commit once tiered

A tier move to remote and a vacuum compaction of the same volume could
interleave and leave the volume unreadable:

- A compaction committing while the upload ran swapped .dat/.idx under
  the transfer, which reopens the .dat by path per part. The move then
  published an object holding the old (or a mixed) layout against the
  compacted .idx, and with keep_local_dat_file=false deleted the only
  compacted .dat.
- A tier move finishing while the compaction copy ran (or between the
  copy and the commit) let the commit swap in the compacted .idx while
  the reload served the pre-compaction remote object through it.

The tier move now refuses to start while the volume is compacting, and
re-checks the compaction revision under the store write lock before it
records the remote file; on a mismatch it deletes the uploaded object
and fails with FailedPrecondition, leaving the volume local. Committing
a compaction on a volume that has a remote file is refused and its
.cpd/.cpx removed, since the reload would read the remote object
through the compacted index.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: abort a tier move whose volume was replaced or removed

The tier-up bookkeeping looked the volume up by id only and compared the
compaction revision. A delete and re-create of the same id during the upload
yields a fresh volume at the same revision, so the move recorded the old
volume's object on the new one and, without keep_local_dat_file, removed the
new .dat. An unmounted volume was skipped and the move reported success,
leaving the uploaded object referenced by nothing.

Capture the volume instance (its data-file access control Arc, as the scan
and read plans do) with the revision, and require both under the store write
lock. A replaced volume fails with FailedPrecondition, a missing one with
NotFound; either way nothing is recorded and the object is deleted after the
lock is released. Go fails in both cases because deleting or unmounting closes
the descriptor its copy reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: refuse to compact a volume tiered to remote storage

Committing a compaction of a tiered volume is refused, since the reload
would read the remote object through the compacted index. The compaction
itself still started: a tiered volume's data backend is the remote object
(the local .dat is dropped or deleted on tier-up), so an explicit vacuum
streamed the whole .dat out of remote storage into a .cpd that the commit
then discarded.

Refuse at the start of the compaction instead, before the .cpd is created,
at the point where Go's copy opens the local .dat. The truncated-index
test now uses a read-only local volume for its sorted index, since a
tiered one no longer reaches the copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: match the ReadOnly(VolumeId) variant in write_volume_needles

#11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-01 23:11:16 +08:00
deepnemesisGitHubcoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>Subhadeep Maitycoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
7c7834c98e Helm: Add Configurable Security Contexts for Chart-Managed Workloads (#11558)
* security context changes

Signed-off-by: Subhadeep Maity <smaity@slb.com>

* extending examples

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

* updating examples

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

* examples

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

* Update k8s/charts/seaweedfs/values.yaml

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Update k8s/charts/seaweedfs/values.yaml

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* review comments

Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>

---------

Signed-off-by: Subhadeep Maity <smaity@slb.com>
Signed-off-by: Subhadeep Maity <322813880+deepnemesis@users.noreply.github.com>
Co-authored-by: Subhadeep Maity <smaity@slb.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-10-01 22:39:54 +08:00
2a42d56437 ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst (#11553)
* ecbalancer: honour total-shards-per-rack cap in Place / PlaceDurabilityFirst

Worker auto-EC encode places via Topology.Place, which capped each shard
type independently (ceil(data/racks), ceil(parity/racks)). On an 8-rack
topology that permits 3 total shards on one rack, so losing two racks
strands 6/14 and a 10+4 volume becomes unreadable.

- tryPlace caps the total shards (data + parity) per rack in both modes,
  whether or not ReplicaPlacement is set.
- rackTotalCap picks the smallest per-rack total the racks' real room
  (free slots, bounded by the per-disk cap and node free slots, counting
  shards already placed) can satisfy. On a uniform cluster it is
  ceil(shards/racks); a nearly full rack raises it just enough that the
  cap alone never fails an encode.
- PlaceDurabilityFirst gets a last rung that drops the rack cap
  ("rack-total-cap" in Relaxed), so it fails only when no disk has room.
  PlaceStrict keeps the cap as a hard limit.
- chooseShardDest tries the next rack when the chosen one has no node
  that fits, and room checks count the per-disk cap, so a rack whose
  disks are all at the cap is no longer picked and then failed on
  (pre-existing: 3-node rack + single-disk rack failed at shard 9).
- Docs no longer claim the cap guarantees surviving rack loss; the
  placement error names the caps in effect; the encode warning no longer
  says replica placement when other constraints were relaxed.

place_rack_cap_test.go covers 10+4 over 8 racks (max 2/rack, 3/rack on
master), a starved rack, nearly full racks, the preferred-tag tier, the
full-disk rack, and rackTotalCap directly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ecbalancer: size the rack total cap from room left under SameRackCount

The rack total cap counted each rack's free disk room, but attempts that
enforce ReplicaPlacement also stop a node at SameRackCount shards. With
SameRackCount=1, four one-node racks and four three-node racks got cap 2,
which fits only 12 of 14 shards: strict placement failed and
durability-first relaxed replica placement although 1 per small rack and
up to 3 per large rack fits.

Attempts that enforce ReplicaPlacement now use a cap sized from each
node's remaining SameRackCount allowance; attempts that relax it keep the
disk-room cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 22:22:34 +08:00
164c3db606 s3: return 403, not 500, when an over-quota bucket refuses a write (#11552)
* s3: return 403, not 500, when an over-quota bucket refuses a write

Filer AssignVolume flattened ErrReadOnly into the free-text
AssignVolumeResponse.Error string, so S3 PutObject / PutObjectPart via
UploadReaderInChunks could not match it with errors.Is and fell through
to 500 InternalError: retryable, and it hides the quota.

Add FilerError READ_ONLY and AssignVolumeResponse.error_code, set it
alongside the unchanged error text, and rebuild the sentinel with
filer_pb.AssignVolumeResponseError. weed_server.ErrReadOnly now aliases
filer_pb.ErrReadOnly so errors.Is matches on both sides, and
mapChunkedUploadErrorToS3Error maps it to ErrAccessDenied. There is no
"read only" substring matching, so a volume server's "volume N is read
only" stays retryable.

Carrying the verdict as a response code rather than a gRPC status keeps
clients from treating it as a transport failure: the S3 gateway does not
fail over across filers and the Java client does not retry it.

Wrap per-chunk copy errors with %w so CopyObject keeps the sentinel, and
map UploadPartCopy chunk errors through mapCopyErrorToS3Error instead of
always returning 500.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: re-run integration tests (PyPI download timeout)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 22:21:03 +08:00
6f9becaa37 volume server: answer BatchDelete on EC needles as Go does (#11541)
* volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go

For an EC volume the handler read only locally mounted shards, so a node
that did not hold the shard with the needle's bytes answered Internal
"ec shard N not available locally". Go's ReadEcShardNeedle fetches the
interval from a peer or reconstructs it. It also mapped every regular
volume read error, including a tombstone, to NotFound "needle not found",
which fs.verify treats as a missing needle; Go returns ErrorDeleted as a
plain error ("already deleted"), which fs.verify skips.

The EC branch now drops the store guard and uses the distributed EC read
the HTTP GET path uses. Errors map like Go: needle absent -> NotFound
"needle not found <decimal id>", tombstoned (regular or EC .ecx/.ecj) ->
Unknown "already deleted", anything else -> Unknown with the error text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: tell EC deletions and vanished volumes apart in VolumeNeedleStatus

The distributed EC reader returned Ok(None) for an absent needle, a needle
a peer reported deleted, and a volume unmounted after the handler's own
existence check. VolumeNeedleStatus answered all three NotFound "needle not
found", which fs.verify -pruneEntries counts as lost data. A reported
deletion was also lost when an earlier interval failed.

The reader now says why it has no needle (EcMiss: NotFound, Deleted,
VolumeNotFound), classifying the local tombstone itself and letting a
reported deletion outrank other interval errors, as Go's ReadEcShardNeedle
does. VolumeNeedleStatus maps Deleted to Unknown "already deleted" and
VolumeNotFound to "volume not found", and drops its separate EC pre-check.
read_ec_shard_needle_distributed keeps its Ok(None) for every miss, so the
other callers are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: answer BatchDelete on EC needles as Go does

With skip_cookie_check, which every weed/ client sends, an EC needle that
was already deleted came back 404 "ec needle <fid> not found". Go's
DeleteEcShardNeedle gets ErrorDeleted from its read and BatchDelete
answers 304 with no error; the filer's deletion classifier only forgives
"already deleted" or an exact "not found", so it booked the repeat delete
as a permanent failure. The same mode also compared the fid cookie and
refused chunk manifests with 406, while Go never reads the needle before
those checks when skipping, so the filer's delete of a manifest chunk's
own fid failed permanently too.

The EC branch now reads with read_ec_shard_needle_or_miss and answers as
Go: skipping, a deletion is 304 and any other miss is 500 with Go's text;
checking, every miss is 404 with Go's text ("already deleted",
"locate in local ec volume: FindNeedleFromEcx: needle not found",
"ec shard <vid> not found"). The cookie and manifest checks run only when
the caller asked for the cookie check, which leaves the non-EC path as it
was.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: match the ReadOnly(VolumeId) variant in write_volume_needles

#11543 matched VolumeError::ReadOnly as a unit variant in Store::write_volume_needles, and #11544 changed it to ReadOnly(VolumeId) in the same merge window. Each passed CI on its own, but master no longer compiles the Rust volume server. Carry the volume id through.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-10-01 22:17:18 +08:00
Chris Lu 3fd5e4c329 volume-rust: match the tuple variant VolumeError::ReadOnly in write_volume_needles 2026-10-01 06:24:49 -07:00
9d3907e36c volume: say "volume N is read only" like Go, so filer retries deletes (#11544)
VolumeError::ReadOnly displayed "volume is read-only". Go's store and
volume say "volume %d is read only", and the filer's deletion classifier
requeues a failed delete only when the error contains "is read only".
Against a Rust volume server a BatchDelete on a read-only volume (tier
move, maintenance) was booked as a permanent failure and the chunk was
never deleted.

ReadOnly now carries the volume id and displays Go's text. The text
reaches clients through BatchDelete results, the HTTP write and delete
error bodies, and gRPC statuses; the gRPC code (FailedPrecondition) and
the HTTP/BatchDelete status codes are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 21:00:04 +08:00
bd953b0f84 volume server: group-commit fsync writes in the write queue (#11543)
* volume: split the write path into reusable steps

do_write_request ran its pre-append checks, the append, the sync
rollback, the index publish and the post-write bookkeeping inline, so a
batched write could only reuse it one needle at a time. Pull the steps
out (check_writable, prepare_write, undo_unsynced_append, publish_write,
finish_write) and the store's volume lookup plus disk-space check
(writable_volume_mut). do_write_request composes them in the same order
with the same early returns; no behaviour change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: group-commit fsync writes in the write queue

The write queue holds one store lock for a batch of up to 128 needles
but wrote them one at a time, so every fsync needle paid its own .dat
sync and its own .idx sync: 2N syncs per batch.

Add Volume::write_needles_grouped, after Go's processBatch. A volume's
entries are split into runs of distinct needle ids (a repeated id starts
a new run, so its dedup and cookie checks see the earlier write). A run
with a durable entry appends everything with append_at_ns chained
through a local, syncs the .dat once, and only then publishes the
entries and syncs the .idx once. A failed .dat sync truncates the .dat
back to the run start (marking the volume unavailable if that fails),
leaves last_append_at_ns and last_modified untouched, and fails every
entry of the run. Runs with no durable entry go through the unchanged
per-needle path. Store::write_volume_needles is the queue's entry point;
the handlers' non-queue path is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: fail closed when a failed append's rollback fails

append_needle discarded the truncate-back result, so a partial write
that could not be rolled back left unindexed bytes on the .dat while
the volume stayed writable; the next append would bury them mid-file,
past the load-time tail check. Route the rollback through
undo_unsynced_append, which marks the volume unavailable when the
truncate fails, so nothing more is appended over an unverified tail.

* volume server: keep a grouped run's I/O error streak from later appends

A synced run stages every append before any entry finishes, so the
success reset in finish_write ran after the failed appends queued
behind the last write to land and erased their media-error streak.
Sent one at a time, those errors would have counted and quarantined
the volume. Skip the reset when an append after the last landed write
added to the streak.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: replay a grouped run's I/O error streak in queue order

Skipping the run's success reset whenever an append after the last
landed write failed kept the errors from before that write as well,
so a run like [EIO, EIO, landed, EIO] reached the quarantine count
that the same writes one at a time (one error) do not. Mark the
streak where each entry is staged and record the run's success at
the last landed write's mark: errors before it are cleared, the ones
after it still count.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: replay a run's I/O error streak in one atomic step

Reads record their outcomes on the tracker without the volume's write
lock, so record_success_at's separate load and store could drop an
error a read counted in between, or restore errors a read had just
cleared. Keep the count and the clear counter in one atomic word and
apply the replay with a single fetch_update.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 20:59:40 +08:00
7944cb4ba2 volume server: keep repeated response headers when proxying a read, like Go (#11542)
In readMode=proxy, proxy_request copied the target's response headers
with HeaderMap::insert, so a header the target sent more than once
(several Set-Cookie, Vary, Link, ...) reached the client with only its
last value. Go's proxyReqToTargetServer adds every value with
w.Header().Add. Append instead of insert; the Server header is still
dropped and status and body handling are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:58:50 +08:00
15d3c65e0c volume server: take the heartbeat's remaining store reads off the runtime (#11540)
* volume server: collect EC heartbeats and adjust volume max off the runtime

The volume pass moved to the blocking pool, but the heartbeat task still
called collect_ec_heartbeat and the following EC shard snapshot, and
Store::maybe_adjust_volume_max, directly on its tokio worker.
maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat
of every writable volume under the store read lock. All of them block the
worker on the node-wide RwLock<Store> whenever a writer holds it or is
queued, and every task sharing that worker stalls with it.

Run the adjustment, on the pulse and after the master changes volume
options, and the EC tick's heartbeat plus shard snapshot through
off_runtime, like the volume pass. apply_master_volume_options now only
reports whether the options changed; the loop adjusts off the runtime.
What is collected and sent, and in what order, is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: keep EC shard deltas out of the volume heartbeat's snapshot

EC shard deltas are the diff between the store's shards and last_ec_shards,
taken when volume_state_notify fires. But the volume tick and the
options-changed heartbeat re-took last_ec_shards from the store too, and a
volume heartbeat carries no shard list: a mount or unmount that landed while
the notify was pending or the volume pass was collecting was absorbed into
the baseline and never sent. The master only learned of it at the next EC
tick, 17 pulses later. The EC tick likewise built its heartbeat and its
baseline under two separate store guards, so a mount between them was lost
the same way.

A volume heartbeat now only takes out of the baseline the expired EC shards
it reports deleted itself, so the next delta does not repeat them. The EC
tick, and the initial EC heartbeat, build the full list and the baseline
under one read guard, still on the blocking pool.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: take the heartbeat's remaining store reads off the runtime

The heartbeat's volume and EC passes and the volume max adjustment already
run on the blocking pool, but several reads of the node-wide RwLock<Store>
were still taken directly on the heartbeat's tokio worker: the digest
report reset before the first heartbeat, the duplicate-UUID directory
lookup and the volume options a master response carries, the EC shard list
a state notification is diffed against, and the deregistration heartbeat
sent on stop and shutdown. The lock is writer-preferring, so with a writer
holding or queued for it each of these parks the worker, and every task
sharing that worker stalls with it.

Run each through off_runtime, which now takes a closure so a pass can carry
what it needs from the master's response. The notify branch's volume
snapshot and EC read become one blocking pass, still under two guards in
the same order. What is collected and sent, and in what order, is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 20:58:01 +08:00
fea14c01a7 volume server: refuse a tier move while compacting, and a commit once tiered (#11539)
* volume server: refuse a tier move while compacting, and a commit once tiered

A tier move to remote and a vacuum compaction of the same volume could
interleave and leave the volume unreadable:

- A compaction committing while the upload ran swapped .dat/.idx under
  the transfer, which reopens the .dat by path per part. The move then
  published an object holding the old (or a mixed) layout against the
  compacted .idx, and with keep_local_dat_file=false deleted the only
  compacted .dat.
- A tier move finishing while the compaction copy ran (or between the
  copy and the commit) let the commit swap in the compacted .idx while
  the reload served the pre-compaction remote object through it.

The tier move now refuses to start while the volume is compacting, and
re-checks the compaction revision under the store write lock before it
records the remote file; on a mismatch it deletes the uploaded object
and fails with FailedPrecondition, leaving the volume local. Committing
a compaction on a volume that has a remote file is refused and its
.cpd/.cpx removed, since the reload would read the remote object
through the compacted index.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: abort a tier move whose volume was replaced or removed

The tier-up bookkeeping looked the volume up by id only and compared the
compaction revision. A delete and re-create of the same id during the upload
yields a fresh volume at the same revision, so the move recorded the old
volume's object on the new one and, without keep_local_dat_file, removed the
new .dat. An unmounted volume was skipped and the move reported success,
leaving the uploaded object referenced by nothing.

Capture the volume instance (its data-file access control Arc, as the scan
and read plans do) with the revision, and require both under the store write
lock. A replaced volume fails with FailedPrecondition, a missing one with
NotFound; either way nothing is recorded and the object is deleted after the
lock is released. Go fails in both cases because deleting or unmounting closes
the descriptor its copy reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:20:16 +08:00
f2498e122a volume: read EC shards fully, like Go's ReadAt (#11537)
EcVolumeShard::read_at and the scrub plan's EcLocalShard::read_at were a
single pread/seek_read. That may legally return fewer bytes than asked
mid-file (FUSE/NFS/CIFS mounts, a signal, very large requests), and an
Interrupted error was not retried. Callers treat a short count as end of
file or corruption: verify_ec_shards compared a zero tail and reported a
parity mismatch, local scrub reported a broken shard, VolumeEcShardRead
ended the stream early, and decode/rebuild/local needle reads failed.

Add storage::io::read_full_at, which loops until the buffer is full or a
read returns 0 and retries Interrupted, so a short count means EOF. Route
both shard read_at methods through it, replace the encoder's private
read_at_most with it, and reuse it for the Windows read_exact_at loop.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:19:43 +08:00
62481f1673 volume server: ec.decode reads the .ecx from the index dir it was copied to (#11536)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:19:17 +08:00
b8f074b7d3 volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go (#11535)
* volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go

For an EC volume the handler read only locally mounted shards, so a node
that did not hold the shard with the needle's bytes answered Internal
"ec shard N not available locally". Go's ReadEcShardNeedle fetches the
interval from a peer or reconstructs it. It also mapped every regular
volume read error, including a tombstone, to NotFound "needle not found",
which fs.verify treats as a missing needle; Go returns ErrorDeleted as a
plain error ("already deleted"), which fs.verify skips.

The EC branch now drops the store guard and uses the distributed EC read
the HTTP GET path uses. Errors map like Go: needle absent -> NotFound
"needle not found <decimal id>", tombstoned (regular or EC .ecx/.ecj) ->
Unknown "already deleted", anything else -> Unknown with the error text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: tell EC deletions and vanished volumes apart in VolumeNeedleStatus

The distributed EC reader returned Ok(None) for an absent needle, a needle
a peer reported deleted, and a volume unmounted after the handler's own
existence check. VolumeNeedleStatus answered all three NotFound "needle not
found", which fs.verify -pruneEntries counts as lost data. A reported
deletion was also lost when an earlier interval failed.

The reader now says why it has no needle (EcMiss: NotFound, Deleted,
VolumeNotFound), classifying the local tombstone itself and letting a
reported deletion outrank other interval errors, as Go's ReadEcShardNeedle
does. VolumeNeedleStatus maps Deleted to Unknown "already deleted" and
VolumeNotFound to "volume not found", and drops its separate EC pre-check.
read_ec_shard_needle_distributed keeps its Ok(None) for every miss, so the
other callers are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:18:34 +08:00
cc1ec48151 volume server: collect EC heartbeats and adjust volume max off the runtime (#11532)
The volume pass moved to the blocking pool, but the heartbeat task still
called collect_ec_heartbeat and the following EC shard snapshot, and
Store::maybe_adjust_volume_max, directly on its tokio worker.
maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat
of every writable volume under the store read lock. All of them block the
worker on the node-wide RwLock<Store> whenever a writer holds it or is
queued, and every task sharing that worker stalls with it.

Run the adjustment, on the pulse and after the master changes volume
options, and the EC tick's heartbeat plus shard snapshot through
off_runtime, like the volume pass. apply_master_volume_options now only
reports whether the options changed; the loop adjusts off the runtime.
What is collected and sent, and in what order, is unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:14:15 +08:00
5ece8dd63c volume server: drop the unused unmount result in VolumeCopy validation, and test-only EC helpers (#11531)
mount_and_reply ignored the Result of store.unmount_volume when a copied
replica failed record count validation, tripping unused_must_use. The Err
branch is unreachable there: the volume was mounted under the same store
write guard, mount_volume refuses an already loaded vid so it is a fresh
Volume with is_compacting false, and a compaction claim needs &mut Volume,
i.e. the store lock. Ignore the result explicitly with a one-line reason.

Store::delete_expired_ec_volumes and Store::remove_ec_volume are called
only from test modules (the heartbeat uses the split
find_expired_ec_volumes / remove_expired_ec_volumes halves), so mark them
#[cfg(test)].

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:13:45 +08:00
github-actions[bot] 67f7fad987 docs: regenerate star history chart 2026-10-01 00:56:45 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2b057eccbf build(deps): bump com.fasterxml.jackson.core:jackson-databind from 2.22.2 to 2.22.3 in /test/java/spark (#11548)
build(deps): bump com.fasterxml.jackson.core:jackson-databind

Bumps [com.fasterxml.jackson.core:jackson-databind](https://github.com/FasterXML/jackson-databind) from 2.22.2 to 2.22.3.
- [Commits](https://github.com/FasterXML/jackson-databind/compare/jackson-databind-2.22.2...jackson-databind-2.22.3)

---
updated-dependencies:
- dependency-name: com.fasterxml.jackson.core:jackson-databind
  dependency-version: 2.22.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 01:43:45 +08:00
Javier GarciaandGitHub 8fdcf69eb0 s3api: report the stored checksum in GetObjectAttributes (#11529)
GetObjectAttributes accepted the Checksum attribute but never filled it
in, as its comment said SeaweedFS did not store S3 checksums. PutObject
and CompleteMultipartUpload store them now, and HeadObject returns them.
Fill in Checksum from the same entry fields, with the ChecksumType and
ChecksumCRC64NVME members the response did not have.

Also run ceph/s3-tests' test_get_checksum_object_attributes in CI.
2026-10-01 01:11:19 +08:00
Javier GarciaandGitHub 988fc4f7ba s3api: do not store aws-chunked in an object's Content-Encoding (#11528)
* s3api: do not store aws-chunked in an object's Content-Encoding

aws-chunked in Content-Encoding names the SigV4 streaming framing of
the request body, which the gateway decodes on upload. PutObject and
CreateMultipartUpload stored the header as sent, so an object uploaded
with "gzip, aws-chunked" was served with that Content-Encoding, and one
uploaded with "aws-chunked" alone was served as aws-chunked. S3 drops
aws-chunked and keeps the other encodings.

Also run ceph/s3-tests' test_object_content_encoding_aws_chunked in CI.

* s3api: read every Content-Encoding field, and drop aws-chunked on copy

A client can send aws-chunked and the object's own encoding as separate
Content-Encoding fields. Only the first was read, so "aws-chunked"
followed by "gzip" left the object without its gzip. Combine all the
fields before dropping aws-chunked. CopyObject with the REPLACE
directive stored the requested Content-Encoding as sent: drop
aws-chunked there too.
2026-10-01 01:10:38 +08:00
ihnokimandGitHub 11e8c4c288 master: follow heartbeat read-only changes in the layout's replica flag (#11527)
A replica's read-only flag in the volume layout only moved on registration
and on volume.mark. A change that arrived in the regular heartbeat updated
the node's record, which the writable list follows, but not the layout
flag, which the vacuum sweep reads. So the sweep kept trying volumes on a
disk that had gone read-only while the server ran, and after a restart it
skipped volumes that had since become writable again until the next
restart (issue #11516).

Apply the reported state to the flag for every changed volume. Only the
flag: the writable list stays with EnsureCorrectWritables and its
capacity guards.
2026-10-01 01:09:09 +08:00
ihnokimandGitHub d3cd061c22 shell: say which read-only volumes volume.vacuum leaves alone (#11525)
* shell: say which read-only volumes volume.vacuum leaves alone

volume.vacuum without -volumeId runs the same sweep as the automatic
vacuum, which skips read-only volumes, and the master's response carries
no result. An operator whose disk filled up runs the command, sees it
return, and watches nothing change (issue #11516).

Before issuing the request, list the read-only volumes whose garbage is at
or above the threshold and point at -volumeId, which is the explicit path
PR #9861 opened for them. The help text says the same.

* shell: volume.vacuum hint survives a failed listing and looks at every replica

Review follow-ups: a failed topology listing no longer stops a sweep
without -volumeId, it only drops the hint; a volume counts as read-only
when any replica is, with the garbage ratio taken from the replica that
reports the most, which is what the sweep itself does; a converted index
that reports deletes without sizes is listed rather than hidden; and the
threshold is printed as given instead of rounded to two decimals.

* shell: do not guess a garbage ratio for a converted index

The master cannot compute one for a volume that reports deletes without
their sizes, and a guess of 1 would send the operator to -volumeId for a
volume the server may decline at that threshold. Leave it out and say so.
2026-10-01 01:08:48 +08:00
895d49b55b s3/iam: manage roles through the IAM API, with an opt-in persistent role store (#11522)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store

Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.

Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
  IAM config file the default stays memory; with none it is the filer, as
  for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
  which outlives the file and may be shared by S3 servers with different
  files. They are served from memory beneath the store, as OIDC providers
  are: a stored role of the same name takes precedence, and deleting it
  restores the file's. A config-file role cannot be changed or deleted
  through the API (UnmodifiableEntity), and removing one from the file
  removes it at the next start. An in-memory store holds them as records,
  as before. They have no creation time, so CreateDate is omitted rather
  than reporting when this server started. SetRoleStore installs a store
  the same way, so a store set after startup keeps the config-file roles,
  as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
  cached filer store otherwise serves a peer's stale role for up to its 5m
  TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
  used to report any failed lookup as "role not found". CreateRole proceeds
  only on a confirmed absence, so an unreadable store cannot let it write
  over an existing role.

IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
  AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
  are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
  changing or deleting it through the API is refused (UnmodifiableEntity)
  rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
  <name>.json in the filer, so this also keeps a name from leaving the role
  store's directory. At most 10 managed policies per role (AWS's default
  quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
  policy, as it already is for users and groups: roles attach policies by
  name, so a policy created later under the deleted one's name would
  otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
  refused rather than dropped.

Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
  kept across updates; a config-file role gets a stable ID derived from its
  name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
  and AssumeRole carry the role's ID (claim "rid"), and a request under a role
  whose current ID differs is denied. Resolving a session's policies by role
  name let a session outlive its role: once a role was deleted, a role later
  created under the same name — with a different trust policy and different
  policies — revived every unexpired session of the old one with the new
  role's permissions. Sessions issued before this change carry no ID and are
  unaffected until they expire.

Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: bind every role session to its role, and change roles atomically

Review follow-ups.

Session binding
- The role-ID check ran only when a session carried no policy names, and
  AssumeRole embeds the role's attached policies, so those sessions kept
  their permissions after the role was deleted or recreated. The check
  now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
  instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
  the name alone: a different role put in the file under the same name
  gets a new ID, while an unchanged role keeps its sessions across restarts.

Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
  role is unchanged since the read, and otherwise re-reads and retries. The
  filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
  role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
  Attach/DetachRolePolicy all go through it, so two gateways no longer
  overwrite each other's changes, a change racing a delete no longer
  writes the role back, and of two concurrent creates one gets
  EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
  broken stream instead of returning what arrived, so DeletePolicy's
  attachment check sees every role. ListRoles skips a role deleted between
  listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
  InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
  comment now says so. Role and OIDC provider actions write their own stores.

Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path

Review follow-ups.

- Authorization evaluates the policies of the role definition the session's
  binding was checked against, instead of reading the role again: a role
  replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
  definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
  and take its ID, duration cap and embedded policies from that same
  definition. A role replaced after the caller's trust check by one that does
  not trust the caller now yields AccessDenied, not a session bound to the
  replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
  write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
  routed and locked like the conditional CreateEntry. DeleteRole decides
  against the role it deletes, so a policy attached meanwhile on another
  server is a DeleteConflict, and a delete never removes a role written
  after its check.
- S3 servers watch the role store's configured basePath, not only
  /etc/iam/roles, so a custom path also drops peers' cached roles on change.

Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: state which roles DeletePolicy's attachment check can see

RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: note that a role store set after startup is not watched for peer changes

S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:45:47 +08:00
Khris RichardsonandGitHub 95e0b74fb6 s3/iam: retry a failed OIDC provider refresh until the store answers (#11521)
* s3/iam: retry a failed OIDC provider refresh until the store answers

RefreshOIDCProvidersFromStore reports a failure and nothing retries it. Its
callers can't: a metadata-subscription event reports each change once, so a
refresh that found the filer unreachable on it (the filer restarting, say)
left a peer's new provider untrusted, or a deleted one trusted, until some
unrelated later change. The refresh after a local IAM API mutation has the
same shape. Only the startup load retried.

A failed refresh now retries in the background with the startup load's
backoff until the store answers. At most one retry runs, however many
refreshes fail meanwhile, and installing another store cancels it. The
startup load uses the same path instead of its own.

Seen on a SeaweedFS operator cluster whose filer restarted while an
S3OIDCProvider was created: the gateway logged "OIDC provider refresh after
/etc/iam/oidc-providers change failed: ... fail to dial". The operator's
periodic re-apply happened to recover it; an IAM API client would not.

* s3/iam: never retry or apply a superseded OIDC provider store, and never drop a failure during a retry

Review of the retry (#11521) found two ways to lose the state it protects.

A refresh of store A that failed as store B was installed could start a
retry for A after B's install had cancelled retries. Nothing cancelled it,
and when A answered it replaced B's providers in STS. The installed store
now changes under the retry lock, a store that is no longer current gets no
retry, and a snapshot of a replaced store is never handed to STS, even when
the refresh listed it just before the swap.

A refresh that failed while a retry ran was dropped by the at-most-one
guard, though the retry might already have listed an older snapshot, so the
change the failed refresh would have loaded stayed unloaded. The retry now
runs once more after its success when a failure arrived meanwhile.

Each has a test that fails without its guard.
2026-09-30 17:33:40 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
0978e7f833 vacuum: keep disk-full read-only volumes reclaimable (#11519)
* storage/topology: keep disk-full read-only volumes vacuumable

The vacuum sweep skipped every read-only replica, so a volume that went
read-only because its disk filled could never reclaim its garbage — the
exact situation compaction exists for. The volume server now reports
disk_space_low in VacuumVolumeCheckResponse, and the sweep skips a
read-only replica only when the flag is clear. An explicit volumeId
vacuum is unaffected: it already bypassed the read-only rule.

The field takes number 4: 2 and 3 are downstream-allocated for tombstone
retention, keeping the wire merge clean.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: measure vacuum free space against live bytes

The pre-compaction space check required the current .dat + .idx size
free, which includes the garbage being reclaimed — on a nearly full disk
that estimate can never fit, so the volume stayed garbage-bound forever.
Measure against the estimated compacted output instead: superblock plus
live index entries plus live content bytes, with the existing ten
percent buffer unchanged. Mirrors the same check in the Rust volume
server.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: count per-needle framing in the compacted-size estimate

The live-bytes estimate covered each live needle's content and index
entry but not its .dat framing (header, checksum, timestamp, padding —
~32 bytes on version 3). For small-needle volumes that is more than the
10% headroom, so a disk with space between the estimate and the real
output still ran out mid-compaction. Rust side mirrors the same formula.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: report disk_space_low only when it is the sole read-only cause

Review feedback (ihnokim, greptile, devin): a volume read-only for low
disk space AND an operator mark or I/O quarantine was still eligible for
the automatic sweep, rewriting a copy meant to stay protected. The flag
now reports only the benign sole-cause case in both servers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* topology: fail closed when the read-only lookup misses in the sweep

A heartbeat can drop the volume from the DataNode cache between the
location-list copy and VacuumVolumeCheck; a lookup error previously
skipped the read-only check entirely. Review feedback (coderabbit).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 17:32:41 +08:00
Alex HuandGitHub 38ce95d960 s3api: always write XML timestamps with three fractional digits (#11520)
CopyObject responses carried LastModified values such as
"2026-09-29T20:30:04.56Z": trailing zeros of the fractional seconds were
trimmed, and a whole-second value had no fraction at all. AWS S3 always
writes exactly three digits ("...04.560Z"), and clients that parse with a
fixed-width pattern reject anything else. minio-java 8.6.0
(yyyy-MM-dd'T'HH:mm:ss.SSS'Z') throws DateTimeParseException, so roughly
one CopyObject in ten fails on the client even though the copy succeeded.

Two causes:

- xsdDateTime marshalled with "2006-01-02T15:04:05.999999999", which
  drops trailing zeros. It now writes UTC with ".000Z".
- CopyObjectResult.MarshalXML had a pointer receiver, but the handlers
  pass the result by value, so encoding/xml never called it and fell back
  to time.Time's RFC 3339 encoding. It now has a value receiver.
  CopyPartResult had no custom marshaller at all; it now uses xsdDateTime.

Follow-up to #8394 / #8398, which truncated these timestamps to
milliseconds but kept the trimmed format.
2026-09-30 16:22:31 +08:00
github-actions[bot] 9a454bbe2b docs: regenerate star history chart 2026-09-30 00:53:03 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
757917f564 filer: evict remote-cached objects under storage pressure (#11515)
* filer: identify remote-mounted entries safe to drop under disk pressure

ListEvictableRemoteEntries walks every mounted directory directly on the
filer store (no lazy remote listing) and returns entries that hold local
chunks fully synchronized with remote, ordered oldest-cached first.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: evict remote-cached chunks oldest-first and vacuum the garbage

uncacheRemoteEntry applies the same transition remote.uncache does -
cleared chunks plus a reset LastLocalSyncTsNs under the entry path lock -
and evictRemoteCachedEntries serializes passes over all mounts until a
byte target is met. Aged victims are preferred; a second pass accepts any
synchronized cached entry when aged ones cannot cover the request, since
a failed read is worse than a dropped hot object.

Cleared chunks only become disk space after compaction, so
reclaimRemoteCacheSpace pairs each pass with a rate-limited VacuumVolume
call that also picks up orphaned partial fills.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: trigger remote cache eviction under storage pressure

A periodic check (30s) reads disk usage from master topology and evicts
remote-mounted cached chunks once any disk crosses
-filer.remoteCacheEvictThreshold (default 0.9; 0 disables), with a vacuum
pass to reclaim the tombstoned needles.

The cold-read cache path also kicks the same reclaim when a fill fails on
exhausted volumes - the request still falls back to streaming from the
remote, but the cache stops being permanently wedged full.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: flush deletion queue before remote cache vacuum

Vacuum ran immediately after eviction while evicted file IDs still sat
in the asynchronous deletion queue, so compaction saw no garbage and the
cache stayed wedged. Flush the queue synchronously first and shorten the
vacuum cooldown so sustained pressure does not wait five minutes between
reclaim passes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover remote cache eviction under capacity pressure

Unit tests pin the eligibility filter and oldest-first ordering; the
integration test runs a constrained two-node setup that saturates the
cache, verifies the oldest synced entry is evicted and vacuumed, and
that a later read re-caches it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: coalesce remote cache reclaim passes

A failed cache fill used to queue behind any in-flight eviction,
stacking full mount traversals during a write-failure storm. Skip the
pass when one is already running; the caller falls back to streaming
from remote regardless.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: stop the remote cache janitor on shutdown

The eviction ticker kept running after Shutdown closed the metadata
store and could traverse a closed store. Give the janitor a context
cancelled from Shutdown and propagate it into its master RPCs and
traversals.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: vacuum only tombstoned volumes and retry deferred passes

VacuumVolume with no volume id swept every collection, compacting
volumes unrelated to the cache fill that failed. Now the reclaim path
collects the vids of file ids actually flushed from the deletion queue
and compacts only those. Vids that land inside the vacuum cooldown stay
in a pending set the janitor retries on each tick, so chunks evicted
just after a sweep are not stranded until the next pressure event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: count only pressured disks when evicting remote cache

The janitor measured the largest excess on one disk but let bytes on
healthy disks satisfy the reclaim target. Split the topology disk view
per physical disk and count only chunk bytes whose volumes sit on an
over-threshold disk; entries contributing nothing there are skipped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: compare remote cache sync time at nanosecond precision

Second-precision mtime comparisons let a local write in the same second
as the last sync still qualify as evictable, discarding unsynced
changes. Compare LastLocalSyncTsNs against full-precision mtime
(mtime_ns round-trips through the entry codec), and apply the same fix
to remote.uncache's inline check.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: invalidate remote sync stamp on local content change

A local overwrite that keeps the remote entry's LastLocalSyncTsNs looks
evictable even though the remote copy no longer matches, and some write
paths stamp mtime at second precision so a timestamp comparison cannot
catch it. UpdateEntry now clears the stamp when chunks change without a
fresh stamp, leaving replicated updates authoritative.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound remote cache master rpcs and vacuum all evicted garbage

VolumeList and VacuumVolume now run under a 30s context so a stalled
master cannot wedge the eviction janitor. The targeted vacuum drops the
garbage threshold so volumes with under 10% deleted bytes still compact.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tolerate straggler fills in remote cache eviction test

Detached fills from the concurrent wave keep racing the final checks:
live chunks legitimately fill both volumes, and a re-cached object can
be evicted again before its commit is observed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: start remote cache eviction loop after filer init

The janitor's first tick dereferences fs.filer; starting the goroutine
before NewFiler assigns it could panic when startup exceeds an interval.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep remote cache vacuum intent across retries

Evicted entries now record their chunk volumes for vacuum directly, so
the intent survives whoever consumes the shared deletion queue first.
A pending volume keeps several vacuum attempts so tombstones that land
late are still compacted, and the janitor retries pending volumes under
the reclaim mutex instead of flushing unrelated deletes every tick.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound each remote cache vacuum request independently

A shared 30s deadline across pending volumes let one slow compaction
cancel the rest. Each VacuumVolume now gets its own context, and pending
volumes keep more attempts since the master reports request acceptance
rather than compaction.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tighten remote cache reclamation bound

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: treat chunk timestamp changes as content changes

chunksEqual now also compares ModifiedTsNs so an update that rewrites a
chunk record still invalidates the remote sync stamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: run remote cache queue flush under the reclaim context

BatchDelete for flushed file ids now uses the caller's context instead of
context.Background(), so a reclaim pass bounded by shutdown or timeout
stops its deletes too. Other callers keep their existing behavior.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote cache vacuum to evicted volumes

The flush no longer feeds the shared deletion queue's ids into the
pending set — only evicted chunks' volumes are tracked, so ordinary
deletions no longer pick up repeated vacuum attempts. The flush also
runs under a shutdown-immune bounded context and is skipped when no
volume is pending.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retry remote cache vacuums even after unmount

Pending volumes were only retried while a remote mount existed; removing
the last mount skipped every later pass and left evicted bytes allocated.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: reclaim partial cache fills that run out of capacity

A fill that fails midway queues its written chunks for deletion, but
when no entries remain evictable the reclaim pass found no pending
volumes and skipped the flush and vacuum entirely, leaving the partial
garbage to the slow periodic vacuum while the disk stayed full. Mark
the failed fill's chunk volumes pending so the pass tombstones and
compacts them even when nothing was evicted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 22:05:42 +08:00
Sascha KühndelandGitHub a033ab76e5 helm: add base dns names to admin cert (#11518) 2026-09-29 21:05:24 +08:00
Ishwar KanseandGitHub 5fcf12c080 docker: publish release and latest images for linux/ppc64le and linux/s390x (#11517)
Add ppc64le and s390x to the normal and large_disk release image
matrices, next to arm/v7 and 386, and to the manual container_latest
fallback so that rebuilding `latest` does not drop them from the
manifest.

The Go binary is cross-compiled without cgo on the build host, and
Dockerfile.go_build already falls back to empty Rust volume server and
worker placeholders on platforms without pre-built Rust binaries. Only
the small final-stage RUN steps run emulated, so these two platforms use
QEMU on ubuntu-latest like arm/v7.

Refs #5558
2026-09-29 21:04:19 +08:00
Khris RichardsonGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Opus 5.5Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
5da137233d s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart (#11510)
* s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart

The S3 server's IAM config loader never read the documented
`oidcProviderStore` key, so the OIDC provider store was always in memory:
a provider created with CreateOpenIDConnectProvider lived in one gateway's
process, was lost on restart, and was never seen by peers. The
/etc/iam/oidc-providers metadata subscription refreshed from that empty
in-memory store.

- Read `oidcProviderStore` and pass it to the IAM manager. With an IAM
  config file the default stays memory. With no config file (zero-config
  IAM, as `weed filer -s3` and operator-managed clusters run) it defaults
  to the filer: there is nothing static to shadow, and providers created at
  runtime otherwise vanish on restart.
- With a store that outlives the process, load the STS runtime view from it
  at startup, so providers created on an earlier boot or on a peer are
  trusted without waiting for the next mutation.
- If the store cannot be read at startup (a filer not up yet), the load is
  retried in the background with backoff until it succeeds: the metadata
  subscription reports only later changes, so providers already stored would
  otherwise stay unknown to STS until one of them changed.
- Mark records mirrored from STS.Providers as `source: static-config`, and
  at startup delete such records whose provider has left the config, so
  removing a provider from the config file still revokes it. Records created
  through the IAM API are never pruned.
- The filer store reported every failed lookup, an unreachable filer
  included, as ErrOIDCProviderNotFound, which CreateOIDCProvider reads as
  "free to create". Only a confirmed absence is now not-found.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: keep config-file OIDC providers out of a persistent store

Review of the previous commit found that mirroring the IAM config file's
providers into a persistent store, and pruning them when they leave the
file, breaks as soon as S3 servers share a filer:

- a server prunes stored config-file providers its own file does not list,
  including ones a peer's file still defines (a zero-config server prunes
  them all);
- mirroring overwrites an API-created provider with the same ARN and marks
  it config-owned, so a later prune deletes it;
- a failed mirror write or a failed prune leaves a stale record trusted;
- a mirrored record is loaded into STS at startup as an IAM-managed provider
  and shadows the config-file provider, dropping the settings a record does
  not carry (jwksUri, roleMapping, policyClaim, ...).

A persistent store now never receives the config file's providers. STS keeps
serving them from its static configuration, as it always has; the IAM API
lists and returns them from memory, refuses to change or delete them
(UnmodifiableEntity; change them in the file) and to create another provider
with their ARN (EntityAlreadyExists). The store holds only providers created
through the IAM API, and those are what startup loads into STS. There is
nothing to prune, so the source marker is gone. An in-memory store keeps its
behaviour: the config file's providers are records in it, as before.

buildOIDCProviderFromRecord also carries PolicyClaim and
AllowedPrincipalTagKeys now; they were dropped whenever an API-created
provider was loaded into STS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: send UnmodifiableEntity as a 400, not an internal error

The IAM API's error writer had no case for UnmodifiableEntity, which the
previous commit returns for a change to a config-file provider, so it went
out as a 500 ServiceFailure that clients retry. AWS sends it as a 400.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: document stored-over-config precedence, drop invented CreateDate, cancel superseded retries

Follow-ups from review of b881982d2:

- A provider stored under the same ARN as a config-file provider takes
  precedence in the IAM API, matching STS, which already prefers
  IAM-managed providers so that an API call can shadow a bootstrap entry.
  Deleting the stored provider brings the config-file one back. This was
  already the behaviour; it is now documented and tested.
- A config-file provider no longer reports its server's start time as
  CreateDate, which changed on every restart; GetOpenIDConnectProvider now
  omits the date for it. An in-memory store still stamps its copies at load,
  as before.
- The startup retry runs under a cancellable context, is cancelled when
  another store is installed, and retries the store it was started for
  rather than reading the manager's field, so replacing the store neither
  leaves the old retry running nor races with it (go test -race).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: serialize OIDC provider refreshes so an older snapshot cannot restore a deleted provider

Refreshes run concurrently: after an IAM API change, on a peer's change
and in the startup retry. Each lists the store and then hands STS the
result, so a refresh that listed before a DeleteOIDCProvider could finish
after that call's own refresh and keep the deleted provider trusted until
the next change. Refreshes now hold a lock from the read to the hand-off,
and a startup retry cancelled by installing another store drops its
snapshot instead of applying it.

The retry-cancellation test waits for the retry by polling instead of a
fixed sleep.

* s3/iam: route SetOIDCProviderStore through installOIDCProviderStore

A store installed after Initialize skipped the static-provider overlay
and startup hydration: config-file providers disappeared from the IAM
API, ErrOIDCProviderStatic no longer protected them, and stored
providers were never trusted until the next mutation or peer event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:51:39 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
62d4f9152a iam: evaluate trust policies deny-by-default (#11513)
* iam: evaluate trust policies deny-by-default

EvaluateTrustPolicy seeded its result with the engine's DefaultEffect,
so a non-matching trust-policy statement set still resolved to Allow
when the IAM config sets policy.defaultEffect=Allow. A caller holding
a validly signed token from a registered provider could then assume a
role its trust policy does not admit.

Trust policies now start from implicit deny, matching AWS semantics and
the pre-d751623 behavior of evaluateTrustPolicy; DefaultEffect still
governs identity-policy evaluation.

Upgrade note: deployments on defaultEffect=Allow whose trust policies do
not match their callers will see those assumptions refused.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: cover trust policy implicit deny under DefaultEffect=Allow

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:36:01 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a901c1a5e2 filer: compare IF_ENTRY_EQUAL chunks by fid, not file_id (#11514)
The stored entry came through FindEntry, which restores chunk file ids
from their fid form, while an expected entry built from a metadata-log
event still carries the serialized form (file_id moved into fid). The
proto.Equal saw file_id "" against the restored id and refused every
stamp, so remote.sync re-uploaded each entry and the RemoteEntry stamp
never landed.

Clone both sides and run BeforeEntrySerialization before comparing, so
chunks match on their fid and the file_id spelling is ignored; the stored
entry and the request's ExpectedEntry are left untouched.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:34:07 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
67b19b2db5 build(deps): bump com.fasterxml.jackson.core:jackson-databind from 2.22.1 to 2.22.2 in /test/java/spark (#11511)
build(deps): bump com.fasterxml.jackson.core:jackson-databind

Bumps [com.fasterxml.jackson.core:jackson-databind](https://github.com/FasterXML/jackson-databind) from 2.22.1 to 2.22.2.
- [Commits](https://github.com/FasterXML/jackson-databind/compare/jackson-databind-2.22.1...jackson-databind-2.22.2)

---
updated-dependencies:
- dependency-name: com.fasterxml.jackson.core:jackson-databind
  dependency-version: 2.22.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:12:55 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c54f1fe69c build(deps): bump github/codeql-action from 4.38.1 to 4.38.2 (#11509)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.1 to 4.38.2.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.38.1...v4.38.2)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:10:07 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
818ff69270 build(deps): bump cloud.google.com/go/kms from 1.34.0 to 1.35.0 (#11508)
---
updated-dependencies:
- dependency-name: cloud.google.com/go/kms
  dependency-version: 1.35.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:55 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
db89ed198b build(deps): bump github.com/twmb/avro from 1.8.0 to 1.9.0 (#11507)
Bumps [github.com/twmb/avro](https://github.com/twmb/avro) from 1.8.0 to 1.9.0.
- [Commits](https://github.com/twmb/avro/compare/v1.8.0...v1.9.0)

---
updated-dependencies:
- dependency-name: github.com/twmb/avro
  dependency-version: 1.9.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:39 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f7b5ca826b build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity from 1.14.0 to 1.14.1 (#11506)
build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity

Bumps [github.com/Azure/azure-sdk-for-go/sdk/azidentity](https://github.com/Azure/azure-sdk-for-go) from 1.14.0 to 1.14.1.
- [Release notes](https://github.com/Azure/azure-sdk-for-go/releases)
- [Commits](https://github.com/Azure/azure-sdk-for-go/compare/sdk/azcore/v1.14.0...sdk/azidentity/v1.14.1)

---
updated-dependencies:
- dependency-name: github.com/Azure/azure-sdk-for-go/sdk/azidentity
  dependency-version: 1.14.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:24 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
878d23b95a build(deps): bump github.com/dustin/go-humanize from 1.0.1 to 1.1.0 (#11505)
Bumps [github.com/dustin/go-humanize](https://github.com/dustin/go-humanize) from 1.0.1 to 1.1.0.
- [Commits](https://github.com/dustin/go-humanize/compare/v1.0.1...v1.1.0)

---
updated-dependencies:
- dependency-name: github.com/dustin/go-humanize
  dependency-version: 1.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:01 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1e0b243aed build(deps): bump golang.org/x/oauth2 from 0.36.0 to 0.37.0 (#11504)
Bumps [golang.org/x/oauth2](https://github.com/golang/oauth2) from 0.36.0 to 0.37.0.
- [Commits](https://github.com/golang/oauth2/compare/v0.36.0...v0.37.0)

---
updated-dependencies:
- dependency-name: golang.org/x/oauth2
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:08:43 +08:00
github-actions[bot] 79e2abdfbf docs: regenerate star history chart 2026-09-29 00:50:14 +00:00
github-actions[bot] 530be3e373 4.48 2026-09-28 15:53:40 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
9b3b12c607 filer: pin-aware reader cache eviction and stream release (#11503)
* filer: synchronize stream pins and release them on transitions

Guard chunkStream.cacher with the ReaderCache lock everywhere: mount
sections share one ChunkReadAt across concurrent reads, and unsynchronized
release could double-unpin. Reads served from the chunk cache now detach
the stream's pin instead of retaining the previous chunk. Eviction prefers
unpinned downloaders so a pinned buffer is not dropped mid-stream. A new
ReleaseStream lets callers drop their pin without destroying the shared
cache; S3 and WebDAV readers use it. lastChunkFid becomes atomic since
concurrent mount reads can update it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep eviction bounded when every downloader is pinned

Both eviction paths still fall back to a pinned victim when no unpinned
one exists, so abandoned stream pins cannot bypass the downloader limit
or stall the memory budget. Budget eviction also rechecks the pin under
the ReaderCache lock at removal time: a stream that pinned the selected
victim in between keeps it mapped and the selection retries.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: restore budget bookkeeping when a victim gets pinned mid-eviction

removeUnpinned losing the pin race left the victim out of the idle list
while still holding its reservation, making it unevictable even as the
pinned fallback. Push it back when the reservation is still live.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 22:37:04 +08:00
Tobias GurtzickandGitHub 3abdef3202 filer: keep shared chunk buffers pinned while another stream reads them (#11502)
The ReaderCache is shared by all streams of a process (every S3 GET, for
instance), but a ChunkReadAt released chunks as if it owned them:

- moving on to the next chunk called UnCache on the previous one,
  destroying the buffer even when other streams were still inside it;
- since #11384 a buffer is dropped once any reader has consumed it to the
  end and no read call is in flight. Streams copy out in slices (256 KiB
  in the S3 gateway), so between two calls a slower stream is not
  attached and loses the buffer to a faster one.

Either way the slower stream refetches the whole chunk from the volume
servers. With many clients downloading the same popular object at once,
each chunk is fetched over and over; in production we saw the S3 gateway
pull ~10 Gbit/s from volume servers while serving ~1 Gbit/s to clients.

A ChunkReadAt now pins the chunk it is positioned in. The pin is taken
and released only under the ReaderCache lock, since concurrent ReadAt
calls on one ChunkReadAt (as in mount) share it. It is released when the
stream reads the chunk to its end, moves to another chunk (including one
served from the chunk cache), or falls back to random reads. A buffer is
dropped once no stream pins it and no read is in progress, if it was
consumed or its last stream left it; a read still in flight when the
stream leaves drops it on detach, as UnCache did via destroy. Eviction by
slot limit and memory budget is unchanged.

lastChunkFid is now guarded as well: concurrent ReadAt calls raced on it.

Tests: two ChunkReadAt instances streaming one object in interleaved
slices fetch each chunk exactly once (2-3 times before); leaving a chunk
for a chunk-cache hit or while another read is in flight releases it;
concurrent ReadAt calls on one ChunkReadAt leave no pins behind under
-race.
2026-09-28 21:56:04 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
43fd5b8d82 volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete

The 2PC generation switch stages each run as <base>.ecNN.v<N> plus
versioned .ecx/.ecj/.vif files. Nothing on the volume server removes
them: isEcDataShardFile only recognises the exact .ecNN name, so the
staged files are invisible to every bookkeeping pass, and even
full_teardown's wipe-all path left them behind. Each re-encode therefore
leaks a full shard set per shard-holding disk.

RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>,
optionally keeping generations at or above a threshold; teardown and the
reconcile wipe remove every generation, and a per-shard delete removes
that shard's staged generations too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: delete staged EC generations older than N via VolumeEcShardsDelete

After a 2PC generation switch commits, the superseded generation's
<base>.*.v<N> files sit on disk with no cleanup path: teardown removes
everything, and a per-shard delete only touches the named shards, so the
executor had no RPC that reclaims just the staged leftovers.

delete_generations_older_than removes staged generation files strictly
below the threshold on every disk. Versioned files are never mounted, so
nothing is unloaded first; the committed generation and the canonical
files are preserved.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: mirror staged EC generation cleanup

Parity with the Go volume server: remove_ec_generation_files sweeps
<base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by
remove_ec_volume_files (which covers both teardown paths) and the new
delete_generations_older_than request field; delete_ec_shards removes a
shard's staged generations along with the canonical file.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: match staged generation filenames literally

filepath.Glob interprets metacharacters in the collection part of the
base name, so a collection like a[bc] could match another volume's
staged files (or miss its own). Scan the directory and compare names
literally instead, mirroring the Rust read_dir implementation.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: report generation-sweep errors and drop the store lock first

- snapshot the location base names under the read lock and run the
  filesystem sweep after dropping it, so a slow disk cannot stall the
  store;
- record per-entry read_dir errors in remove_ec_generation_files and
  propagate them from remove_ec_shard_generations instead of flatten()
  skipping them;
- warn when a staged-shard generation fails to delete rather than
  reporting success with files left behind.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: fail shard delete when the staged-generation listing fails

A transient ReadDir failure fell back to removing canonical shard names
only: staged .v<N> files survived while the RPC still reported success,
leaving the leak invisible to retrying callers. ENOENT still means the
disk simply has no such directory; other listing errors now propagate.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: propagate staged-generation removal failures

delete_ec_shards logged remove_ec_shard_generations errors and the RPC
returned success while staged .v<N> files remained, diverging from the
Go handler which surfaces the failure. The sweep keeps processing the
remaining shards, retains the first error, and volume_ec_shards_delete
maps it to Status::internal so callers can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: notify state change even when the shard sweep errors

delete_ec_shards already deletes and unmounts the shards before
returning a staged-generation failure, so returning early skipped
volume_state_notify and the master kept routing to them until the next
heartbeat. Notify before propagating the error.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:55:25 +08:00
150a69fe11 master: make volume capacity reservation timeout configurable (#11426) (#11497)
* master: make volume capacity reservation timeout configurable (#11426)

* master: expire reservations on reads, fix int timeout units

- AvailableSpaceForReservation now expires reservations too: a node that
  is full of reservations is filtered out before TryReserveCapacity can
  clean them, which stranded expired capacity indefinitely.
- Drop TryReserveCapacityWithTimeout: a per-call timeout lets one caller
  expire another's live reservations, and the Node interface stays
  stable for implementations outside this tree.
- parseReservationTimeout no longer routes integer values through
  GetDuration, which read them as nanoseconds; bare numbers are
  seconds. The 5m fallback is now the shared DefaultReservationTimeout.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 18:40:11 +08:00
4fec65d949 filer: demote client-cancelled directory listing log from error (#11495) (#11496)
* filer: demote client-cancelled directory listing log from error (#11495)

* filer: quote path in canceled listing log

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 14:35:41 +08:00
github-actions[bot] f564918685 docs: regenerate star history chart 2026-09-28 00:52:53 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f9289f0570 s3: do not promote ?prefix into the object for non-List actions (#11494)
* s3: do not promote ?prefix into the object for non-List actions

authRequestWithAuthType mapped an empty object to the prefix parameter for
every action, so PUT /bucket?versioning&prefix=x authorized as Write:bucket/x.
An object-scoped grant (Write:bucket/*) could then change bucket versioning,
lifecycle, cors, and object-lock configuration, and the promoted object also
made ResolveS3Action report s3:PutObject to attached IAM policies.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object

resolveS3AuthTarget fed the promoted prefix to ResolveS3Action, so a
bucket-level ?uploads request resolved as s3:GetObject on the prefix ARN
in the admin explicit-deny check. Resolve both action and resource
against the object the bucket listing actually scopes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object in AuthorizeAction

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: drop the unreachable object-level uploads case from the resolver test

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 07:17:58 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
4303b3aa4c s3: keep a listing's start position inside the requested prefix (#11493)
* s3: a list marker that sorts past the prefix leaves nothing to list

AWS scopes a listing to keys under Prefix; StartAfter, Marker and
continuation tokens only reposition inside that range. A marker that
diverges from the prefix at a larger byte is after every key the prefix
can match, so the page is empty. normalizePrefixMarker used to keep such
a marker as the walk cutoff at the bucket root, where the walk descends
into the marker's own directory and returns keys the prefix never names.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the listing variant's action when a prefix is promoted to object

authRequestWithAuthType promotes ?prefix= into the object argument for
the legacy CanDo path. ResolveS3Action treats a non-empty object as
object-level, so a bucket-level ?versions or ?uploads request carrying a
prefix missed its specific action and fell back to the base List action:
an s3:ListBucket grant then covered s3:ListBucketVersions, and an
explicit Deny on the specific action was skipped on the same path.

Resolve the action against the same bucket-level object the resource
ARN already uses.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Update weed/s3api/auth_credentials.go

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-28 06:58:44 +08:00
a0ee7ba314 s3: ignore empty intermediate directories in bucketHasUserObjects (#11490) (#11491)
* s3: ignore empty intermediate directories in bucketHasUserObjects (#11490)

* s3: keep nested reserved-named dirs from hiding user objects

Reserved folders (.uploads, *.versions) are internal only at the bucket
root; deeper entries with those names are user key prefixes and must be
walked. Also treat a missing subdirectory as empty via isFilerNotFound
(list errors cross gRPC as status errors, not the sentinel), let names
containing backslashes count as objects, and walk iteratively so empty
chains deeper than the old scan depth no longer report non-empty.

* s3: treat reserved-named directories as internal at every level

Object listing interprets .uploads and *.versions directories as
internal storage wherever they appear, so walking them during the
emptiness check would report invisible version remnants as user objects
and block deletion. A reserved name on a file still counts, matching
listing which only special-cases directories.

* s3: count explicit directory objects under reserved names

A directory object created by PutObject (MIME or prefix-object marker
set) is user data even when named .uploads or *.versions; only a plain
directory with a reserved name is internal storage.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 06:56:46 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a976b21010 s3: require dedicated object-lock permissions for x-amz-object-lock-* headers (#11492)
* s3: require dedicated object-lock permissions for x-amz-object-lock-* headers

PutObject, CreateMultipartUpload, and PostPolicy honor the retention and
legal-hold headers after only the route's s3:PutObject check, so a
write-only principal could pin a version under COMPLIANCE retention that
nobody can remove before its retain-until date. On AWS these headers
require s3:PutObjectRetention / s3:PutObjectLegalHold. validateObjectLockHeaders
is the shared funnel for all four call sites; it now authorizes the
corresponding dedicated action when each header is present.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: record the verified POST-policy signer as the request identity

The handler authenticated the form policy signature but stored only the
signer's name, so downstream authorization (the object-lock header check)
re-authenticated the form-signed request as anonymous and evaluated the
wrong principal.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 20:17:12 +08:00
e57f8c4d87 volume server: read GET/HEAD needles off the store lock, and only once (#11487)
The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-27 20:13:15 +08:00
67691a1eea volume server: split volume_copy into phases and type the delete-after-status gate (#11485)
volume_copy was one ~400-line handler, and the rule that an existing
local replica is deleted only after the source's ReadVolumeFileStatus
succeeded was held by statement order alone. The keep_remote_data=true
that the pre-copy delete and the failed-copy rollback must share was
kept in sync by a comment pointing from one to the other.

The handler is now a ~60-line orchestrator over connect_to_copy_source,
SourceVolumeStatus::fetch, delete_existing_replica,
plan_copy_destination and a VolumeCopyJob whose run() drives
preallocate_dat, transfer_files, finish_copied_files and
mount_and_reply, with cleanup_failed_copy on error.
delete_existing_replica takes a &SourceVolumeStatus, which only fetch
can construct (private field in a child module), so the delete cannot
be called before the status RPC. Both deletes go through
delete_replica_keep_remote.

Pure refactor: call order, status codes and messages, cancellation
checks, throttling, progress reports and cleanup are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-27 20:12:55 +08:00
hsdfatandGitHub 02353444ac fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts (#11483)
* fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts

Port of the Go VolumeCopy hardening in #11238 and #11252.

- Pick the destination disk before deleting the existing replica, counting
  the slot that replica holds as free. If no disk qualifies, the healthy
  replica is kept instead of being deleted.
- Read the source's VolumeStatus before and after the copy. When both
  succeed and the counts did not change, the mounted replica's file and
  deleted counts must match; on mismatch it is unmounted and its files
  removed. A failed "before" read skips the check; a failed "after" read
  fails the copy.

* fix(volume-rust): let a departing caller cancel VolumeCopy's post-copy status read

Go reads the source's status after the copy with stream.Context(), so the
call ends when the caller leaves. The Rust call had no such link: a
source that stalled there held the copied, unmounted files after the
caller was gone. Race it against the response channel, like the other
blocking steps, so the usual error cleanup removes the partial copy.
2026-09-27 19:40:07 +08:00
68944e83a3 volume: typed tier errors so a missing remote object answers NotFound (#11484)
remote_storage/s3_tier.rs returned Result<_, String> from every
transfer (upload_file, download_file, read_range[_blocking],
delete_file[_blocking]) and from the tier runtime helpers. The tier
move handlers could only wrap that in Status::internal, so a .dat whose
remote object is gone was indistinguishable from an I/O failure to
weed shell.

Add TierError { NotFound, Io, RuntimeUnavailable, Aborted }. Each
variant carries the existing message verbatim. NotFound follows the
rules remote_storage/s3.rs already uses: raw 404 status on HEAD,
NoSuchKey code on GET; a bare 404 on GET stays Io. A progress-callback
Err becomes Aborted. VolumeError gains a transparent Tier variant and
From<VolumeError> for Status maps Tier(NotFound) to NotFound; the tier
move handlers go through status_with_context, so their message text is
unchanged. Every other tier failure is still Internal.

The remote needle read path keeps io::Error::other, so its error kind
and vacuum's handling of it do not change.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:39:25 +08:00
00310f6588 volume server: run the vacuum compaction copy without the store lock (#11482)
* volume server: run the vacuum compaction copy without the store lock

VacuumVolumeCompact held the store write lock for the whole live-needle
copy, including every progress blocking_send on the 16-deep stream. On a
large volume that is minutes with every read, write and heartbeat on the
node parked behind it, long enough for the master to unregister the node.

Split compaction the way Go's CompactByIndex runs it. A short locked step
claims the volume's compacting flag, records the makeup_diff watermark
(index size and compaction revision) and opens fresh .dat/.idx handles.
The copy then replays .idx up to the watermark and copies from those
handles with the store lock released; writes that land meanwhile are
replayed by makeup_diff at commit, as before. The flag is an Arc<AtomicBool>
released when the job is dropped, so every exit path clears it.

Because the flag is now visible to other callers, the operations that would
pull the files out from under the copy refuse while it is set: unmount
(and VolumeConfigure, which unmounts and remounts), delete (checked before
the volume is removed from the map, which a refused destroy used to leave
unmounted), cleanup, and index relocation. A second compact and a commit
stay no-ops, as in Go. The pre-copy fsync is dropped: the copy reads its
own handles through the page cache and .cpd/.cpx are fsynced before commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: keep a read-only in-memory index's size for the compaction copy

The unlocked copy replays .idx up to index_file_size(). A read-only
volume whose .sdx could not be built loads its index into memory without
a writer, so that size stayed 0: the copy came out empty and the commit
replaced the volume with it. CompactNeedleMap::load_from_idx now records
the rows it loaded, which is also what Go's IndexFileSize reports for a
read-only index.

The copy's index replay now stops reading at the recorded size instead
of walking rows appended since, which makeup_diff replays anyway.

Adds tests for compacting a read-only volume on both the sorted index
and the in-memory fallback, and for VolumeConfigure stopping when the
unmount is refused during a copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a vacuum copy as soon as its client is gone

The progress callback only noticed a closed response stream when a report
was due, every 128 MiB. With the copy now running outside the store lock,
a copy nobody waits for keeps the volume marked compacting and so keeps
refusing unmount, delete and cleanup until that next report. Check the
stream on every callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:59 +08:00
5c9c424a84 volume server: stream ReadAllNeedles without holding the store lock (#11481)
* volume server: stream ReadAllNeedles without holding the store lock

read_all_needles held store.read() while Volume::read_all_needles read
every live needle of the volume into a Vec, and kept holding it through
the whole blocking_send loop. Memory grew with the volume, and a slow
client parked the scan in a send with the guard held; needle writes and
the heartbeat take store.write() on a writer-preferring lock, so the node
stopped serving until the client caught up.

Take a DatScanPlan (fresh .dat open, end bound) under a short guard and
walk it with the guard released, sending one needle at a time. Each
record is checked against the live needle map under a brief read guard,
as the scan reaches it, and only a live record is parsed, so a damaged
stale copy does not fail the stream. Records appended while a pass ran
are walked by a follow-up plan, so a needle overwritten during the scan
is streamed once, as its new copy. A vacuum commit or re-create of the
volume during the scan fails the stream, since the map's offsets no
longer describe the pinned file; the plan carries the volume instance
and compaction revision for that check.

DatScanPlan::scan_records yields records unparsed; scan keeps its
behaviour on top of it. Volume::read_all_needles has no caller left and
is removed; its tests move to the RPC.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: check ReadAllNeedles liveness only once the send can proceed

The per-record liveness check ran before blocking_send, so a scan parked
on a full channel held a record it had already judged live. An overwrite
landing during that park left the old copy in the stream, and the
continuation over appended records then streamed the new copy as well.

Reserve channel space first, then take the store read guard, check the
record against the needle map and enqueue it through the permit before
releasing the guard. The wait for space still happens without the lock;
the record is parsed before the guard is taken, and its parse error only
counts if the record turns out to be live.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:43 +08:00
5218e68554 volume server: collect heartbeats under the store read lock, off the runtime (#11480)
* volume server: collect heartbeats under the store read lock, off the runtime

Every pulse tick, options change and volume-state notification took
store.write() for a whole heartbeat pass, directly on the async heartbeat
task. The pass fstats every volume's .dat twice and hashes its report, so
on a server with many volumes it held the store exclusively for the whole
scan: reads and writes stalled, and with the writer-preferring RwLock a
pending pass parked every new reader too.

The pass only needs to mutate the store for a few rare actions: removing
expired EC volumes, deleting expired volumes past their removal delay, and
setting no-write on IO-quarantined volumes. It now runs under store.read(),
records those as (disk, volume id) actions, and applies them afterwards
under a short store.write() that is only taken when there is something to
do. Each action re-checks its target under the write lock, so a volume
written to, replaced or removed in between is left alone. Expired EC
volumes are still removed before the volume pass, as before, because the
EC shard count feeds the disk-space-low max volume count. Every pass runs
on the blocking pool via spawn_blocking.

The heartbeat message is unchanged for the same store state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: take has_no_ec_shards with the heartbeat's volume list

The heartbeat pass took has_no_ec_shards from the EC phase's read lock,
then built the volume list under a second one. An EC shard mounted in
between went out as "no EC shards" beside a volume list taken after the
mount, and the master clears a server's EC registrations on that flag.

has_no_ec_shards is now computed under the same read lock as the volume
list, with the EC phase's filter: not expired, not quarantined, at least
one shard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:23 +08:00
github-actions[bot] 00cffa028c docs: regenerate star history chart 2026-09-27 00:53:30 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a261f90e18 vacuum: let the sweep release volumes that stay empty and quiet (#11477)
* vacuum: let the sweep release volumes that stay empty and quiet

Vacuuming reclaims bytes but not slots: a fully emptied volume stays
registered to its collection forever, and since growth is gated only on
slot count a store at 99% free disk can still refuse writes to other
collections (#11429). volume.deleteEmpty exists but is manual-only.

With -vacuumDeleteEmptyAfterSeconds (or master.vacuumDeleteEmptyAfterSeconds
under weed server/mini; default 0, off) the automatic sweep now deletes
replica copies that have stayed empty and quiet for that long, the same
rule volume.deleteEmpty applies on demand: remote-backed copies are
skipped, and every delete carries the volume server's onlyEmpty /
onlyGarbage guards so a copy written since the last report is refused
rather than removed. Copies that still hold data or were written
recently stay; only a volume whose every copy is deleted leaves the
sweep's work map, sparing a compaction of bytes that are all deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: harden empty-volume sweep against partial and racing deletes

Review follow-up on #11477:

- delete a volume only when every replica copy is a verifiable
  empty-and-quiet candidate; deleting the empty copy of a volume whose
  sibling holds live files would silently cut its replica count
  (greptile P1).
- drain the volume out of the writable list before deleting, the same
  drain the compact pass uses, so PickForWrite stops assigning it and
  pending writes settle (devin).
- bound the VolumeDelete RPC so one stalled server cannot hold the
  vacuum lock indefinitely (greptile P1, reusing allocateVolumeTimeout).

The vid2location panic scenario raised in review does not exist:
VolumeLocationList methods are nil-receiver safe and a missing vid just
fails enoughCopies, so a partially deleted volume skips compaction
instead of crashing the sweep.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: unregister deleted empty replicas and prune the sweep list

A successful VolumeDelete only updates the volume server; the master
still tracked the replica and kept it in the sweep's location list for
the compaction pass (coderabbit on #11477). Unregister the replica right
after its delete succeeds and drop it from the sweep copy, so a partially
deleted volume only compacts copies that still exist.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: pin deleting volumes out of the writable list across heartbeats

Review follow-up on #11477 (greptile): DrainAndRemoveFromWritable only
removed the volume once; a heartbeat landing between the drain and the
replica deletes re-evaluated writability and re-added it, so a client
write could reach a replica whose siblings were already gone and leave
the volume under-replicated when the last copy refused its onlyEmpty
delete.

MarkDeleting records the vid in deletingVolumes — checked inside
setVolumeWritable so heartbeat, capacity-recovery, and admin re-add
paths all hold it out — and UnmarkDeleting releases it once the sweep
finishes the copy pass. A partially deleted volume's surviving replicas
then return to writable through the normal heartbeat path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: restore writability when a sweep delete survives

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:05:22 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
be29f44d87 s3: record requester identity before the authz verdict (#11479)
* s3: record requester identity before the authz verdict for audit

Identity was only stored in request context on the success branch, so
denied requests reached WriteErrorResponse without requester attribution
and audit entries had empty requester/requester_arn/requester_identity.
Authentication failures still resolve no identity, so unauthenticated
denials stay unattributed.

Fixes #11474

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the resolved identity through authz denial in Auth

Review follow-up on #11479 (devin): authRequest discarded the identity
on every error, so a request that authenticated fine but failed the
action check still reached handleAuthResult with no identity and the
deny path could not audit a requester. Auth now calls
authRequestWithAuthType directly, the same entry AuthPostPolicy uses,
so the resolved identity reaches the error writer; a failed authN
still resolves no identity and stays unattributed. The regression test
now signs a denied request end to end through iam.Auth.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:03:01 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2864bc0fe8 s3: honor configured session bounds on AssumeRole and LDAP identity (#11478)
* sts: export CalculateSessionDuration

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: honor configured session bounds on AssumeRole and LDAP identity

prepareSTSCredentials hardcoded a one-hour session when the caller
omitted DurationSeconds, so sts.tokenDuration was ignored and
sts.maxSessionLength only clamped explicit requests: asking for 3600s
against a 20m ceiling was rejected while omitting the parameter was
granted a full hour (#11473). The two affected handlers now use the
same default-then-cap calculation as AssumeRoleWithWebIdentity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through role store copies

copyRoleDefinition rebuilt RoleDefinition field by field and dropped
MaxSessionDuration, so memory-backed role stores silently discarded the
per-role session bound on every write and read (devin on #11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: apply per-role MaxSessionDuration to resolved session durations

Review follow-up on #11478 (devin): the role bound only ever applied to
explicit DurationSeconds values — an omitted duration resolved to the
configured default and sailed past a shorter role max on every assume
path.

- capDurationByRole now resolves min(requested||tokenDuration, roleMax),
  so AssumeRoleWithWebIdentity and AssumeRoleWithCredentials cap
  defaults the same way they cap explicit values.
- prepareSTSCredentials caps the calculated duration at the named
  role's MaxSessionDuration, covering the AssumeRole and LDAP handlers;
  self-assumption has no role definition to consult.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through the cached role store

genericCopyRoleDefinition drops MaxSessionDuration the same way
copyRoleDefinition did, so the cached filer role store reads back a zero
maximum and every downstream duration cap is skipped (greptile on
#11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: only materialize defaults that pass session duration validation

Review follow-up on #11478 (greptile): materializing an omitted
DurationSeconds into an explicit value could exceed the service's own
input bound (a configured tokenDuration above maxSessionLength) and turn
a previously working request into a validation error.

capDurationByRole now leaves nil anything the service can resolve
better itself, clamps a tightened default at maxSessionLengthSeconds,
and floors a role bound below 900s to the tightest issuable value.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:01:51 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
ab95d58b7c s3: keep dedicated object-lock actions pinned during action resolution (#11475)
* s3: keep dedicated object-lock actions pinned during action resolution

A coarse action that already names a dedicated operation (governance
bypass, retention, legal hold, bucket object-lock config) now resolves to
itself before request shape is consulted. Previously a synthetic
DELETE ?versionId authorization request re-resolved to
s3:DeleteObjectVersion, so the bypass check was satisfied by the
delete-version grant alone; with the pin it evaluates
s3:BypassGovernanceRetention as intended.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: cover pinned object-lock actions against competing query params

Locks in the resolution for every dedicated action in the pin set, incl.
the retention and legal-hold shapes carrying versionId.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 21:42:34 +08:00
2f641a63d6 filer: honor is_moved only from ring member connections (#11456)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks

updateLocalEntry records the RemoteEntry stamp after an upload by writing the
event's entry back with UpdateEntry. The filer deletes every stored chunk
absent from an updated entry, so when the file was rewritten while its upload
was in flight (or the event is a replay), the stale snapshot deletes the
rewrite's chunks: the entry then points at the new fid with no needle behind
it, and the rewrite's own upload fails and is skipped as superseded.

The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's
chunk fids, evaluated by the filer under the path lock. A refused stamp means
the filer moved past this event; the superseding event follows in the log and
stamps the current entry, so the refusal is logged and skipped like a
superseded upload.

Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount,
filer.remote.sync; hold the remote (docker pause) so one upload stays in
flight, rewrite the file through the filer, unpause. Before: the entry's chunk
is 404 on every volume server. After: the stale stamp is refused, the rewrite's
chunk stays live and reads back after a vacuum.

* filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored

The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a
rewrite that touched inline content or metadata alone still compared
equal and the stale snapshot overwrote the live entry. The new clause
compares the whole stored entry against the event's entry under the
same path lock.

* filer: route conditional UpdateEntry to the entry's owner filer

Two filers locking the same path locally could still pass a stale
condition on the non-owner while the owner's entry had moved on. When a
condition or expected_extended precondition is set, forward the request
to the entry's owner the same way conditional CreateEntry does, with
is_moved bounding the hop.

* filer: compare IF_ENTRY_EQUAL against the normalized expected entry

FindEntry grows FileSize to the chunk extent, so a raw event entry with
FileSize still zero failed the condition on an unchanged file and the
stamp was skipped, letting a replay upload the object again.

* filer.remote.sync: classify refused stamps by gRPC status only

A FailedPrecondition substring in an unrelated error would have been
swallowed as a skipped stamp; status.FromError already unwraps.

* remote sync: keep the event entry intact for IF_ENTRY_EQUAL

* filer: honor is_moved only from ring member connections

is_moved is caller-controlled, so a request could set it to skip owner
routing and run a conditional check under a non-owner's lock. Verify the
marker against the peer's connection address and the lock ring members;
an unverified marker is ignored and the request routes like a fresh one.

* filer: refuse unverifiable is_moved at a non-owner, cache ring IPs

Follow-up fixes from review on the is_moved provenance check:

- checkMovedMarker replaces "ignore and re-forward" for markers that did
  not arrive on a ring member's connection. Re-forwarding a claimed hop
  could cycle while rings disagree; instead the request is refused with
  FailedPrecondition unless this filer is the key's owner, in which case
  applying locally is correct anyway.
- ringMemberIPs caches resolved member addresses per ring membership so
  hostname-advertising deployments do not pay a DNS lookup per forwarded
  request; failed lookups are not cached so a DNS blip self-heals.
- DistributedUnlock no longer dereferences the nil response of a failed
  next-hop RPC.

* filer: refuse unverifiable is_moved with PermissionDenied, not FailedPrecondition

A routing refusal is different in kind from a write-condition mismatch:
remote sync treats FailedPrecondition as a stale stamp and skips it, so
reusing that code let a routing failure pass as synced. Owner checks now
also run before the peer-IP lookup so the common accept path does no DNS.

* filer: expire resolved ring member IPs after 5 minutes

A member's hostname can re-resolve to a new IP while its ring address
stays unchanged; caching forever would reject its genuine forwards until
a membership change or restart.

* filer: deduplicate concurrent ring member DNS lookups

At cache expiry, parallel forwarded requests would each resolve every
member hostname serially; singleflight collapses them into one lookup
per ring membership.

* filer: detach the shared ring lookup from the caller's context

The singleflight winner's ctx is cancelled when its request ends; the
shared result would then be an incomplete member list and genuine
forwards denied. The lookup now runs on a detached context with its
own deadline so a canceled caller cannot poison it.

* filer: resolve ring member hostnames in parallel

The shared lookup gave every member one serial budget, so a few slow
resolutions could leave later members out of the cached list and reject
their genuine forwards. Each member now resolves concurrently under its
own detached deadline.

* filer: gather literal member IPs before spawning lookups

A ring mixing IP literals and hostnames raced: the literal appends ran
unlocked alongside the resolver goroutines' locked appends. Split into
two passes so only hostname results share the mutex.

---------

Co-authored-by: jsas <1351492+jsas@users.noreply.github.com>
2026-09-26 19:42:23 +08:00
Mohd Quamar TyagiGitHubcoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>Chris Lucoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
80fd3635d2 volume: skip TTL last-write scan when it cannot fit its budget (#11472)
* volume: skip TTL last-write scan when it cannot fit its budget

* Update weed/storage/volume_checking.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-26 19:41:59 +08:00
Chris LuandGitHub 7129e1178e s3: evaluate bucket policy before ACL public-read for anonymous requests (#11471)
* s3: evaluate bucket policy before ACL public-read for anonymous requests

AuthWithPublicRead granted anonymous access on a public-read ACL before
consulting the bucket policy, so an explicit Deny (e.g. s3:ListBucket)
was skipped for anonymous callers while still enforced for authenticated
ones. Run the policy engine first: a matching Deny or Allow is honored,
otherwise fall through to the ACL grant as before.

* s3: defer object-level anonymous requests to the handler's policy recheck

Evaluating the bucket policy with a nil entry at middleware time makes
tag conditions like s3:ExistingObjectTag/<key> resolve against missing
values, so a conditional Deny could wrongly block anonymous Get/Head on
a public bucket whose handler recheck would permit it. Object requests
now take the ACL grant and let Get/HeadObjectHandler re-evaluate with
the fetched entry; only bucket-level requests (List, HeadBucket), which
have no such recheck, are decided by the middleware policy verdict.

Reading the bucket config first also refreshes the compiled policy on a
cache miss, so a remotely deleted policy cannot leave a stale verdict
in the engine for nonresident buckets.

* s3: recheck bucket policy before serving directory objects

handleDirectoryObjectRequest runs before the object handlers' policy
recheck, so directory content on a public-read bucket was served to
anonymous callers without any policy evaluation. Evaluate the policy
with the directory entry, matching the recheck the file path performs.
2026-09-26 17:59:57 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
3c1e8ca7a8 volume: never finish serving a needle whose data fails its CRC (#11467)
* storage: hold back last chunk until CRC verifies on whole-needle reads

Above PagedReadLimit the needle is streamed: headers and body go out
before the checksum is computed, so a corrupted needle was served as
200 with bad bytes and readers could not fall back to a replica. The
final chunk is now written only after the checksum verifies; on a
mismatch the response ends short of Content-Length and the client sees
a failed transfer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: verify needle checksum on streamed reads (parity)

Mirror the Go fix: carry the needle checksum in NeedleStreamInfo and
have StreamingBody accumulate the CRC and verify it before emitting the
last frame; a mismatch ends the body with an error so the client sees
the transfer fail rather than receiving corrupt bytes that look
complete.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* server: abort the transfer when a committed stream fails

A writeFn failure after any byte or WriteHeader call leaves the declared
status and Content-Length already sent; http.Error's text then joins the
body and can exactly fill the withheld tail of a corrupted needle read —
the client sees a complete 200 instead of a failed transfer to retry.

Track whether the response is committed (headers sent, or bytes buffered
for the deferred flush) and panic with http.ErrAbortHandler instead of
appending an error body; pre-commit failures keep the 500 path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* server: drop the response writer wrapper from the committed-response check

Counting buffered writes is enough: with no bytes buffered the status and
headers cannot have gone out, and the range branches commit via the
explicit WriteHeader call before writeFn runs. The extra ResponseWriter
wrapper added a new Write sink site that CodeQL flags.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:50:41 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
0f3ba98e11 volume: make volume.scrub report a live needle whose stored id is damaged (#11468)
* storage: scrub live needles' stored id against the index key

scrubVolumeData only compared the needle's stored id for tombstones, so
header damage on a live needle — where the data CRC cannot see it —
passed every scrub mode while reads of that needle kept failing or
serving the wrong key's data. Compare the id for every indexed needle.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: scrub live needles' stored id against the index key (parity)

Mirror the Go scrub fix: compare the stored needle id with the index
key for live needles too, not only for deleted ones.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: cover damaged live needle id in scrub test

The tombstone test proved the index-key check fires for deleted entries;
add the live-needle mirror of Go's TestScrubVolumeDataChecksLiveNeedleId
so a regression in the live path is caught in Rust too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:10:14 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
80a26020d7 util: serialize all ViperProxy access so startup cannot hit concurrent map read/write (#11470)
* util: serialize every ViperProxy method; stop promoting unlocked viper calls

ViperProxy embedded *viper.Viper, so only the five declared methods took
the mutex while every promoted call — GetStringMap in backend.LoadConfiguration
was the reported crash — touched viper's maps unsynchronized. `weed server`
starts the volume server (SetDefault writer) and the master (GetStringMap
reader) back to back, and a race build reports the pair on a plain start.

The wrapped viper is now a named field: a method must be declared here to
exist on the proxy, so unsynchronized access fails at compile time rather
than at runtime. Every promoted use in the tree (GetStringMap, GetUint32,
GetFloat64, GetDuration, IsSet, AllKeys, Set) gets a locked wrapper;
NewViperProxy replaces struct literals for local vipers. GetStringMap
deep-copies its result — viper hands back the internal subtree, so
iterating it after the lock is released would race the next writer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: take the shared lock while LoadConfiguration merges a config file

viper.MergeInConfig rewrites the same maps the proxy serializes; without
the lock a merge can race a concurrent SetDefault or reader exactly like
the reported startup crash.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: deep-copy slice elements in the GetStringMap snapshot

A slice of maps inside the returned subtree still shared the inner maps —
copy elements recursively so nothing the caller mutates is viper's
internal state.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: add the missing AutomaticEnv wrapper used by tests

sse_reader_test reaches it through GetViper(); without the wrapper the
call no longer exists once the viper field stopped being embedded.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: return a fresh slice from GetStringSlice

A stored []string comes back uncast from viper — the backing array is
shared internal state like the GetStringMap subtree, so copy it while
holding the lock.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:09:29 +08:00
yi111GitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
5389f61cef volume server: do not finish a GET when the needle CRC mismatches (#11464)
* volume server: do not finish a GET when the needle CRC mismatches

A streamed full-needle read compared the CRC only after every page had been written. Once the response buffer flushed, the client already had a completed 200 and the corrupt bytes. Hold the last page until the checksum matches, and if an earlier page has already been flushed, abort the connection instead of calling http.Error.

Fixes #11459

* volume server: abort partial-content bodies on write error too

The non-Range path drops the unflushed tail and aborts on a mid-body
error; the single-range and multi-range paths still flushed it after
WriteHeader(206) was committed, delivering corrupt bytes as a complete
body.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: assert the started 200 is aborted in the write-error test

The test previously returned on any request error, so it passed without
verifying the abort. It now asserts the client got the committed 200
headers and then a failed body read. Also trims comments.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:00:22 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
8ad2f29e3e shell: let volume.deleteEmpty drop volumes with no live needles (#11437)
* shell: let volume.deleteEmpty drop volumes with no live needles

The candidate check only accepted a .dat at superblock size, so a volume
whose every needle was deleted still had to be vacuumed first — minutes
of compaction to rewrite bytes that were all garbage anyway. FileCount
counts every indexed entry and DeleteCount every entry made garbage by
overwrite or delete, so FileCount <= DeleteCount means nothing live
remains and the volume can be unlinked directly. The quietFor guard is
unchanged.

* volume server: add only_garbage VolumeDelete guard

VolumeDelete(only_empty) refuses every volume that ever held data, so a
volume whose needles are all deleted could only be removed after a
vacuum rewrote it. The new only_garbage flag deletes only when the byte
counters show nothing live: DeletedSize covering all of ContentSize, the
same all-garbage state vacuum measures. Byte counters are used because
the file/delete counts drift on index reload.

* rust volume: mirror only_garbage VolumeDelete guard

Same check as the Go server: a volume deletes under only_garbage when
its deleted bytes cover all content bytes. The grpc handler rejects
before the store drops the volume from its map, since destroy errors
after removal would still unmount it.

* volume delete: let either enabled check pass, keep onlyEmpty on the wire

An upgraded shell sending only_garbage to a pre-upgrade server would be
read as an unconditional delete (field ignored, only_empty false). The
request now keeps only_empty set so old servers check emptiness and
refuse, while new servers delete when either check passes.

* volume.deleteEmpty: skip remote-backed and protected read-only volumes

A remote-tiered replica shares its cloud object with the other replicas,
so keepRemoteData=false on one delete removes data they still reference.
Protected read-only volumes are quarantined or under maintenance, which
is exactly when a replica should not be dropped.

* volume delete: validate guarded copies across disks before deleting

* volume delete: hold copy locks across guarded validate-and-delete

CheckVolumeDeletable released each copy's locks before Destroy ran, so a
write landing on a later copy between the two passes refused its destroy
after earlier copies were already removed. Pin every copy's
dataFileAccessLock (and its location's volumesLock) across validation and
removal so a refused delete leaves all copies intact.

* volume delete: send deleted-volume notices after releasing locks

A blocking send on a full DeletedVolumesChan under volumesLock can stall
the heartbeat loop that drains it while it waits on the same locks.
Collect the notices under the lock span and send after release.

* pb: restore generated-file cosmetics to match the repo's protoc version

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 15:53:39 +08:00
Ilia DemianenkoandGitHub c58bd0dfd3 s3: honor assignment fsync in UploadWithRetry (#11449)
* fix: honor assignment fsync in UploadWithRetry

* Tests feedback
2026-09-26 12:01:52 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f7680cf812 volume: one long-lived tokio runtime for blocking tiered S3 reads (#11448)
* volume: one long-lived tokio runtime for blocking tiered S3 reads

block_on_tier_future, behind read_range_blocking and delete_file_blocking,
spawned an OS thread and built a fresh current-thread tokio runtime on every
call, then tore the runtime down. On an S3-tiered volume that is once per
needle read, per streamed 64 KiB chunk, per DatScanPlan record and per
destroy. The SDK client's pooled HTTPS connections are driven by tasks on the
runtime a request ran on, so each teardown dropped the pool and every call
re-dialed and re-handshook TLS. A panic inside the SDK was also flattened to
the fixed string "tier runtime thread panicked".

Now one process-wide runtime (OnceLock, multi_thread, 2 workers named
tier-io) drives all tier I/O; block_on_tier_future spawns onto it and parks
the caller on an mpsc channel for the JoinHandle result. Blocking the caller
is unavoidable (the storage layer is synchronous) and is what the old code
did through thread::spawn().join(). Handle::block_on is not used because the
wrappers are also reached from inside another runtime's worker, where it
panics with "Cannot start a runtime from within a runtime". JoinError panics
are downcast to &str/String and the payload is kept in the error.

Tests cover runtime reuse (Handle::id equal across calls, thread name
tier-io), calls from a std thread, from spawn_blocking, and directly from
current-thread and multi-thread runtime contexts, and the panic payload.
Against the old body 7 of 9 fail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: return a tier runtime build failure instead of panicking

Review follow-up. tier_runtime() expect'ed the runtime build, so an OS
refusing threads panicked inside Volume::destroy (after the volume left
the in-memory map, before its files were removed) and inside needle
reads, bypassing their error paths. Keep the runtime in a
Mutex<Option<Runtime>> behind tier_handle() -> Result<Handle, String>:
a failed build is returned to the caller through block_on_tier_future's
existing Result and is not cached, so a later call retries once the
pressure is gone. The lock is held only while building.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the shared tier I/O runtime

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:50 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
4299fdf578 volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest (#11446)
* volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest

Go's VolumeEcShardsDelete full teardown calls vs.store.UnloadEcVolume in
the blanket path (weed/server/volume_grpc_erasure_coding.go:488) and
location.UnloadEcVolume in the generation-fenced path (:511): each disk
that had the volume registered drops it, closes its shard descriptors
and gives back its ec_shards gauge before the artifacts are unlinked.

The Rust handler used Store::remove_ec_volume / DiskLocation::remove_ec_volume
instead, which only remove the map entry. Store::remove_ec_volume also
stops at the FIRST disk holding the vid, so on a split-disk volume
(shards on several disks) the blanket teardown left the sibling disks'
EcVolume registered with open fds while the unlink loop deleted their
files underneath it: the heartbeat kept advertising shards whose files
were gone, the inodes stayed pinned by the open descriptors, and the
VOLUME_GAUGE{collection,"ec_shards"} never came back down. The fenced
path leaked the gauge and the descriptors the same way on the one disk
it wiped.

Both paths now use the unload_ec_volume helpers from #11413 (every disk
for the blanket teardown, the strictly-older disk for the fenced one),
and the two Status::internal messages name the disk directory like Go's
"... on %s: %w".

Regression tests build a two-disk store with the same vid mounted on
each disk (the SplitDiskEcFixture, which gains a collection knob so the
gauge read is isolated from parallel tests mounting under "") and assert
that a blanket teardown leaves no EcVolume registered on any disk and
returns the gauge to its pre-mount value, and that a fenced teardown
decrements the gauge for the older disk's shard while preserving the
newer disk. Both fail against the previous handler.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments on the EC full-teardown unload path

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:23 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
3e9fc9e75b volume: reject invalid -minFreeSpace and duration flags at startup like Go (#11447)
* volume: reject invalid -minFreeSpace and duration flags at startup like Go

Symptom: the Rust volume server swallowed invalid values for three
startup flags and ran with a limit the operator did not ask for.
-minFreeSpace=150 became 150 bytes (Go: fatal, percent must be 0..=100),
-minFreeSpace=10GiBx and -minFreeSpace=50B fell through to the 1%
default (Go: fatal), and -inflightUploadDataTimeout=30sec, =abc or =30
silently became the 60 s default (Go: flag.Duration usage error, exit 2).
Go's byte-size units (1,024MB, 42 mib, 100Ki, 1eb) were also not
understood and mapped to 1%.

Go reference: weed/util/minfreespace.go:73-106 (MustParseMinFreeSpace /
ParseMinFreeSpace), weed/util/bytes.go:212-243 (ParseBytes and
bytesSizeTable), weed/command/volume.go:226-227 (flag.Duration, i.e.
time.ParseDuration grammar).

Fix: the two inflight timeout Cli fields are now std::time::Duration
with a clap value_parser that ports time.ParseDuration (units ns/us/µs/
μs/ms/s/m/h, fractions, bare 0, int64-nanosecond overflow), so clap
prints a usage error naming the value and exits 2 at the same layer as
Go's flag package. parse_min_free_spaces returns Result and is built
from Go-parity parse_min_free_space / parse_bytes; resolve_config_with_env
aborts with Go's "The value specified in -minFreeSpace not a valid
value" message, the fail-fast pattern it already uses for --max, --index
and --readMode. No new dependency. Negative durations and NaN percent
are rejected explicitly (Go quirks with no sane Rust representation).

Tests cover the Go fixtures from minfreespace_test.go and bytes_test.go,
the duration grammar, and flag-level rejection through Cli::try_parse_from
and resolve_config_with_env; the flag-level tests were confirmed to fail
against the previous parser.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: say that -minFreeSpace commas separate directories, never thousands

Review follow-up. parse_bytes keeps Go's ParseBytes thousands-comma
support, but the flag is split on commas into per-directory entries
before any entry is parsed (Go's MustParseMinFreeSpace does the same),
so -minFreeSpace=1,024MB is the two entries 1 and 024MB. The doc
comment and the single-entry tests advertised the comma form as if it
were reachable from the flag. Document the split and add a flag-level
test: two entries from parse_min_free_spaces, and resolve_config
aborting on the directory-count mismatch with one -dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the strict flag parsing

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:06 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
afce0a3dd3 docs: restore the HTTP REST API reference (#11454)
* docs: restore the HTTP REST API reference

The REST API documentation was lost in the README refactor, and the wiki
only covers the master server API. Add a complete reference for the three
HTTP surfaces — filer paths, master file-id/topology endpoints, and
volume-server content endpoints — generated from the actual handlers and
query parameters.

* docs: correct REST API details flagged in review

Tagging uses Seaweed- headers not query params, the filer recursive
delete option changes the DELETE default, omitted resize mode does not
mean fit, default file mode is 0660, the master redirect is 308, the
listing flag is -dirListLimit, TUS is enabled by default at /.tus, and
-port.public opens the separate read-only listener.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: match the tagging delete example to canonical header names

Go canonicalizes Seaweed-k1 to Seaweed-K1 on write, and the delete list
is compared case-sensitively, so ?tagging=k1,k2 would not match.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:59:43 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>franchbChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
83d44be0f3 volume: detect S3 not-found by typed SDK errors, not the "service error" string (#11444)
* volume: detect S3 not-found by typed SDK errors, not the "service error" string

remote_storage/s3.rs decided ObjectNotFound by matching the Display
output of an aws_sdk_s3 SdkError against "NoSuchKey" / "404" /
"NotFound". In the locked SDK (aws-smithy-runtime-api 1.11.6,
src/client/result.rs:487-497) that Display is a fixed string per
variant, "service error" for every S3 error, so ObjectNotFound was
unreachable: every missing remote object surfaced as
Other("s3 get object: service error") with the real cause discarded.

Go (weed/remote_storage/s3/s3_storage_client.go) uses typed checks:
HEAD (373-374): awserr.RequestFailure with StatusCode() == 404;
GET (436-437): awserr.Error with Code() == s3.ErrCodeNoSuchKey.

read_file now matches SdkError::ServiceError whose GetObjectError
is_no_such_key(); a bare 404 on GET stays a generic error, as in Go.
stat_file matches HeadObjectError::is_not_found() or a raw HTTP 404
status, Go's actual condition. Non-service errors fall through to
Other unchanged. Every SdkError message in s3.rs and s3_tier.rs is
formatted with DisplayErrorContext so the S3 error code and message
survive instead of "service error".

Six network-free unit tests drive the client through a canned
HttpClient (404 NoSuchKey, bare 404 on GET and HEAD, 404 with a
foreign body on HEAD, 403 AccessDenied on GET and HEAD). They need
aws-smithy-runtime-api as a dev-dependency; it is already in the lock
at a single version, so no new crates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: HEAD not-found is the raw 404 status alone, as in Go

Review follow-up. The HEAD arm also accepted the SDK's NotFound error
code on any status, so a 400 carrying <Code>NotFound</Code> became a
missing object. Go's stat looks only at RequestFailure.StatusCode() ==
404 (weed/remote_storage/s3/s3_storage_client.go:373); do the same. The
raw status still covers the body-less 404 the SDK turns into NotFound
and a 404 whose body names a foreign code. Regression test for the
non-404 NotFound body, which failed against the previous arm.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the typed S3 not-found checks

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: franchb <ru.balloo@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:59:09 +08:00
975cec9228 s3api: exclude marker part in listObjectParts pagination (#11463)
* s3api: exclude marker part in listObjectParts pagination

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: guard listObjectParts marker boundary and enhance pagination test

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: fold in review feedback from the parallel #11462 fix

Same core fix; this adds the explanatory comment, tightens the overflow
guard to math.MaxInt64, makes the fake filer sort entries like a real
listing, and adds the marker-exclusivity assertions alongside the
pagination walk.

Co-authored-by: yi111 <yi111@users.noreply.github.com>

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: yi111 <yi111@users.noreply.github.com>
2026-09-26 11:58:23 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f31a026b2a master,filer: fix lock ring poisoning after leader change (#11453)
* cluster: never broadcast an empty lock ring

An empty member list is never a usable ring state, but a delayed
RemoveServer on a former leader can fire after the new leader already
broadcast the recovered ring. That late broadcast carries a newer
wall-clock version, so clients accept the empty ring and permanently
reject the good one.

Skip the broadcast entirely when the member list is empty, keeping the
last non-empty snapshot for reconnecting clients.

* cluster: periodically rebroadcast the lock ring

Ring updates are purely event-driven, so one lost or poisoned update is
permanent until the next membership change — with a single filer that may
never come. Re-arm a per-group timer after every broadcast so the current
leader keeps re-sending the ring; clients reject nothing newer than their
last accepted version, so a re-sent snapshot always heals a stale view.

* filer,s3api: reset the lock ring on master change

Ring versions are per-master monotonic — each master stamps wall-clock
nanoseconds — so a late high-version update accepted from a former leader
makes the new leader's snapshot look stale forever. Detect a leader
change across the reconnect gap (currentMaster is cleared between
attempts, so remember the last served master) and reset the ring to
bootstrap state so the new leader's view always applies.

* cluster: fail lock acquisition when no lock server exists

retryUntilLocked loops forever, so a filer reporting an empty lock ring
wedges every append write indefinitely. Bound only the "no lock server
found" case — ordinary contention is still waited out since the holder
releases eventually. The constructors now return nil on failure: the
filer append path and S3 object writes fail fast, while mounts degrade
to their existing lockless mode.

* cluster: reset only the ring version on master change

Ring versions are per-master monotonic, so a version gate reset is all a
leader change needs. Clearing the whole ring made every filer its own
write owner until the next update and dropped the prior-owner window for
keys the new leader remaps; the last ring now keeps routing until the
new leader's snapshot transitions off it.

* cluster: skip redundant ring installs and defer rebroadcasts

An unchanged member list now only bumps the accepted version instead of
installing a snapshot: periodic rebroadcasts no longer fire the
topology-change callback or restart the prior-owner window. And a
rebroadcast that lands inside a membership stabilization window yields
to the pending timer rather than publishing an intermediate ring.

* cluster,mount: bound lock unavailability, fail ops that cannot lock

Only 'lock already owned' contention retries without bound now; every
other failure — no lock server, or a dead ring member refusing
connections — shares the same unavailability budget, so a ring naming
departed filers can no longer hang a lock forever. Mount open-write,
create, and rename fail with EAGAIN when the required lock cannot be
acquired instead of proceeding without cross-mount serialization.

* cluster: check pending stabilization inside the broadcast critical section

rebroadcast released the mutex between the pending-timer check and
nextBroadcastUpdate, so a membership change arriving in the gap could arm
a stabilization timer while the rebroadcast emitted an intermediate ring.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: acquire path locks before mutating create/rename state

Create took the DLM lock only after the filer create, so a lock failure
returned EAGAIN with an eagerly persisted file left behind. Rename marked
source handles renamed before acquiring locks, so a failed acquisition
left them suppressing old-path flushes for a rename that never happened.
Both now take the locks first; the create's lock is released again if the
entry race loses to another creator and AcquireHandle takes over.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: keep the old-path lock when rename lock migration fails

The migration stopped the handle's lock before acquiring the replacement,
so a nil result left the handle writing with no lock at all. Acquiring the
new-path lock first means failure keeps the existing lock instead of
reporting success with serialization dropped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: skip new-path rename lock when a handle already holds it

A target file open for write on this mount already carries a lock on
newPath; the lock manager does not grant a second lock to the same
owner, so the rename would wait on itself until the handle closed.
Also avoid locking twice when old and new paths coincide.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: hand the rename's target lock to the migrating handle

The rename holds a lock on newPath for its duration, so the response
migration's fresh acquisition waited on that same lock until the handle
released — under fhLockTable, blocking the handle's own close. Adopt the
rename's lock directly; nested move responses still acquire their own.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: move the replaced target's lock to the renamed handle

When the target path was already locked by an open handle on this
mount, the migrated source handle kept only its stale old-path lock —
the target's close would then release the last lock on the new path
while the renamed handle was still open. Adopt the replaced handle's
lock instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: stop the handle lock inside the fh lock on release

ReleaseHandle stopped fh.dlmLock before taking the fhLockTable slot, so
a rename migration holding that slot could still observe and adopt a
lock that was already stopping. Stopping under the fh lock makes the
transfer serialize against the release.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the replaced target's lock for the renamed handle

When the target path is already locked by an open handle on this mount,
adopting it at migration time keeps the renamed path protected after
that handle closes, without waiting on a lock this mount already holds.
If the handle was released mid-migration the claimed lock is stopped,
and a fresh acquire covers the case where it was already gone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the target handle's lock before the rename runs

Skipping the new-path lock when a handle already holds it let that
handle's close release the lock mid-rename, leaving the path unguarded
until the response migrated it. Take over the lock at check time and
hold it for the rename's duration: the response adopts it for the
migrating handle, or it returns to the target handle / is released on
failure. The target handle lookup also falls back to the entry's stored
inode for a forgotten path mapping.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: read handle locks only under the fh lock during rename

The loose dlmLock reads raced ReleaseHandle, which now mutates the lock
inside the handle lock; check and claim it under the same hold.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:57:00 +08:00
github-actions[bot] 4914c14982 docs: regenerate star history chart 2026-09-26 00:41:51 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2f6c237238 filer: keep lazy remote reads from resurrecting deleted paths (#11452)
* filer: keep lazy remote reads from resurrecting deleted paths

Under a remote mount with filer.remote.sync as write-back, a path that
was deleted or renamed away could come back as a chunkless remote-only
entry: between the local delete and the daemon's remote delete, a store
miss made maybeLazyFetchFromRemote trust a bucket that was behind the
filer. The ghost then outlived the remote object -- HEAD answered 200,
GET failed, and nothing cleaned it up.

The filer now tombstones paths it deletes under a remote mount, learned
both synchronously from its own delete path and from peer metadata
events. The lazy fetch and the lazy listing skip a tombstoned path until
the path is written again, until the mount's persisted write-back sync
offset has passed the delete event (the remote delete has landed), or
until a generous TTL covers a mount without a daemon.

Fixes #11440

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: cover recursive remote deletes with an ancestor tombstone

A recursive delete now records the directory tombstone before walking
children, so a partial traversal or a store that drops the subtree
without listing it still leaves every descendant covered. Directory
tombstones also subsume older descendant entries on add, descendant
adds covered by a standing ancestor are skipped, and an existing
tombstone can be refreshed even at capacity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote tombstones to the deleted object's generation

A remote object whose own mtime postdates the local delete is a new
generation, not the one the tombstone hides, so a recreated directory
can surface remote writes made after its delete while old-generation
objects stay hidden. Lazy fetch now stats the remote object before
deciding, listings pass each child's remote mtime, and a sync offset
releases a tombstone once it reaches the delete's own timestamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: rebuild remote deletion tombstones after restart

In-memory tombstones are lost on restart while remote write-back
offsets persist, so a filer boot replays the persisted metadata log
from the oldest mount offset and folds deletes back into the tombstone
set through the same event handler. Lazy remote reads hold off while
the replay runs so a pending delete cannot resurrect in the gap.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: release remote tombstones only after their delete event lands

The write-back offset orders against event timestamps, but the synchronous
delete path recorded tombstones with the local clock before its event was
emitted — a later unrelated event could already have pushed the mount's
watermark past that guess, releasing the tombstone before the daemon
applied the delete. Tombstones recorded ahead of their event are now
marked pending and can only be lifted by the event confirming them or by
TTL; event-stamped tombstones release through the offset as before.

The remote-mtime generation bypass is dropped: remote and filer clocks
are independent, and a pending remote delete removes whatever object sits
at the path, so a "newer" remote object would only resurrect as a
phantom. Tombstoned lookups now skip the remote stat entirely.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: drop dir tombstone when recursive delete fails before listing

The ancestor tombstone is recorded before the child listing; if that
listing fails nothing was deleted, and the leftover tombstone would hide
still-existing remote children for the whole TTL. Tombstones for children
already deleted stay, since their remote deletes are still owed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: block lazy remote reads on startup tombstone rebuild

The rebuild gate is now a done-channel set synchronously before the
replay goroutine starts, so no lazy read can slip through in between.
Reads wait on it with context cancellation instead of returning an
empty miss that makes remote-only objects look deleted.

The replay start is floored at now-TTL: mounts without a recorded
write-back offset previously replayed the whole persisted history, and
events older than the TTL would only build already-expired tombstones.
The gate check now runs after the mount lookup so replaying the meta
log's own directory listings does not deadlock on the gate, and the
replay retries with backoff until it succeeds instead of failing open.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: mark restamped tombstone pending until its delete event lands

When a local delete raises an existing tombstone's timestamp, the new
value is only a local clock guess ahead of that delete's event. Leaving
the tombstone un-pending lets a write-back offset release it before the
event is actually consumed, reopening the resurrection window.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound tombstone replay to the tombstone TTL

Persisted-log replay retried forever, keeping lazy remote reads gated
indefinitely when the log cannot be read. Cap retries at the tombstone
TTL measured from replay start: past that point every tombstone would
have expired anyway, so opening the gate loses no protection.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: re-check deletion tombstone before persisting lazy fetch

A delete landing while StatFile is in flight passed the earlier
tombstone check but still persisted the fetched entry, resurrecting a
path whose remote delete is pending. Re-check right before CreateEntry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract a lazily persisted entry when a delete raced the insert

The pre-insert tombstone check still leaves a window between the check
and the store insert. Since deletes always record the tombstone before
removing the entry, a tombstone visible right after a successful insert
means the delete already ran: delete the entry back out so the
tombstoned path stays deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: note why the replay deadline can safely open the gate

Deletes made after startup are captured by the live delete and event
paths, so a stalled replay can only be missing pre-restart deletes, all
of which are past the tombstone TTL by the deadline.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract only the entry a lazy remote read materialized

Deleting by path after a raced delete could remove a legitimate rewrite
that replaced the fetched entry. Verify the stored entry still matches
the remote object (or the just-created directory shape) before deleting,
and apply the same post-insert check to lazy listing children.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: require full-entry equality before retracting a lazy entry

Remote-only matching still removed a write that had updated the fetched
entry, e.g. appended chunks. Compare the persisted entry against what
this read materialized; any change means a real update owns the path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 08:00:39 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
df4995b894 volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock (#11443)
* volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock

make_volume_readonly looked the volume up, notified the master (step 1),
then took the store write lock (step 2) and marked the volume only `if
let Some(..)`. When the volume left the store during step 1 -- a master
round trip, during which an unmount or a heartbeat expiry can land --
the missing else meant the RPC reported success for a volume the server
no longer has, and step 3 told the master again that it is read-only.

Go's Store.MarkVolumeReadonly (weed/storage/store.go) returns
"volume %d not found" when findVolume comes back nil, and
makeVolumeReadonly (weed/server/volume_grpc_admin.go) returns that error
before the step-3 notification. The Rust step 2 now does the same:
find_volume_mut(vid) -> Status::not_found("volume {vid} not found"), and
the `?` skips step 3, as it already did for a set_read_only_persist
failure. The scrub caller already matches NotFound to skip such a
volume instead of failing the whole report; it now actually gets it.
volume_mark_writable already returns NotFound under its write lock.

The regression test opens the step-1 window deterministically: step 1
awaits the current_master_url read lock, so the test holds its write
guard, lets make_volume_readonly park there after its own lookup
succeeded, unmounts the volume, then releases the guard. With no master
configured the notification is a no-op, so the write lock in step 2 is
the only place left that can notice the volume is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments on the vanished-volume mark-readonly path

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:05:53 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
b750853c42 shell: refuse s3.bucket.create on an existing bucket (#11455)
* shell: refuse s3.bucket.create on an existing bucket

CreateEntry without o_excl replaces the bucket entry, dropping every
extended attribute: lifecycle configuration, owner, versioning and the
irreversible Object Lock flag. Send o_excl so a re-run fails with
'bucket already exists' instead of silently resetting the bucket.

* filer: fail exclusive creates when the lookup itself fails

CreateEntry discards FindEntry errors, so an o_excl create hitting a
transient store failure would take the insert path and upsert over the
entry it was meant to preserve. Propagate the lookup error when o_excl
is set; non-exclusive creates keep their existing semantics.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* shell: test s3.bucket.create requests an exclusive create

Exercises the command end to end through a fake filer gRPC server and
asserts the OExcl flag reaches the wire along with the already-exists
error path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* shell: synchronize captured requests and assert the exact bucket error

The fake filer records CreateEntry requests on the gRPC server goroutine,
so reads need the same mutex; the test also now checks for the exact
"bucket my-bucket already exists" message rather than any error that
mentions existence.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:04:54 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
317e756b9a volume server: validate ext and collection in gRPC CopyFile/ReceiveFile (Rust) (#11451)
* volume server: validate ext and collection in gRPC CopyFile

Port the Go-side checks (checkVolumeFileExtension, checkVolumeCollection)
to the Rust volume server so a client-supplied collection or ext carrying
a separator or ".." cannot fold a path outside the volume directory.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: validate ext and collection in gRPC ReceiveFile

Same port on the write path: the file ReceiveFile creates is built from
client-supplied fields, so reject traversal there too. Reported through
the response error field, matching Go's SendAndClose.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 09:47:16 +08:00
5b79f51e3c filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks (#11435)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks

updateLocalEntry records the RemoteEntry stamp after an upload by writing the
event's entry back with UpdateEntry. The filer deletes every stored chunk
absent from an updated entry, so when the file was rewritten while its upload
was in flight (or the event is a replay), the stale snapshot deletes the
rewrite's chunks: the entry then points at the new fid with no needle behind
it, and the rewrite's own upload fails and is skipped as superseded.

The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's
chunk fids, evaluated by the filer under the path lock. A refused stamp means
the filer moved past this event; the superseding event follows in the log and
stamps the current entry, so the refusal is logged and skipped like a
superseded upload.

Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount,
filer.remote.sync; hold the remote (docker pause) so one upload stays in
flight, rewrite the file through the filer, unpause. Before: the entry's chunk
is 404 on every volume server. After: the stale stamp is refused, the rewrite's
chunk stays live and reads back after a vacuum.

* filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored

The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a
rewrite that touched inline content or metadata alone still compared
equal and the stale snapshot overwrote the live entry. The new clause
compares the whole stored entry against the event's entry under the
same path lock.

* filer: route conditional UpdateEntry to the entry's owner filer

Two filers locking the same path locally could still pass a stale
condition on the non-owner while the owner's entry had moved on. When a
condition or expected_extended precondition is set, forward the request
to the entry's owner the same way conditional CreateEntry does, with
is_moved bounding the hop.

* filer: compare IF_ENTRY_EQUAL against the normalized expected entry

FindEntry grows FileSize to the chunk extent, so a raw event entry with
FileSize still zero failed the condition on an unchanged file and the
stamp was skipped, letting a replay upload the object again.

* filer.remote.sync: classify refused stamps by gRPC status only

A FailedPrecondition substring in an unrelated error would have been
swallowed as a skipped stamp; status.FromError already unwraps.

* remote sync: keep the event entry intact for IF_ENTRY_EQUAL

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-25 09:02:34 +08:00
635f69a821 test(s3tables): add an OLake Iceberg catalog integration test (#11441)
OLake (github.com/datazip-inc/olake) is a CDC and ingestion engine that writes
Apache Iceberg. It covers two paths none of the existing catalog suites reach.

It is a strict Java Iceberg client. OLake does not write Iceberg from Go — its
Go process spawns a Java sidecar over gRPC and writes through the official
Apache Iceberg library, because the Go library has no equality deletes and CDC
needs them. That makes it the client class weed/s3api/iceberg's metadata
compliance backfill exists to serve: the one that fails with "Cannot parse
missing long current-snapshot-id" when spec-required keys are omitted.

And it produces equality deletes. Its upsert path commits operation=overwrite
with an equality-delete file and a delete manifest. ClickHouse, Doris, Trino,
Spark and DuckDB all only append, so nothing else in this directory exercises
a delete manifest at all.

Six subtests: the destination check reaches SUCCEEDED and actually loads the
REST catalog; discover enumerates the source; a full sync commits a snapshot;
PyIceberg reads back what the Java writer committed; an update plus a re-sync
records an overwrite carrying equality deletes with a delete manifest in the
current snapshot; and the catalog does not rewrite manifests the official Java
writer produced, which gives the manifest-repair path a negative test to go
with the ClickHouse positive one.

What it deliberately does not assert is a delete-applied read. PyIceberg
refuses to scan a table carrying equality deletes (apache/iceberg#6568) while
reading its metadata fine, and an engine that can apply them costs a
multi-gigabyte image. Recording the commit correctly is the catalog's
contract; applying deletes on read is the engine's. The README says so, says
the read half was verified by hand once with StarRocks, and warns against
later "upgrading" this to a PyIceberg rows read — which would either fail or,
if PyIceberg ever starts skipping deletes instead of raising, pass by not
looking.

The workflow job asserts the suite actually ran rather than trusting a green
exit, for the same reason: at least one top-level PASS and zero SKIP.

No product change — the destination config is the generic catalog_type=rest
with the standard OAuth2 client-credentials flow, and s3_path_style is not even
set, since OLake turns it on itself whenever s3_endpoint is non-empty.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 07:39:27 +08:00
Chris LuandGitHub 11791fad6a filer: resolve the collection a bucket delete drops (#11439)
* filer: resolve the collection a bucket delete drops

A bucket delete dropped the collection named after the bucket, which
assumes bucket name is collection name. With a collection rule the
write path honors, deleting the bucket either orphaned its collection
or, when a bucket was named after a shared collection, removed volumes
other buckets still write to.

Resolve the collection through the same rule chain the write path uses
and drop it only when no other bucket resolves there too. A listing
failure keeps the collection, the safe side of an unknown.

* filer: prove collection exclusivity across all paths before dropping it

The sibling-bucket scan missed every non-bucket writer: a broad rule like
'/' or '/buckets/', a rule under a surviving bucket, or a rule on an
unrelated path can route into the same collection. Check every storage
rule's prefix instead, and mirror the grouped gateway's explicit
<group>_<bucket> collection, which otherwise resolves a rule-named
collection the bucket never wrote to.

* s3: let the filer own the collection decision on bucket delete

Both entry points deleted a name-derived collection around the filer's
own resolved delete, bypassing its exclusivity check and wiping sibling
data. The filer now resolves the collection a bucket actually used,
including the grouped form.

* filer: keep a collection the default write route also uses

Rule-less writes outside buckets land in the filer's default collection,
so a bucket resolving there shares it with them.
2026-09-25 07:30:42 +08:00
Chris LuandGitHub 56d2f05ccd topology: wake the vacuum dispatcher when a worker frees quota (#11436)
* topology: wake the vacuum dispatcher when a worker frees quota

The dispatch loop slept a fixed 10s whenever every pending volume was
waiting for a per-server quota slot, so a sweep took volumes x 10s
regardless of how fast the compactions were. Workers now signal on a
buffered channel after crediting quota; the dispatcher waits on it with
the 10s sleep kept only as a timeout.

* master: add -vacuumIntervalSeconds to tune the automatic sweep interval

The 14-minute base interval was a literal inside the refresh loop while
every neighbouring vacuum knob was already a flag. Defaults to 840s,
unchanged.

* topology: keep the 14 minute floor on the vacuum interval

A zero-valued MasterOption or a negative -vacuumIntervalSeconds left the
sweep sleeping only its jitter, so treat non-positive intervals as the
previous default.
2026-09-25 02:14:59 +08:00
Chris LuandGitHub 8c1be63c92 ecbalancer: let a non-overflow parity shard leave a data-bearing rack (#11438)
The parity pass only queued shards past the per-type cap, so a single
parity shard sharing a rack with data was never a move candidate even
when an empty data-free rack existed (2+1 over 3 DCs settled 2/1/0).
Non-overflow candidates now move too, but only to a rack without data;
overflow shards keep the existing data-rack fallback.
2026-09-25 02:14:23 +08:00
bb9942c646 volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes (#11413)
* volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes

The Rust VolumeEcShardsGenerate went straight into write_ec_files: no unload
of an already-mounted EC volume and no stale-artifact sweep. Only .ec00..ecNN
on the encoding disk were truncated, so a retry could mix two encode runs. A
stale N.ec03 left on a sibling disk survived, reconcile later mounted it
against the new .ecx, and the new .vif made the encode_ts_ns identity guard
pass, so reads served old-run bytes at new-run offsets.

Mirror Go's VolumeEcShardsGenerate (#9880 / #9953): UnloadEcVolume on every
disk, then removeStaleEcArtifacts on every disk location before encoding.
remove_ec_volume_files_full_teardown already has removeStaleEcArtifacts'
semantics (.ec00..ec31, .ecx/.ecj/.ecsum[.vN] in both the data and idx dirs,
.vif only on a shard-only disk; never the source .dat/.idx), so reuse it. Add
Store::unload_ec_volume, which unlike remove_ec_volume does not stop at the
first disk and closes the descriptors so the unlink frees the inodes. The
store write lock covers only unload + sweep, not the encode.

The failure arm now also drops the generation-0 .ecsum, as Go's defer does.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: wake the heartbeat after VolumeEcShardsGenerate unloads shards

The pre-encode unload drops mounted EC shards from memory, but unlike every
other unmount path it did not wake the heartbeat, so the master kept routing
reads to shards this server no longer serves until the next pulse. Notify
once the store lock is released, and before the sweep error propagates: a
failed sweep has unloaded the shards too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: clean up encode artifacts when the .vif write fails too

Go's shouldCleanup defer covers every error before the .vif commits,
not just a failed encode. A serialize or write failure on the .vif left
the fresh .ecNN/.ecx/.ecsum behind, which the next generate would have
to rely on the new sweep to remove. Extract the cleanup and run it on
the .vif error paths as well.

* volume server: write the EC .vif atomically

Go's SaveVolumeInfo writes a temp file, syncs it, and renames it over
the target, so a failed write leaves the previous metadata intact and a
read-only .vif fails the save. The direct fs::write truncated the file
first, so a write or sync failure could leave an empty .vif even after
cleanup_encode removed the generated shards.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:10:54 +08:00
f0afcf904d volume: an EC volume needs an .ecx to mount, and a 0-byte stub never outranks a real index (#11415)
* volume: an EC volume needs a non-empty .ecx to mount

Two gaps against Go in how the Rust volume server treats the .ecx.

EcVolume::new mounted with no index at all. The per-shard
VolumeEcShardsMount path picks the disk by shard file alone, so a shard
whose .ecx was on no local directory still registered and was
advertised to the master; every VolumeEcShardRead then failed with
"ecx file not open", and add_shard's 0-byte guard was neutralised
because ecx_file_size stayed 0. Go's NewEcVolume returns an error
wrapping os.ErrNotExist. EcVolume::new now fails with NotFound, and
Store::mount_ec_shard looks up the .ecx owner across all disks first
(findEcxIdxDirForVolume) so a shard on a sibling disk of its index
still mounts instead of turning into a hard failure.

A 0-byte .ecx stub, as left by a failed EC distribute copy, counted as
a valid index. Go requires Size() > 0 wherever the file steers a
decision: HasEcxFileOnDisk, findEcxIdxDirForVolume, indexEcxOwners
(shared by reconcile and mirror), and VolumeEcShardsCopy removes a
copied 0-byte .ecx and fails the copy. Mirror each through one
is_usable_ecx_file helper. NewEcVolume itself still accepts a lone
0-byte .ecx as a legitimate empty index, but prefers a non-empty copy,
local directory first, over a stub in the other directory; the
resolution in EcVolume::new now follows the same order.

Tests that mounted EC volumes without any .ecx get a real fixture.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: mount_ec_shard tries every disk; reconcile ignores a 0-byte local .ecx

mount_ec_shard returned the first disk's error, so an unusable shard copy
(a 0-byte .ecNN left by an interrupted move) hid a good copy on the next
disk. Like Go's MountEcShards, keep scanning: NotFound means "not this
disk", any other failure is collected, and an all-disks-fail error names
every disk tried. "No .ecx on any local disk" is now told apart from
"shard not on this server".

The orphan-shard reconcile took its locally-mirrored fast path whenever a
local .ecx existed at all. A 0-byte stub there registered the shards against
an empty index while the owner index skipped that same stub. Go gates the
fast path on HasEcxFileOnDisk; do the same. ec_local_ecx_path loses its last
production caller and becomes test-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: match Go's mount error text and skip the owner stat on the owning disk

MountEcShards in Go skips the HasEcxFileOnDisk stat when the disk's own
directories already hold the .ecx, dedups a shared -dir.idx across
locations in findEcxIdxDirForVolume, and reports "load failures" with
the same wording. Also drop two issue-number references from comments.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:08:56 +08:00
c1ccbcda13 volume server: the write queue answers uploads with the needle's real ETag (#11414)
With SEAWEED_WRITE_QUEUE=1 every upload came back with ETag "00000000".
The upload handler built the needle with Needle::default(), so its
checksum was CRC(0), and handed a clone of it to the queue. The CRC was
only computed in the write path, on the worker's clone, and WriteResult
carries no checksum back, so n.etag() in the handler formatted the zero
checksum. The direct path writes through &mut n and was correct.

Compute the checksum in the handler while building the needle, the way
Go's CreateNeedleFromRequest does, over the same bytes the write path
hashes (the stored data, gzipped or not). The ETag and the has-name flag
are read before the write, so the needle is moved into the queue instead
of cloned, which also drops a full payload copy per queued upload.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 07:08:36 +08:00
94a68fa9b9 volume: walk_index_file keeps row alignment across short reads (#11412)
* volume: walk_index_file keeps row alignment across short reads

walk_index_file issued one Read::read per batch and decoded whatever came
back. Read::read may legally return a short count that is not a multiple
of the 17-byte entry size (FUSE and network filesystems, a BufReader whose
capacity is not a multiple of 17). The split entry at the end of the batch
was dropped with no carry and the next read started mid-entry, so every
later row was decoded from misaligned bytes and fed to the index as a
garbage key/offset/size. This function backs every in-memory index load.

Go's WalkIndexFile is immune because it reads through io.ReaderAt, which
returns a full buffer or an error. Fill the batch buffer until it is full
or the reader reports EOF, retrying ErrorKind::Interrupted, and only then
decode whole entries. Reads stay batched at ROWS_TO_READ entries.

EOF semantics are unchanged and match Go: on io.EOF Go decodes the whole
entries in the final buffer, ignores a trailing partial entry and returns
nil. A torn final entry is still skipped without an error here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim walk_index_file comments

The batch-fill loop and the ShortReader test helper each carried a
paragraph where a sentence suffices.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:08:26 +08:00
Chris LuandGitHub b3a8701989 lance: authenticate the catalog with Bearer tokens and x-api-key (#11431)
* lance: accept OAuth2 bearer tokens for catalog auth

Lance and LanceDB clients can only send OAuth2 / Bearer / API-Key
headers on catalog calls, never SigV4, so behind an auth-enabled S3
gateway every namespace request failed with 403 Access Denied.

Mirror the Iceberg catalog's OAuth2 support: POST /oauth/token accepts
an S3 access key / secret key as client_id / client_secret, validates
them against IAM, and returns a signed JWT. The Auth middleware accepts
that token as a Bearer credential before falling through to SigV4.

Closes #11430

* lance: accept x-api-key header carrying an S3 credential

The Lance namespace spec's third auth scheme maps api_key onto the
x-api-key header. Accept "access_key:secret_key" there and validate it
against IAM, so clients that only hold static headers can authenticate
without minting a token first.

* lance: answer invalid_client with the Basic challenge

RFC 6749 5.2 requires a 401 from the token endpoint to carry
WWW-Authenticate matching the scheme the client used, so it knows how
to retry.

* lance: cap the token endpoint request body

/oauth/token is unauthenticated, so ParseForm needs the same size
bound decodeBody applies to every other catalog request.

* lance: keep query strings out of request logs

/oauth/token rejects a client_secret sent in the query, but the
logging middleware and the catch-all wrote RequestURI to the log
before that rejection ran. Log the path alone so a mis-sent secret
never reaches the log.

* lance: log the escaped path, not the decoded one

URL.Path decodes percent escapes, so a request like /%0aFORGED could
split log lines. EscapedPath keeps the encoding while still dropping
the query string.
2026-09-24 07:02:59 +08:00
b9ad62fc16 [Volume] Keep DAT and index state consistent after async batch Sync failure (#11425)
* fix 11400

* persist failed-recovery quarantine and harden rollback

- record the unavailable state in a .unavailable marker, fsync it, and
  re-arm it on load so a restart cannot serve an unverified pair
- quarantine the volume so heartbeats stop advertising it
- block MarkVolumeWritable while unavailable, rechecked under noWriteLock
- fail every request of a failed batch, not only the succeeded ones
- restore the needle map and truncate .dat on inline fsync rollback failure
- add truncateIndex for the sorted-file needle map
- mirror the fail-closed semantics in the Rust volume server

* volume: erase rolled-back mappings instead of leaving tombstones

A rolled-back batch or failed inline write used Delete() to undo a
needle that did not exist beforehand, leaving a tombstoned map entry
whose stale offset makes the next write to that needle fail reading a
header that no longer exists. Add removeMapping/restoreMapping to the
mappers so recovery erases entries that were absent before the batch
and reinstates the exact prior offset/size for ones that were,
including tombstones. The index row still goes through Delete so a
replay forgets the needle.

* volume: gate bulk readers on unavailable and fsync the marker's dir

- fsync_dir(&self.dir) synced the volume dir's parent, not the dir
  holding .unavailable; pass the marker path so the create survives
  a host crash
- export UnavailableError and check it in ReadAllNeedles,
  VolumeTailSender, VolumeIncrementalCopy, and IncrementalBackup so
  replica-sync paths cannot stream or append data from an unverified
  .dat/.idx pair; mirror on the Rust side via read_dat_slice,
  read_all_needles, dat_scan_plan, and the incremental-copy handler

* volume: drop issue references from comments near touched code

* volume: stop active scans when the volume becomes unavailable

The stream entry-point checks ran once per RPC, so a volume quarantined
by a failed recovery mid-scan kept serving data. Recheck availability
per needle/chunk on the detached read paths: tail scan and heartbeat,
read-all, incremental copy, incremental backup writes, and the Rust
StreamingBody chunk reads. Rust incremental copy also rejects a
quarantined volume before sync_to_disk touches the backend.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-24 06:57:44 +08:00
Chris LuandGitHub d848b8ed00 rust volume: regression test for makeup_diff replay across a 32 GiB offset boundary (#11410) (#11427)
* rust volume: test makeup_diff replay across a 32 GiB offset boundary

Issue #11410 corrupted a replayed write's index offset in Go's makeupDiff
by patching only four of the five offset bytes. The Rust makeup_diff
already encodes the whole offset through idx_entry_to_bytes and
Offset::from_actual_offset; this adds the mirror of
TestConcurrentWriteCrossesOffsetBoundary so a regression would fail here
the same way it does under -tags=5BytesOffset on the Go side.

Sparse-truncate the .dat to 64 GiB, compact, write, commit: the index
offset must equal the .cpd size and the needle must stay readable
through a second vacuum. Gated on the 5bytes feature since a 64 GiB
.dat exceeds the 32 GiB range of 4-byte offsets.

* rust volume: skip the offset-boundary replay test on Windows

Windows set_len allocates the full 64 GiB extension instead of a sparse
range, so the test fails with StorageFull on CI runners. Gate it to unix,
where set_len leaves the extension unallocated.
2026-09-23 18:08:20 +08:00
github-actions[bot] 196c71b613 docs: regenerate star history chart 2026-09-23 00:46:51 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
559ec33498 build(deps): bump github/codeql-action from 4.38.0 to 4.38.1 (#11423)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.0 to 4.38.1.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.38.0...v4.38.1)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:47 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d6397fdc50 build(deps): bump dorny/paths-filter from 3 to 4 (#11422)
Bumps [dorny/paths-filter](https://github.com/dorny/paths-filter) from 3 to 4.
- [Release notes](https://github.com/dorny/paths-filter/releases)
- [Changelog](https://github.com/dorny/paths-filter/blob/master/CHANGELOG.md)
- [Commits](https://github.com/dorny/paths-filter/compare/v3...v4)

---
updated-dependencies:
- dependency-name: dorny/paths-filter
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:38 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
264030c08c build(deps): bump docker/setup-qemu-action from 4.3.0 to 4.4.0 (#11421)
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 4.3.0 to 4.4.0.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](https://github.com/docker/setup-qemu-action/compare/v4.3.0...v4.4.0)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: 4.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:30 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
8099e71934 build(deps): bump golang.org/x/crypto from 0.56.0 to 0.57.0 (#11419)
Bumps [golang.org/x/crypto](https://github.com/golang/crypto) from 0.56.0 to 0.57.0.
- [Commits](https://github.com/golang/crypto/compare/v0.56.0...v0.57.0)

---
updated-dependencies:
- dependency-name: golang.org/x/crypto
  dependency-version: 0.57.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:21 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
63eb67c9da build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts from 1.50.0 to 1.51.0 (#11420)
build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts

Bumps [github.com/aws/aws-sdk-go-v2/service/sts](https://github.com/aws/aws-sdk-go-v2) from 1.50.0 to 1.51.0.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.50.0...service/s3/v1.51.0)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/service/sts
  dependency-version: 1.51.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:37 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
93efc64af6 build(deps): bump go.uber.org/atomic from 1.11.0 to 1.12.0 (#11416)
Bumps [go.uber.org/atomic](https://github.com/uber-go/atomic) from 1.11.0 to 1.12.0.
- [Release notes](https://github.com/uber-go/atomic/releases)
- [Changelog](https://github.com/uber-go/atomic/blob/master/CHANGELOG.md)
- [Commits](https://github.com/uber-go/atomic/compare/v1.11.0...v1.12.0)

---
updated-dependencies:
- dependency-name: go.uber.org/atomic
  dependency-version: 1.12.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:29 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f2b08e47e5 build(deps): bump github.com/tarantool/go-tarantool/v3 from 3.0.1 to 3.0.2 (#11417)
build(deps): bump github.com/tarantool/go-tarantool/v3

Bumps [github.com/tarantool/go-tarantool/v3](https://github.com/tarantool/go-tarantool) from 3.0.1 to 3.0.2.
- [Release notes](https://github.com/tarantool/go-tarantool/releases)
- [Changelog](https://github.com/tarantool/go-tarantool/blob/master/CHANGELOG.md)
- [Commits](https://github.com/tarantool/go-tarantool/compare/v3.0.1...v3.0.2)

---
updated-dependencies:
- dependency-name: github.com/tarantool/go-tarantool/v3
  dependency-version: 3.0.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:20 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
5079d926e4 build(deps): bump cloud.google.com/go/kms from 1.33.0 to 1.34.0 (#11418)
Bumps [cloud.google.com/go/kms](https://github.com/googleapis/google-cloud-go) from 1.33.0 to 1.34.0.
- [Release notes](https://github.com/googleapis/google-cloud-go/releases)
- [Changelog](https://github.com/googleapis/google-cloud-go/blob/main/documentai/CHANGES.md)
- [Commits](https://github.com/googleapis/google-cloud-go/compare/kms/v1.33.0...dlp/v1.34.0)

---
updated-dependencies:
- dependency-name: cloud.google.com/go/kms
  dependency-version: 1.34.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:13 +09:00
github-actions[bot] a9e1b57dcd docs: regenerate star history chart 2026-09-22 00:46:32 +00:00
Chris LuandGitHub e2608edda4 volume: fix 5-byte index offset corruption in makeupDiff (#11411)
* volume: encode all offset bytes when makeupDiff replays a write

makeupDiff patched only bytes 8:12 of the index entry, so under the
5BytesOffset build the fifth byte kept the old offset's high bits and the
replayed needle's index pointed 32 GiB-aligned ranges away from its body.
A later vacuum then dropped the entry as unreadable. Rebuild the entry with
needle_map.ToBytes, the same encoder the tombstone branch just below uses.

* volume: test makeupDiff replay across a 32 GiB offset boundary

Sparse-file test: truncate the .dat to 64 GiB after one write, compact,
write a second needle, commit, and assert the index offset matches the
.compacted size and the needle stays readable through a second vacuum.
Only runs under -tags=5BytesOffset.
2026-09-21 00:59:35 -07:00
0f2ecb766f volume server: reject non-ASCII input instead of panicking (#11406)
* volume server: reject non-ASCII input instead of panicking

Three parsers sliced attacker-supplied strings by byte offset, so a
multi-byte character split inside itself and panicked the task:

  - parse_needle_id_cookie took the last 8 bytes as the cookie and the
    rest as the needle id. Reachable from VolumeServer.BatchDelete,
    whose file_ids come straight off the wire as protobuf strings;
    that handler already answers 400 per bad fid, so the guard turns a
    panicked RPC into the error it was already written to return.

  - TTL::read took the unit as the last byte and the count as
    everything before it, so "?ttl=5<multi-byte>" split mid-character.
    The HTTP upload path does TTL::read(..).ok() and drops an invalid
    TTL; AllocateVolume maps the Err to InvalidArgument.

Both now reject non-ASCII up front. Hex and a digits-plus-unit TTL are
ASCII by definition, so no accepted input changes -- covered by tests
alongside the rejection cases.

The six response-* header overrides were inserted with
parse().unwrap(). They come from the query string, so
"?response-cache-control=%0Aevil" decodes to a value HeaderValue
rejects and the unwrap panicked the connection task,
unauthenticated. They now skip the override, matching the if-let the
chunked-response path in the same file already uses.

ReplicaPlacement::from_string was reported as a fourth site but is not
one: reaching chars[2] requires chars[0] and chars[1] to be ASCII
digits, which forces the padded string to be three single-byte
characters, so a multi-byte character always lands on a to_digit()
None first. Kept as a regression test rather than a change.

Each fix was confirmed against the unfixed code first: the parser
tests panic with "byte index N is not a char boundary", and the
integration tests panic at handlers.rs:1413 and ttl.rs:88.

Not a vector, contrary to the report: the HTTP request line. The path
is not percent-decoded before parsing, so "%C3%A9" stays ASCII and
fails the length check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: fall back to needle MIME when response-content-type is invalid

Skipping an unparseable override left the response without any
Content-Type because the override had already bypassed the normal MIME
selection. Also correct a test comment that described a chars[2] panic
which cannot be reached.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
2026-09-20 23:35:05 -07:00
Chris Lu 6848cdf9e1 s3: close seaweedfs-quota policy-confusion gap (#11409)
* s3: count seaweedfs-quota as an operation subresource

PUT /bucket?policy&seaweedfs-quota was not rejected by
hasAmbiguousSubresource because operationSubresources omitted the
seaweedfs-quota key. The router then picks the policy route
(registered first) while the IAM action resolver may resolve the
request to s3:PutBucketQuota, letting a quota-only identity write a
bucket policy. Reject the combination before routing, matching the
fix for policy&tagging (#10987).

* s3: resolve seaweedfs-quota after other bucket subresources

The quota routes are registered last among the bucket subresource
routes, but the action resolver found seaweedfs-quota inside the
unordered bucketQueryActions map, so a request carrying it alongside
another selector could be authorized as the quota operation while the
router served the earlier-registered handler. Resolve it explicitly at
the end so the resolver agrees with the router, mirroring how
list-type is handled.

* s3: count resolver subresources in the ambiguity guard

hasAmbiguousSubresource only counted operationSubresources, so adding
a query parameter to the action resolver without updating that list
reopened the authorize-one-serve-another gap. Count bucketQueryActions
keys as operation selectors too, and add a test that walks the
registered routes and fails on any query key that is neither an
operation subresource nor a known modifier.
2026-09-20 23:34:57 -07:00
ca62d4297b volume: load the .ecj deletion journal in chunks, and repair a torn tail (#11408)
* volume: load the .ecj deletion journal in chunks, and repair a torn tail

Two independent defects in the EC deletion journal's load path.

1. The loader issued one NEEDLE_ID_SIZE-byte positional read per entry.

   That is fine for a healthy journal -- kilobytes -- and pathological for a
   large one. A `.ecj` is semantically a SET of deleted needle ids but is
   written as an append-only log that nothing dedupes, and several paths append
   a peer's ENTIRE journal onto the local one (VolumeEcShardsCopy with
   copy_ecj_file, EC index recovery, and ec_decode's deliberate cross-holder
   merge), so a volume whose shards are repeatedly balanced between two servers
   grows the file without bound.

   Observed in production: 1.51 TB and 1.30 TB on the two holders of one 10+4
   volume containing ~100 distinct ids. At that size the per-entry loop is
   ~188e9 syscalls, run synchronously while holding the deleted_needles write
   lock and before the HTTP port opens. The process sits at 100% of one core
   with a small RSS -- the set stays tiny because the ids repeat -- reading at a
   few MiB/s because 8-byte reads defeat readahead, logs nothing after "Adding
   storage location", and ignores SIGTERM. The master then unregisters every
   volume it holds and reads of them fail. 4.46 and 4.47 are both affected.

   Read in 1 MiB chunks and build into a local set, merging once at the end so
   the write lock is not held for the whole scan. Measured on a 256 MiB journal
   of 100 distinct ids: 33,554,500 syscalls -> 257, identical resulting set.

2. A torn tail silently corrupted later deletes.

   The journal handle is in append mode, so writes land at the physical end
   regardless of alignment. A trailing partial record therefore pushed every
   later append out of alignment: the loader skipped the partial bytes, but the
   next mount decoded them together with the leading bytes of the following
   entry, producing one garbage id and dropping the delete that came after the
   tear -- after acknowledging it.

   Truncate to a whole number of records at mount, before anything can append.
   The repair uses its own read+write (non-append) handle: on Windows,
   append(true) requests FILE_APPEND_DATA without FILE_WRITE_DATA (and
   .write(true) is subsumed by .append(true)), so SetEndOfFile through the
   journal handle fails with ERROR_ACCESS_DENIED.

   The same trap exists in journal_delete's recovery path, which calls set_len
   on the append handle to roll back a partial write whose sync failed. It is
   error-handled rather than fatal, so on Windows that rollback silently does
   not happen. Untouched here; worth a separate fix.

Bounding the journal's growth needs compaction, which is deliberately not in
this change: replacing the file under a store that can hold several EcVolume
instances for one volume id requires coordinating with the other holders, and
that belongs at the store layer. Sent separately.

Tests: a journal spanning several read chunks loads every entry; a trailing
partial record is ignored rather than panicking; a torn tail is truncated at
mount and a delete taken afterwards survives a remount.

* volume: roll back a failed .ecj append through a dedicated write handle

The append handle lacks FILE_WRITE_DATA on Windows, so the set_len
rollback after a failed sync silently did nothing and the journal could
drift one record past deleted_needles. Same trap as the torn-tail repair
in this file; fix it the same way. Also format the new tests.

* volume: mirror chunked .ecj load and torn-tail repair in Go

---------

Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:41 -07:00
4bb40732bb volume server: ReceiveFile loses bytes and hides fsync failures (#11407)
* volume server: ReceiveFile loses bytes and hides fsync failures

Three defects in one handler, all on the path that receives a pushed
.dat/.idx/.vif or EC shard:

  - `f.write(&content)` never compared the return to content.len().
    A short write (ENOSPC, NFS) counted only the bytes that landed,
    so every later chunk was written at a shifted offset and the RPC
    answered error: "" with a byte count that looked right. Go's
    os.File.Write loops. Now write_all.

  - `let _ = f.sync_all();` discarded EIO and answered success with
    the full byte count. Go omits the check too, but
    ReceiveFileResponse carries an `error` field and the caller
    renames the staged file into place on success -- so a silent
    fsync failure publishes a file whose data never reached the
    platter. Flush and fsync failures are now reported.

  - Both the per-chunk write and the final fsync were blocking
    std::fs calls inside the async fn, on the runtime worker that is
    also driving the stream. Switched to tokio::fs + BufWriter, the
    shape `drain_copy_stream_to_file` in this same file already uses
    and documents. The partial-file cleanup on the error path moves
    to tokio::fs::remove_file for the same reason.

The handler had no test at all, which is how the short-write bug
survived. Added a round-trip over a real connection with ragged chunk
boundaries, asserting the bytes on disk and not only the reported
count -- a dropped or reordered chunk changes the file even when
bytes_written still adds up.

That test guards the rewrite; it does not reproduce the original
faults. ENOSPC and EIO need fault injection that this suite has no
harness for, so the short-write and fsync paths are argued from the
code, not demonstrated by a failing test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: remove the staged file on every ReceiveFile error reply

Flush and fsync failures returned early and left the partial .copying or
shard file behind, as did the pre-existing write-error path. Route all
response-level errors through one cleanup block, matching Go's
close-and-remove on a failed write.

* volume server: tighten ReceiveFile comments

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:37 -07:00
8ff2e0777e volume server: HTTP DELETE on a distributed EC volume (#11405)
* volume server: HTTP DELETE on a distributed EC volume

The delete handler validated the cookie with EcVolume::read_ec_shard_needle,
which reads only locally-mounted shards and errors "ec shard N not available
locally" for any interval held by a peer. Every Err was mapped to 500 and no
.ecj tombstone was appended, so on a standard 10+4 spread over 14 servers an
HTTP delete of an EC needle could not succeed. The GET path already goes
through read_ec_shard_needle_distributed.

Route the delete's read through the same distributed reader. It does a
local-first pass in its snapshot phase, so the all-shards-local case costs
what it did before, and no store guard is held across the await (the reader
takes its own; RwLockReadGuard is !Send).

Two smaller corrections fall out of the new return type:

  - the reader reports both "needle not in the index" and "volume vanished
    between the has_ec check and the snapshot" as Ok(None), which collapses
    the old Some(Ok(None)) and None arms into one 404;
  - an io::ErrorKind::NotFound now answers 404 rather than 500, matching the
    GET path. Telling a caller to retry a delete that can never succeed was
    half the bug.

The cookie check and its ordering before the journal append are unchanged.

Not addressed here: Rust journals the tombstone locally while Go routes it to
the primary shard holder. That is a separate behaviour change and belongs in
its own PR against the same issue-10 checkbox.

The regression test mounts 13 of 14 shards, leaving out the one holding the
needle's interval. The distributed reader seeds its Reed-Solomon buffers from
locally mounted siblings, so with >= 10 survivors it reconstructs with no peer
fan-out -- which makes the bug reproducible on a single node. Against the
unfixed handler the test fails with 500 vs 202.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: fail the delete when the EC volume unmounts mid-request

find_ec_volume_mut returning None used to fall through to a 202 with no
.ecj tombstone written, reporting success for a delete that did not
happen. Answer 404 like the other volume-vanished arms so the caller can
retry after a remount.

* volume server: forward EC needle deletes to a primary-shard holder

Mirror Go's doDeleteNeedleFromAtLeastOneRemoteEcShards: the tombstone is
journaled on one holder of the needle's primary data shard via
VolumeEcBlobDelete (or the local journal when this server holds the
shard), falling back to any other shard holder when the primary has
none. Journaling only on the node that received the DELETE scattered
tombstones across whichever server took the request.

* volume server: route BatchDelete EC deletes through the same forwarding

BatchDelete had the same local-journal divergence as HTTP DELETE, plus a
gap the old code admitted in a comment: the .ecx index cannot supply the
needle's cookie, so EC deletes ran with no cookie check at all. A
distributed read now fills the needle for every EC entry — matching Go's
DeleteEcShardNeedle, which reads and compares the fid cookie even when
skip_cookie_check is set — and the tombstone forwards via
delete_ec_shard_needle_distributed. A needle deleted between read and
journal reports 304 like Go's ErrorDeleted; a vanished volume reports
500 so the filer retries.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:31 -07:00
ac876eef21 rust volume: build outgoing gRPC clients in one place and give every dial a connect timeout (#11354)
server/grpc_client.rs stopped at build_grpc_endpoint() -> Endpoint, so all 13
production call sites hand-wrote the same .connect() + X::with_interceptor()
+ two max_*_message_size() lines. Four of them -- VolumeCopy,
VolumeTailReceiver, VolumeEcShardsCopy and the HTTP chunk batch-delete fan-out
-- dialed with no timeout at all, so an unreachable peer whose TCP handshake
never completes (SYN dropped, blackholed route, host behind a silent firewall)
left the operation waiting on the kernel's own retry budget, minutes long.

Add GrpcDialOptions (unary / long / stream presets), connect_channel(), and
volume_server_client() / master_client() / filer_client() constructors that
attach the request-id interceptor and lift both message-size limits, then
route all 13 sites through them. build_grpc_endpoint is private again, so
connect_channel is the only way out of the module and no call site can dial
without picking up a bound. Each site's existing timeouts are preserved
exactly; the four bare dials gain a 5 s connect timeout and nothing else. No
per-request deadline was added to any streaming call: Endpoint::timeout is a
per-request bound on time-to-first-response-headers for every request the
channel carries, so a value picked for one short call would also be the header
deadline for the whole-volume transfer sharing the dial.

The new bound covers the TCP handshake only -- tonic hands connect_timeout to
HttpConnector::set_connect_timeout. A peer that completes the handshake and
then stalls in the TLS or HTTP/2 exchange is still unbounded at those four
sites, as are the RPCs themselves. That is why the three ping_* helpers keep
their outer tokio::time::timeout: replacing it with connect_timeout would have
narrowed a whole-connect bound they already had.

main.rs no longer re-declares GRPC_MAX_MESSAGE_SIZE and the three
keepalive/window constants; it imports them from grpc_client.rs so the
inbound server and the outgoing clients cannot drift apart.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 19:41:37 -07:00
github-actions[bot] ddf009ffe7 docs: regenerate star history chart 2026-09-21 00:50:25 +00:00
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Fable 5.1Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
818f3bb71b rust volume: share the I/O-error tracker between Volume and EcVolume (#11351)
* rust volume: share the I/O-error tracker between Volume and EcVolume

Volume and EcVolume each carried the same three fields - a mutex-held
last error, a consecutive count and a sticky quarantine flag - and the
same four methods over them, identical except for the path qualifier on
is_storage_io_error. The tolerance the count is compared against was a
fourth copy: heartbeat.rs held VOLUME_IO_ERROR_TOLERANCE for volumes,
ec_volume.rs held IO_ERROR_TOLERANCE for EC, and the volume test helper
open-coded the same 3, so the two paths could drift apart silently.

Go keeps this in one place already: weed/storage/io_error.go holds
IoErrorTracker, IoErrorTolerance and isStorageIoError, and Volume embeds
the tracker. Go's EcVolume has to re-implement it only because those
fields are unexported and EC lives in another package.

storage::io_error::IoErrorTracker now owns that state, with record /
state / should_quarantine / mark_quarantined / reset and the single
IO_ERROR_TOLERANCE. is_storage_io_error moves into the same file, so it
sits with the tracker that is now its only caller, the way io_error.go
is laid out. Both volume kinds embed one tracker and keep their existing
method names as delegates, so the ~16 internal call sites and the
readers in heartbeat.rs, store.rs and grpc_server.rs change only where
the two threshold comparisons become should_quarantine().

Volume::last_io_error and EcVolume::reset_io_error_state had no callers
and are gone.

Unchanged: what counts as a storage-media error - is_storage_io_error
changed file, not body, and is still the single predicate both volume
kinds share, where Go's EcVolume tests EIO directly and so misses the
Windows codes. Also unchanged: the tolerance value, the metric increment
on every counted error, and the sticky quarantine - a success clears the
count and the last error but never the flag, which only reset lifts. In
the heartbeat the state read moved inside the quarantine branch, so the
common path no longer takes the tracker's mutex or clones the last-error
string; should_quarantine's two relaxed loads run either way.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: hoist absolute_display_path into server

handlers.rs and ui.rs each held a byte-identical copy of the helper that
turns a configured -dir into an absolute path for display. The status
JSON and the status page are meant to show the same directory, so the
two copies had to be edited together to stay that way.

The helper now lives in server/mod.rs as pub(crate) and both callers use
it. No behaviour change: same body, same call sites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: keep EcVolume::reset_io_error_state

Moving both volume types onto the shared IoErrorTracker dropped
EcVolume's public reset while Volume kept its own, so the two sides of
the tracker drifted apart.

mark_quarantined is sticky: a later successful read clears the error
count through record(), but the quarantine flag only comes down through
reset(). Without the delegate an EC volume that hit sustained media
errors could not be returned to service in place once the storage was
repaired. Go exposes the same method as EcVolume.ResetIoErrorState
(weed/storage/erasure_coding/ec_volume.go:114).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* rust volume: name the shared tracker after Go's IoErrorTracker

- check_read_write_error, get_io_error_state, mark_io_quarantined,
  reset_io_error_state match weed/storage/io_error.go one to one
- io_error module is pub(crate) like the io module beside it
- restore EcVolume::reset_io_error_state so both volume kinds expose the
  same recovery surface
- trim comments that restate the code

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 16:57:23 -07:00
7643f4f541 rust: a seaweed-common crate for the address and TLS helpers both crates carry (#11358)
* rust: a seaweed-common crate for the address and TLS helpers both crates carry

seaweed-volume and seaweed-worker are separate cargo trees with separate
lockfiles and no root manifest, so anything both of them need has had to be
written twice. Two of those copies are a correctness risk rather than a typing
cost, and this crate is where they stop being copies.

address.rs is the HTTP<->gRPC port rule: `host:port` means gRPC on port+10000,
`host:port.grpcPort` names it outright. The two copies had already drifted —
the worker's bracketed IPv6 literals, the volume server's did not — so the rule
lives here once, returning a typed AddressError whose Display text is the volume
server's original wording, with join_host_port public beside it. A test asserts
two of those messages in full rather than by substring, because the wording is
the contract its callers hand to a Status or an io::Error; the other three end
in a std ParseIntError message, which is std's to reword. The enum is
#[non_exhaustive] so a future variant is not a breaking change for either
consumer. The tests are both crates' cases together, plus the IPv6,
already-bracketed and normalisation cases neither copy covered on its own.

tls.rs is install_default_crypto_provider. Both binaries link aws-lc-rs and ring
transitively, so rustls cannot auto-select and tonic's client TLS panics on
first use; each binary has to pin one and it has to be the same one, which is
exactly the kind of choice that should not exist twice. It is safe to share
because `cargo tree -i rustls` resolves a single rustls in each tree (0.23.37 in
seaweed-volume, 0.23.43 in seaweed-worker) and cargo unifies all
semver-compatible `rustls = "0.23"` requirements into one crate per binary, so
this crate writes the same process-wide static its consumer reads. rustls is
already in both graphs — directly in the volume server, through tonic's
tls-aws-lc in seaweed-worker-core — so the dependency adds no crate to either.

rust-version is 1.91.1, the lower of the two consumers' floors, so depending on
this crate cannot raise either tree's MSRV; verified with
`cargo +1.91.1 check --all-targets`. The lockfile is committed even though this
is a library: CI builds it directly, so a committed lock is what makes those
runs reproducible and their caches stable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust: take the address and TLS helpers from seaweed-common

Both public signatures are kept, so no caller outside the two wrapper files
changes. parse_grpc_address stays `Result<String, String>` and maps the typed
error through Display; server_to_grpc_address stays `Option<String>` and drops
it with .ok(). Their doc comments and the volume server's 13 call sites are
otherwise untouched.

Three behaviours change, each in the direction of the copy that was already
right:

- The volume server now brackets IPv6 literals. `::1:19333` used to come back as
  `::1:29333`, which build_grpc_endpoint rejects with "invalid gRPC endpoint
  http://::1:19333: invalid authority" — an IPv6 master or EC peer could not be
  dialled at all. Two tests in grpc_client.rs pin it, one on the string and one
  on the endpoint the string builds.
- The volume server now emits the *parsed* gRPC port of the dotted form instead
  of the original text it had just validated, so `host:8080.018080` and
  `host:8080.+18080` come back as `host:18080` rather than as authorities the
  URI parser rejects. Same port either way; only malformed spellings change.
- The worker's dotted form now validates the HTTP port it discards.
  `server_to_grpc_address("host:abc.18080")` used to answer Some("host:18080");
  it now answers None, which is what the volume server's copy has always done.

install_default_crypto_provider becomes a re-export in both trees, so
`crate::security::tls::install_default_crypto_provider` and
`weed_lance_worker::tls::install_default_crypto_provider` still resolve. The
lance crate's `rustls = "0.23"` was its only direct use of rustls and goes away
with the body; seaweed-common states the same requirement, so neither the
resolved version nor the enabled features move in either lockfile.

The PEM test fixtures stay where they are. The two tests that use them are not
duplicates: the volume server's exercises build_grpc_endpoint, and the lance one
exists precisely because aws-lc-rs and ring are both linked in that crate's
graph. Only the literals are shared, and exporting test fixtures from a library
to dedupe two constants costs more than it saves.

A path dependency outside both trees means every build context that copies one
crate directory has to copy the other. The repo has one: the Rust source-build
stage of docker/Dockerfile.go_build, which now copies seaweed-common beside
seaweed-volume. Every workflow whose `paths:` filter keys on a crate directory
gains `seaweed-common/**` — the two Rust test workflows, rust_binaries_dev,
container_dev and performance. The tag- and dispatch-triggered ones
(rust_binaries_release, container_release_unified, container_latest) have no
`paths:` filter and need nothing.

The two Rust test workflows also run `cargo test` in seaweed-common, from their
unit-test job, because a path dependency is not a workspace member and neither
tree's own `cargo test` reaches it. Each step builds into its job's cached
target directory, and both cache keys now hash seaweed-common/Cargo.lock as well
so a change there invalidates the cache it would otherwise silently reuse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docker: keep go_build working for BRANCH revisions without seaweed-common

The rust_builder stage copies seaweed-common unconditionally now that seaweed-volume path-depends on it, but BRANCH can name any revision — including ones that predate the crate. Create the directory in the builder stage so the COPY always has a source; an empty dir beside an old seaweed-volume is harmless.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-20 13:30:58 -07:00
f1ed270942 rust volume: one S3 tier registry instead of two kept in sync by hand (#11357)
* rust volume: one S3 tier registry instead of two kept in sync by hand

`VolumeServerState.s3_tier_registry` and `global_s3_tier_registry()` held
the same S3 tier backends. `apply_storage_backends` — the only production
writer — registered every backend into both, and each half of the tiering
code then read a different one: the gRPC tier-move handlers resolved the
backend from the per-server field, while `Volume`'s remote mount and
destroy paths resolved it from the global registry, because a `Volume` has
no handle to the server state. Two registries that must agree, kept in
agreement by a duplicated `register_s3_backend` call and a comment in a
test constructor explaining the hand-sync.

Delete the field and let both tier-move handlers resolve from the global
registry, so `apply_storage_backends` registers once and no longer needs
the server state at all. Injecting a registry handle through `VolumeSpec`
instead was considered and rejected here: it would touch every `Volume`
constructor for no functional gain, and the process-wide registry is what
`Volume` already uses.

Behaviour is unchanged: the same names were registered in both registries,
so every lookup resolves exactly as before. The tier-down test now
registers its backend only in the global registry — before this change it
fails with `remote storage s3.tier_down_delete not found from supported:
[]`. The tier-up handler had no test at all, so it gets a cheap probe:
register a backend only in the global registry, ask for that destination,
and check the call gets past the lookup — the response is dropped straight
away, so the transfer sees a departed caller and never opens a connection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: await the tier-up probe terminal error instead of racing it

Dropping the response left it to chance whether the detached transfer saw the closed channel before its initial check; if it won that race it went on to attempt the multipart upload with no one waiting on the outcome. Hold the stream and read until the dead endpoint fails the upload — the terminal error proves the task ran and finished, so no background network work outlives the test.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-20 13:30:40 -07:00
7dbbdac030 rust volume: keep the EC shard-location map, its refresh time and stale mark under one lock (#11356)
The per-EcVolume shard-location cache was three fields under three locks:
an RwLock<HashMap> for the map, a Mutex<Option<Instant>> for the time it
was last refreshed, and a Mutex<bool> for the stale mark. Nothing tied
them together. merge_shard_locations published the merged map, released
the write lock, and only then stamped the refresh time; both readers
(scrub_ec_volume_distributed's snapshot and build_snapshot) took the two
guards one after the other. A reader landing between the two writes
paired a freshly merged map with the previous lookup's timestamp -- and
that pair is exactly what needs_refresh judges, so a read went back to
the master for a map that had just been refreshed. Go keeps the same
state in one struct behind one ShardLocationsLock. replace_shard_locations
documented itself as "a single observable step" while being two.

Fold the three fields into one ShardLocationCache behind a single RwLock.
merge_shard_locations upserts and stamps in one write section,
shard_locations_snapshot returns the map and its time from one read
section, and mark_shard_locations_stale / claim_shard_locations_refresh
move the mark's read-and-consume onto the cache. The three zero-caller
accessors -- set_shard_locations, replace_shard_locations,
get_shard_locations -- are deleted, and the field is now private, so the
invariant cannot be sidestepped from outside the module. The two test
seeding sites go through merge_shard_locations, which already produces
the state they were writing by hand.

Unchanged: the freshness rule. needs_refresh keeps its thresholds and
still judges the caller's snapshot -- the map that caller will actually
read from, not whatever is cached by the time the claim runs -- so only
the stale mark is read from under the new lock. The master lookup, the
completeness guard in write_back_shard_locations and the per-shard upsert
semantics are untouched.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:30:15 -07:00
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Fable 5.1Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
6d676eda67 rust volume: typed errors for store compaction so gRPC can answer NotFound (#11355)
* rust volume: typed errors for store compaction so gRPC can answer NotFound

The vacuum entry points on `Store` returned `Result<_, String>`, so the
gRPC layer had nothing to branch on and answered `Status::internal` for
every failure. A vacuum loop that races a volume being moved or deleted
saw the same code as a disk going bad, and `weed shell` could only tell
the two apart by matching on the message text.

`VolumeError` gains `VolumeNotFound(VolumeId)` — the existing `NotFound`
is needle-level and carries no payload — and `InsufficientSpace`, and
`compact_volume`, `commit_compact_volume`, `cleanup_compact_volume` and
`delete_collection` return it. `impl From<VolumeError> for tonic::Status`
in `server/mod.rs` maps not-found to `not_found`, read-only to
`failed_precondition`, insufficient space to `resource_exhausted`,
already-exists to `already_exists`, and everything else to `internal`;
the four RPCs prefix their own context with `status_with_context`, so a
message reads "commit compact volume 7: volume id 7 is not found". The
store-side "during compact" / "during commit compact" / "during cleaning
up" suffixes are gone, and the free-space message drops the volume id the
prefix already supplies.

`check_compact_volume` had no callers — `VacuumVolumeCheck` computes the
garbage level from its own `find_volume` — and is deleted. `compact_volume`
folded the size estimate into its first lookup, dropping the `unwrap()`
re-lookup that only existed to dodge a borrow.

`ascending_visit` on `CompactNeedleMap`, `RedbNeedleMap`,
`SortedFileNeedleMap` and the `NeedleMap` dispatch is now generic over the
visitor's error type, like `CompactMap::ascending_visit` already was. The
three signatures that can fail on their own bound `E: From<String>` to
carry those failures; the in-memory walk in `iter_entries` names
`Infallible`, which says in the type what its comment used to say in prose.

No Go shell command matches on the old error text: the strings exist only
in weed/storage/store_vacuum.go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments and answer the same codes from Go

- vacuum_volume_check reports VolumeError::VolumeNotFound like the other
  vacuum RPCs instead of its own "not found volume id" wording
- drop doc comments that restate what the code says
- Go volume server wraps ErrVolumeNotFound/ErrInsufficientSpace from
  store_vacuum.go so VacuumVolumeCheck/Compact/Commit/Cleanup and
  DeleteCollection answer NotFound/ResourceExhausted, matching the Rust
  volume server; volumeDeleteStatusError generalized to volumeStatusError

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: prefix operation context on vacuum errors

Lower-level errors forwarded by CompactVolume, CommitCompactVolume,
CommitCleanupVolume and DeleteCollection carry no volume id or operation
name. Wrap with %w so the status mapping still sees the sentinel chain,
matching the context the Rust server's status_with_context adds.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: map NotEmpty to FailedPrecondition, share mapper in VolumeDelete

Go's volumeStatusError maps ErrVolumeNotEmpty to FailedPrecondition; the
Rust Status conversion was missing it and volume_delete kept a hand-rolled
match. Route it through status_with_context like the vacuum handlers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 13:30:02 -07:00
d002481037 rust volume: derive has_remote_file instead of mirroring it (#11353)
Volume carried `pub has_remote_file: bool` next to `pub volume_info`,
and the bool was only ever the answer to `!volume_info.files.is_empty()`:
outside the two constructors, `refresh_remote_write_mode` was the single
writer. Both fields being public made the pair a convention rather than
an invariant. Every caller that touched `volume_info.files` — load_vif
twice, the tier-up handler, the tier-down handler and its rollback — had
to remember to call `refresh_remote_write_mode` afterwards, and a caller
that forgot would leave the volume advertising a write mode its .vif
contradicts, or serving a remote .dat through a writable needle map.

The bool becomes `has_remote_file()`, computed from the list, so it
cannot drift. `volume_info` becomes private with a `volume_info()`
reader, and edits to the reference list go through
`update_remote_files(|files| ...)`, which applies the closure and then
refreshes the derived write mode and the needle map. With no caller left
outside the module, `refresh_remote_write_mode` is private.

Unchanged: the refresh logic itself, the order of operations in both
tier handlers, and the tier-down rollback semantics. The rollback still
snapshots the removed reference before the refresh runs, restores it on
failure, and re-refreshes unconditionally on the error path — the second
`update_remote_files` call runs with a no-op closure when there was
nothing to restore, exactly as the old code re-ran the refresh whether
or not it had re-inserted a reference.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:29:47 -07:00
1df8c05bc3 rust volume: document every unsafe block and stop mutating the process env in tests (#11352)
Three production `unsafe` blocks carried no `// SAFETY:` comment at all
(`libc::fallocate`, `libc::sysinfo`, `libc::statvfs`), and nothing made
that an error: `clippy::undocumented_unsafe_blocks` is a `restriction`
lint, allow-by-default, and appeared nowhere in either crate. Turn it on
in `seaweed-volume`'s `[lints.clippy]` and in the worker workspace's
`[workspace.lints.clippy]`, then document what each block relies on.

`memory_status.rs` and `disk_location.rs` get their blocks narrowed to
the `zeroed()` and the libc call, so each comment sits next to the
operation it justifies and the arithmetic is outside the block. Both
turn the success test into an early return on failure; the casts, the
multiplication order and the values returned on either path are
unchanged.

The bigger problem was in `config.rs`'s tests. `with_temp_env_var` and
`with_cleared_security_env` called `std::env::set_var`/`remove_var`,
claiming soundness because every caller holds `process_state_lock()`.
That mutex only serialises the fourteen annotated tests in this module.
The same lib test binary runs the `grpc_server.rs` tests, which bind a
`TcpListener`, dial loopback and drive a multi-thread tokio runtime, and
tonic/hyper/rustls/aws-sdk all read the environment lazily on those
threads — which is exactly the race Rust 2024 made these calls unsafe
for. `restore_env_var` had no SAFETY comment at all. `#[serial]` would
not have helped: it serialises annotated tests, which the mutex already
did.

So the config layer no longer reads the environment implicitly. An
`EnvLookup<'a> = &'a dyn Fn(&str) -> Option<OsString>` is threaded from
the public entry points down to every reader — `HOME`, `USERPROFILE`,
the twenty-four `WEED_*` keys and `SEAWEED_WRITE_QUEUE`. `parse_cli` and
`parse_security_config` keep their signatures and pass `process_env`, a
thin wrapper over `std::env::var_os`; `resolve_config` becomes
`resolve_config_with_env` (private, one caller). Tests build one with
`fake_env` instead, so no test touches the real environment and every
`unsafe` in the module is gone.

`process_state_lock()` stays, with a smaller job: `set_current_dir` is
safe but still process-global, so the tests that move the working
directory are still serialised against the ones that read it. Tests
naming an explicit config file never reach that search and no longer
take the lock.

No production behaviour changes: the same keys are read in the same
order with the same precedence, and `env_string` reproduces
`std::env::var(key).ok()` — absent and non-UTF-8 both read as unset.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:29:07 -07:00
Chris LuandChris Lu 44ba070d83 helm: document admin.allowInsecureBind and drop stale comment
- README and values.yaml now describe the allowInsecureBind escape
  hatch alongside the non-loopback bind guard
- remove a PR reference from the CI test comment
2026-09-20 12:28:53 -07:00
062238bb5c helm: let admin.allowInsecureBind satisfy the admin auth render guard
PR #11185 made `weed admin` refuse to bind a non-loopback address
without -adminPassword or mTLS. PR #11228 added -allowInsecureNoAuth
as an explicit opt-out for operators who restrict admin access some
other way (e.g. a NetworkPolicy plus an authenticating reverse proxy).

The chart's render-time guard added by #11236 (admin-statefulset.yaml,
seaweedfs.admin.authEnabled) predates -allowInsecureNoAuth and only
recognizes password-based auth, so there was no values.yaml path to
express that choice: the chart would fail(...) even though the binary
itself would start fine with a warning.

Add admin.allowInsecureBind (default false) to the seaweedfs.admin.authEnabled
helper's checks; when true it renders -allowInsecureNoAuth on the admin
command and satisfies the render guard alongside the existing
password-based checks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 12:28:53 -07:00
Chris LuandChris Lu b77c42ff32 test: use BaseEndpoint for the checksum test client 2026-09-20 01:15:11 -07:00
Chris LuandChris Lu 26fc90187e s3: accept x-amz-checksum-mode from the query string, case-insensitively
Presigned HeadObject/GetObject requests hoist x-amz-checksum-mode into the
signed query string, so a strict header-only check would withhold stored
checksums on presigned reads that AWS honors.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu ef463fe1af s3: read complete-request checksum values from headers or query
Presigned CompleteMultipartUpload requests hoist x-amz-checksum-type and
the full-object checksum header into the signed query string, so a
header-only lookup would skip BadDigest validation for them.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu 17ad5a1419 s3: accept FULL_OBJECT checksums without per-part checksums at complete
COMPOSITE uploads must still carry every part checksum in the complete
request, but FULL_OBJECT uploads may instead supply the whole-object
checksum in an x-amz-checksum-* request header. Compare that header
against the computed object checksum and return BadDigest on mismatch,
matching AWS.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu b03419ee92 s3: reject UploadPart checksum algorithms conflicting with the upload
An UploadPart that explicitly selects a different checksum algorithm than
the one declared at CreateMultipartUpload would store a checksum
CompleteMultipartUpload could never accept. Reject the conflict up front
with InvalidRequest, matching AWS.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu f849b7c823 s3: validate per-part checksums in CompleteMultipartUpload
Parse the Checksum* elements of each completed part and enforce what AWS
does for uploads created with x-amz-checksum-algorithm: every part must
carry a checksum in the complete request (InvalidRequest when missing,
BadDigest when it differs from the stored part checksum), and an
x-amz-checksum-type header must match the upload resolved checksum type
(BadDigest). Add the issue-11401 reproduction as a regression test.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu f15b980976 s3: UploadPart inherits the checksum algorithm of its multipart upload
AWS computes a checksum for every part of an upload created with
x-amz-checksum-algorithm, even when the part request carries no checksum
headers. Mirror that: when the part request specifies no algorithm, apply
the one stored on the upload entry so the part entry keeps a checksum
CompleteMultipartUpload can fold into the object checksum.
2026-09-20 01:15:11 -07:00
hsdfatandGitHub 110b485bae fix(volume): stop ScanVolumeFileFrom at a header it cannot advance past (#11398)
fix(volume): stop scans at a header they cannot advance past

A corrupt .dat header with a very negative size gives a record length
(NeedleHeaderSize + NeedleBodyLength) of zero or less: v3 sizes -43..-36
and v2 sizes -35..-28 give exactly zero, and smaller sizes give a
negative length. ScanVolumeFileFrom advanced by that length, so it
re-read the same header forever or stepped back into the record before
it. weed fix, weed export, weed compact, incremental weed backup and the
tail sender behind volume.move and volume.merge could hang on such a
volume, and weed compact could also finish with a .cpx that had dropped
every needle after the header.

Return an error wrapping needle.ErrorCorrupted instead. The check runs
after the visitor has seen the record, so the rebuild scanner still
stops quietly with io.EOF. Smaller negative sizes whose record length is
positive are still stepped over, preserving the salvage behavior
compaction relies on.

Mirror the guard into the Rust volume scans: DatScanPlan::scan and
read_all_needles fail on a non-positive record length, as does
scan_dat_head, so a corrupt header cannot stall a tail pass or leave the
repair scan walking stale offsets.
2026-09-19 22:26:56 -07:00
06dda12e4b fix(volume): validate sizes in ReadNeedleBlob and WriteNeedleBlob (#11399)
* fix(volume): reject negative sizes in ReadNeedleBlob and WriteNeedleBlob

A ReadNeedleBlob RPC with a size of -44 or below (-36 on v2 volumes)
panics in makeslice inside needle.ReadNeedleBlob. The volume gRPC server
has no recovery interceptor, so one request kills the process. Smaller
negative sizes return bytes that are not a record.

WriteNeedleBlob accepted a negative size whenever the blob header
carried the same value: it appended the blob to .dat and indexed the
needle with that size, which reads as deleted.

Reject size < 0 in both Volume methods. Size 0 still passes, since
delete records carry it. The Rust volume server got the same storage
guards in #11345.

* fix(volume): reject needle blobs whose length does not match their size

WriteNeedleBlob appends the blob as is. A blob that is not the length
its size implies leaves .dat off the 8-byte grid, and every later
ordinary write to the volume is indexed at a truncated offset and reads
back as EOF. A blob off by 8 bytes keeps the grid but leaves bytes that
a .dat scan reads as the next record.

The in-tree callers already send exact lengths. The one case this newly
refuses is a copy between volumes of different needle versions, and
that case already writes a broken record: a v3 record lands on a v2
volume with 8 extra bytes, and a v2 record on a v3 volume either fails
the timestamp check or lands 8 bytes short.

This is separate from the negative-size guards, whose Rust counterpart
is #11345. The Rust server does not check the length yet.

* fix(volume): guard the blob buffer allocation in needle.ReadNeedleBlob

Volume.ReadNeedleBlob rejected negative sizes, but needle.ReadNeedleBlob
still sized its buffer from the size and is called directly by vacuum and
other paths. Reject a deletion marker before make() there too, and use
size.IsDeleted() in the volume-level checks.

* fix(volume): mirror the blob length check in the rust volume server

write_needle_blob_and_index checked the size against the blob header but
appended the blob verbatim, so a blob that is not the length its size
implies still leaves .dat off the record grid. Match the Go check.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 21:28:37 -07:00
a93a1ab2eb fix(volume): return an error instead of 201 when a write lands on no volume (#11397)
* fix(volume): return an error instead of 201 when a write lands on no volume

ReplicatedWrite only writes locally when this server holds the volume.
For a volume id no server holds, the master lookup returns no locations,
so the write went nowhere and the upload still got 201 Created. The same
happened for a type=replicate write to a server without the volume, so
the primary, or the S3 chunk fan-out, counted a replica that was never
written.

A server without the volume still forwards the write to the replicas the
master lists. When there is nothing to forward to, fail with "volume N
not found on host:port". PostHandler returns that as 500, the status the
Rust volume server already returns here, and uploaders re-assign on 5xx.

Fixes #6609

* volume: reuse Store.HasVolume, drop issue ref from test comment

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 19:03:21 -07:00
cd1e738422 [Volume] Scrub local deletion tombstones during FULL scrub (#11396)
* fix 11388

* fix(volume): scrub validates local deletion tombstones

TombstoneFileSize (-1) is an .idx-only sentinel; the physical record it
points at carries a zero-sized body. Normalize deleted index sizes to 0
via onDiskSize before computing disk usage and calling ReadData, so
corrupted or truncated tombstone records are detected instead of
skipped. Offset-zero entries (remote logical deletes, no .dat record)
remain skipped, and the physical needle id is checked against the index
key. Mirror the behavior in the Rust volume server.

* fix(volume): scrub preserves physical size of deleted non-tombstone entries

Size.Raw()/raw() already encodes the index-to-disk mapping: tombstone
(-1) -> 0, other negative sizes -> their absolute value (the offset then
points at the original record, per the ReadDeleted path). Use it instead
of mapping every deleted size to 0.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 18:55:27 -07:00
github-actions[bot] 3dec359d6b docs: regenerate star history chart 2026-09-20 00:50:27 +00:00
Chris LuandGitHub f6a3286b32 fix(volume): derive needle body tail bound from the version layout (#11395)
* fix(volume): derive needle body tail bound from the version layout

The size guard in ReadNeedleBodyBytes computed the tail length as
checksum, plus timestamp only for Version3. Forks and future on-disk
formats whose tail carries more fields would silently under-check and
still panic in readNeedleTail on a truncated body. Derive the tail from
NeedleBodyLength minus data and padding so the bound stays exact for
every version.

Iterate IsSupportedVersion in the new tests instead of hardcoding
v1-v3 so downstream formats get covered automatically, and skip
versions the build cannot write rather than failing on them.

* test: skip needle write only on the unsupported-version error

A blanket skip would hide a real writer regression. Skip the version
subtest only when the writer reports the version is not supported in
this build (the error text differs between builds), and fail on any
other write error.
2026-09-19 12:20:01 -07:00
Chris LuandGitHub 01bb3b3053 s3api: add Snowflake s3compat API integration tests (#11394)
* s3api: add Snowflake s3compat API integration tests

Run the upstream snowflakedb/snowflake-s3compat-api-test-suite against a
local SeaweedFS server in CI. test/s3/snowflake/run.sh starts weed server
with S3 (-s3.autoCreateBucket=false so missing-bucket PUTs return
NoSuchBucket), prepares the fixtures the suite needs (versioned bucket,
deny-all-policy bucket, >1000-object prefix), clones the suite, patches
it to path-style addressing, and runs mvn -Dtest=S3CompatApiTest.

The suite also exposed that GetBucketLocation returned 404 NoSuchBucket
for a malformed bucket name; validate the name first and return
400 InvalidBucketName like AWS.

* test: harden snowflake s3compat runner per review

- Pin the upstream suite to a tested commit (SUITE_REV) instead of the
  moving default branch
- Bind the test server to loopback only
- Require the AccessDenied error code when verifying the denied bucket
- Fix README so go install runs in a subshell
- checkout with persist-credentials: false
- Make the concurrency group unique per PR, and widen path filters to
  the storage/operation/wdclient/cluster/pb packages the S3 stack uses

* test: advertise loopback ip for snowflake test server

-ip.bind 127.0.0.1 alone left the volume server advertising the host's
primary address, so chunk uploads were refused. Also set -ip 127.0.0.1
and disable the Iceberg/Lance listeners so the harness is loopback-only
and does not collide with other local services.
2026-09-19 03:29:51 -07:00
hsdfatandGitHub 5769057af3 fix(volume): return an error instead of panicking on a corrupt needle size (#11393)
ReadNeedleBodyBytes sliced the needle body with the size from the needle
header without checking it. A corrupted .dat header carrying size -1 still
gets a positive body length (16 bytes on v3), so vacuum compaction read
that body and panicked with "slice bounds out of range [:-1]".

Writers never put a negative size in a .dat header: a delete appends a
size-0 record, and TombstoneFileSize only lives in the .idx. Reject a size
that is negative or leaves no room for the checksum/timestamp tail with an
error wrapping ErrorCorrupted. ScanVolumeFileFrom already logs body read
errors and moves on, so compaction now skips the record like any other
corrupt needle.

Fixes #6763
2026-09-18 23:00:06 -07:00
github-actions[bot] 4160b92864 docs: regenerate star history chart 2026-09-19 00:45:06 +00:00
Chris LuandGitHub 37bf1cd91d volume: validate copy/tail source addresses before dialing (#11390)
* pb: stop exiting the process on malformed server addresses

ServerToGrpcAddress and GrpcAddressToServerAddress called glog.Fatalf
when hostAndPort could not parse the port, which os.Exit(255)ed the whole
process. A caller-supplied copy or tail source address reached this path
synchronously in the serving goroutine, so one anonymous VolumeCopy with
a non-numeric port terminated the volume server.

Log the parse error and return the input unchanged instead: the dial or
request that consumes the address then fails as an ordinary error.

* volume: validate copy and tail source addresses before dialing

VolumeCopy, VolumeEcShardsCopy and VolumeTailReceiver dial a
caller-supplied source address (SourceDataNode / SourceVolumeServer)
with no endpoint validation, so an anonymous caller could aim the volume
server at loopback, link-local (cloud metadata) or other unintended
destinations and read dial behavior back as a connectivity oracle.

Apply the same peer-target deny list FetchAndWriteNeedle uses for
replica targets: the source must be a bare host:port whose host is not
loopback, link-local or unspecified; cluster peers stay reachable on
private networks, and -volume.allowUntrustedRemoteEndpoints opts out.
The loopback-using copy tests set the flag to keep exercising the copy
path in process.

* rust volume: validate copy and tail source addresses before dialing

Mirror the Go guard on the Rust volume server: volume_copy,
volume_ec_shards_copy and volume_tail_receiver dial a caller-supplied
source address, so run it through validate_replica_target first (bare
host:port; no loopback, link-local or unspecified hosts; private peers
stay allowed). --volume.allowUntrustedRemoteEndpoints opts out; the test
fixture and the Rust test-cluster launcher set it so loopback sources in
tests keep working.

* volume: pin validated copy/tail source addresses at dial time

validateReplicaTarget resolves the source hostname once, but the gRPC
client resolved it again at connect, leaving a DNS-rebinding window for
hostname sources. The copy and tail source dials now run through the
same guardedDialerPolicy the remote-storage path uses, so every resolved
address is re-checked against the replica deny list (private peers
allowed) immediately before the TCP connect. guardedDialerPolicy also
moves to util.OutboundDialContext so the guarded path keeps the -ip.bind
source binding the default gRPC dialer had.

The Rust volume server mirrors this with connect_guarded, a tonic
connector that resolves, re-checks each address, and connects to the
first passing IP; handlers use it whenever the untrusted-endpoint
opt-out is off. A handler-level test now exercises the enabled
validation branches for all three source-taking RPCs.

* pb: return empty server address for malformed grpc addresses

GrpcAddressToServerAddress used to return the unparseable input on a
hostAndPort failure, so a malformed raft address (e.g. "host:abc")
flowed into admin dashboard master maps unchanged. Return an empty
string instead, skip empty conversions at the two raft-cluster merge
sites, and drop the now-stale comment about the fatal exit the earlier
commit removed.

* test: opt erasure-coding loopback clusters out of the remote endpoint guard

The erasure-coding suites drive VolumeEcShardsCopy / VolumeCopy between
volume servers bound to 127.0.0.1, which the copy/tail source guard now
rejects by default. Pass -volume.allowUntrustedRemoteEndpoints to the
test volume launches, matching what the volume_server framework
harnesses already do.

* admin: only claim fallback master leadership on an empty raft response

A nonempty RaftListClusterServers response whose entries were all
rejected left masterMap empty, so the fallback marked the reachable
current master as leader the same way a genuinely empty (non-raft)
response does. Track whether the successful response returned zero
servers and only promote the fallback master then.
2026-09-18 12:55:47 -07:00
Chris LuandGitHub a6d72bc272 s3api: delete orphaned chunks only when the entry is confirmed absent (#11389)
* s3api: test for chunks deleted under an entry the filer committed

Issue #11387: the filer can report a create failure after inserting the
entry (e.g. a parent-directory creation failing post-insert). The error
arrives in the response rather than as a transport status, so it maps
to a definitive error and putToFiler deletes the chunks of the live
entry.

* s3api: confirmCreateLanded also reports a confirmed-absent entry

The verification a failed create runs can answer both directions: the
entry matching the uploaded chunks proves the write landed, and an
authoritative not-found proves the uploaded chunks are orphaned. Return
both outcomes so the cleanup path can gate on the fact rather than the
error class. An empty upload can never prove a landing, so a zero-chunk
entry match no longer upgrades the outcome.

* s3api: delete orphaned chunks only when the entry is confirmed absent

A failed create no longer skips verification based on the error class: the filer can fail after inserting the entry (issue #11387) and a partially-applied routed transaction can leave it behind too, both surfacing as definitive errors. Every failed create now resolves the entry's fate, and the uploaded chunks are deleted only when the entry is confirmed absent; anything unverifiable keeps them for vacuum.

* s3api: confirm absence on every filer the create could have committed on

A lock-path create fails over across filers, so the entry can live on a replica the routed owner has not caught up to; one not-found does not prove absence. The confirmation now queries the owner, the prior owner, and the failover set, declaring absent only when none of them has the entry.

* s3api: bound the reconciliation lookups confirmCreateLanded runs

The lookups ran on context.Background() under the object write lock, so a connected filer that never replies could stall the write path. One timeout now covers the whole enumeration; an expired budget fails the remaining lookups as uncertain, which keeps the chunks.
2026-09-18 12:30:07 -07:00
Chris LuandGitHub c72eda50a8 s3: drop implicit reader cache budget that throttled S3 GETs (#11384)
* fix(filer): leave reader cache unbounded without an explicit budget

NewReaderCache silently installed a 256MiB ReaderCacheBudget when the
caller passed none. Only weed mount opts into a budget; every other
caller (S3 gateway, WebDAV, query engine, mq logstore) inherited the
cap. Under ~90 concurrent S3 GETs of medium objects, prefetch wants
far more than 64 chunk buffers, so reserve() serialized chunk fetches,
clients timed out and retried, and the retry re-downloaded chunks the
cancelled request had already fetched.

A nil budget now means unbounded, restoring the pre-4.47 behavior for
callers that never asked for a memory cap; reserve/complete/release
are nil-safe. The mount path is unchanged and still enforces
-readerCacheSizeMB.

Fixes #11380

* feat(s3): expose -s3.readerCacheSizeMB reader buffer budget

Operators who want the S3 gateway read path memory-bounded can now
opt in: -s3.readerCacheSizeMB on weed filer/server/mini and
-readerCacheSizeMB on standalone weed s3, matching the mount flag.
The default 0 keeps the unbounded pre-4.47 behavior; a positive value
installs a shared ReaderCacheBudget across in-flight and retained
chunk buffers for all S3 GETs.

* fix(filer): validate chunk size before consulting the reader budget

A nil budget returned early and skipped the negative chunkSize check,
letting a corrupted size reach mem.Allocate and panic. Also drop the
command-specific flag prefix from the S3 validation error since
standalone weed s3 exposes the option as -readerCacheSizeMB.

* filer: drop chunk buffers once fully consumed

ReaderCache retained every completed chunk buffer in the downloaders
map until the slot limit evicted it, so buffers lingered after all
readers finished with them.

Track attached readers on each SingleChunkCacher and remove the cacher
when the last reader consumes the buffer to its end. In-flight download
deduplication and the prefetch handoff are unchanged: a buffer always
survives until fully read, partial reads keep it available, and an
attached reader pins a consumed buffer until it detaches. Repeat reads
now go through the chunk cache where enabled, or refetch.

* filer: drop consumed buffers on last detach, rechecked under cache lock

Two review findings on the drop-on-consume change:

- Removal only fired when the detaching reader itself reached the chunk
  end. If the end-reaching reader finished first and the last remaining
  reader did a partial read or cancelled, the consumed buffer and its
  budget reservation lingered until eviction. Track a persistent
  consumed flag instead, so any end-reaching read marks the buffer and
  the last detach drops it.
- remove() checked only map identity, so a reader attaching between the
  reader count hitting zero and removal could attach to a cacher that
  was then deleted underneath it. removeConsumed() re-checks identity,
  readers == 0, and consumed under the ReaderCache lock; a raced attach
  keeps the cacher and its own detach retries the removal.
2026-09-18 01:02:32 -07:00
Chris LuandGitHub 87ee3b63a2 s3: abort completed multipart uploads metadata-only (#11385)
* s3: abort a completed upload's leftover directory metadata-only

A .uploads/<id> directory can outlive the object it completed into when
the commit's metadata-only removal failed or the gateway died in between;
the restored part entries then share chunks with the published object.
AbortMultipartUpload deleted the directory recursively, chunks and all,
so aborting such a leftover destroyed a committed object (#11382).

Run the same check s3.clean.uploads gained in #11375 before deleting:
when the object entry or a version file under <key>.versions carries the
upload id, remove .uploads/<id> metadata-only and answer the abort; when
the lookup cannot decide, refuse with InternalError rather than risk
live chunks.

* s3: apply the completed-upload check to lifecycle MPU abort

lifecycleAbortMPU ran the same destructive recursive delete on
.uploads/<id>. Reuse uploadCompleted so a leftover whose object entry or
version file carries the upload id is removed metadata-only, and an
undecidable lookup retries later instead of freeing live chunks.

* s3: serialize abort's upload-dir delete with the object's commit

The completed check alone leaves a race: abort can read completed=false,
then an in-flight completion publishes the object over the same part
chunks before the recursive delete frees them.

Run the check and delete inside the object write lock, which non-routed
completions hold for their whole finalize. With an owner, send the data
delete as an ObjectTransaction on the object's lock key — a routed
commit then either loses its upload-exists precondition after our delete
or has already stamped the object, which the transaction's
IF_EXTENDED_NOT_EQUAL condition detects and falls back to a
metadata-only remove. lifecycleAbortMPU shares removeUploadDir so both
callers get the same ordering.

* s3: check for an empty object before resolving its write owner

* s3: check completion at the abort's resolved object key

An upload record missing ExtMultipartObjectKey skipped the completed
check entirely even though the request's Key names the object.
2026-09-18 01:01:52 -07:00
Chris LuandGitHub 0ca1c19821 s3api: unify auth error handling across s3tables, iceberg and lance (#11381)
* s3api: fail closed when S3 Tables signature verification fails

* s3api: avoid nil Account dereference in S3 Tables auth log

* iceberg: return auth error instead of falling back to DefaultAllow

* lance: return auth error instead of falling back to DefaultAllow

* s3api: stop trusting client-supplied s3-account-id

The header is set by the server after successful authentication; scrub
inbound values alongside the other internal headers, and apply the same
admin guard to the header fallback branch of getAccountID that the
identity branch already has.

* test: cover table-catalog auth wrappers and principal resolution

* test: configure anonymous identity where catalog clients do not sign

* s3api: scrub s3-account-id after signature verification
2026-09-18 01:01:04 -07:00
Chris LuandGitHub f40687b34e s3: tighten STS session token handling (#11383)
* s3api: test that a session token must not reveal its credential

* sts: derive secret access key with HMAC keyed on the signing key

* s3api: stop accepting STS session tokens as bearer credentials

* security: reject STS session tokens on filer and admin gRPC auth

* test: sign s3/iam framework requests with the session credential

* s3api: exercise the real auth pipeline in the end-to-end harness
2026-09-17 23:50:16 -07:00
Chris LuandGitHub 15520f601f s3: commit multipart upload and remove .uploads atomically; purge completed uploads metadata-only (#11375)
* s3: commit versioned multipart upload in one transaction

CompleteMultipartUpload wrote the version file, flipped the .versions
pointer, then removed .uploads/<id> metadata-only as a best-effort
post-commit step. A filer error or gateway crash in that window left the
upload directory referencing the same chunks as the published object, and
the next s3.clean.uploads run purged it with data -- corrupting a
committed object.

Put the version file, remove the upload directory metadata-only (its
chunks are the object's chunks), and recompute the latest pointer in one
ObjectTransaction under the object's per-path lock on the owner filer.
The mutation order keeps every partial state safe: the chunks stay
referenced at all times, and a published object never coexists with the
upload directory the cleaner would purge.

Unused part entries are freed before the transaction, since the
metadata-only directory delete would otherwise leak their chunks.

* s3: remove upload directory inside the multipart object PUT

The same committed-object/stranded-upload window existed on the
suspended and non-versioned paths: writeMultipartObject committed the
object, then a best-effort rm dropped .uploads/<id>. Ride the
metadata-only removal on the routed PUT itself so the two land in one
transaction; the unrouted mkFile fallback keeps post-commit cleanup.

* shell: purge completed uploads metadata-only in s3.clean.uploads

A leftover .uploads/<id> can outlive a committed object when the
completion's metadata-only delete fails or the gateway dies in between;
its part entries then share chunks with the live object, and a recursive
purge frees them out from under it.

Before purging a stale upload, check whether it completed: the object
entry or any version file under <key>.versions carrying the upload id.
If so, delete with skipChunkDeletion. If the lookup fails, skip the
upload for this run rather than risk live chunks.

* s3: abort multipart completion when unused part cleanup fails

Deleting the upload directory metadata-only erases the only metadata
pointing at part entries whose deletion failed, orphaning their chunks.
Propagate the error so the completion fails while the upload directory
still exists and the request remains retriable.

* s3: require the upload directory to exist at multipart commit

A delete that does not take the object lock (abort, lifecycle,
s3.clean.uploads) can remove .uploads/<id> and its chunks between the
prepare step and the commit transaction. The commit now carries an
IF_EXISTS precondition on the upload directory so the race fails the
request with NoSuchUpload instead of publishing an object over freed
chunks.

* s3: keep the version file when the upload directory is gone

The finalize transaction has no rollback, so a failure at the
latest-pointer recompute leaves the version written and .uploads/<id>
removed. Deleting the version then destroys the only remaining record of
the upload, making a retried CompleteMultipartUpload return NoSuchUpload
while the version's chunks leak. Roll back only while the upload
directory survives; otherwise keep the version, which a retry resolves
through SeaweedFSUploadId and the version reconciler promotes.

* s3: keep manifests when a routed object write partially commits

For non-versioned and suspended completions the object PUT precedes the
upload-directory DELETE, so an error can mean the object entry exists
while the response reports failure. Freeing this attempt's manifest
chunks then destroys the committed object. Keep them when the object
entry survived, and after a failed null-marker finalize which always
follows a committed write.

* s3: skip the keep-version path on precondition failure

A rejected precondition means no mutation ran, so there is no version
file to preserve and this attempt's manifests are orphans the error
cleanup should free.

* s3: keep manifests when the object-existence check itself fails

A transient lookup error previously read as absent, letting the error
cleanup free manifest chunks a committed object still references.

* s3: keep the upload directory when post-commit part cleanup fails

Removing it metadata-only after a failed entry delete erases the only
reference to the leftover chunks. Leave the directory so the entries
keep their chunk references for s3.clean.uploads or manual recovery.

* pb: fix filer list entry counting on 32-bit

int(limit) wraps to -1 on 386 when limit is math.MaxUint32, so the
beyond-limit check discarded every streamed entry. Compare in uint64
instead; the semantics are unchanged on 64-bit platforms.

* shell: resolve trailing-slash object keys in s3.clean.uploads

Completion stores a key ending in / inside the directory it names
(<bucket>/dir/dir), but FullPath+DirAndName on the normalized key
looked one level too high. Deriving dir and name with path.Dir and
path.Base mirrors getEntryNameAndDir so the completed-upload check
finds the entry instead of purging its chunks.

* s3: heal a suspended completion hidden behind a delete marker

Removing .uploads/<id> inside the commit transaction means a failed
finalizeSuspendedNullWrite leaves nothing to retry against: the object
entry is committed but the marker still makes the key read as deleted,
and a retried CompleteMultipartUpload can only report NoSuchUpload.
When the upload directory is gone, check the regular path for an entry
carrying the upload id and re-run the marker finalize, so the retry
both succeeds and repairs the key. Only suspended buckets can hold
this state; anything newer owns the key.

* s3: report store errors when resuming a committed multipart object
2026-09-17 21:09:21 -07:00
Chris LuandGitHub bdc37a1e86 mount/shell: bucket allow-empty-folders toggle, mount keeps explicit false (#11370)
* mount: keep a deliberate bucket allow-empty-folders setting

* shell: s3.bucket.allowEmptyFolders toggles the empty folder cleaner

* shell: guard allow-empty-folders toggle with expected extended attrs

* filer: drop cached empty-folder policy on bucket entry update

* mount: guard allow-empty-folders write with expected extended attrs

* filer: skip caching a stale cleanup policy read across an update

* mount, shell: snapshot the full extended attributes for update preconditions

* filer: fail closed and invalidate on all bucket entry events for cleanup policy

* filer: key the cleanup policy generation by bucket

* filer: requeue cleanup when the bucket policy cannot be loaded

* filer: expire idle cleanup policy generations

* filer: skip requeueing cleanup after the cleaner stops

* filer: bound cleanup retries on repeated policy failures

* filer: cover cleanup requeue on repeated policy failures

* filer: keep cleanup policy generations while reads are in flight

* filer: exercise the cleanup queue lifecycle in the retry-cap test
2026-09-17 20:30:53 -07:00
Chris LuandGitHub 2d2619f0b4 ci: make telemetry deploy work on Oracle Linux 7 (#11377)
* telemetry/server: tidy module dependencies

* ci: make telemetry deploy work on Oracle Linux 7

* ci: install telemetry unit and logrotate on every deploy

* ci: abort telemetry deploy on install failures
2026-09-17 20:00:25 -07:00
Chris LuandGitHub ce1e0dc30a s3api: don't delete chunks when CreateEntry outcome is ambiguous (#11376)
* s3api: map ambiguous filer transport errors to retryable 503

Canceled, DeadlineExceeded and Unavailable can be returned after the
filer applied the write, so the outcome is ambiguous. Reporting them as
a 4xx tells the client not to retry; report ServiceUnavailable instead.

* s3api: verify entry existence before deleting orphaned chunks

A failed CreateEntry can still have landed on the filer when the error
is a transport failure, and entryCreated=false would tombstone chunks a
live entry references, leaving a dangling pointer that survives only
because reads pass readDeleted=true until vacuum reclaims the needle.

Before deleting, look the entry up: if it is stored with the same
chunks, the write succeeded; if the lookup cannot be answered, keep the
chunks for vacuum to reclaim; only a confirmed absence still cleans up.

* s3api: regression tests for ambiguous CreateEntry outcomes

Covers the three post-create-failure cases in putToFiler: the entry
landed despite the error (treat as success, keep chunks), the entry is
confirmed absent (delete orphans), and the outcome is unverifiable
(keep chunks, return error).

* volume: count reads served from deleted needles

A readDeleted read succeeding on a tombstoned needle is the signal that
metadata still points at deleted data. Count it under a
readDeletedNeedle handler label in both the Go and Rust volume servers
so the condition is visible before vacuum turns it into a 404.

* s3api: never delete chunks on an ambiguous create error

Review feedback on the first fix showed verification could still go
wrong in both directions: a stale or lagged lookup could report
not-found for a committed entry, a prefix object stores its chunks on a
directory entry, and filer-side manifestization rewrites the top-level
chunk ids the comparison relied on.

Rework the rule so the outcome classes are asymmetric:

- A transport-level error (anything filerErrorToS3Error maps to a
  retryable 503) is ambiguous and never deletes chunks; the lookup can
  only upgrade the write to success.
- Any other error is a definitive filer refusal and still cleans up.

confirmCreateLanded asks the write owner first, resolves the stored
entry through chunk manifests, requires an exact match of the uploaded
file ids, and on success runs the finalize callback the failed create
skipped (under the object write lock, with the same rmObject undo the
create path uses). Zero-chunk writes stay ambiguous since they cannot
be told apart by chunks.

* s3api: cover definitive refusals and stale entries in put tests

The confirmed-failure case now uses a definitive refusal so it still
exercises orphan cleanup, and a new case keeps chunks when the stored
entry belongs to an older object rather than this PUT.

* volume: count deleted-needle reads once per request

Streamed Go reads ran the deleted check in readNeedle and again in
readNeedleDataInto, and non-streamed Rust reads in stream_info and the
full-read fallback, double-counting one request. Count at the single
entry probe each implementation takes per GET: readNeedle in Go,
read_needle_stream_info in Rust.

* s3api: run recovered-write rollback under the object lock

Two follow-ups from review: ResolveChunkManifest returns traversed
manifest blobs in its manifestChunks output, so requiring it empty
rejected every manifestized landing; and the rmObject undo ran after
the object write lock was released, so a concurrent newer write could
be deleted between finalize failure and rollback. Compare only the
resolved data chunks and keep the undo inside the lock.

* s3api: verify, finalize and roll back recovered creates in one lock

A lookup done before the object write lock let a concurrent PUT replace
the entry between the chunk comparison and the finalize/rollback
section, so a failed afterCreate could rmObject a newer write. Run the
owner lookup, manifest resolution, chunk comparison, afterCreate and
the conditional undo inside a single withObjectWriteLock section.
2026-09-17 19:58:49 -07:00
github-actions[bot] d4e11a471d docs: regenerate star history chart 2026-09-18 01:41:46 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
08d5daf0c1 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc from 1.44.0 to 1.45.0 (#11371)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:16:38 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
8d34433308 build(deps): bump go.opentelemetry.io/otel/exporters/zipkin from 1.36.0 to 1.45.0 (#11373)
build(deps): bump go.opentelemetry.io/otel/exporters/zipkin

Bumps [go.opentelemetry.io/otel/exporters/zipkin](https://github.com/open-telemetry/opentelemetry-go) from 1.36.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.36.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/zipkin
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:10:51 -07:00
Chris LuandGitHub 4fd67001d9 security: require go 1.26.6 and bump vulnerable deps (#11374)
* security: require go 1.26.6 and bump vulnerable deps

A dependency scan of the 4.47 release flagged the bundled toolchain
and modules:

- github.com/golang/go < 1.26.6 (CVE-2026-39821, CVE-2026-56853,
  CVE-2026-56859, CVE-2026-56862, CVE-2026-56864, CVE-2026-56865,
  CVE-2026-33818, CVE-2026-46600): raise the go directive to 1.26.6
  so every built artifact requires the fixed toolchain.
- google.golang.org/grpc (CVE-2026-84445, CVE-2026-84304): move to
  the fixed dev pseudo-version; released tags through v1.85.0-dev
  remain in the affected range.
- github.com/pelletier/go-toml/v2 <= v2.4.2 (unbounded parser
  recursion): v2.4.3.
- alpine libcrypto3/libssl3 < 3.5.8-r0 (CVE-2026-75803,
  CVE-2026-63073, CVE-2026-63075, CVE-2026-63076, CVE-2026-63072,
  CVE-2026-54874, CVE-2026-18798, CVE-2026-14456, CVE-2026-14457):
  the release images already apk-upgrade the final stage; extend the
  same to the telemetry and admin-integration images.

Same bumps applied to the test/kafka, test/sftp,
kafka-client-loadtest, and telemetry/server modules.

* telemetry: send integration test report above the 10 GiB floor

The collect endpoint keeps reports only when TotalDiskBytes >=
proto.MinDiskBytes, but the integration test still sent 1 GiB, so the
server counted the report and skipped storing it. No cluster_id series
was ever created and /metrics lacked seaweedfs_telemetry_volume_servers.
Send just above the floor (via proto.MinDiskBytes so it cannot silently
drift again) so the expected per-cluster metrics are exported.
2026-09-17 15:09:55 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c6a3280595 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace from 1.44.0 to 1.45.0 (#11372)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-17 14:36:15 -07:00
799c495226 rust volume: one positional read helper; never seek a dup'd handle on Windows (#11342)
* rust volume: one positional read helper; never seek a dup'd handle on Windows

Positional read-exact was hand-rolled four times: the complete
cross-platform version in needle_map/sorted_file.rs, a Windows-only half
in volume.rs whose unix half was inlined as a
cfg(unix)/cfg(windows)/compile_error! triple at three call sites, a
byte-identical Windows-only copy in ec_volume.rs, and read_full_at in
ec_bitrot.rs. Three more sites -- EcVolumeShard::read_at,
EcLocalShard::read_at and ec_encoder::read_at_most -- hand-rolled the
short-read-permitted variant with a cfg(not(unix)) arm that
try_clone()s the handle and seeks it.

That last arm is wrong. A duplicated descriptor shares one kernel file
offset with the original, so seek-then-read is two syscalls against
state another thread can move in between: a concurrent reader or an
append repositions the offset and the read returns bytes from somewhere
else entirely. EcLocalShard::read_at documents that it must never seek,
one line above the seek. Windows seek_read carries its own offset in a
single call, so that window does not exist.

All seven now go through storage::io::{read_exact_at, read_at}, whose
module doc records why duplicating a handle is not a way to get a
private file position -- opening the file again is, as
Volume::dat_scan_plan already does. read_at_most keeps its own
fill-until-EOF loop; only the per-iteration positional read changes.

Behaviour on unix is unchanged: every unix arm was already
FileExt::read_exact_at or FileExt::read_at. The one exception is
ec_bitrot::verify_shard_blocks, which now retries on EINTR (std's
read_exact_at does; the loop it replaces did not) and, on unix, reports
the standard "failed to fill whole buffer" text instead of "short read
on shard block". The Windows arm still says "unexpected EOF in
seek_read"; both carry ErrorKind::UnexpectedEof, as before.

NeedleStreamSource::read_exact_at and Volume::read_exact_at_backend keep
their signatures; only their bodies shrink.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: retry Interrupted in Windows read_exact_at

Unix std's FileExt::read_exact_at ignores ErrorKind::Interrupted and
retries, but the Windows seek_read loop propagated it, so the shared
exact-read contract differed by platform. seek_read can surface
ERROR_OPERATION_ABORTED, which std maps to Interrupted.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 11:47:38 -07:00
Chris LuandGitHub 4ec564469a s3tables: hide inaccessible catalog resources (#11365)
* s3tables: hide inaccessible table buckets

* s3tables: hide inaccessible namespaces

* s3tables: hide inaccessible tables

* s3tables: hide inaccessible resources in rename and namespace delete

RenameTable/RenameView denied on the source now report the same
not-found as a missing source, and the destination name conflict is
checked only after destination authorization so a denied caller cannot
distinguish an existing destination namespace or name from a missing
one. DeleteNamespace denials use the same formatted message as a
missing namespace.
2026-09-17 11:43:52 -07:00
Chris LuandGitHub 66f1754896 s3: enforce dedicated Object Lock actions (#11362)
s3: enforce dedicated object lock actions
2026-09-16 20:34:06 -07:00
Chris LuandGitHub 994e1f7d64 admin: replace Font Awesome with MIT-licensed icons (#11364)
admin: replace Font Awesome with MIT icons
2026-09-16 20:29:14 -07:00
github-actions[bot] 74eeac6b66 docs: regenerate star history chart 2026-09-17 00:46:37 +00:00
0eb638f503 fix(ec): BatchDelete cookie fail-closed via locate_data geometry (#11348)
* fix(ec): BatchDelete cookie fail-closed via locate_data geometry

* fix(ec): honor skip_cookie_check, require full cookie header

* fix(ec): retry short cookie header reads, still fail closed on EOF

* chore(ec): trim cookie validation comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-16 16:25:03 -07:00
David ChristopherandGitHub 1a285c1334 filer: join shutdown paths before closing metadata store (#11363)
Serve can return when its listener closes while HTTP requests are still draining. The main path could then close the metadata store before those requests finish.

Make signal, context, and Serve-exit paths join one shutdown sequence. Drain gRPC and HTTP concurrently with 15-second default limits, then close the store. Test both completion orders.
2026-09-16 16:20:39 -07:00
Eliah RusinandGitHub caf3d157e6 fix(ec): encode drops tombstoned needles, last-wins replay (#11347)
* fix(ec): encode drops tombstoned needles, last-wins replay

* fix(ec): drop zero-offset rows in encode, match readNeedleMap
2026-09-16 15:40:22 -07:00
Chris LuandGitHub 3ebc05930d s3: separate Object Lock configuration permission (#11361)
* s3: separate object lock configuration permission

* test: synchronize manifest cancellation setup
2026-09-16 15:27:09 -07:00
Chris LuandGitHub 0c7beec697 server: add filer-specific disableHttp flag (#11360) 2026-09-16 14:18:57 -07:00
David ChristopherandGitHub a859f0a019 filer: preserve accepted metadata log records on shutdown (#11359)
fix: flush metadata log before closing filer store

Serialize sealed-batch handoffs with shutdown, reject late appends, and wait for log-buffer workers before closing the filer metadata store.

Cover queued writes, interval and explicit flushes, late-write rejection, and pending persistence with shutdown tests.
2026-09-16 12:23:59 -07:00
Eliah RusinandGitHub def25ca84d fix(ec): validate ShardId at gRPC boundary, reject >=32 (#11346) 2026-09-16 08:33:22 -07:00
Eliah RusinandGitHub 701e397337 fix(volume): reject negative Size, recover poisoned store lock (#11345) 2026-09-16 08:12:11 -07:00
Eliah RusinandGitHub 4fc9ada2ec ci: run seaweed-volume unit tests on Windows (#11349) 2026-09-16 08:06:21 -07:00
a73ba3adbb rust volume: parse vid/fid paths once; the proxy redirect drops the extension like Go (#11341)
handlers.rs split needle URLs in three places and the three disagreed.
Go does it once, in parseURLPath (weed/server/common.go:218-249), and
dispatches on the slash count: /vid/fid/filename takes the extension off
the filename and leaves the fid whole, /vid/fid takes it off the fid, and
the comma form splits the last segment on its last comma and dot.

Two of the Rust copies got that wrong:

- extract_file_id returned the path unchanged when it found no comma, so
  a JWT fid claim, which Go compares against vid + "," + fid for every URL
  form (volume_server_handlers.go:361-364), could never match a slash-form
  request. With a JWT key configured, every read, write or delete of
  /3/01637037d6 was a 401.
- build_proxy_request_info's slash branch had no extension handling, so a
  redirect for /3/01637037d6.jpg sent the client to /3,01637037d6.jpg.
  Go's proxyReqToTargetServer formats "%s/%s,%s" from the already-stripped
  fid (volume_server_handlers_read.go:128-137) and so emits
  /3,01637037d6. The peer still serves either form, since the comma form
  strips the extension again, so this one is parity rather than breakage.

Replace all three with one parse_needle_path returning vid, fid, ext and
filename borrowed from the path. The fid keeps its _delta suffix, as in
Go: parse_needle_id_cookie applies it and the JWT check strips it. The
leading slash stays optional, so chunk manifest fids still parse.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 01:13:06 -07:00
71f8128d75 shell: fs.verify -pruneEntries deletes entries whose needles are lost (#11338)
* shell: fs.verify -pruneEntries deletes entries whose needles are lost

* shell: harden fs.verify -pruneEntries guards; VolumeNeedleStatus returns NotFound for absent needles

* shell: resolve chunk manifests in fs.verify metadata path; require confirmed deletion before counting prunes

* shell: anchor fs.verify legacy missing-needle error matching

* shell: keep fs.verify metadata scan alive on manifest resolution failures

* shell: classify EC missing needles and keep manifest failures unverified

VolumeNeedleStatus now canonicalizes erasure_coding.NotFoundError to
codes.NotFound, so absent needles in EC volumes reach the prune path
through the same stable contract as regular volumes. The client-side
isNeedleMissingError keeps recognizing the legacy wrapped EC shape
("locate in local ec volume: ... needle not found") for mixed-version
clusters.

A chunk manifest that fails to resolve is now an entry-level
verification failure even when the raw top-level chunks are healthy:
the file is not fully readable without the manifest. Raw chunks are
still verified on a resolution failure so a missing top-level manifest
needle is classified and can be pruned. The per-entry logic is
extracted into resolveAndVerify for testability.

* shell: trim fs.verify prune comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 20:42:18 -07:00
github-actions[bot] 01545fc4ff docs: regenerate star history chart 2026-09-16 00:50:47 +00:00
1f037e48f9 s3: a list marker that sorts before the prefix excludes nothing (#11322)
* s3: a list marker that sorts before the prefix excludes nothing

ListObjects `marker` and ListObjectsV2 `start-after` are a plain key
cutoff: list the keys that sort after it. A marker that sorts before the
prefix and is not under it therefore excludes no key carrying the prefix,
and the listing must equal the one with no marker at all.

normalizePrefixMarker treated every marker that does not start with the
prefix as "something wrong" and the listing came back empty. Clients send
this shape routinely: docker/distribution's S3 storage driver walks
prefix "<root>/<path>/" with start-after "<root>" (its rootdirectory), so
on SeaweedFS a registry walk saw an empty bucket. zot read that as "no
repositories": /v2/_catalog was empty, GC/scrub/retention never saw a
repo, and on restart its storage parse deleted every repository's
metadata as "no longer in storage".

listFilerEntries now lists as if no marker were given when the marker
sorts before the prefix; the response still echoes the marker the client
sent. A marker that sorts after the prefix's subtree is left alone: it may
legitimately sit inside a partial-name prefix's match set, which
normalizePrefixMarker already handles, and otherwise correctly lists
nothing.

Reproduce on 4.44 and 4.47:

  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot/zot/"  # all keys
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot"       # KeyCount 0
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=a"         # KeyCount 0

* s3: keep the prefix's own key excluded by a marker that names it

Fold the before-prefix marker rule into normalizePrefixMarker, which now also
derives prefixEndsOnDelimiter from the effective marker instead of each cursor
rebuilding the expression.

A marker equal to the prefix is no longer trimmed to a subtree cutoff:
start-after "a/b/" with prefix "a/b/" excludes only the "a/b/" key, so the walk
starts inside that directory and its children still list.

Adds a listing-level test that walks the whole path for both start-after shapes
a registry sends, and covers the new normalization cases.

* s3: leading slashes do not hide a marker that names the prefix

* s3: echo the V1 marker the client sent, not the walk's cutoff

* s3: filter only the walk's cutoff from the V1 page, not the echoed marker

* s3: skip the key an exclusive marker names as it streams

---------

Co-authored-by: Zuse <be9c90a8-c104-4be2-b7a4-9f92eb833ac8@forge.local>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 16:50:58 -07:00
563c729e70 rust volume: stream the tail scan and release the store lock (#11275)
* rust volume: add a .dat scan plan that runs without the store lock

DatScanPlan captures a fresh .dat handle, the version, the start offset
and an end bound while the caller holds a store guard, then visits one
record at a time with positional reads that never touch the Volume, the
way Go's ScanVolumeFileFrom feeds a scanner.

The handle pins the inode the offset was resolved against: a vacuum
commit renames .cpd over .dat and destroy unlinks it, and neither
rewrites the pinned bytes. The end bound is read while no writer can
hold store.write(), so the scan never meets a partial append. It is a
fresh open, not try_clone, because on Windows read_exact_at uses
seek_read, which moves a cursor a clone shares with the writer.

A header whose size is negative, or does not fit before the end bound,
ends the pass before the body length is computed or anything is
allocated. In today's scan a negative size reaches needle_body_length
and either overflows the buffer size or walks the scan from a wrong
offset. A size near i32::MAX overflows padding_length's i32 arithmetic,
which panics in debug builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: stream the tail scan with the store lock released

volume_tail_sender read every needle from the start offset to EOF into a
Vec while holding store.read(). volume.merge tails from zero, so that was
the whole volume in memory. And because needle writes and the heartbeat
take store.write() on a lock that prefers writers, the whole node stopped
serving until the scan finished: the failure #11235 fixed for EC scrub.

Each pass now runs on a blocking thread. Under one store guard it
resolves the start offset and captures a DatScanPlan, then drops the
guard and sends each needle as it is read, as Go's
VolumeFileScanner4Tailing does. This replaces the one-guard-across-
search-and-scan rule from the previous commit with a stronger invariant:
the offset, the handle and the end bound come from the same guard, and
the handle pins the inode, so a vacuum commit mid-scan cannot point the
offset into the compacted file.

A scan error now ends the stream with Status::internal instead of a
clean EOF, as Go's `streamFollow: %w` does. Once needles stream, a clean
EOF after a partial pass would let volume.move treat a truncated tail as
complete. A panic in the pass is reported the same way. A receiver that
hangs up is also noticed between skipped needles, not only on a send.

Unchanged: the append_at_ns filter, the header on every 2MB chunk, the
caught-up heartbeat without a scan, and the draining countdown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: fail the tail pass on a short read below the snapshot end

DatScanPlan::scan treated an UnexpectedEof on the header or body read as
the end of the data and returned Ok. Every byte below the captured end
existed when the plan was taken, so a short read there can only mean the
inode was truncated under the plan: an unmount followed by a VolumeCopy
of the same volume id reopens .dat with truncate(true). The pass then
reported Scanned, the next pass found the volume gone, and the stream
ended cleanly after a prefix of the planned records, which volume.move
would take as a complete tail.

Both short-read arms now fail the scan with an I/O error that names the
offset and the snapshot end, so tail_pass reports Status::internal as it
does for every other read failure. The break arms were carried over from
scan_raw_needles_from, where the whole scan ran under the store guard and
nothing could truncate the file.

Found by the Devin and Greptile reviews on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: sum the needle padding in i64 so a corrupt size cannot overflow

padding_length added the header, checksum and timestamp widths to the
needle size in i32. A size read from a corrupt header can sit near
i32::MAX, and that sum then overflows: a panic with overflow checks, a
wrapped padding without. DatScanPlan::scan bounds the size against the
bytes left before computing the body length, but that only keeps such a
size out of the arithmetic while under 2 GiB of the file remains, so on a
large volume the scan could still reach the overflow and, in release,
size a buffer from garbage.

Sum in i64 in both version branches. The result is at most
NEEDLE_PADDING_SIZE, so it still fits Size. The scan comment no longer
claims the bound check prevents the overflow.

Found by the CodeRabbit review on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: propagate dat scan parse failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:40:06 -07:00
Chris LuandGitHub 55367afded Update README with performance details of 'weed' binary
Clarified the performance characteristics of the 'weed' binary, emphasizing O(1) complexity for read and write operations.
2026-09-15 14:24:27 -07:00
a9ecfeef45 helm: roll master pods when master config changes (#11331)
* helm: roll master pods when master config changes

The master loads master.toml once at startup (startAdminScripts reads
master.maintenance.scripts and sleep_minutes via viper with no config
watching), and the master ConfigMap is mounted with subPath, which
kubelet never refreshes in a running pod. So a change to
.Values.master.config today updates the ConfigMap but running masters
keep executing the old configuration until something else restarts
them.

Add a checksum/config annotation on the master pod template, following
the existing checksum/s3config pattern on the filer and s3 pods, so a
master config change triggers a rolling restart of the masters.

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Guard against duplicate keys

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Add checksum to deployment as well

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Always ensure the annotation is set

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Update comments

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Soften stance

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* helm: merge pod annotations before checksums

---------

Signed-off-by: Evans Mungai <mbuevans@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:11:58 -07:00
Chris LuandGitHub beaf96a51d s3: cover object lock retention on version deletes (#11335)
* s3: cover WORM guarded version deletes

* s3: trim version delete comments
2026-09-15 13:31:32 -07:00
Chris LuandGitHub 93d4a6aefd s3: drain request body before error response (#11334)
* s3: drain request body before error response

* s3: keep oversized request bodies drainable
2026-09-15 13:09:07 -07:00
166af06a2b rust: cargo fmt both crates, with a commented-out fmt --check CI step (#11329)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-worker: run cargo fmt

Layout only; no token in the workspace changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: run cargo fmt

Layout only; no token in the crate changes. Every earlier Rust PR here
formatted only the blocks it touched so as not to drown its diff in
this one, and this commit is that debt paid in a single place. rustfmt
needed two passes to settle one block in handlers.rs; the committed
form is the fixed point, so `cargo fmt --check` is clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out cargo fmt --check step to both Rust workflows

Same shape as the commented clippy step from #11312: the check is
written out so that making formatting a gate is a one-line uncomment,
and whether to do that stays a maintainer call.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-15 09:29:22 -07:00
517f60e875 rust-volume: fold the 8–15-argument functions into parameter structs (#11328)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 09:04:16 -07:00
Eliah RusinandGitHub 49a680dd64 rust: tonic 0.14 / prost 0.14, drop the result_large_err allow (#11327) 2026-09-15 04:40:46 -07:00
87332eb60b Cloud/remote storage & tiering: configurable multipart upload/download concurrency (#11319)
* pb: add multipart concurrency fields to RemoteConf and tier move requests

RemoteConf gains upload_concurrency/download_concurrency (0 = client
default); VolumeTierMoveDatToRemote/FromRemote requests gain a
concurrency field (0 = backend default).

* remote storage: honor RemoteConf upload/download concurrency in s3 and azure clients

s3 client: ReadFile passes conf download_concurrency to the downloader,
WriteFile uses upload_concurrency for the uploader; previously
hard-coded 1 upload / 5 download parts. 0 keeps defaults. Same for
azure client.

* storage: plumb concurrency through backend interface and tier upload/download

BackendStorage.CopyFile/DownloadFile take a concurrency hint (<=0 =
backend configured default); s3 backend reads
upload_concurrency/download_concurrency from scaffold config with
parseConcurrency fallback, rclone updated to the new signature. Tier
move gRPC handlers forward the request concurrency to the backend.

* shell: -upload_concurrency/-download_concurrency for remote.configure, -concurrent for volume.tier

remote.configure exposes upload/download concurrency persisted into
RemoteConf; volume.tier move/evict commands forward -concurrent to the
tier move requests. Documented in master-cloud.toml scaffold.

* test: cover concurrency propagation in remote tier integration test

* remote.configure: merge existing config on partial update

Load the stored RemoteConf before saving so a partial update (e.g. only
-upload_concurrency) preserves credentials, endpoints, and type instead
of replacing them with new-config defaults. Only treat a confirmed
ErrNotFound as a new configuration; propagate all other load errors so a
transient filer failure does not overwrite stored settings.

On a type transition, reset backend-specific fields to the destination
type's new-config defaults rather than inheriting the old backend's
empty values. Bound configured concurrency to a sane maximum.

* remote storage: honor configured download concurrency in S3 and Azure

ReadFileWithConcurrency now resolves a zero request override against the
client's configured download_concurrency (new downloadConcurrency()
helpers), so the remote-mount/cache read path honors
RemoteConf.DownloadConcurrency instead of the hard-coded default.

Azure also clamps the resolved value to math.MaxUint16 regardless of
whether the fallback was used, preventing uint16 wraparound when a
configured value exceeds 65535.

* shell: rename -concurrent to -concurrency and validate tier transfer bounds

Rename the -concurrent flag to -concurrency across volume.tier.upload,
volume.tier.download, and volume.tier.compact to match the proto field and
RemoteConf field names. Add validateTierConcurrency to reject values that
would wrap int32 or exceed a 1024 cap before constructing the request.

* server: clamp tier move concurrency in gRPC handlers

Add clampTierConcurrency to both VolumeTierMoveDatToRemote and
VolumeTierMoveDatFromRemote handlers so a direct gRPC caller cannot spawn
an unbounded number of network workers.

* trim verbose comments added with concurrency feature

Remove redundant doc comments on the backend interface, rclone backend,
s3_backend parseConcurrency, and test helpers that restated the obvious.

* remote.configure: apply type defaults before re-parse so explicit flags win

applyTypeDefaults ran after the second flag parse, overwriting explicit
destination flags (e.g. -s3.region=eu-west-1) with new-config defaults.
Move the type-transition default reset before the re-parse so user-supplied
flags override the destination defaults.

* remote.configure: only treat explicit -type as a type transition

The first parse defaults -type to s3, so a concurrency-only update on an
existing non-S3 config captured requestedType=s3 and wrongly triggered a
type transition, resetting the stored backend to S3. Use fs.Visit to
detect whether -type was explicitly supplied; an omitted -type keeps the
stored backend.

---------

Co-authored-by: Jack Meredith <9480542+jackusm@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 22:09:08 -07:00
github-actions[bot] e735c12869 docs: regenerate star history chart 2026-09-15 00:48:17 +00:00
e4ca0d09e7 s3: preserve versions for POST policy uploads (#11316)
* s3: preserve versions for POST policy uploads

Route POST policy uploads through the existing version-aware write helpers
and validate promoted Object Lock headers before writing.

Return the generated version ID when versioning is enabled, return
x-amz-version-id: null when versioning is suspended, and omit the header
when versioning has never been enabled.

* s3: reuse versioning helpers in POST policy handler

Route the POST policy handler through the existing getVersioningState
and isObjectLockEnabled helpers instead of open-coding the object-lock
forces-versioning-enabled rule, matching the PUT path.

Drop the x-amz-version-id: null response header for suspended
versioning; the PUT handler omits it and the S3 PutObject sample
response for suspended buckets does not include it. Trim the moved
fileSize comment.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:36:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
01433e801d build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0 (#11306)
* build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0

Bumps [github.com/redis/go-redis/v9](https://github.com/redis/go-redis) from 9.21.0 to 9.22.0.
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/redis/go-redis/compare/v9.21.0...v9.22.0)

---
updated-dependencies:
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.22.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* test(redis_conf): track go-redis 9.22.0 default read timeout of 5s

go-redis 9.22.0 raised the default ReadTimeout from 3s to 5s (part of the
cross-SDK configuration alignment). Update TestUnsetKeepsGoRedisDefaults to
expect the new default so the bump in #11306 stops failing CI.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:16:35 -07:00
1f61097d4d helm: grant List to the generated read-only S3 identity (#11318)
* helm: grant List to the generated read-only S3 identity

The chart's anvReadOnly identity only carried the Read action, so its
credentials could GetObject and HeadObject but every ListObjects request
was denied: List is a separate action and the identity check is an exact
match. Add List so the read-only credentials can list buckets and objects.
Writes stay denied. Update the README example to match.

Bump the chart to 4.47.1. The chart label is part of the s3 and
all-in-one pod templates, so the upgrade rolls the gateways and they
reload the identity config, which is only read at startup.

Fixes #11317

* helm: roll standalone S3 and all-in-one on s3 config changes

Mirror the filer checksum/s3config pod annotation in the standalone S3
and all-in-one deployments so a changed generated S3 secret triggers a
rollout during a normal helm upgrade without relying on a chart version
bump.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:11:22 -07:00
Chris LuandGitHub 0e82b4e351 s3: populate Initiated timestamp in ListMultipartUploads (#11313)
* s3: populate Initiated timestamp in ListMultipartUploads

ListMultipartUploads returned each upload with only Key and UploadId,
omitting the Initiated timestamp. Clients such as GeeseFS rely on this
field to expire stale uploads and crash on its absence. Set Initiated
from the upload directory entry creation time so repeated listings
preserve the original initiation time.

* test/s3: verify Initiated timestamp in ListMultipartUploads

Add an integration test that initiates a multipart upload, lists it,
and asserts the Initiated field is populated and preserved across
repeated listings rather than reflecting the listing time.
2026-09-14 16:03:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0bd048b76f build(deps): bump google.golang.org/api from 0.296.0 to 0.297.0 (#11307)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.296.0 to 0.297.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.296.0...v0.297.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.297.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 15:09:14 -07:00
Chris LuandGitHub c997e54096 admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback (#11314)
* admin: extract isFlagExplicitlySet helper from applyViperFallback

No behavior change; the inline flag-visit check becomes a reusable
helper so the upcoming bind-address default can share it.

* admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback

PR #11185 made the admin HTTP server default to 127.0.0.1 to stop
exposing the unauthenticated admin API on the network by accident.
That also locked out operators who already authenticated with
-adminPassword: their UI became unreachable from the network after
upgrade unless they added -ip=0.0.0.0 (see #11303).

An authenticated deployment is safe to expose, so auto-upgrade the -ip
default to 0.0.0.0 when -adminPassword or [https.admin] mTLS is
configured. The loopback default stays for the unauthenticated case, so
the unauthenticated API is never exposed on the network. An explicit
-ip is always honored.

The worker gRPC control plane has no password auth (only mTLS), so it
must not follow the HTTP upgrade. Give it a separate bind address that
stays on loopback unless -ip is explicit, so adminPassword no longer
re-exposes the unauthenticated worker stream.

* admin: hint loopback-only bind in startup banner

When the admin server binds to loopback (the default for the
unauthenticated case), print a one-line hint that it is not reachable
from other hosts and how to expose it. This helps operators who, after
the #11185 loopback default, can no longer reach the UI from another
machine quickly see the cause and the fix without reading the docs.

* admin: keep worker gRPC on loopback, decouple from https.admin mTLS

The worker gRPC auto-upgrade to 0.0.0.0 was gated on hasMTLS, which
reads the https.admin (HTTP) mTLS config. The worker gRPC mTLS comes
from grpc.admin + grpc.ca, a separate config, so:

- https.admin mTLS without grpc.admin mTLS widened the worker gRPC to
  0.0.0.0 unauthenticated (re-exposing the control plane), and
- grpc.admin mTLS without https.admin mTLS left the worker gRPC on
  loopback, blocking authenticated remote workers.

Drop the worker gRPC auto-upgrade entirely. The worker gRPC keeps the
raw -ip value (loopback by default), matching the pre-existing
behavior; an operator who wants remote workers sets -ip explicitly.
Only the HTTP admin listener auto-upgrades to 0.0.0.0 when
authenticated.

Addresses review feedback on #11314 from Devin and Greptile.
2026-09-14 14:04:35 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2aa6af033d build(deps): bump github.com/go-sql-driver/mysql from 1.10.0 to 1.10.1 (#11308)
Bumps [github.com/go-sql-driver/mysql](https://github.com/go-sql-driver/mysql) from 1.10.0 to 1.10.1.
- [Release notes](https://github.com/go-sql-driver/mysql/releases)
- [Changelog](https://github.com/go-sql-driver/mysql/blob/master/CHANGELOG.md)
- [Commits](https://github.com/go-sql-driver/mysql/compare/v1.10.0...v1.10.1)

---
updated-dependencies:
- dependency-name: github.com/go-sql-driver/mysql
  dependency-version: 1.10.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 14:04:10 -07:00
Chris LuandGitHub 02749c1192 s3api: configurable trusted-proxy allowlist for aws:SourceIp (#11302) (#11315)
* s3api: add TrustedProxies allowlist helper for aws:SourceIp extraction

Introduces a policy_engine.TrustedProxies type that parses a
comma-separated list of bare IPs and CIDRs (mirroring Guard.UpdateWhiteList)
and extracts the client IP for aws:SourceIp condition evaluation.

When the direct TCP peer is in the allowlist, X-Forwarded-For is walked
right-to-left skipping trusted hops (then X-Real-Ip); otherwise the direct
peer address is returned. This is the building block for restoring
configurable forwarded-header trust removed in b88156f (#11231), as
proposed in #11302.

* s3api: honor trusted-proxy allowlist in bucket/IAM policy engine

Make ExtractConditionValuesFromRequest a method on *PolicyEngine so it
can use the engine TrustedProxies when resolving aws:SourceIp. With no
allowlist configured the behavior is unchanged from b88156f: the direct
TCP peer is used and forwarded headers are ignored. When an allowlist is
configured via SetTrustedProxies, requests from a trusted peer honor
X-Forwarded-For (right-to-left) then X-Real-Ip.

Update the two call sites (auth_credentials.go, s3api_bucket_policy_engine.go)
and the engine tests to the method form, and add a regression test for the
trusted-proxy path.

* s3api: honor trusted-proxy allowlist in IAM role/session policies

Make extractRequestContext and extractSourceIP methods on
*S3IAMIntegration so they can use the integration TrustedProxies when
resolving aws:SourceIp. With no allowlist configured the behavior is
unchanged from b88156f: the direct TCP peer is used and forwarded
headers are ignored. When an allowlist is configured via
SetTrustedProxies, requests from a trusted peer honor X-Forwarded-For
(right-to-left) then X-Real-Ip.

Update the call site in isActionExplicitlyDeniedByIAM to type-assert
the integration and use the method, and add a regression test for the
trusted-proxy path.

* s3api: load [s3.trusted_proxies] from security.toml and wire to engines

Read s3.trusted_proxies.white_list (comma-separated IPs/CIDRs) from
security.toml and propagate the allowlist to the bucket policy engine,
the IAM policy engine (persisted across rebuilds via
IdentityAccessManagement.SetTrustedProxies), and the IAM integration.
Reloaded on SIGHUP alongside the JWT signing keys. Document the new
section in the scaffold security.toml.

Closes #11302.

* s3api: harden TrustedProxies parsing and X-Forwarded-For traversal

Canonicalize bare IP entries (via net.ParseIP + String) so non-canonical
IPv6 allowlist entries such as 2001:0db8::1 match peers rendered as
2001:db8::1, and log+skip unparseable bare entries instead of storing
them inertly.

When walking X-Forwarded-For right-to-left, stop at the first malformed
(non-empty, unparseable) entry instead of skipping it, and only fall
back to the leftmost valid IP when the chain was well-formed. This
prevents a malformed hop from masking a forged IP to its left.

Addresses review feedback on #11315.

* s3api: make TrustedProxies reload race-free via atomic.Pointer

Store the trusted-proxy allowlist behind sync/atomic.Pointer in
PolicyEngine and S3IAMIntegration so SIGHUP reloads (which swap the
allowlist) cannot race with concurrent request handlers reading it.
This mirrors the existing Guard guardState pattern. The
IdentityAccessManagement copy is already protected by iam.m.

Addresses review feedback on #11315.
2026-09-14 13:54:26 -07:00
ac03d3fd78 shell: warn when fs.mergeVolumes source holds only orphan needles (#11310)
* shell: warn when fs.mergeVolumes source holds only orphan needles

fs.mergeVolumes traverses filer entries, so a source volume whose
needles are all orphans — filer entries lost to a crashed write or a
wiped filer store — produces only the plan header and exits 0: no move,
no skip, no error. Operators read that as a successful merge while the
real cleanup (volume.fsck) never runs, and dat>idx volumes keep coming
back read-only after restarts.

Count the source-volume needles seen during traversal and, when a plan
source was never seen but its index still reports needles, print a
warning pointing at volume.fsck. Dry-run warns too.

* shell: make needle counting concurrency-safe and count manifest sub-chunks

TraverseBfs runs its callbacks from five workers, so the plain
needlesSeen map raced between source-heavy merges (fatal concurrent
map writes). All increments now funnel through a mutex-guarded
recordSeen closure.

Manifest sub-chunks that live on planned source volumes are now
recorded too — rewriteManifestChunk visits them (including dry-run
and capacity-skipped ones) but previously never marked their source,
which produced false 'orphan needles' warnings for sources whose
chunks were all reached through manifests.

* shell: extract sourceNeedleCounter so the concurrency test covers the production path

The orphan-warning recording was a closure local to Do, so
TestWarnUnreferencedSources_ConcurrentRecording could only exercise a
test-local copy of it — a regression in the production mutex would pass
the test. Lift the map and mutex into a sourceNeedleCounter type with
record/count methods and use it from Do and the test, so the -race test
now drives the actual recording path. Trim the verbose comments added
with the warning while here.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:48 -07:00
adaf3534fa rust: clippy-clean both crates and adopt the std APIs the 1.91 MSRV allows (#11312)
* rust: apply clippy --fix to both crates

The mechanical part of a clippy sweep: `cargo clippy --all-targets --fix`
on seaweed-volume and the seaweed-worker workspace, hand-reviewed. Both
manifests declare their MSRV (1.91.1 and 1.94.1), so every suggestion
clippy applied is within it: the collapsible_if sites become let chains
(1.88, edition 2024), `% n == 0` becomes is_multiple_of (1.87),
chunks_exact with a constant becomes as_chunks (1.88), repeat().take()
becomes repeat_n (1.82), and io::Error::new(Other, ..) becomes
io::Error::other (1.74). The rest is redundant clones, borrows, casts,
closures and field names.

Nothing here changes behaviour. The three let_and_return sites in
needle_map.rs and store_ec.rs deserve a note: the `let result = ..;
result` shape was a deliberate edition-2021 workaround to drop a redb
guard before the table it borrows. Edition 2024 drops tail-expression
temporaries before locals, which is why clippy now flags it, and the
two comments that described the workaround say so instead.

Manual edits on top of the tool output: the blocks clippy rewrote are
re-indented the way rustfmt lays them out (only those blocks — the
crate is not rustfmt-clean and a whole-crate fmt would bury this diff),
the blank lines let_and_return left behind are removed, and the CRC
legacy_value test compares against a literal worked out from the
original shift formula rather than restating rotate_right.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: clear the clippy warnings --fix cannot apply, and say why the rest stay

Hand fixes for the lints clippy only reports. Behaviour is unchanged
throughout; each rewrite is the one clippy names.

- needless_range_loop (7): index loops over shard vectors become
  iterator loops. Where the old code indexed `v[..n]` the new loop
  iterates `v[..n]` so an undersized vector still panics the same way.
- field_reassign_with_default (6): struct literals with `..Default`.
- redundant_pattern_matching (3): `if let Err(_) = guard.check()` becomes
  `.is_err()`, which also releases the read guard at the end of the
  condition instead of at the end of the block.
- manual_strip (2), manual_checked_ops, format_in_format_args,
  redundant_locals, wrong_self_convention (to_vif takes self by value,
  so it is into_vif; CompactEntry is Copy, so to_needle_value takes self).
- type_complexity (2): `OrphanShardLoad` and `RawNeedleEntry` name two
  tuples that were spelled out inline.
- new_without_default: CompactNeedleMap gets a Default that calls new().
- suspicious_open_options: a test helper spells out `.truncate(false)`,
  which is what `.create(true).write(true)` already did.

What stays, and the attribute that says so:

- too_many_arguments (10): `#[expect]` on each function. Folding 8–15
  parameters into a struct is a design change, not a lint fix.
- await_holding_lock / readonly_write_lock: one test holds the store
  write guard across a sleep on purpose, as a barrier that parks the
  copy task at the mount block. `#[expect(.., reason = ..)]` records it.
- module_inception: needle/needle.rs mirrors the Go package layout.

Two lints become crate-wide policy in `[lints.clippy]`, with the reason
next to each: result_large_err, because every RPC path returns
tonic::Status (176 bytes) and boxing it would change every handler
signature; and needless_update, because `..Default::default()` on a
protobuf message literal is what lets a proto gain a field without
touching every constructor (all 11 sites are pb messages). The worker
workspace gets the same table and its members opt in with
`lints.workspace = true`; its generated plugin.rs also allows
large_enum_variant on prost's oneof enums.

Both crates are now clean under `cargo clippy --all-targets -- -D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: use the std APIs the 1.91 MSRV already pays for

The crate declares rust-version 1.91.1, so a few things the code still
worked around are plain std now. All of them come from the 1.85–1.91
release notes; nothing here needs a newer toolchain than the manifest
already requires.

- std::sync::LazyLock (1.80) replaces the lazy_static! block in
  metrics.rs, and the lazy_static dependency goes. Every use site reads
  the same through Deref, so no caller changes.
- Duration::from_mins / from_hours (1.91) replace `from_secs(v * 60)`
  and `from_secs(v * 3600)` in the option parser and the shard-location
  refresh TTLs. One difference for the parser: an absurd count that
  overflows u64 seconds now panics in release builds too, where the
  multiplication used to wrap.
- Result::flatten (1.89) replaces `.and_then(|r| r)` on the replication
  join handle.
- OsStr::display (1.87) replaces `to_string_lossy()` where the name was
  only being formatted; the output is byte-identical.
- `#[allow]` becomes `#[expect]` (1.81) on the suppressions that are
  meant to be permanent, so a suppression that stops being needed
  becomes a warning rather than lingering. Doing that found four that
  already had: dead_code on ChunkManifest, base_name and last_io_error,
  and too_many_arguments on read_from_data_shards, which is down to
  seven parameters. Those attributes are deleted. The three allows that
  depend on cfg (a unix-only mutation, a linux-only field set, a
  profiling-only parameter) stay as allow, because expect would be
  unfulfilled on the other platforms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out clippy step to both Rust workflows

Both crates are warning-free under `cargo clippy --all-targets
-D warnings` now. Whether that becomes a gate is a policy call, so the
step is present but commented out; uncommenting it is the whole change.
The comment points at the `[lints.clippy]` table where crate-wide
exceptions are recorded, so the gate does not become a reason to
sprinkle allows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: guard parse_duration against overflow panics

Duration::from_mins/from_hours panic when the count overflows u64
seconds. Use checked_mul so an oversized CLI value falls back to the
parser default instead of crashing volume startup.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:29 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
49f20489e4 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.20.1 to 1.20.4 (#11305)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.20.1 to 1.20.4.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.20.1...service/mq/v1.20.4)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.20.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:13:14 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bdec508da9 build(deps): bump golang.org/x/image from 0.45.0 to 0.46.0 (#11304)
Bumps [golang.org/x/image](https://github.com/golang/image) from 0.45.0 to 0.46.0.
- [Commits](https://github.com/golang/image/compare/v0.45.0...v0.46.0)

---
updated-dependencies:
- dependency-name: golang.org/x/image
  dependency-version: 0.46.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:45:16 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fd33c07843 build(deps): bump github/codeql-action from 4.37.9 to 4.38.0 (#11311)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.9 to 4.38.0.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.9...v4.38.0)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:41:36 -07:00
Chris LuandGitHub cf38c01978 admin: bind worker gRPC listener to -ip instead of wildcard (#11300)
* admin: bind worker gRPC listener to -ip instead of wildcard

The worker/plugin gRPC control plane called net.Listen("tcp", ":port")
directly, so it wildcard-bound every interface and ignored the -ip setting.
A cluster bound to loopback still exposed the unauthenticated
WorkerService/PluginControlService streams on 0.0.0.0. Bind through
util.JoinHostPort(bindIp, port) so the listener honors -ip like the
master, filer, and volume gRPC listeners.

* admin: warn when worker gRPC is exposed off loopback without mTLS

The worker gRPC stream has no password auth, so grpc.admin mTLS is the
only effective control once the listener leaves loopback. An operator who
sets -adminPassword and binds -ip=0.0.0.0 authenticates the HTTP API but
still exposes the unauthenticated worker control plane. Log a startup
warning naming the port and the mTLS knobs so the exposure is not silent.

* admin: address review on worker gRPC bind fix

- mini: reserve the admin gRPC port with util.JoinHostPort so an IPv6
  bindIp (e.g. ::1) does not form an invalid unbracketed address and
  lose the reservation.
- worker gRPC: track whether grpc.admin mTLS credentials actually loaded
  rather than only whether they were configured, and gate the
  non-loopback exposure warning on that. A cert/key that fails to load
  now still warns instead of silently suppressing.
2026-09-13 21:48:14 -07:00
Chris LuandGitHub f4bad510c9 test/ec: pin rack in seedAndSpread volume.grow to stop silent no-ops (#11299)
seedAndSpread() calls `volume.grow -dataNode X` without pinning the rack.
The master's grow picks the rack by weighted-random when -rack is unset,
and only one of the three racks holds the requested data node, so an
unpinned grow lands on the wrong rack two times out of three. The
VolumeGrow RPC swallows the "No matching data node" failure for
non-cache collections, so those grows count as success without creating
a volume. The per-server cap (maxGrowsPerServer=4) is then exhausted by
silent no-ops before the volumes ever spread, and seedAndSpread times
out with "volumes never spread across >=2 disks on all 3 nodes".

Pin -dataCenter dc1 and -rack rack{i} alongside -dataNode so every grow
reaches the target node. This removes the timing-sensitive assumption
that made TestECVacuumDuplicateShardClaimAcrossDisks flaky.
2026-09-13 21:05:55 -07:00
Chris LuandGitHub 15d9f6c6fe rust-volume: fix Windows build of find_needle_from_ecx (#11298)
* rust-volume: fix Windows build of find_needle_from_ecx

The .ecx binary-search fallback path used on non-Unix targets
(Seek + Read, both &mut self receivers) requires the ecx_file
binding to be mutable. On Unix the read_exact_at path takes &self,
so the mut would be unused there — gate that warning with
#[cfg_attr(unix, allow(unused_mut))].

Without this the build-rust-volume-windows CI job fails with
E0596 at ec_volume.rs:1033, breaking the weed-volume_windows_amd64
release asset.

* rust-volume: use positional seek_read for .ecx lookups on Windows

The previous fix (making ecx_file mut) compiled but left the Windows
fallback using Seek + Read on the shared .ecx file cursor. Concurrent
find_needle_from_ecx calls could interleave seek/read and read the
wrong index entry, corrupting the binary search (raised by Devin and
Greptile review on the PR).

Switch the Windows path to std::os::windows::fs::FileExt::seek_read,
which is positional (offset passed via OVERLAPPED, cursor untouched)
and takes &self — so the binding no longer needs mut, and concurrent
callers on the cached handle can't interfere. Mirrors the existing
read_exact_at helper in storage::volume.

Add a compile_error fallback for non-unix/non-windows targets to
match the convention in storage::volume.
2026-09-13 20:08:01 -07:00
ea179963c0 filer: clean up manifest resolve error propagation and add webdav tes… (#11297)
filer: clean up manifest resolve error propagation and add webdav test (#78)

Drop GitHub issue references from comments and trim verbose comments.
Replace the viewFromChunksOrErr helper with the existing
NonOverlappingVisibleIntervals + ViewFromVisibleIntervals at the stream
call sites, and add a WebDavFile.Read regression test for the manifest
resolution failure path.

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-13 18:34:09 -07:00
693 changed files with 68668 additions and 13379 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.9
uses: github/codeql-action/init@v4.38.2
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.9
uses: github/codeql-action/autobuild@v4.38.2
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.9
uses: github/codeql-action/analyze@v4.38.2
+2 -1
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'weed/**'
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'seaweed-worker/**'
- 'docker/**'
- 'go.mod'
@@ -152,7 +153,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
run: |
@@ -129,7 +129,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
+10 -6
View File
@@ -156,7 +156,7 @@ jobs:
runs-on: ubuntu-latest
strategy:
matrix:
platform: [amd64, arm64, arm, 386]
platform: [amd64, arm64, arm, 386, ppc64le, s390x]
variant: ${{ fromJSON(needs.setup.outputs.variants) }}
steps:
@@ -236,7 +236,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
if: matrix.platform != 'amd64'
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
run: |
cat > /tmp/buildkitd.toml <<EOF
@@ -405,7 +405,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.2
if: always()
with:
sarif_file: trivy-results.sarif
@@ -505,7 +505,9 @@ jobs:
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-amd64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-ppc64le \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-s390x
# The copy and the signature below use the digest this run pushed, not whatever the tag points at by then.
DIGEST=$(jq -er '."containerimage.descriptor".digest' /tmp/manifest.json)
echo "digest=${DIGEST}" >> "$GITHUB_OUTPUT"
@@ -549,13 +551,15 @@ jobs:
echo "Using skopeo to copy..."
retry_with_backoff skopeo copy --all docker://ghcr.io/chrislusf/seaweedfs@${DIGEST} docker://chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
else
echo "Using docker buildx imagetools (pulling 4 images from Docker Hub)..."
echo "Using docker buildx imagetools (pulling 6 images from Docker Hub)..."
# Fallback: create manifest directly on Docker Hub (pulls from Docker Hub - rate limited)
retry_with_backoff docker buildx imagetools create -t chrislusf/seaweedfs:${BASE_TAG}${SUFFIX} \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-amd64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-ppc64le \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-s390x
fi
- name: Sign
@@ -46,7 +46,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
-
name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
-
name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
@@ -149,12 +149,16 @@ jobs:
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/arm64, arch: arm64, runner: ubuntu-24.04-arm, qemu: false }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/arm/v7, arch: armv7, runner: ubuntu-latest, qemu: true }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/386, arch: i386, runner: ubuntu-latest, qemu: false }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/ppc64le, arch: ppc64le, runner: ubuntu-latest, qemu: true }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/s390x, arch: s390x, runner: ubuntu-latest, qemu: true }
# Large disk - multi-arch
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/amd64, arch: amd64, runner: ubuntu-latest, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/arm64, arch: arm64, runner: ubuntu-24.04-arm, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/arm/v7, arch: armv7, runner: ubuntu-latest, qemu: true }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/386, arch: i386, runner: ubuntu-latest, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/ppc64le, arch: ppc64le, runner: ubuntu-latest, qemu: true }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/s390x, arch: s390x, runner: ubuntu-latest, qemu: true }
# Full tags - multi-arch
- { variant: full, tag_suffix: _full, dockerfile: ./docker/Dockerfile.go_build, build_args: "TAGS=elastic,gocdk,rclone,sqlite,tarantool,tikv,ydb", rust_variant: normal, platform: linux/amd64, arch: amd64, runner: ubuntu-latest, qemu: false }
@@ -231,7 +235,7 @@ jobs:
- name: Set up QEMU
if: (github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant) && matrix.qemu
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -456,7 +460,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.2
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
@@ -85,7 +85,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@1f40c72289eff860ee54a304f1438e3cff362e0a # v1
uses: docker/setup-qemu-action@99012661954931238ded8c8b007157a8430204e1 # v1
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v1
+45 -35
View File
@@ -35,11 +35,48 @@ jobs:
cd telemetry/server
go mod tidy
echo "Building telemetry server..."
GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
CGO_ENABLED=0 GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
cd ../..
ls -la telemetry-server
echo "Build completed successfully"
- name: Generate Service Configuration
if: github.event_name == 'workflow_dispatch' && (inputs.setup || inputs.deploy)
env:
REMOTE_USER: ${{ secrets.TELEMETRY_USER }}
run: |
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/bin/sh -c 'exec /home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353 >>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log 2>>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log'
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
- name: First-time Server Setup
if: github.event_name == 'workflow_dispatch' && inputs.setup
env:
@@ -61,40 +98,6 @@ jobs:
touch ~/seaweedfs-telemetry/logs/telemetry.log ~/seaweedfs-telemetry/logs/telemetry.error.log && \
chmod 644 ~/seaweedfs-telemetry/logs/*.log"
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353
Restart=always
RestartSec=5
StandardOutput=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log
StandardError=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
# Copy configuration files
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
@@ -137,11 +140,18 @@ jobs:
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Copy updated service and logrotate files
scp -i ~/.ssh/deploy_key telemetry.service telemetry_logrotate $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Check if service exists and deploy accordingly
ssh -i ~/.ssh/deploy_key $REMOTE_USER@$REMOTE_HOST "
if systemctl list-unit-files telemetry.service >/dev/null 2>&1; then
echo 'Service exists, performing update...'
set -e
sudo systemctl stop telemetry.service
sudo mv ~/seaweedfs-telemetry/telemetry.service /etc/systemd/system/
sudo mv ~/seaweedfs-telemetry/telemetry_logrotate /etc/logrotate.d/seaweedfs-telemetry
sudo systemctl daemon-reload
mkdir -p ~/seaweedfs-telemetry/bin
mv ~/seaweedfs-telemetry/tmp/telemetry-server ~/seaweedfs-telemetry/bin/
chmod +x ~/seaweedfs-telemetry/bin/telemetry-server
+263
View File
@@ -116,6 +116,35 @@ jobs:
grep -q "security-config" /tmp/security.yaml
echo "Security configuration renders correctly"
echo "=== Testing secure bucket-creation hook certificate mounts ==="
helm template test $CHART_DIR \
--show-only templates/shared/post-install-bucket-hook.yaml \
--set s3.enabled=true \
--set 's3.createBuckets[0].name=data' \
--set global.seaweedfs.enableSecurity=true \
> /tmp/security-bucket-hook.yaml
test "$(grep -cE '^[[:space:]]*- name: ca-cert$' /tmp/security-bucket-hook.yaml)" -eq 2
test "$(grep -cE '^[[:space:]]*- name: client-cert$' /tmp/security-bucket-hook.yaml)" -eq 2
grep -q 'mountPath: /usr/local/share/ca-certificates/ca/' /tmp/security-bucket-hook.yaml
grep -q 'mountPath: /usr/local/share/ca-certificates/client/' /tmp/security-bucket-hook.yaml
grep -q 'secretName: test-seaweedfs-ca-cert' /tmp/security-bucket-hook.yaml
grep -q 'secretName: test-seaweedfs-client-cert' /tmp/security-bucket-hook.yaml
echo "Secure bucket-creation hook mounts its CA and client certificate"
echo ""
echo "=== Testing admin.allowInsecureBind satisfies the admin auth render guard ==="
helm template test $CHART_DIR --set admin.enabled=true --set admin.allowInsecureBind=true \
> /tmp/admin-allow-insecure-bind.yaml
grep -q -- "-allowInsecureBind" /tmp/admin-allow-insecure-bind.yaml
echo "admin.allowInsecureBind renders -allowInsecureBind and passes the render guard"
if helm template test $CHART_DIR --set admin.enabled=true > /tmp/admin-no-auth.yaml 2>/tmp/admin-no-auth.err; then
echo "FAIL: admin.enabled=true with no auth configured should fail to render"
exit 1
fi
grep -q "admin.allowInsecureBind" /tmp/admin-no-auth.err
echo "admin with no auth configured still fails the render guard, and the guard mentions admin.allowInsecureBind"
echo ""
echo "=== Testing JWT expiration overrides ==="
helm template test $CHART_DIR \
@@ -704,6 +733,168 @@ jobs:
helm template test $CHART_DIR --set cosi.enabled=true > /tmp/cosi.yaml
grep -q "seaweedfs-cosi" /tmp/cosi.yaml
echo "COSI driver renders correctly"
echo ""
echo "=== Testing configurable pod and container security contexts ==="
helm template test $CHART_DIR > /tmp/security-context-defaults.yaml
security_context_args=()
for component in master volume filer s3 sftp admin worker allInOne cosi; do
security_context_args+=(
--set "$component.podSecurityContext.enabled=true"
--set "$component.podSecurityContext.seccompProfile.type=RuntimeDefault"
--set "$component.containerSecurityContext.enabled=true"
--set "$component.containerSecurityContext.privileged=false"
--set "$component.containerSecurityContext.allowPrivilegeEscalation=false"
--set "$component.containerSecurityContext.readOnlyRootFilesystem=true"
--set "$component.containerSecurityContext.capabilities.drop[0]=ALL"
--set "$component.containerSecurityContext.seccompProfile.type=RuntimeDefault"
)
done
helm template test $CHART_DIR \
"${security_context_args[@]}" \
--set s3.enabled=true \
--set s3.createBuckets[0].name=test \
--set sftp.enabled=true \
--set admin.enabled=true \
--set admin.secret.adminPassword=ci-admin-password \
--set worker.enabled=true \
--set volume.idx.type=hostPath \
--set volume.idx.hostPathPrefix=/tmp \
--set cosi.enabled=true \
--set global.seaweedfs.tmpDir.sizeLimit=64Mi > /tmp/security-contexts.yaml
helm template test $CHART_DIR \
"${security_context_args[@]}" \
--set allInOne.enabled=true \
--set master.enabled=false \
--set volume.enabled=false \
--set filer.enabled=false \
--set global.seaweedfs.tmpDir.sizeLimit=64Mi > /tmp/security-contexts-aio.yaml
python3 - /tmp/security-context-defaults.yaml /tmp/security-contexts.yaml /tmp/security-contexts-aio.yaml <<'PYEOF'
import sys
import yaml
errors = []
workloads = 0
containers = 0
chart_managed_init_containers = 0
components = set()
def validate_container(workload_name, container):
context = container.get("securityContext", {})
prefix = f"{workload_name}/{container['name']}"
if "enabled" in context:
errors.append(f"{prefix}: internal enabled flag leaked into container securityContext")
if context.get("privileged") is not False:
errors.append(f"{prefix}: privileged is not false")
if context.get("allowPrivilegeEscalation") is not False:
errors.append(f"{prefix}: allowPrivilegeEscalation is not false")
if context.get("readOnlyRootFilesystem") is not True:
errors.append(f"{prefix}: readOnlyRootFilesystem is not true")
if context.get("capabilities", {}).get("drop") != ["ALL"]:
errors.append(f"{prefix}: capabilities.drop is not exactly [ALL]")
if context.get("seccompProfile", {}).get("type") != "RuntimeDefault":
errors.append(f"{prefix}: seccompProfile is not RuntimeDefault")
mounts = {mount["name"]: mount for mount in container.get("volumeMounts", [])}
if mounts.get("seaweedfs-tmp", {}).get("mountPath") != "/tmp":
errors.append(f"{prefix}: writable temporary volume is not mounted at /tmp")
with open(sys.argv[1]) as stream:
default_documents = [document for document in yaml.safe_load_all(stream) if document]
for document in default_documents:
if document.get("kind") not in ("Deployment", "StatefulSet", "Job"):
continue
name = document["metadata"]["name"]
pod = document["spec"]["template"]["spec"]
if "securityContext" in pod:
errors.append(f"{name}: pod securityContext should be absent by default")
if any(volume["name"] == "seaweedfs-tmp" for volume in pod.get("volumes", [])):
errors.append(f"{name}: writable /tmp volume should be absent by default")
for container in pod.get("containers", []):
if "securityContext" in container:
errors.append(
f"{name}/{container['name']}: container securityContext "
f"should be absent by default"
)
if any(
mount["name"] == "seaweedfs-tmp"
for mount in container.get("volumeMounts", [])
):
errors.append(
f"{name}/{container['name']}: writable /tmp mount "
f"should be absent by default"
)
for path in sys.argv[2:]:
with open(path) as stream:
documents = [document for document in yaml.safe_load_all(stream) if document]
for document in documents:
if document.get("kind") not in ("Deployment", "StatefulSet", "Job"):
continue
workloads += 1
name = document["metadata"]["name"]
pod = document["spec"]["template"]["spec"]
volumes = {volume["name"]: volume for volume in pod.get("volumes", [])}
temporary = volumes.get("seaweedfs-tmp", {}).get("emptyDir")
if temporary is None:
errors.append(f"{name}: writable /tmp emptyDir is missing")
elif temporary.get("sizeLimit") != "64Mi":
errors.append(f"{name}: temporary volume sizeLimit is not 64Mi")
component = document["spec"]["template"]["metadata"]["labels"].get("app.kubernetes.io/component")
if component:
components.add(component)
else:
errors.append(f"{name}: app.kubernetes.io/component label is missing")
pod_context = pod.get("securityContext", {})
if "enabled" in pod_context:
errors.append(f"{name}: internal enabled flag leaked into pod securityContext")
if pod_context.get("seccompProfile", {}).get("type") != "RuntimeDefault":
errors.append(f"{name}: pod seccompProfile is not RuntimeDefault")
for container in pod.get("containers", []):
containers += 1
validate_container(name, container)
for container in pod.get("initContainers", []):
if container["name"] != "seaweedfs-vol-move-idx":
continue
chart_managed_init_containers += 1
validate_container(name, container)
expected_components = {
"master", "volume", "filer", "s3", "sftp", "admin", "worker",
"objectstorage-provisioner", "bucket-hook", "seaweedfs-all-in-one",
}
if components != expected_components:
errors.append(
f"security context workload coverage is incomplete: "
f"expected {sorted(expected_components)}, got {sorted(components)}"
)
if workloads != 10 or containers != 12:
errors.append(
f"expected 10 workloads and 12 containers, got "
f"{workloads} workloads and {containers} containers"
)
if chart_managed_init_containers != 1:
errors.append(
f"expected one chart-managed seaweedfs-vol-move-idx init container, "
f"got {chart_managed_init_containers}"
)
if errors:
print("\n".join(f"FAIL: {error}" for error in errors), file=sys.stderr)
sys.exit(1)
print(f"Validated security contexts on {workloads} workloads and {containers} containers")
PYEOF
# The resize hook depends on lookup finding a live StatefulSet and
# therefore cannot render during helm template. Keep static coverage
# for both security-context blocks.
grep -Fqx ' securityContext: {{- omit .Values.volume.podSecurityContext "enabled" | toYaml | nindent 8 }}' \
"$CHART_DIR/templates/volume/volume-resize-hook.yaml"
grep -Fqx ' securityContext: {{- omit .Values.volume.containerSecurityContext "enabled" | toYaml | nindent 12 }}' \
"$CHART_DIR/templates/volume/volume-resize-hook.yaml"
echo "Volume resize hook security contexts are covered"
echo ""
echo "=== Testing long release name: service names match DNS references ==="
@@ -1591,6 +1782,78 @@ jobs:
- name: Create kind cluster
uses: helm/kind-action@v1.15.0
- name: Verify volume resize hook security contexts
run: |
set -e
CHART_DIR="k8s/charts/seaweedfs"
NS="resize-hook-security"
kubectl create namespace "$NS"
kubectl apply -n "$NS" -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: data1-resize-seaweedfs-volume-0
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
EOF
helm template resize "$CHART_DIR" -n "$NS" --dry-run=server \
--set volume.dataDirs[0].name=data1 \
--set volume.dataDirs[0].type=persistentVolumeClaim \
--set volume.dataDirs[0].size=2Gi \
--set volume.podSecurityContext.enabled=true \
--set volume.podSecurityContext.seccompProfile.type=RuntimeDefault \
--set volume.containerSecurityContext.enabled=true \
--set volume.containerSecurityContext.privileged=false \
--set volume.containerSecurityContext.allowPrivilegeEscalation=false \
--set volume.containerSecurityContext.readOnlyRootFilesystem=true \
--set volume.containerSecurityContext.capabilities.drop[0]=ALL \
--set volume.containerSecurityContext.seccompProfile.type=RuntimeDefault \
--set global.seaweedfs.tmpDir.sizeLimit=64Mi \
> /tmp/security-context-resize-hook.yaml
python3 - /tmp/security-context-resize-hook.yaml <<'PYEOF'
import sys
import yaml
with open(sys.argv[1]) as stream:
documents = [document for document in yaml.safe_load_all(stream) if document]
jobs = [
document for document in documents
if document.get("kind") == "Job"
and document["metadata"]["name"].endswith("-volume-resize-hook")
]
if len(jobs) != 1:
raise AssertionError(f"expected one volume resize hook Job, got {len(jobs)}")
pod = jobs[0]["spec"]["template"]["spec"]
assert pod["securityContext"] == {
"seccompProfile": {"type": "RuntimeDefault"},
}
assert len(pod["containers"]) == 1
assert pod["containers"][0]["securityContext"] == {
"allowPrivilegeEscalation": False,
"capabilities": {"drop": ["ALL"]},
"privileged": False,
"readOnlyRootFilesystem": True,
"seccompProfile": {"type": "RuntimeDefault"},
}
assert pod["containers"][0]["volumeMounts"] == [
{"mountPath": "/tmp", "name": "seaweedfs-tmp"},
]
assert pod["volumes"] == [
{"emptyDir": {"sizeLimit": "64Mi"}, "name": "seaweedfs-tmp"},
]
PYEOF
kubectl delete namespace "$NS"
echo "Volume resize hook security contexts render correctly"
- name: Run chart-testing (install)
run: |
ct install --target-branch ${{ github.event.repository.default_branch }} --all --chart-dirs k8s/charts \
+1
View File
@@ -8,6 +8,7 @@ on:
- 'go.mod'
- 'go.sum'
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/perf/**'
- '.github/workflows/performance.yml'
workflow_dispatch:
+88 -3
View File
@@ -5,6 +5,7 @@ on:
branches: [ master ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/volume_server/**'
- 'weed/pb/volume_server.proto'
- 'weed/pb/volume_server_pb/**'
@@ -13,6 +14,7 @@ on:
branches: [ master, main ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/volume_server/**'
- 'weed/pb/volume_server.proto'
- 'weed/pb/volume_server_pb/**'
@@ -27,6 +29,29 @@ permissions:
jobs:
changes:
name: Detect changed paths
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
outputs:
rust: ${{ steps.filter.outputs.rust }}
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
fetch-depth: 0
- name: Filter changed paths
id: filter
uses: dorny/paths-filter@v4
with:
filters: |
rust:
- 'seaweed-volume/**'
- '.github/workflows/rust-volume-server-tests.yml'
rust-unit-tests:
name: Rust Unit Tests
runs-on: ubuntu-22.04
@@ -52,13 +77,73 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
- name: Build Rust volume server
run: cd seaweed-volume && cargo build --release
# The crate is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[lints.clippy]` in
# seaweed-volume/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-volume && cargo clippy --all-targets -- -D warnings
# The crate is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-volume && cargo fmt --check
# seaweed-common is a path dependency of this crate, not a member of its
# workspace, so the run below does not reach its own tests. It builds into
# this job's cached target directory, and the cache key above covers the
# shared crate's lock, so the aws-lc-sys that rustls pulls in is restored
# with the cache instead of compiled from scratch on every run.
- name: Run shared-crate unit tests
env:
CARGO_TARGET_DIR: ${{ github.workspace }}/seaweed-volume/target
run: cd seaweed-common && cargo test
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
- name: Run Rust unit tests (redb experimental cursor)
run: cd seaweed-volume && cargo test --features redb-experimental-cursor --lib storage::needle_map
rust-unit-tests-windows:
name: Rust Unit Tests (Windows)
runs-on: windows-latest
timeout-minutes: 30
needs: [changes]
if: needs.changes.outputs.rust == 'true'
defaults:
run:
shell: bash
steps:
- name: Checkout code
uses: actions/checkout@v7
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# No glibc on Windows: key the cache on the toolchain and OS only.
- name: Fingerprint build toolchain
id: toolchain
run: echo "fingerprint=windows-rustc-$(rustc -V | awk '{print $2}')" >> "$GITHUB_OUTPUT"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-windows-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-windows-${{ steps.toolchain.outputs.fingerprint }}-
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
@@ -95,7 +180,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
@@ -177,7 +262,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
+25 -1
View File
@@ -5,12 +5,14 @@ on:
branches: [ master ]
paths:
- 'seaweed-worker/**'
- 'seaweed-common/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
push:
branches: [ master, main ]
paths:
- 'seaweed-worker/**'
- 'seaweed-common/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
@@ -49,7 +51,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/release
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-worker-${{ steps.toolchain.outputs.fingerprint }}-
@@ -73,6 +75,28 @@ jobs:
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The workspace is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[workspace.lints.clippy]`
# in seaweed-worker/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-worker && cargo clippy --workspace --all-targets -- -D warnings
# The workspace is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-worker && cargo fmt --all --check
# seaweed-common is a path dependency of core and lance, not a member of
# this workspace, so `--workspace` below does not reach its own tests.
# Release and this job's cached target directory, and the cache key above
# covers the shared crate's lock. That lock pins the same rustls and
# aws-lc-sys this workspace resolves, so the release build above has
# already paid for them.
- name: Run shared-crate unit tests
env:
CARGO_TARGET_DIR: ${{ github.workspace }}/seaweed-worker/target
run: cd seaweed-common && cargo test --release
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
+1
View File
@@ -5,6 +5,7 @@ on:
branches: [ master ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- '.github/workflows/rust_binaries_dev.yml'
permissions:
+3 -2
View File
@@ -139,7 +139,8 @@ jobs:
unit-tests:
name: Go Unit Tests (Implicit Directory)
runs-on: ubuntu-latest
timeout-minutes: 10
# Leave time for setup and the focused run before the nine-minute suite.
timeout-minutes: 20
steps:
- name: Checkout code
@@ -159,5 +160,5 @@ jobs:
- name: Run all S3 API tests
run: |
cd weed/s3api
go test -v -timeout 5m
go test -v -timeout 9m
+106
View File
@@ -0,0 +1,106 @@
name: "Snowflake S3Compat API tests"
on:
push:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/iam/**'
- 'weed/command/**'
- 'weed/storage/**'
- 'weed/operation/**'
- 'weed/wdclient/**'
- 'weed/cluster/**'
- 'weed/pb/**'
- 'test/s3/snowflake/**'
- 'go.mod'
- 'go.sum'
- '.github/workflows/s3-snowflake-tests.yml'
pull_request:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/iam/**'
- 'weed/command/**'
- 'weed/storage/**'
- 'weed/operation/**'
- 'weed/wdclient/**'
- 'weed/cluster/**'
- 'weed/pb/**'
- 'test/s3/snowflake/**'
- 'go.mod'
- 'go.sum'
- '.github/workflows/s3-snowflake-tests.yml'
concurrency:
group: ${{ github.event.pull_request.number || github.ref }}/s3-snowflake-tests
cancel-in-progress: true
permissions:
contents: read
jobs:
snowflake-s3compat-tests:
name: Snowflake S3Compat API tests
runs-on: ubuntu-22.04
timeout-minutes: 30
env:
WORK_DIR: /tmp/seaweedfs-snowflake-tests
steps:
- name: Check out code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Set up Java
uses: actions/setup-java@v6
with:
java-version: '17'
distribution: 'temurin'
cache: 'maven'
- name: Install SeaweedFS
run: |
cd weed
go install -buildvcs=false
weed version
- name: Run Snowflake S3Compat API tests
timeout-minutes: 20
run: |
# Starts weed server, creates the buckets/objects the suite needs,
# clones the upstream suite, and runs mvn -Dtest=S3CompatApiTest.
bash test/s3/snowflake/run.sh
- name: Show logs on failure
if: failure()
run: |
echo "=== SeaweedFS Server Log ==="
tail -200 "$WORK_DIR/weed.log" || echo "No server log"
echo ""
echo "=== Surefire results ==="
cat "$WORK_DIR"/snowflake-s3compat-api-test-suite/s3compatapi/target/surefire-reports/*.txt 2>/dev/null || echo "No surefire reports"
- name: Upload test results
if: always()
uses: actions/upload-artifact@v7
with:
name: snowflake-s3compat-surefire-reports
path: /tmp/seaweedfs-snowflake-tests/snowflake-s3compat-api-test-suite/s3compatapi/target/surefire-reports/
retention-days: 14
- name: Cleanup
if: always()
run: |
pkill -9 -f "weed server" || true
rm -rf "$WORK_DIR" || true
+111
View File
@@ -439,6 +439,117 @@ jobs:
path: test/s3tables/catalog_clickhouse/test-output.log
retention-days: 3
olake-iceberg-catalog-tests:
name: OLake Iceberg Catalog Integration Tests (${{ matrix.tag }})
runs-on: ubuntu-22.04
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
include:
# Pinned baseline, and latest so new OLake releases are exercised
# without a code change. OLake's Iceberg writer is a Java sidecar
# whose Iceberg version moves independently of the Go release, so
# the latest leg is the one that catches library drift.
- olake-image: olakego/source-postgres:v0.10.1
tag: "v0.10.1"
- olake-image: olakego/source-postgres:latest
tag: latest
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull ${{ matrix.olake-image }}
pull postgres:16
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Install SeaweedFS
run: |
go install -buildvcs=false ./weed
- name: Run OLake Iceberg Catalog Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_olake
env:
OLAKE_IMAGE: ${{ matrix.olake-image }}
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting OLake Iceberg Catalog Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "OLake Iceberg catalog integration tests failed"
exit 1
}
# The suite skips itself when Docker is unavailable, so a green job is not
# by itself evidence that anything ran. Assert execution explicitly.
- name: Assert the suite actually ran
working-directory: test/s3tables/catalog_olake
run: |
log=test-output.log
if [ ! -f "$log" ]; then
echo "::error::no test-output.log; the suite did not run"
exit 1
fi
passes=$(grep -c '^--- PASS' "$log" || true)
skips=$(grep -c '^--- SKIP' "$log" || true)
echo "top-level PASS=$passes SKIP=$skips"
if [ "$skips" -gt 0 ]; then
echo "::error::the OLake suite skipped $skips top-level test(s); the environment it needs was not provisioned, so this job proves nothing"
grep '^--- SKIP' "$log" | head -20
exit 1
fi
if [ "$passes" -lt 1 ]; then
echo "::error::the OLake suite recorded no passing top-level test"
exit 1
fi
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_olake
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|olake|postgres)" || true
echo "=== Containers ==="
docker ps -a | head -30 || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: olake-iceberg-catalog-test-logs-${{ matrix.tag }}
path: test/s3tables/catalog_olake/test-output.log
retention-days: 3
polaris-integration-tests:
name: Polaris Integration Tests
runs-on: ubuntu-22.04
+4
View File
@@ -289,6 +289,7 @@ jobs:
s3tests/functional/test_s3.py::test_object_write_check_etag \
s3tests/functional/test_s3.py::test_object_write_cache_control \
s3tests/functional/test_s3.py::test_object_write_expires \
s3tests/functional/test_s3.py::test_object_content_encoding_aws_chunked \
s3tests/functional/test_s3.py::test_object_write_read_update_read_delete \
s3tests/functional/test_s3.py::test_object_metadata_replaced_on_put \
s3tests/functional/test_s3.py::test_object_write_file \
@@ -311,6 +312,7 @@ jobs:
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_good \
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_object_ifunmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_checksum_object_attributes \
s3tests/functional/test_s3.py::test_bucket_head \
s3tests/functional/test_s3.py::test_bucket_head_notexist \
s3tests/functional/test_s3.py::test_object_raw_authenticated \
@@ -1149,6 +1151,7 @@ jobs:
s3tests/functional/test_s3.py::test_object_write_check_etag \
s3tests/functional/test_s3.py::test_object_write_cache_control \
s3tests/functional/test_s3.py::test_object_write_expires \
s3tests/functional/test_s3.py::test_object_content_encoding_aws_chunked \
s3tests/functional/test_s3.py::test_object_write_read_update_read_delete \
s3tests/functional/test_s3.py::test_object_metadata_replaced_on_put \
s3tests/functional/test_s3.py::test_object_write_file \
@@ -1171,6 +1174,7 @@ jobs:
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_good \
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_object_ifunmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_checksum_object_attributes \
s3tests/functional/test_s3.py::test_bucket_head \
s3tests/functional/test_s3.py::test_bucket_head_notexist \
s3tests/functional/test_s3.py::test_object_raw_authenticated \
+2 -3
View File
@@ -17,10 +17,11 @@ SeaweedFS is a simple and highly scalable distributed file system. There are two
1. to store billions of files!
2. to serve the files fast!
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently.
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently. Both read and write operations have O(1) complexity and can run at the full speed supported by the underlying hardware.
- [Download Binaries for different platforms](https://github.com/seaweedfs/seaweedfs/releases/latest)
- [Wiki Documentation](https://github.com/seaweedfs/seaweedfs/wiki)
- [HTTP REST API](REST_API.md) for the filer, master, and volume servers
- Community: [Slack](https://join.slack.com/t/seaweedfs/shared_invite/enQtMzI4MTMwMjU2MzA3LTEyYzZmZWYzOGQ3MDJlZWMzYmI0OTE4OTJiZjJjODBmMzUxNmYwODg0YjY3MTNlMjBmZDQ1NzQ5NDJhZWI2ZmY), [Twitter](https://twitter.com/SeaweedFS), [Telegram](https://t.me/Seaweedfs), [Reddit](https://www.reddit.com/r/SeaweedFS/), [Mailing List](https://groups.google.com/d/forum/seaweedfs)
- [SeaweedFS White Paper](https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS_Architecture.pdf) and introduction slides: [2025.5](https://docs.google.com/presentation/d/1tdkp45J01oRV68dIm4yoTXKJDof-EhainlA0LMXexQE/edit?usp=sharing), [2021.5](https://docs.google.com/presentation/d/1DcxKWlINc-HNCjhYeERkpGXXm6nTCES8mi2W5G0Z4Ts/edit?usp=sharing), [2019.3](https://www.slideshare.net/chrislusf/seaweedfs-introduction)
@@ -400,8 +401,6 @@ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
The text of this page is available for modification and reuse under the terms of the Creative Commons Attribution-Sharealike 3.0 Unported License and the GNU Free Documentation License (unversioned, with no invariant sections, front-cover texts, or back-cover texts).
[Back to TOC](#table-of-contents)
# Sponsors #
+344
View File
@@ -0,0 +1,344 @@
# SeaweedFS HTTP REST API
SeaweedFS exposes three HTTP surfaces:
| Service | Default port | Addressing |
|---------|--------------|------------|
| Filer | 8888 | File system paths (`/dir/name`) |
| Master | 9333 | File id assignment and cluster topology |
| Volume server | 8080 | File content by file id (`vid,fid`) |
Most clients only need the filer API (paths) or the S3 API. The master and
volume APIs are the lower-level blob store interface.
Conventions applying to all three:
- Responses are JSON unless noted otherwise. Append `&pretty=y` to pretty-print.
- A file id (`fid`) has the form `volumeId,fileKeyCookie`, e.g. `3,01637037d6`.
An optional suffix selects a reserved id from a `count` assignment
(`3,01637037d6_1`, `_2`, ...), and an optional extension
(`3,01637037d6.jpg`) sets the content type on reads.
- `replication` is a 3-digit replica placement `xyz`: `x` copies in other
data centers, `y` on other racks in the same data center, `z` on other
volume servers on the same rack. `000` = no replication, `001` = one copy
on the same rack, `010` = one copy on a different rack, `100` = one copy in
another data center, `200` = two copies in two other data centers, `110` =
one copy in another data center plus one on another rack.
- `ttl` units: `m` minute, `h` hour, `d` day, `w` week, `M` month, `y` year.
## Filer API (port 8888)
The filer presents a POSIX-like namespace over the volume servers.
### Upload a file
```bash
# PUT the raw body to the target path
curl -T /home/chris/myphoto.jpg "http://localhost:8888/dir/myphoto.jpg"
# or POST as multipart form (the part filename becomes the entry name)
curl -F file=@/home/chris/myphoto.jpg "http://localhost:8888/dir/"
```
Response `201 Created`:
```json
{"name":"myphoto.jpg","size":43234,"eTag":"0x6c656...","mtime":"...","chunks":[...]}
```
Query parameters:
| Parameter | Description | Default |
|-----------|-------------|---------|
| `collection` | collection name | empty |
| `replication` | replica placement code | filer default |
| `ttl` | file expiration, e.g. `3d` | never |
| `disk` | disk type to store on | filer default |
| `fsync` | `true` fsyncs on the volume server | false |
| `dataCenter` | preferred data center | empty |
| `rack` | preferred rack | empty |
| `dataNode` | preferred volume server | empty |
| `saveInside` | store small content inside the metadata instead of a volume | false |
| `maxMB` | split the upload into chunks of this many MB | filer `-maxMB` |
| `mode` | unix permission bits, e.g. `0644` | `0660` |
| `op` | `append` appends to an existing file | overwrite |
| `skipCheckParentDir` | `true` skips the parent-directory existence check | false |
### Create a directory
```bash
curl -X POST "http://localhost:8888/dir/newdir/"
```
A POST to a path ending in `/` with no content creates the directory,
including missing parents.
### Read a file
```bash
curl "http://localhost:8888/dir/myphoto.jpg"
```
Supports `Range` requests (`Accept-Ranges: bytes`), `ETag`, and the
`If-None-Match` / `If-Modified-Since` conditional headers. `HEAD` returns
headers only. Entry headers stored as extended attributes are echoed back,
minus internal `Seaweed-` and `xattr-` keys.
Entry metadata instead of content:
```bash
curl "http://localhost:8888/dir/myphoto.jpg?metadata=true"
```
`metadata=true&resolveManifest=true` additionally resolves chunked-manifest
entries into their real chunk list.
### List a directory
```bash
curl -H "Accept: application/json" "http://localhost:8888/dir/?limit=10&lastFileName=a.jpg"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `limit` | max entries per page | filer `-dirListLimit` |
| `lastFileName` | resume listing after this entry name | empty |
| `namePattern` | include only names matching the wildcard | empty |
| `namePatternExclude` | exclude names matching the wildcard | empty |
The JSON response carries `Path`, `Entries`, `Limit`, `LastFileName`,
`ShouldDisplayLoadMore`, and `EmptyFolder`. Without the `Accept` header the
filer renders its HTML browser.
### Move and copy
```bash
curl -X POST "http://localhost:8888/dir/newname.jpg?mv.from=/dir/myphoto.jpg"
curl -X POST "http://localhost:8888/dir/copy.jpg?cp.from=/dir/myphoto.jpg"
```
`mv.from` renames or moves the source to the request path (`204 No Content`).
`cp.from` copies it.
### Append
```bash
curl -T chunk2.bin "http://localhost:8888/dir/file.bin?op=append"
```
### Delete
```bash
curl -X DELETE "http://localhost:8888/dir/myphoto.jpg"
curl -X DELETE "http://localhost:8888/dir/?recursive=true"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `recursive` | delete a non-empty directory tree | false; when the filer runs with `filer.options.recursive_delete=true`, deletes are recursive unless `recursive=false` |
| `ignoreRecursiveError` | keep deleting remaining entries after an error | false |
| `skipChunkDeletion` | remove only the metadata, keep volume data | false |
### Tagging
Tags are carried as `Seaweed-`-prefixed request headers, not query
parameters; `?tagging` selects the tagging handler and `?tagging=K1,K2`
lists the keys to remove. Header names are canonicalized on write
(`Seaweed-k1` is stored as `Seaweed-K1`), and the delete list is matched
case-sensitively against the stored names.
```bash
curl -X PUT -H "Seaweed-k1: v1" -H "Seaweed-k2: v2" "http://localhost:8888/dir/file.jpg?tagging"
curl -X DELETE "http://localhost:8888/dir/file.jpg?tagging=K1,K2"
```
### Read by file id
```bash
curl "http://localhost:8888/?proxyChunkId=3,01637037d6"
```
The filer proxies the chunk read to the right volume server, so only the
filer port needs to be exposed.
### Resumable uploads
The filer serves the [TUS protocol](https://tus.io/) for resumable uploads
(`POST`, `PATCH`, `HEAD` on upload URLs). It is enabled by default at
`/.tus`; `-tusBasePath` changes the endpoint base path.
### Health
`GET /healthz` and `GET /readyz` return `200 OK`.
## Master API (port 9333)
Write-affecting endpoints are automatically proxied to the current leader, so
any master in the quorum can serve them.
### Assign a file id
```bash
curl "http://localhost:9333/dir/assign?count=1&replication=001&collection=turbo&dataCenter=dc1&ttl=3d&disk=ssd"
{"count":1,"fid":"3,01637037d6","url":"127.0.0.1:8080","publicUrl":"localhost:8080"}
```
Upload the file content to `http://<url>/<fid>` afterwards. With `count>1`,
use `<fid>_1`, `<fid>_2`, ... for the additional ids.
| Parameter | Description | Default |
|-----------|-------------|---------|
| `count` | file ids to reserve | 1 |
| `collection` | collection name | empty |
| `dataCenter` | preferred data center | empty |
| `rack` | preferred rack | empty |
| `dataNode` | preferred volume server | empty |
| `replication` | replica placement | master `-defaultReplication` |
| `ttl` | file expiration, e.g. `3d` | never |
| `disk` | disk type | empty |
| `dataSize` | expected file size in bytes | 0 |
| `preallocate` | bytes to preallocate for new volumes | master `-volumePreallocate` |
| `writableVolumeCount` | grow this many volumes when none are writable | master default |
| `memoryMapMaxSizeMb` | memory-mapped file size (Windows) | 0 |
### Look up a volume or file id
```bash
curl "http://localhost:9333/dir/lookup?volumeId=3"
{"locations":[{"url":"localhost:8080","publicUrl":"localhost:8080"}]}
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `volumeId` | volume id; a full `vid,fid` is accepted too | required |
| `fileId` | like `volumeId`, but also returns a write JWT when security is on | empty |
| `collection` | speeds up the lookup | empty |
| `read` | `yes` generates a read JWT instead of a write JWT | empty |
### Store a file in one call
```bash
curl -F file=@/home/chris/report.pdf "http://localhost:9333/submit?collection=turbo&replication=001"
{"fileName":"report.pdf","fid":"3,01637037d6","fileUrl":"localhost:8080/3,01637037d6","size":43234,"eTag":"0x6c656..."}
```
`POST /submit` accepts multipart file data plus the `dir/assign` placement
parameters (`count`, `collection`, `dataCenter`, `rack`, `replication`,
`ttl`, `disk`), assigns a file id, uploads to the volume server, and returns
the result.
### Redirect to a file
```bash
curl -v "http://localhost:9333/3,01637037d6"
```
`GET /{fileId}` answers `308 Permanent Redirect` to a volume server holding
the file, preserving the query string (e.g. image-resize parameters).
### Cluster status
```bash
curl "http://localhost:9333/dir/status?pretty=y" # full topology tree
curl "http://localhost:9333/vol/status?pretty=y" # every volume on every node
curl "http://localhost:9333/collection/info?collection=turbo"
curl "http://localhost:9333/collection/info?collection=turbo&detail=true"
```
`collection/info` returns aggregated `TotalSize`, `FileCount`, `UsedSize`,
`VolumeCount`; `detail=true` splits them per volume layout.
### Grow volumes
```bash
curl "http://localhost:9333/vol/grow?count=4&replication=001&collection=turbo&ttl=5d&disk=ssd&dataCenter=dc1&rack=rack1"
{"count":4}
```
`count` is required; the placement parameters match `dir/assign`. One volume
serves one write at a time, so pre-allocated volumes raise write concurrency.
### Vacuum deleted space
```bash
curl "http://localhost:9333/vol/vacuum?garbageThreshold=0.4"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `garbageThreshold` | minimum deleted-bytes ratio before a volume is compacted | master `-garbageThreshold` (0.3) |
Vacuuming makes a volume read-only, copies live needles to a new volume, and
swaps it in.
### Delete a collection
```bash
curl "http://localhost:9333/col/delete?collection=benchmark"
```
Deletes all volumes of the collection, including erasure-coded shards.
`204 No Content` on success.
### Health
```bash
curl -I "http://localhost:9333/healthz" # liveness
curl -I "http://localhost:9333/readyz" # readiness
curl "http://localhost:9333/" # web UI
```
## Volume server API (port 8080)
The volume server stores file content by file id. Clients normally get the
volume URL from `dir/assign` or `dir/lookup`.
### Upload
```bash
curl -F file=@/home/chris/myphoto.jpg "http://127.0.0.1:8080/3,01637037d6"
{"name":"myphoto.jpg","size":43234,"eTag":"0x6c656...","mime":"image/jpeg","contentMd5":"..."}
```
PUT or POST the body (or a multipart `file` part) to `/{vid},{fid}`.
`204 No Content` is returned when the content is unchanged. `?ts=<unix>`
sets the stored modification time.
### Read
```bash
curl "http://127.0.0.1:8080/3,01637037d6"
curl "http://127.0.0.1:8080/3,01637037d6.jpg" # sets Content-Type from the extension
```
Supports `Range` and `HEAD`. Image files can be resized server-side:
| Parameter | Description |
|-----------|-------------|
| `width`, `height` | resize bounds in pixels |
| `mode` | `fit` (contain) or `fill` (cover); omitted resizes to `width`/`height` |
| `crop_x1`, `crop_y1`, `crop_x2`, `crop_y2` | explicit crop rectangle |
| `cm` | `false` returns the chunk-manifest blob instead of resolving it |
| `readDeleted` | `true` reads soft-deleted needles |
| `collection` | passed through redirects for the right volume |
### Delete
```bash
curl -X DELETE "http://127.0.0.1:8080/3,01637037d6"
{"size":43234}
```
`?ts=<unix>` sets the deletion timestamp. Replicated volumes propagate the
delete to every replica.
### Status
```bash
curl "http://localhost:8080/status?pretty=y" # disk and volume inventory
curl -I "http://localhost:8080/healthz" # liveness/readiness
```
`OPTIONS` preflights answer CORS headers. When `-port.public` differs from
`-port`, the volume server opens a separate read-only public listener on
that port; `-publicUrl` sets the address it advertises to clients.
+6
View File
@@ -14,6 +14,9 @@ RUN cd /go/src/github.com/seaweedfs/seaweedfs && \
git checkout $BRANCH) || \
(echo "ERROR: Branch/commit $BRANCH not found in repository" && \
echo "Available branches:" && git branch -a && exit 1))
# seaweed-common only exists on revisions that have it; a BRANCH predating it
# still needs the directory so the COPY into rust_builder below never fails.
RUN mkdir -p /go/src/github.com/seaweedfs/seaweedfs/seaweed-common
ARG TARGETOS TARGETARCH TARGETVARIANT
RUN cd /go/src/github.com/seaweedfs/seaweedfs/weed \
&& export LDFLAGS="-X github.com/seaweedfs/seaweedfs/weed/util/version.COMMIT=$(git rev-parse --short HEAD)" \
@@ -31,6 +34,9 @@ ARG TAGS
COPY weed-volume-prebuilt/ /prebuilt/
COPY weed-worker-prebuilt/ /prebuilt-worker/
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-volume /build/seaweed-volume
# seaweed-common is a path dependency of seaweed-volume that lives beside it,
# so the source build below needs it in the same relative position.
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-common /build/seaweed-common
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/weed /build/weed
WORKDIR /build/seaweed-volume
RUN if [ -f "/prebuilt/weed-volume-${TARGETARCH}" ]; then \
+5 -2
View File
@@ -1,11 +1,14 @@
FROM alpine:latest
# Install required packages
RUN apk add --no-cache \
RUN apk upgrade --no-cache && \
apk add --no-cache \
ca-certificates \
fuse \
curl \
jq
jq \
libcrypto3 \
libssl3
# Copy our locally built binary
COPY weed-local /usr/bin/weed
+2
View File
@@ -30,3 +30,5 @@ sleep_minutes = 17 # sleep minutes between each script execution
bucket = "volume_bucket" # an existing bucket
endpoint = "http://server2:8333"
storage_class = "STANDARD_IA"
# upload_concurrency = 5 # concurrent multipart part uploads per volume (volume.tier.upload -concurrent overrides)
# download_concurrency = 5 # concurrent multipart part downloads per volume (volume.tier.download -concurrent overrides)
+47 -47
View File
@@ -1,6 +1,6 @@
module github.com/seaweedfs/seaweedfs
go 1.26.0
go 1.26.6
require (
cloud.google.com/go v0.123.0 // indirect
@@ -14,7 +14,7 @@ require (
github.com/coreos/go-semver v0.3.1 // indirect
github.com/coreos/go-systemd/v22 v22.7.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/dustin/go-humanize v1.0.1
github.com/dustin/go-humanize v1.1.0
github.com/eapache/go-resiliency v1.6.0 // indirect
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 // indirect
github.com/eapache/queue v1.1.0 // indirect
@@ -25,7 +25,7 @@ require (
github.com/facebookgo/subset v0.0.0-20200203212716-c811ad88dec4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-redsync/redsync/v4 v4.17.0
github.com/go-sql-driver/mysql v1.10.0
github.com/go-sql-driver/mysql v1.10.1
github.com/go-zookeeper/zk v1.0.4 // indirect
github.com/golang/protobuf v1.5.4
github.com/golang/snappy v1.0.0
@@ -70,7 +70,7 @@ require (
github.com/spf13/afero v1.15.0 // indirect
github.com/spf13/cast v1.10.0 // indirect
github.com/spf13/viper v1.21.0
github.com/stretchr/testify v1.11.1
github.com/stretchr/testify v1.12.1
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203
github.com/syndtr/goleveldb v1.0.1-0.20190318030020-c3a204f8e965
github.com/tidwall/gjson v1.18.0
@@ -90,18 +90,18 @@ require (
gocloud.dev v0.46.0
gocloud.dev/pubsub/natspubsub v0.46.0
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.56.0
golang.org/x/crypto v0.57.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.45.0
golang.org/x/image v0.46.0
golang.org/x/net v0.58.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/oauth2 v0.37.0
golang.org/x/sys v0.48.0
golang.org/x/text v0.42.0 // indirect
golang.org/x/tools v0.49.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.296.0
google.golang.org/api v0.297.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7
google.golang.org/protobuf v1.36.12
gopkg.in/inf.v0 v0.9.1 // indirect
modernc.org/b v1.0.0 // indirect
@@ -111,7 +111,7 @@ require (
)
require (
cloud.google.com/go/kms v1.33.0
cloud.google.com/go/kms v1.35.0
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0
github.com/DATA-DOG/go-sqlmock v1.5.2
github.com/Jille/raft-grpc-transport v1.6.1
@@ -122,9 +122,9 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2 v1.47.0
github.com/aws/aws-sdk-go-v2/config v1.32.35
github.com/aws/aws-sdk-go-v2/credentials v1.20.1
github.com/aws/aws-sdk-go-v2/credentials v1.20.4
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
@@ -146,21 +146,21 @@ require (
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rclone/rclone v1.75.1
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/redis/go-redis/v9 v9.22.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.7
github.com/tarantool/go-option v1.1.0
github.com/tarantool/go-tarantool/v3 v3.0.1
github.com/tarantool/go-tarantool/v3 v3.0.2
github.com/testcontainers/testcontainers-go v0.44.0
github.com/tikv/client-go/v2 v2.0.7
github.com/twmb/avro v1.8.0
github.com/twmb/avro v1.9.0
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
go.uber.org/atomic v1.12.0
golang.org/x/sync v0.23.0
golang.org/x/tools/godoc v0.1.0-deprecated
google.golang.org/grpc/security/advancedtls v1.0.0
)
@@ -185,7 +185,7 @@ require (
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
@@ -283,34 +283,34 @@ require (
github.com/xeipuuv/gojsonreference v0.0.0-20180127040603-bd5ef7bd5415 // indirect
github.com/xo/terminfo v0.0.0-20220910002029-abceb7e1c41e // indirect
github.com/zeebo/xxh3 v1.1.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/mock v0.5.2 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/mod v0.38.0 // indirect
go.yaml.in/yaml/v3 v3.0.5 // indirect
golang.org/x/mod v0.41.0 // indirect
gonum.org/v1/gonum v0.17.0 // indirect
)
require (
cel.dev/expr v0.25.2 // indirect
cel.dev/expr v0.25.3 // indirect
cloud.google.com/go/auth v0.23.2 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
cloud.google.com/go/iam v1.12.0 // indirect
cloud.google.com/go/monitoring v1.30.0 // indirect
filippo.io/edwards25519 v1.2.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.8.0
github.com/Azure/azure-sdk-for-go/sdk/storage/azfile v1.7.0 // indirect
github.com/Azure/go-ntlmssp v0.1.1 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0 // indirect
github.com/Files-com/files-sdk-go/v3 v3.3.194 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/resourcemapping v0.57.0 // indirect
github.com/IBM/go-sdk-core/v5 v5.23.1 // indirect
@@ -327,20 +327,20 @@ require (
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0
github.com/aws/smithy-go v1.28.1
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
@@ -363,7 +363,7 @@ require (
github.com/elastic/gosigar v0.14.3 // indirect
github.com/emersion/go-message v0.18.2 // indirect
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.37.0 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be // indirect
github.com/envoyproxy/protoc-gen-validate v1.3.3 // indirect
github.com/fatih/color v1.18.0 // indirect
github.com/felixge/httpsnoop v1.1.0 // indirect
@@ -388,7 +388,7 @@ require (
github.com/gogo/protobuf v1.3.2 // indirect
github.com/golang-jwt/jwt/v4 v4.5.2 // indirect
github.com/google/s2a-go v0.1.9 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.20 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.21 // indirect
github.com/gorilla/schema v1.4.1 // indirect
github.com/gorilla/securecookie v1.1.2 // indirect
github.com/gorilla/sessions v1.4.0
@@ -438,7 +438,7 @@ require (
github.com/oracle/oci-go-sdk/v65 v65.121.0 // indirect
github.com/panjf2000/ants/v2 v2.12.1 // indirect
github.com/patrickmn/go-cache v2.1.0+incompatible // indirect
github.com/pelletier/go-toml/v2 v2.4.1 // indirect
github.com/pelletier/go-toml/v2 v2.4.3 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.29
@@ -489,9 +489,9 @@ require (
go.etcd.io/bbolt v1.5.0 // indirect
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 // indirect
go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.45.0 // indirect
@@ -499,9 +499,9 @@ require (
go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
golang.org/x/term v0.46.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
+92 -92
View File
@@ -6,8 +6,8 @@ atomicgo.dev/keyboard v0.2.9 h1:tOsIid3nlPLZ3lwgG8KZMp/SFmr7P0ssEN5JUsm78K8=
atomicgo.dev/keyboard v0.2.9/go.mod h1:BC4w9g00XkxH/f1HXhW2sXmJFOCWbKn9xrOunSFtExQ=
atomicgo.dev/schedule v0.1.0 h1:nTthAbhZS5YZmgYbb2+DH8uQIZcTlIrd4eYr3UQxEjs=
atomicgo.dev/schedule v0.1.0/go.mod h1:xeUa3oAkiuHYh8bKiQBRojqAMq3PXXbJujjb0hw8pEU=
cel.dev/expr v0.25.2 h1:K6j46C81hXtZQfuX60cVWQFBJahKSE2gfRbNuvr5bFs=
cel.dev/expr v0.25.2/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cel.dev/expr v0.25.3 h1:A2jO8jwOugrrovveCWfj0KEZOfqiLgAcwjpHPhzIGw0=
cel.dev/expr v0.25.3/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cloud.google.com/go v0.26.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.34.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.38.0/go.mod h1:990N+gfupTy94rShfmMCWGDn0LpTmnzTp2qbd1dvSRU=
@@ -298,8 +298,8 @@ cloud.google.com/go/kms v1.4.0/go.mod h1:fajBHndQ+6ubNw6Ss2sSd+SWvjL26RNo/dr7uxs
cloud.google.com/go/kms v1.5.0/go.mod h1:QJS2YY0eJGBg3mnDfuaCyLauWwBJiHRboYxJ++1xJNg=
cloud.google.com/go/kms v1.6.0/go.mod h1:Jjy850yySiasBUDi6KFUwUv2n1+o7QZFyuUJg6OgjA0=
cloud.google.com/go/kms v1.9.0/go.mod h1:qb1tPTgfF9RQP8e1wq4cLFErVuTJv7UsSC915J8dh3w=
cloud.google.com/go/kms v1.33.0 h1:pG0X78m212b2pv9N4fdMoUO69LuZGQ9kSvn8sHBOFAo=
cloud.google.com/go/kms v1.33.0/go.mod h1:CSGvW6GnMQbY+1nOHcIzhMtHSbExXlOmCKjWtYVjcpA=
cloud.google.com/go/kms v1.35.0 h1:nJ/ktaqspx1nPM9vIcO0SHbhqCAm8nvAxL1siuVgKm0=
cloud.google.com/go/kms v1.35.0/go.mod h1:0++71pIHvJL+GmMa8K4jOWFq7gNOX3jm2PRMSJwTKJw=
cloud.google.com/go/language v1.4.0/go.mod h1:F9dRpNFQmJbkaop6g0JhSBXCNlO90e1KWx5iDdxbWic=
cloud.google.com/go/language v1.6.0/go.mod h1:6dJ8t3B+lUYfStgls25GusK04NLh3eDLQnWM3mdEbhI=
cloud.google.com/go/language v1.7.0/go.mod h1:DJ6dYN/W+SQOjF8e1hLQXMF21AkH2w9wiPzPCJa2MIE=
@@ -553,10 +553,10 @@ gioui.org v0.0.0-20210308172011-57750fc8a0a6/go.mod h1:RSH6KIUZ0p2xy5zHDxgAM4zum
git.sr.ht/~sbinet/gg v0.3.1/go.mod h1:KGYtlADtqsqANL9ueOFkWymvzUvLMQllU5Ixo+8v3pc=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6 h1:He8afgbRMd7mFxO99hRNu+6tazq8nFF9lIwo9JFroBk=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6/go.mod h1:8o94RPi1/7XTJvwPpRSzSUedZrtlirdB3r9Z20bi2f8=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0 h1:aokoqcHvaGjiM3VpjKDfMMnF/8epJ+Q1HLJ7CudztqE=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0/go.mod h1:/WYEx9pcM9Y+Dd/APJaNlSvVSvzl54rrMdZT5+Oi2LM=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0 h1:CU4+EJeJi3TKYWEcYuSdWsjzw0nVsK/H0MSQOiPcymU=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0/go.mod h1:q0+UTSRvShwUCrR/s5HtyInYphN7Wvxb7snFM3u+SLA=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1 h1:zvXfGJCWvywnCA814d8ZiVyt+fm9nnTE8xSb99zRyfo=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1/go.mod h1:iptorS+VYKFL2N6PnebpS91dubG35eAOEERnT4PJbQU=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1 h1:u93s+zU2JD62im61Bm5CZIc1ZrOJaIAWEg0WOrMVkEo=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1/go.mod h1:oXtinPO4OLj9d1DOTrqrL1oRwGhcqadvAmrl6wTeGlk=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.4.0 h1:xFaZZ+IubdftrDHnGGwZ6QvQ3KHTtWl2MCK+GMt2vxs=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.4.0/go.mod h1:mCBhUhlMjLLJKr5aqw2TNS/VqJOie8MzWq3DAMJeKso=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 h1:fhqpLE3UEXi9lPaBRpQ6XuRW0nU7hgg4zlmZZa+a9q4=
@@ -577,8 +577,8 @@ github.com/Azure/go-ntlmssp v0.1.1 h1:l+FM/EEMb0U9QZE7mKNEDw5Mu3mFiaa2GKOoTSsNDP
github.com/Azure/go-ntlmssp v0.1.1/go.mod h1:NYqdhxd/8aAct/s4qSYZEerdPuH1liG2/X9DiVTbhpk=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1 h1:WJTmL004Abzc5wDB5VtZG2PJk5ndYDgVacGqfirKxjM=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1/go.mod h1:tCcJZ0uHAmvjsVYzEFivsRTN00oz5BEsRgQHu5JZ9WE=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 h1:RHK7bS+HQMslb1sZpAokUt+zTVmue0hKSs2C791hhzU=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2/go.mod h1:HKpQxkWaGLJ+D/5H8QRpyQXA1eKjxkFlOMwck5+33Jk=
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0 h1:Nljr4q1GRA/5vCrMONS+g4u4LRHNgOXVSh3O43J2CnI=
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0/go.mod h1:Y33QHnf0FfdVewFFISOGe20mkZbxX4H839o955/PoeI=
github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU=
github.com/BurntSushi/xgb v0.0.0-20160522181843-27f122750802/go.mod h1:IVnqGOEym/WlBOVXweHU+Q+/VP0lqqI8lqeDx9IjBqo=
github.com/Codefor/geohash v0.0.0-20140723084247-1b41c28e3a9d h1:iG9B49Q218F/XxXNRM7k/vWf7MKmLIS8AcJV9cGN4nA=
@@ -593,8 +593,8 @@ github.com/FilenCloudDienste/filen-sdk-go v0.0.39 h1:tgV5jYL6dsXop9TpDTIQU6UwJjw
github.com/FilenCloudDienste/filen-sdk-go v0.0.39/go.mod h1:0cBhKXQg49XbKZZfk5TCDa3sVLP+xMxZTWL+7KY0XR0=
github.com/Files-com/files-sdk-go/v3 v3.3.194 h1:dtOFxSTWWRpkmvXa6ycNiw8dVDu1wkgzcXyVV1VafNc=
github.com/Files-com/files-sdk-go/v3 v3.3.194/go.mod h1:rl0WumSN9gSo775DgvQv+wMQ8rlb0ES/1hU5jkMtLXg=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 h1:yzIYdwuro811Z27D3T80Wkd3rqZzb0K43nner7Eh1yE=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0/go.mod h1:pJTkW8hEUIIi3Pf65lPZOnn4Y81yCllX6IWk2jNXdkM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 h1:bN1gA3of5bXtbnLsRPrwfmbbe7A5UWFlcTHseujLnpc=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0/go.mod h1:Yj5vHEz/aAepZGliRJsA6uvHAVAQyEwajq9ORCHPxzM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 h1:jLdiS1vO+XJFyDSWRHBx56r4s/NNtcl5J6KyCcWUX/w=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0/go.mod h1:8lmpHY+1VRoteiOwyrQMDt1YGXOrFKCz+1wJW7n3ODY=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/cloudmock v0.57.0 h1:cSjUzZ7KU8hicTgzaSv9NmSyM9fTVK3y5lsBUl3wOis=
@@ -710,48 +710,48 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2 v1.47.0 h1:0jsHallhJCeaU0Ko48c/3FK1ctOQ7NpzggxriJOQ8MQ=
github.com/aws/aws-sdk-go-v2 v1.47.0/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 h1:LAfOuhAH331fmOjTQpAaOlH+Ftn7RzSDJ2VFwjdMMy4=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18/go.mod h1:4e5xhuXHx1e4U9EthvbPP1r/DIMp5c2823OL8karzcM=
github.com/aws/aws-sdk-go-v2/config v1.32.35 h1:UEzXuET8E42lxBPijuACu/tEK7v5lFPlk0Q+GT5WD9E=
github.com/aws/aws-sdk-go-v2/config v1.32.35/go.mod h1:KaMtJpFa2JlL2BStjjHQVwQpzZEmw+ND/EgVrfFoo2g=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1 h1:Z8GRNEx0u9sDkZOq4PUnN8mjGwbUQGRzMSXpvt3d8xQ=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1/go.mod h1:uBIK00kFo95dnemqfFMTWx0X8YRqsh6ecIoCjjOkZqM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 h1:YIEBqcqRnpi4Pfv0YHImtgi6czGCwKHANC7SwmUAVD0=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1/go.mod h1:imEf0oufgAo8KAkCHhrOdqGEC0YWx1PPBQH82shSxGw=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4 h1:hTvrJJseKbvw32kmiE0G+u/9ZqpqscjDrTigHIXP2qs=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4/go.mod h1:gWp9O1ZBWwpcIrgV+mVHk4gZUurAEDkgypu/OXOlIaw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 h1:AM4hHjww+PSFtt6E+UrBrPlZkWsePCLEt9AjkfQX+yM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0/go.mod h1:3x/yXezeQjpOvBb4jEMxrS8SXvpdvJ5abv6l5c1gWM8=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11 h1:eBXB8KZgzQ8A9QB4iJS4aw/u6+4OY3i2hQXPABeAIOg=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11/go.mod h1:N9+5pG27Fy61GUL5YXVLXDTLmUudMrgwsuDbgBMNLxQ=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 h1:pc138gM1CW+XPc60rEwUlwwuwWFQK16CI1T7v1F9Oec=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1/go.mod h1:1+koxpPIbfBdfzP6vojm5/zTpTQ/micYwlxIiNB3TxI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 h1:K0JsbZQj+1h208Ro1zHeA4l7bMp0NvRffHQ91q8Ol1s=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1/go.mod h1:W3/vL6EtCIatICGy9ab29QhMuae+cOKPWcMxv02CO+Q=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 h1:yhw5KD1phVyP9vijxOUzDfEtJx+bt+L63k+VfuiYFAA=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1/go.mod h1:ZW2e0d7DYlRxlS9hEiMXE47gTdX5KRN4byUiNbUpG+Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 h1:Hp/VgjP0BysR3OgLlR057Vz2LcbbVnoWeJ+3qWiS/fY=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3/go.mod h1:nwGV5qw7F1IZPgxCvA/ph8N2TAuz+BkRG/bXn808qMA=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 h1:MUaM4f+kj1ZIBPZfUS8cxP1GKXXZtHJjAthy93AN7SM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3/go.mod h1:6YmVmEVRI5ZZzRjCSsb9SryKH0hAlMRdgA7kG9aDvBU=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 h1:fuSCw4Z2qfRCztMPO3GXJNSiEp6Wee+WOLwrHHUMy9c=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3/go.mod h1:6SxcHheD1pPR5+kWm1wGvjlL/YqUsh267sAfEmN4K7A=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 h1:uZOinZb+h7lZw8IYzP1z1IuEnueB76/EFkcf/fEW4Ag=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31/go.mod h1:NRtwAM/p5VRt03TlEUs0pH3TeWamWdf4YyJpSrzPYLc=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 h1:RmmWQPREQdk9U+PfqeHW3MqZaBaNK7TpV9W3RY+b+7g=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1/go.mod h1:0A3W4F+68ZnNk5XcNL/e9HFMwnP8RlEicFfy6eOEDyw=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 h1:bON1rJf67TSTDCKg816AAIE4xSTtoo9tl0XRkO72R+I=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3/go.mod h1:c5BBpjJcQXpfeq9iASyVKA3T6vX6B6LEXY4mL/gklDY=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 h1:HLPAVrlLDaN2boN0xJx7MgaQDNEO3Q+c9L6kl/8m47Q=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39/go.mod h1:Pg/dVfsNkm1hsIDK/gMvCKtmyNfNTV12mrgHqVE/6Oo=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3 h1:IKoCZqfWfZzSBi16QFQ+QcbQ3LRQ7QgB1S5tDAyPBQQ=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3/go.mod h1:RBpRcXiM4s2pOInVs32GsBonnje+fiAj4mcrStRmlCA=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 h1:mdMtSVKdQ3+mzBh+l0ogrFYZVQUCg6pJZOirA2ARsYE=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1/go.mod h1:9IqUlsJDbUPcg6cgx3WEzXdjrbWzLDQrak0aaSqlTcI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 h1:ZD5qFpWcaOKdTuhBi431pIDkCgrMkMlMT6jlpSPoIRI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0/go.mod h1:8Nuuf+tR346PjJ3MvZPh9pekbLiLQFWJhzMXfwy7alA=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 h1:B6WFn91tobD6gG4724ONHaqrpKsoETGnv98LHe/yIGM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1/go.mod h1:tWuiVBUtPBr8/rgRiYS8Uf85sHcAN+G7XS3D3CEoUh8=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 h1:6yeYCWFvgbI2TI3K6jr9LtBNhXgJ7g4xqD+DEiaDDmM=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1/go.mod h1:naFe83jSMuYkH+QjQPX8n1MLhBkeCFM5Lsnh5m5wz3c=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1 h1:Sv2xPnRHlThSUtVujYuUBPI/Il8si6UPHXL8DMiB/F0=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1/go.mod h1:mKo/CzaCz8qytGW70NG4vIIGAx1HXTlb5lHNkC5k3lk=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 h1:JGeeBcMlhg1xtOXYpeCaTQBZObtXMPQCUqBcmr65NRA=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0/go.mod h1:XwteswG9EOMRFm73UT0t+MbTwyLxMrEXkU6e+v92Lzo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 h1:obhahQXDEdVEv8y5bTKXR30LVaxYe1kyYM0L7l2Iq+k=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0/go.mod h1:6twZZ/aXHNy1vXUO8koUbp++MYzMASkOgEBdkbJYmO0=
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0 h1:Zpnqa6XtrNzXZnwbdCqHOXpXhMsa01ql/pcRQ1sb4hk=
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0/go.mod h1:/8JRcdTt//hG0Q4BTmGbuOplT7ABe+5rdtqUHqXvYIM=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
@@ -953,8 +953,8 @@ github.com/dropbox/dropbox-sdk-go-unofficial/v6 v6.4.0/go.mod h1:gDXhl0OElhzYoDs
github.com/dsnet/try v0.0.3 h1:ptR59SsrcFUYbT/FhAbKTV6iLkeD6O18qfIWRml2fqI=
github.com/dsnet/try v0.0.3/go.mod h1:WBM8tRpUmnXXhY1U6/S8dt6UWdHTQ7y8A5YSkRCkq40=
github.com/dustin/go-humanize v1.0.0/go.mod h1:HtrtbFcZ19U5GC7JDqmcUSB87Iq5E25KnS6fMYU6eOk=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/dustin/go-humanize v1.1.0 h1:dbKTrvD0klcbBV/h4AWJdMuZogJACoMlvWIWZ5b2xWg=
github.com/dustin/go-humanize v1.1.0/go.mod h1:hc1CvRkJMsgxqjmjMQF3QNRAZBwY8AXBAzKYoSX9sFI=
github.com/eapache/go-resiliency v1.6.0 h1:CqGDTLtpwuWKn6Nj3uNUdflaq+/kIPsg0gfNzHton30=
github.com/eapache/go-resiliency v1.6.0/go.mod h1:5yPzW0MIvSe0JDsv0v+DvcjEv2FyD6iZYSs1ZI+iQho=
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 h1:Oy0F4ALJ04o5Qqpdz8XLIpNA3WM/iSIXqxtqo7UGVws=
@@ -987,8 +987,8 @@ github.com/envoyproxy/go-control-plane v0.10.3/go.mod h1:fJJn/j26vwOu972OllsvAgJ
github.com/envoyproxy/go-control-plane v0.11.0/go.mod h1:VnHyVMpzcLvCFt9yUz1UnCwHLhwx1WguiVDV7pTG/tI=
github.com/envoyproxy/go-control-plane v0.14.0 h1:hbG2kr4RuFj222B6+7T83thSPqLjwBIfQawTkC++2HA=
github.com/envoyproxy/go-control-plane v0.14.0/go.mod h1:NcS5X47pLl/hfqxU70yPwL9ZMkUlwlKxtAohpi2wBEU=
github.com/envoyproxy/go-control-plane/envoy v1.37.0 h1:u3riX6BoYRfF4Dr7dwSOroNfdSbEPe9Yyl09/B6wBrQ=
github.com/envoyproxy/go-control-plane/envoy v1.37.0/go.mod h1:DReE9MMrmecPy+YvQOAOHNYMALuowAnbjjEMkkWOi6A=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be h1:SWe0x6yfglnxuvOiYgTTnNq7QD/yvqthOh1RBI8Bj8w=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be/go.mod h1:PYEOlng9XcrulfyWpm49jECTPV0LT4q8cO7fLW/xwgk=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0 h1:/G9QYbddjL25KvtKTv3an9lx6VBE2cnb8wp1vEGNYGI=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0/go.mod h1:Wk+tMFAFbCXaJPzVVHnPgRKdUdwW/KdbRt94AzgRee4=
github.com/envoyproxy/protoc-gen-validate v0.1.0/go.mod h1:iSmxcyjqTsJpI2R4NaDN7+kN2VEUnK/pcBlmesArF7c=
@@ -1109,8 +1109,8 @@ github.com/go-redsync/redsync/v4 v4.17.0 h1:FFJ+uxZs44y4Sq10//IFKic9T94AYl+u3Sog
github.com/go-redsync/redsync/v4 v4.17.0/go.mod h1:CKVA6qwT07S/916i+Yd9h1/8YFQhCCpPYTQhvvYytJo=
github.com/go-resty/resty/v2 v2.17.2 h1:FQW5oHYcIlkCNrMD2lloGScxcHJ0gkjshV3qcQAyHQk=
github.com/go-resty/resty/v2 v2.17.2/go.mod h1:kCKZ3wWmwJaNc7S29BRtUhJwy7iqmn+2mLtQrOyQlVA=
github.com/go-sql-driver/mysql v1.10.0 h1:Q+1LV8DkHJvSYAdR83XzuhDaTykuDx0l6fkXxoWCWfw=
github.com/go-sql-driver/mysql v1.10.0/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-stack/stack v1.8.0/go.mod h1:v0f6uXyyMGvRgIKkXu+yp6POWl0qKG85gN/melR3HDY=
github.com/go-task/slim-sprig v0.0.0-20230315185526-52ccab3ef572 h1:tfuBGBXKqDEevZMzYi5KSi8KkcZtzBcTgAUUtapy0OI=
github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI=
@@ -1262,8 +1262,8 @@ github.com/googleapis/enterprise-certificate-proxy v0.1.0/go.mod h1:17drOmN3MwGY
github.com/googleapis/enterprise-certificate-proxy v0.2.0/go.mod h1:8C0jb7/mgJe/9KK8Lm7X9ctZC2t60YyIpYEI16jx0Qg=
github.com/googleapis/enterprise-certificate-proxy v0.2.1/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.2.3/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.3.20 h1:t/xL64VUoN69MuMRQuJETqYGOw4Z9mSRJK9epIEtwFk=
github.com/googleapis/enterprise-certificate-proxy v0.3.20/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/enterprise-certificate-proxy v0.3.21 h1:OFdQ3tnCX/zaQ0Cedur3D3z7kI6HiLX9g3TiAN4/DFU=
github.com/googleapis/enterprise-certificate-proxy v0.3.21/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/gax-go/v2 v2.1.0/go.mod h1:Q3nei7sK6ybPYH7twZdmQpAd1MKb7pfu6SK+H1/DsU0=
@@ -1652,8 +1652,8 @@ github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144T
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
github.com/patrickmn/go-cache v2.1.0+incompatible/go.mod h1:3Qf8kWWT7OJRJbdiICTKqZju1ZixQ/KpMGzzAfe6+WQ=
github.com/pborman/getopt v0.0.0-20170112200414-7148bc3a4c30/go.mod h1:85jBQOZwpVEaDAr341tbn15RS4fCAsIst0qp7i8ex1o=
github.com/pelletier/go-toml/v2 v2.4.1 h1:j5OMOImsH+j2k7GJ5YO+RxfWwohNiH6t5zB/+h3bagc=
github.com/pelletier/go-toml/v2 v2.4.1/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pelletier/go-toml/v2 v2.4.3 h1:GTRvJQutkOSftxIFD5xw9aepkYNuPWmVJpffdDPYVpY=
github.com/pelletier/go-toml/v2 v2.4.3/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 h1:XeOYlK9W1uCmhjJSsY78Mcuh7MVkNjTzmHx1yBzizSU=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14/go.mod h1:jVblp62SafmidSkvWrXyxAme3gaTfEtWwRPGz5cpvHg=
github.com/peterh/liner v1.2.2 h1:aJ4AOodmL+JxOZZEL2u9iJf8omNRpqHc/EbrK+3mAXw=
@@ -1770,8 +1770,8 @@ github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 h1:N/ElC8H3+5X
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475/go.mod h1:bCqnVzQkZxMG4s8nGwiZ5l3QUCyqpo9Y+/ZMZ9VjZe4=
github.com/rdleal/intervalst v1.5.0 h1:SEB9bCFz5IqD1yhfH1Wv8IBnY/JQxDplwkxHjT6hamU=
github.com/rdleal/intervalst v1.5.0/go.mod h1:xO89Z6BC+LQDH+IPQQw/OESt5UADgFD41tYMUINGpxQ=
github.com/redis/go-redis/v9 v9.21.0 h1:FPBE4hhbAke+TLmcY3WkpbDffJEomdqPn3HYiqAtL9E=
github.com/redis/go-redis/v9 v9.21.0/go.mod h1:v/M13XI1PVCDcm01VtPFOADfZtHf8YW3baQf57KlIkA=
github.com/redis/go-redis/v9 v9.22.0 h1:laDvpYXTJtZLloinw1fA5Kqd6HAEH2XKxOkG/PDq2F0=
github.com/redis/go-redis/v9 v9.22.0/go.mod h1:y2g0Wj8rQvuK0ELM+oxSudcLtC09JScs98I/X9gRWY4=
github.com/redis/rueidis v1.0.76 h1:RdDWuvlYBSp+bTrBvaXqJnNEL3VVzsnjo+0psPFgLc4=
github.com/redis/rueidis v1.0.76/go.mod h1:UsfHPSbomB6QAVMk4iiFkzRy0nh9o7scDGa+SitvBY4=
github.com/redis/rueidis/rueidiscompat v1.0.76 h1:7LikbiqCQqCsZXeZ+akgZMnjIV/J0VHih9PIX4gGZC4=
@@ -1897,8 +1897,8 @@ github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o
github.com/stretchr/testify v1.8.2/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4=
github.com/stretchr/testify v1.8.3/go.mod h1:sz/lmYIOXD/1dqDmKjjqLyZ2RngseejIcXlSw2iwfAo=
github.com/stretchr/testify v1.8.4/go.mod h1:sz/lmYIOXD/1dqDmKjjqLyZ2RngseejIcXlSw2iwfAo=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/stretchr/testify v1.12.1 h1:EuwCh5fleGS7H32xRwO3wRGT7DxrDhLAT6FF8MpWDWE=
github.com/stretchr/testify v1.12.1/go.mod h1:MDEgiDPPsNp5cuIrHPPCyornHKgEVbtFUmoNlxoYthg=
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203 h1:QVqDTf3h2WHt08YuiTGPZLls0Wq99X9bWd0Q5ZSBesM=
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203/go.mod h1:oqN97ltKNihBbwlX8dLpwxCl3+HnXKV/R0e+sRLd9C8=
github.com/subosito/gotenv v1.6.0 h1:9NlTDc1FTs4qu0DDq7AEtTPNw6SVm7uBMsUCUjABIf8=
@@ -1918,8 +1918,8 @@ github.com/tarantool/go-iproto v1.1.0 h1:HULVOIHsiehI+FnHfM7wMDntuzUddO09DKqu2Wn
github.com/tarantool/go-iproto v1.1.0/go.mod h1:LNCtdyZxojUed8SbOiYHoc3v9NvaZTB7p96hUySMlIo=
github.com/tarantool/go-option v1.1.0 h1:ShoOhNsdL41sRpm4hXCRDjV8H0WzPkd4UnKhLKbW//w=
github.com/tarantool/go-option v1.1.0/go.mod h1:hMr9z2JXOWlgdCBpCPSL2nwp8718GKYvNBJ+ZuzJbCo=
github.com/tarantool/go-tarantool/v3 v3.0.1 h1:vaUX4xmVmXh2dIJ/LqlX1MXK3iYqAqV6YiE54Wwl/qg=
github.com/tarantool/go-tarantool/v3 v3.0.1/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/tarantool/go-tarantool/v3 v3.0.2 h1:9ZtHllun80QX7KS9tqd3dXa+QAu2BfqFtX1ZWcNnaW8=
github.com/tarantool/go-tarantool/v3 v3.0.2/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/testcontainers/testcontainers-go v0.44.0 h1:/Fwh6HY1mIikhnm9e7HwoxGycx0lzRAE0f5VQpjFxzI=
github.com/testcontainers/testcontainers-go v0.44.0/go.mod h1:IcnwQrYTO86xHXu5bvMaBH7ATlbS3Qn1M1QWW3c66rE=
github.com/testcontainers/testcontainers-go/modules/compose v0.44.0 h1:8YcW51jhgpkkiRVe10Wj9TCBthJmoNpU2fK5WSf7TQ8=
@@ -1965,8 +1965,8 @@ github.com/tsuna/gohbase v0.0.0-20201125011725-348991136365/go.mod h1:zj0GJHGvyf
github.com/tv42/httpunix v0.0.0-20150427012821-b75d8614f926/go.mod h1:9ESjWnEqriFuLhtthL60Sar/7RFoluCcXsuvEwTV5KM=
github.com/twitchyliquid64/golang-asm v0.15.1 h1:SU5vSMR7hnwNxj24w34ZyCi/FmDZTkS4MhqMhdFk5YI=
github.com/twitchyliquid64/golang-asm v0.15.1/go.mod h1:a1lVb/DtPvCB8fslRZhAngC2+aY1QWCk3Cedj/Gdt08=
github.com/twmb/avro v1.8.0 h1:UMWLg+nH4P3yad5Om7yFSohYLy2RG1s7BcFFiOvmK9Q=
github.com/twmb/avro v1.8.0/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/avro v1.9.0 h1:JSiqewo3AANj7vlVCQXBsIZTA3LiiBble7Xp8T4lbtA=
github.com/twmb/avro v1.9.0/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/murmur3 v1.1.8 h1:8Yt9taO/WN3l08xErzjeschgZU2QSrwm1kclYq+0aRg=
github.com/twmb/murmur3 v1.1.8/go.mod h1:Qq/R7NUyOfr65zD+6Q5IHKsJLwP7exErjN6lyyq3OSQ=
github.com/twpayne/go-geom v1.6.1 h1:iLE+Opv0Ihm/ABIcvQFGIiFBXd76oBIar9drAwHFhR4=
@@ -2103,30 +2103,30 @@ go.opencensus.io v0.24.0 h1:y73uSU6J157QMP2kn2r30vwW1A2W2WFwSCGnAVxeaD0=
go.opencensus.io v0.24.0/go.mod h1:vNK8G9p7aAivkbmorf4v+7Hgx+Zs0yY+0fOtgBfjQKo=
go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64=
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 h1:NmLfL734pJhM0JKaYd2Y28+nY9dPRWYAAbxhRCrKXPw=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0/go.mod h1:tNAsgd8avTGke1+MndXlU5Cru4PQ9Ai/cCNWQv/ZJ/s=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 h1:9jR0ZPRok9ryaOQ2Wx8rg5F7Aon59mxrqbVI60/vlBk=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0/go.mod h1:VSme3o2fvSg5bVg0dRzyHaj4Z5EVhG+g2Fde6LKzmQA=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 h1:2yEATaop1/a1I4psnSLgWVPLWwCzkqWakgJy7xTDVy0=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0/go.mod h1:D7J12YRapIekYyPWgGPlA/23pRmpSEZC5xJC/TTLI9U=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0 h1:MCcYL7J6Vt/X0kjqbMZkekCmwsurbQRbL69vkiye2lk=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0/go.mod h1:3jnStNwSufK+f5ktjL4EPcwtig4rtd81NS70lqHuXl8=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 h1:LMuyCAyfalSjDyjdC65nK6N0zoTT63+E/u95X0JovZI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0/go.mod h1:085m8qbm4hgc8rZWGDEa4vmyyo2c3nPxUslYUKUIU04=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 h1:SUplec5dp06reu1zaXmOXdvqH398taqrDXqUl99jxSc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0/go.mod h1:ho2g4N+ane+swq5I/VBkKWnRDY4kUINH3FuqyZqX/Ug=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 h1:lgh3PiVrRUWMLOVSkQicxzZll5NjF1r+AtsX1XRIHw0=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0/go.mod h1:5Cnhth3m/AgOeTgE3ex12pPmiu/gGtZit03kSzx9X7s=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 h1:hqxVTu/GtBF+vJ8d1fzW7fRxZFvgoDjWcxwwCaFDYpU=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0/go.mod h1:z5fVEF4X5v0ESvlJqBrrFlBVoj5EQuefZpzsu7R+x5Q=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 h1:s0n95ya5tOG03exJ5JySOdJFtwGo4ZQ+KeY7Zro4CLI=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0/go.mod h1:m9wRxtKA2MZ1HcnNC4BKI+9aYe434qRZTCvI7QGUN7Y=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 h1:KN3btaILMTxR4QDHVGAO87lq5ButzK7l+kIfLuxQ1oA=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0/go.mod h1:yNcodmUclM4InyWoOwX/YW4Jri0Gj5FWAlM+NqCrtqY=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/metric/x v0.67.0 h1:PcicCNZFkZ4bXfSooXdo3WN7RBOVOtjVdo1wD358Uns=
@@ -2140,13 +2140,13 @@ go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbi
go.opentelemetry.io/proto/otlp v0.7.0/go.mod h1:PqfVotwruBrMGOCsRd/89rSnXhoiJIqeYNgFYFoEGnI=
go.opentelemetry.io/proto/otlp v0.15.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v0.19.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk=
go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/atomic v1.6.0/go.mod h1:sABNBOSYdrvTF6hTgEIbc7YasKWGhgEQZyfxyTvoXHQ=
go.uber.org/atomic v1.7.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
go.uber.org/atomic v1.9.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
go.uber.org/atomic v1.11.0 h1:ZvwS0R+56ePWxUNi+Atn9dWONBPp/AUETXlHW0DxSjE=
go.uber.org/atomic v1.11.0/go.mod h1:LUxbIzbOniOlMKjJjyPfpl4v+PKK2cNJn91OQbhoJI0=
go.uber.org/atomic v1.12.0 h1:BvcXdFKuviU4fTL/f+SxdQ5qJX/Jix8pAkgdUcb3XOE=
go.uber.org/atomic v1.12.0/go.mod h1:I6c4cg+6HCxRjfjSsYtApoFILnpc0CGUdGkXVqbYVNk=
go.uber.org/goleak v1.1.10/go.mod h1:8a7PlsEVH3e/a/GLqe5IIrQx6GzcnRmZEufDUTk4A7A=
go.uber.org/goleak v1.1.12/go.mod h1:cwTWslyiVhfpKIDGSZEM2HlOvcqm+tG4zioyIeLoqMQ=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
@@ -2163,8 +2163,8 @@ go.uber.org/zap v1.27.1 h1:08RqriUEv8+ArZRYSTXy1LeBScaMpVSTBhCeaZYfMYc=
go.uber.org/zap v1.27.1/go.mod h1:GB2qFLM7cTU87MWRP2mPIjqfIDnGu+VIO4V/SdhGo2E=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
go.yaml.in/yaml/v3 v3.0.5 h1:N6y/pJk8buWs9NY5ERU2HSMfm+IuD/OtfdAnq6kESPw=
go.yaml.in/yaml/v3 v3.0.5/go.mod h1:HVTZu1O7/Vkt2N+BFy8Zza+lnLsABggaTM2ZpNIGuKg=
go.yaml.in/yaml/v4 v4.0.0-rc.6 h1:1h7H1ohdUh93/FyE4YaDa1Zh64K6VVbjF4K6WUxMtH4=
go.yaml.in/yaml/v4 v4.0.0-rc.6/go.mod h1:aZqd9kCMsGL7AuUv/m/PvWLdg5sjJsZ4oHDEnfPPfY0=
gocloud.dev v0.46.0 h1:niIuZwSjMtBx8K+ITB2s5kZullB13PGOS2ZoQPZxQ4Q=
@@ -2193,8 +2193,8 @@ golang.org/x/crypto v0.6.0/go.mod h1:OFC/31mSvZgRz0V1QTNCzfAI1aIRzbiufJtkMIlEp58
golang.org/x/crypto v0.7.0/go.mod h1:pYwdfH91IfpZVANVyUOhSIPZaFoJGxTFbZhFTx+dXZU=
golang.org/x/crypto v0.13.0/go.mod h1:y6Z2r+Rw4iayiXXAIxJIDAJ1zMW4yaTpebo8fPOliYc=
golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf4=
golang.org/x/crypto v0.56.0 h1:GUh5Ii4J5jtcseSMiRqr1jXCNHoxjeV9Fmekc2oLy6Y=
golang.org/x/crypto v0.56.0/go.mod h1:OMW5y6CY9l38uPLmxU6l6pwcXp1obtLo3e6gT7gQR2I=
golang.org/x/crypto v0.57.0 h1:3ZVCjf8Ggz7zneR/EHRVx68Ctf+2pmIMP2UFhh9cC6M=
golang.org/x/crypto v0.57.0/go.mod h1:Fdz0i5U6CoizGwLda9DttjSk6qlZo25zYNtR+ycvuZA=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2225,8 +2225,8 @@ golang.org/x/image v0.0.0-20210607152325-775e3b0c77b9/go.mod h1:023OzeP/+EPmXeap
golang.org/x/image v0.0.0-20210628002857-a66eb6448b8d/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20211028202545-6944b10bf410/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20220302094943-723b81ca9867/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.45.0 h1:FMb1nTbH5H9vF55SriQHgFw5GnNL9Jg6L25BwXKzhB0=
golang.org/x/image v0.45.0/go.mod h1:n62x/7RqlwXDvGsSU4u6IUTUf6KghUZ9Bt7cG/T9Fx4=
golang.org/x/image v0.46.0 h1:b1+oYj0Jbp6K5MDT4i4/eZpYlk3V8SJhhDKh6LBHAyQ=
golang.org/x/image v0.46.0/go.mod h1:3B3W05VGVQyuXucLINLjXKrqISASfi4Xj+iCVkLMwew=
golang.org/x/lint v0.0.0-20181026193005-c67002cb31c3/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
golang.org/x/lint v0.0.0-20190227174305-5b3e6a55c961/go.mod h1:wehouNa3lNwaWXcvxsM5YxQ5yQlVC4a0KAMCusXpPoU=
golang.org/x/lint v0.0.0-20190301231843-5614ed5bae6f/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
@@ -2258,8 +2258,8 @@ golang.org/x/mod v0.8.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.9.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.12.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.13.0/go.mod h1:hTbmBsO62+eylJbnUtE2MGJUyE7QWk4xUqPFrRgJ+7c=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c=
golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o=
golang.org/x/net v0.0.0-20180724234803-3673e40ba225/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180826012351-8a410e7b638d/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
@@ -2354,8 +2354,8 @@ golang.org/x/oauth2 v0.0.0-20221014153046-6fdb5e3db783/go.mod h1:h4gKUeWbJ4rQPri
golang.org/x/oauth2 v0.4.0/go.mod h1:RznEsdpjGAINPTOF0UH/t+xJ75L18YO3Ho6Pyn+uRec=
golang.org/x/oauth2 v0.5.0/go.mod h1:9/XBHVqLaWO3/BRHs5jbpYCnOZVjj5V0ndyaAM7KB4I=
golang.org/x/oauth2 v0.6.0/go.mod h1:ycmewcwgD4Rpr3eZJLSB4Kyyljb3qDh40vJ8STE5HKw=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/oauth2 v0.37.0 h1:JUlcxA8oAtauLfiH8FX2/FkAWHAdi0QtGCGc+hofE98=
golang.org/x/oauth2 v0.37.0/go.mod h1:IxwZNxUULJmpBFf9K/9NTMSIfZZuvuTy1gGxhigP/58=
golang.org/x/sync v0.0.0-20180314180146-1d60e4601c6f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20181108010431-42b317875d0f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
@@ -2374,8 +2374,8 @@ golang.org/x/sync v0.0.0-20220929204114-8fcdb60fdcc0/go.mod h1:RxMgew5VJxzue5/jJ
golang.org/x/sync v0.1.0/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.3.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.4.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.0.0-20180810173357-98c5dad5d1a0/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180830151530-49385e6e1522/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180905080454-ebe1bf3edb33/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
@@ -2477,8 +2477,8 @@ golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.8.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.12.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.13.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210220032956-6a3ed077a48d/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210615171337-6886f2dfbf5b/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
@@ -2491,8 +2491,8 @@ golang.org/x/term v0.6.0/go.mod h1:m6U89DPEgQRMq3DNkDClhWw02AUbt2daBVO4cn4Hv9U=
golang.org/x/term v0.8.0/go.mod h1:xPskH00ivmX89bAKVGSKKtLOWNx2+17Eiy94tnKShWo=
golang.org/x/term v0.12.0/go.mod h1:owVbMEjm3cBLCHdkQu9b1opXd4ETQWc3BhuQGKgXgvU=
golang.org/x/term v0.13.0/go.mod h1:LTmsnFJwVN6bCy1rVCoS+qHT1HhALEFxKncY3WNNh4U=
golang.org/x/term v0.45.0 h1:NwWyBmoJCbfTHpxrWoZ9C6/VxOf7ic219I8xZZFdrf0=
golang.org/x/term v0.45.0/go.mod h1:9aqxs0blBcrm/n0L9QW0aRVD+ktan8ssZromtqJC43w=
golang.org/x/term v0.46.0 h1:3+OXuTbaKDgwk8jTi3aSLHRlmWqHEUDUtxnbFigO4YE=
golang.org/x/term v0.46.0/go.mod h1:+K02xbkittuwc0Am4abfA3Fc+XRGXkvBXNO88NCXPoc=
golang.org/x/text v0.0.0-20170915032832-14c0d48ead0c/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.1-0.20180807135948-17ff2d5776d2/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
@@ -2511,8 +2511,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI=
golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2589,8 +2589,8 @@ golang.org/x/tools v0.6.0/go.mod h1:Xwgl3UAJ/d3gWutnCtw505GrjyAbvKui8lOU390QaIU=
golang.org/x/tools v0.7.0/go.mod h1:4pg6aUX35JBAogB10C9AtvVL+qowtN4pT3CGSQex14s=
golang.org/x/tools v0.13.0/go.mod h1:HvlwmtVNQAhOuCjW7xxvovg8wbNq7LwfXh/k7wXUl58=
golang.org/x/tools v0.14.0/go.mod h1:uYBEerGOWcJyEORxN+Ek8+TT266gXkNlHdJBwexUsBg=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
golang.org/x/tools/godoc v0.1.0-deprecated h1:o+aZ1BOj6Hsx/GBdJO/s815sqftjSnrZZwyYTHODvtk=
golang.org/x/tools/godoc v0.1.0-deprecated/go.mod h1:qM63CriJ961IHWmnWa9CjZnBndniPt4a3CK0PVB9bIg=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -2668,8 +2668,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.296.0 h1:Nn5EHeKdGx70MFClaV/II0gsWUm6xhEjb0xYLylVvaA=
google.golang.org/api v0.296.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/api v0.297.0 h1:WktxTsnnx0yZNnsR6j0q6hR21RnnK81FHTOPy/ux4OE=
google.golang.org/api v0.297.0/go.mod h1:S4m8x0M6OkQpkOzGk1y9JG2sm4fFQrMh6dxzjCTszhE=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2805,8 +2805,8 @@ google.golang.org/genproto v0.0.0-20230222225845-10f96fb3dbec/go.mod h1:3Dl5ZL0q
google.golang.org/genproto v0.0.0-20230306155012-7f2fa6fef1f4/go.mod h1:NWraEVixdDnqcqQ30jipen1STv2r/n24Wb7twVTGR4s=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d h1:C9v1o0/4quuhOAfmRXA2j+we0PqZIp8traLdeogF3Ms=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d/go.mod h1:Wz2wFJntZFmLGo7pLDXZ3wYk5hyc0Mb+SkHhDDXT+lU=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d h1:QwnJwPte4XXAkhPu26LTDIahnsMSUV0kK8HkxbC+Pc4=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d/go.mod h1:WRrQ7/7N19PypuT0fxLOL5Lq0waoiRri4FbtHDEKrGE=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 h1:lrupDmKL3p5kEX1M92oan027eCKcouzjuPbH6YBK+Rs=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1/go.mod h1:q/3oV3jAi5vwelxsVAprMBC8BcM2zmNe+IjRGd+9/ks=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 h1:cYNAzI2sUwhmCcoj9TxvihSrqsxt6uIkj3rDRhSDmW4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
@@ -2849,8 +2849,8 @@ google.golang.org/grpc v1.51.0/go.mod h1:wgNDFcnuBGmxLKI/qn4T+m5BtEBYXJPvibbUPsA
google.golang.org/grpc v1.52.0/go.mod h1:pu6fVzoFb+NBYNAvQL08ic+lvB2IojljRYuun5vorUY=
google.golang.org/grpc v1.53.0/go.mod h1:OnIrk0ipVdj4N5d9IUoFUx72/VlD7+jUsHwZgwSMQpw=
google.golang.org/grpc v1.55.0/go.mod h1:iYEXKGkEBhg1PjZQvoYEVPTDkHo1/bjTnfwTeGONTY8=
google.golang.org/grpc v1.85.0-dev h1:HxkDyKIIZPpFnroC56tQv5gNuKTmVvi0t7TzOf5zt7g=
google.golang.org/grpc v1.85.0-dev/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7 h1:5+EEM1fC0yjOZID0NUZVrE2+8M/+1TclNrSz/l1xMYs=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7/go.mod h1:Ovl0ECo4xx5r4kn/6d4BPSNB7OIFuu6EAjOzjtVAKaM=
google.golang.org/grpc/cmd/protoc-gen-go-grpc v1.1.0/go.mod h1:6Kw0yEErY5E/yWrBtf03jp27GLLJujG4z/JK95pnjjw=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6 h1:ExN12ndbJ608cboPYflpTny6mXSzPrDLh0iTaVrRrds=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6/go.mod h1:6ytKWczdvnpnO+m+JiG9NjEDzR1FJfsnmJdG7B8QVZ8=
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.47"
appVersion: "4.48"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.47.0
version: 4.48.2
+23 -2
View File
@@ -286,7 +286,7 @@ metadata:
app.kubernetes.io/component: s3
stringData:
# this key must be an inline json config file
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read","List"]}]}'
```
#### Source S3 credentials from an existing Secret
@@ -376,7 +376,10 @@ start on a non-loopback address without `-adminPassword`, so the chart fails at
render time if `admin.ip` is non-loopback and authentication is not configured via
`admin.secret.adminPassword`, `admin.secret.existingSecret`, or
`WEED_ADMIN_PASSWORD` supplied through `admin.extraEnvironmentVars` /
`admin.secretExtraEnvironmentVars`. The whole `127.0.0.0/8` range and `::1` are
`admin.secretExtraEnvironmentVars`. Setting `admin.allowInsecureBind` renders
`-allowInsecureBind` and bypasses this guard; it leaves the admin API
unauthenticated on the network, so use it only when access is otherwise
restricted (e.g. network policies). The whole `127.0.0.0/8` range and `::1` are
treated as loopback (matching `weed admin`); `localhost` is treated as
non-loopback. Set `admin.ip` to a loopback address only if you also replace the
httpGet probes (e.g. with an `exec` probe that checks `127.0.0.1`).
@@ -567,6 +570,23 @@ Two things worth knowing before you turn this on:
The DNS selectors default to CoreDNS as kubeadm, kind and the managed offerings from AWS, Google and Azure install it. On OpenShift, override `egress.dnsNamespaceSelector` and `egress.dnsPodSelector` to match `openshift-dns`; see the comment in `values.yaml`.
## Pod and container security contexts
Pod and container security contexts are configurable independently for every built-in workload and remain empty by default for backwards compatibility. The examples in `values.yaml` show how to enable a `RuntimeDefault` seccomp profile, disable privilege escalation and privileged mode, drop all Linux capabilities, and use a read-only root filesystem.
SeaweedFS uses `/tmp` for Unix sockets, temporary uploads, worker task files, and other runtime data. When `readOnlyRootFilesystem` is enabled for a built-in component, the chart mounts a writable `emptyDir` at `/tmp` for its chart-managed containers. Its optional size limit can be configured globally:
```yaml
global:
seaweedfs:
tmpDir:
sizeLimit: 1Gi
```
The chart does not enable `runAsNonRoot` by default because its default `hostPath` storage may be owned by root. To enforce the Kubernetes `restricted` Pod Security Standard, use storage that is writable by a non-root user and configure `runAsNonRoot` or use the OpenShift overrides below.
Security contexts configured for a component also apply to the chart-managed helper containers for that component. User-provided init containers and sidecars must define their own container security context and writable mounts.
## OpenShift Support
SeaweedFS can be deployed on OpenShift or any cluster enforcing the Kubernetes "restricted" Pod Security Standard. By default, OpenShift blocks containers that run as root or use `hostPath` volumes.
@@ -575,6 +595,7 @@ To deploy on OpenShift, use the provided `openshift-values.yaml` which overrides
1. Use `PersistentVolumeClaims` instead of `hostPath`.
2. Enable `runAsNonRoot` and omit hardcoded UIDs to allow OpenShift to assign valid UIDs automatically.
3. Apply appropriate `seccompProfile` and drop capabilities.
4. Use a read-only root filesystem with writable temporary storage at `/tmp`.
Usage:
```bash
@@ -15,6 +15,7 @@
# automatically assign a valid UID from the namespace's allocated range.
# 3. Dropping all Linux capabilities and setting allowPrivilegeEscalation: false
# 4. Enabling RuntimeDefault seccompProfile
# 5. Using a read-only root filesystem with writable temporary storage
#
# Usage:
# helm install seaweedfs seaweedfs/seaweedfs \
@@ -49,6 +50,7 @@ master:
containerSecurityContext:
enabled: true
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
runAsNonRoot: true
@@ -76,6 +78,7 @@ volume:
containerSecurityContext:
enabled: true
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
runAsNonRoot: true
@@ -101,6 +104,7 @@ filer:
containerSecurityContext:
enabled: true
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
runAsNonRoot: true
@@ -125,6 +129,7 @@ s3:
containerSecurityContext:
enabled: true
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
runAsNonRoot: true
@@ -9,7 +9,7 @@
{{- $adminAuthEnabled := include "seaweedfs.admin.authEnabled" . }}
{{- $adminIp := .Values.admin.ip | default "0.0.0.0" }}
{{- if and (not (include "seaweedfs.admin.isLoopbackIp" $adminIp)) (ne $adminAuthEnabled "true") }}
{{- fail (printf "admin.ip is set to %q (non-loopback) but admin authentication is not configured. Since `weed admin` 4.46 refuses to bind a non-loopback address without authentication, the admin container would exit on startup. Set admin.secret.adminPassword or admin.secret.existingSecret, or supply WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars / admin.secretExtraEnvironmentVars, or set admin.ip to a loopback address such as 127.0.0.1 (note: a loopback bind makes the chart's httpGet readiness/liveness probes fail)." $adminIp) -}}
{{- fail (printf "admin.ip is set to %q (non-loopback) but admin authentication is not configured. Since `weed admin` 4.46 refuses to bind a non-loopback address without authentication, the admin container would exit on startup. Set admin.secret.adminPassword or admin.secret.existingSecret, or supply WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars / admin.secretExtraEnvironmentVars, or set admin.ip to a loopback address such as 127.0.0.1 (note: a loopback bind makes the chart's httpGet readiness/liveness probes fail), or set admin.allowInsecureBind to true to opt out via -allowInsecureBind (INSECURE: exposes the admin API unauthenticated on the network)." $adminIp) -}}
{{- end }}
apiVersion: apps/v1
kind: StatefulSet
@@ -176,19 +176,23 @@ spec:
-dataDir={{ .Values.admin.dataDir }} \
{{- end }}
{{- if .Values.admin.masters }}
-masters={{ .Values.admin.masters }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ .Values.admin.masters }} \
{{- else if .Values.global.seaweedfs.masterServer }}
-masters={{ .Values.global.seaweedfs.masterServer }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ .Values.global.seaweedfs.masterServer }} \
{{- else }}
-masters={{ range $index := until (.Values.master.replicas | int) }}${SEAWEEDFS_FULLNAME}-master-{{ $index }}.${SEAWEEDFS_FULLNAME}-master.{{ $.Release.Namespace }}:{{ $.Values.master.port }}{{ if lt $index (sub ($.Values.master.replicas | int) 1) }},{{ end }}{{ end }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ range $index := until (.Values.master.replicas | int) }}${SEAWEEDFS_FULLNAME}-master-{{ $index }}.${SEAWEEDFS_FULLNAME}-master.{{ $.Release.Namespace }}:{{ $.Values.master.port }}{{ if lt $index (sub ($.Values.master.replicas | int) 1) }},{{ end }}{{ end }} \
{{- end }}
{{- if $urlPrefix }}
-urlPrefix={{ $urlPrefix }}{{- if .Values.admin.extraArgs }} \{{ end }}
-urlPrefix={{ $urlPrefix }} \
{{- end }}
{{- if .Values.admin.allowInsecureBind }}
-allowInsecureBind \
{{- end }}
{{- range $index, $arg := .Values.admin.extraArgs }}
{{ $arg }}{{- if lt $index (sub (len $.Values.admin.extraArgs) 1) }} \{{ end }}
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.admin.containerSecurityContext (tpl (.Values.admin.extraVolumeMounts | default "") .) (tpl (.Values.admin.extraVolumes | default "") .)) | nindent 12 }}
{{- if or (eq .Values.admin.data.type "hostPath") (eq .Values.admin.data.type "persistentVolumeClaim") (eq .Values.admin.data.type "emptyDir") (eq .Values.admin.data.type "existingClaim") }}
- name: admin-data
mountPath: /data
@@ -264,6 +268,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.admin.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.admin.containerSecurityContext (tpl (.Values.admin.extraVolumeMounts | default "") .) (tpl (.Values.admin.extraVolumes | default "") .) false) | nindent 8 }}
{{- if eq .Values.admin.data.type "hostPath" }}
- name: admin-data
hostPath:
@@ -37,13 +37,17 @@ spec:
{{- with .Values.allInOne.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.allInOne.podAnnotations | default dict) }}
{{- $existingS3ConfigSecret := or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret }}
{{- if $existingS3ConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace $existingS3ConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.allInOne.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.allInOne.restartPolicy }}
{{- if .Values.allInOne.affinity }}
@@ -302,6 +306,7 @@ spec:
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.allInOne.containerSecurityContext (tpl (.Values.allInOne.extraVolumeMounts | default "") .) (tpl (.Values.allInOne.extraVolumes | default "") .)) | nindent 12 }}
- name: data
mountPath: /data
{{- if and .Values.allInOne.s3.enabled (or .Values.allInOne.s3.enableAuth .Values.s3.enableAuth .Values.filer.s3.enableAuth) }}
@@ -429,6 +434,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.allInOne.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.allInOne.containerSecurityContext (tpl (.Values.allInOne.extraVolumeMounts | default "") .) (tpl (.Values.allInOne.extraVolumes | default "") .) false) | nindent 8 }}
{{- include "seaweedfs.licenseVolume" . | nindent 8 }}
- name: data
{{- if eq .Values.allInOne.data.type "hostPath" }}
@@ -25,6 +25,10 @@ spec:
organizations:
- "SeaweedFS CA"
dnsNames:
- '{{ include "seaweedfs.fullname" . }}-admin'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc.cluster.local'
- '*.{{ include "seaweedfs.fullname" . }}-admin'
- '*.{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}'
- '*.{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc'
@@ -113,6 +113,7 @@ spec:
{{- end -}}
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.cosi.containerSecurityContext (tpl (.Values.cosi.extraVolumeMounts | default "") .) (tpl (.Values.cosi.extraVolumes | default "") .)) | nindent 12 }}
- mountPath: /var/lib/cosi
name: socket
{{- if .Values.cosi.enableAuth }}
@@ -148,6 +149,9 @@ spec:
resources:
{{- toYaml . | nindent 12 }}
{{- end }}
{{- if .Values.cosi.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.cosi.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
- name: seaweedfs-cosi-sidecar
image: "{{ .Values.cosi.sidecar.image }}"
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
@@ -159,6 +163,7 @@ spec:
fieldRef:
fieldPath: metadata.namespace
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.cosi.containerSecurityContext "" "") | nindent 12 }}
- mountPath: /var/lib/cosi
name: socket
{{- with .Values.cosi.sidecar.resources }}
@@ -172,6 +177,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.cosi.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.cosi.containerSecurityContext (tpl (.Values.cosi.extraVolumeMounts | default "") .) (tpl (.Values.cosi.extraVolumes | default "") .) true) | nindent 8 }}
- name: socket
emptyDir: {}
{{- if .Values.cosi.enableAuth }}
@@ -43,19 +43,15 @@ spec:
{{- with .Values.filer.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.filer.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.filer.podAnnotations | default dict) }}
{{- if .Values.filer.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
checksum/s3config: {{ $configSecret | toYaml | sha256sum }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
checksum/s3config: {{ include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.filer.restartPolicy }}
{{- if .Values.filer.affinity }}
@@ -233,6 +229,7 @@ spec:
{{ . }} \
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.filer.containerSecurityContext (tpl (.Values.filer.extraVolumeMounts | default "") .) (tpl (.Values.filer.extraVolumes | default "") .)) | nindent 12 }}
{{- if (or (eq .Values.filer.logs.type "hostPath") (eq .Values.filer.logs.type "persistentVolumeClaim") (eq .Values.filer.logs.type "emptyDir")) }}
- name: seaweedfs-filer-log-volume
mountPath: "/logs/"
@@ -341,6 +338,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.filer.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.filer.containerSecurityContext (tpl (.Values.filer.extraVolumeMounts | default "") .) (tpl (.Values.filer.extraVolumes | default "") .) false) | nindent 8 }}
{{- if eq .Values.filer.logs.type "hostPath" }}
- name: seaweedfs-filer-log-volume
hostPath:
@@ -43,13 +43,10 @@ spec:
{{- with .Values.master.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.master.podAnnotations | default dict) }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.master.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.master.restartPolicy }}
{{- if .Values.master.affinity }}
@@ -182,6 +179,7 @@ spec:
{{ . }} \
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.master.containerSecurityContext (tpl (.Values.master.extraVolumeMounts | default "") .) (tpl (.Values.master.extraVolumes | default "") .)) | nindent 12 }}
- name : data-{{ .Release.Namespace }}
mountPath: /data
{{- if or (eq .Values.master.logs.type "hostPath") (eq .Values.master.logs.type "persistentVolumeClaim") (eq .Values.master.logs.type "emptyDir") }}
@@ -261,6 +259,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.master.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.master.containerSecurityContext (tpl (.Values.master.extraVolumeMounts | default "") .) (tpl (.Values.master.extraVolumes | default "") .) false) | nindent 8 }}
{{- include "seaweedfs.licenseVolume" . | nindent 8 }}
{{- if eq .Values.master.logs.type "hostPath" }}
- name: seaweedfs-master-log-volume
@@ -35,13 +35,15 @@ spec:
{{- with .Values.s3.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.s3.podAnnotations | default dict) }}
{{- if .Values.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.s3.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.s3.restartPolicy }}
{{- if .Values.s3.affinity }}
@@ -155,6 +157,7 @@ spec:
{{ . }} \
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.s3.containerSecurityContext (tpl (.Values.s3.extraVolumeMounts | default "") .) (tpl (.Values.s3.extraVolumes | default "") .)) | nindent 12 }}
{{- if or (eq .Values.s3.logs.type "hostPath") (eq .Values.s3.logs.type "emptyDir") }}
- name: logs
mountPath: "/logs/"
@@ -243,6 +246,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.s3.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.s3.containerSecurityContext (tpl (.Values.s3.extraVolumeMounts | default "") .) (tpl (.Values.s3.extraVolumes | default "") .) false) | nindent 8 }}
{{- if .Values.s3.enableAuth }}
- name: config-users
secret:
@@ -60,7 +60,7 @@ stringData:
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read","List"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
{{ toJson .Values.filer.s3.auditLogConfig | nindent 4 }}
@@ -170,6 +170,7 @@ spec:
-userStoreFile=/etc/sw/seaweedfs_sftp_config \
-filer={{ include "seaweedfs.componentName" (list . "filer-client") }}.{{ .Release.Namespace }}:{{ .Values.filer.port }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.sftp.containerSecurityContext (tpl (.Values.sftp.extraVolumeMounts | default "") .) (tpl (.Values.sftp.extraVolumes | default "") .)) | nindent 12 }}
{{- if or (eq .Values.sftp.logs.type "hostPath") (eq .Values.sftp.logs.type "emptyDir") }}
- name: logs
mountPath: "/logs/"
@@ -249,6 +250,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.sftp.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.sftp.containerSecurityContext (tpl (.Values.sftp.extraVolumeMounts | default "") .) (tpl (.Values.sftp.extraVolumes | default "") .) false) | nindent 8 }}
{{- if .Values.sftp.enableAuth }}
- name: config-users
secret:
@@ -76,6 +76,43 @@ Inject extra environment vars in the format key:value, if populated
{{- end }}
{{- end -}}
{{/*
Writable temporary directory for containers using a read-only root filesystem.
Input: list of the root context, the component container security context, the
rendered extraVolumeMounts and extraVolumes, and whether the pod has secondary
chart-managed containers that mount seaweedfs-tmp. A user-supplied /tmp mount
only covers the main container, so the volume is still emitted for secondaries;
a user-supplied seaweedfs-tmp volume is reused rather than duplicated.
*/}}
{{- define "seaweedfs.tmpDirCovered" -}}
{{- regexMatch `(?m)^\s*-?\s*mountPath:\s*['"]?/tmp/?['"]?\s*(#.*)?$` (index . 2) -}}
{{- end -}}
{{- define "seaweedfs.tmpDirVolume" -}}
{{- $root := index . 0 -}}
{{- $securityContext := index . 1 -}}
{{- if and $securityContext.enabled $securityContext.readOnlyRootFilesystem
(or (index . 4) (ne (include "seaweedfs.tmpDirCovered" .) "true"))
(not (regexMatch `(?m)^\s*-?\s*name:\s*['"]?seaweedfs-tmp['"]?\s*(#.*)?$` (index . 3))) }}
- name: seaweedfs-tmp
{{- with $root.Values.global.seaweedfs.tmpDir.sizeLimit }}
emptyDir:
sizeLimit: {{ . | quote }}
{{- else }}
emptyDir: {}
{{- end }}
{{- end }}
{{- end -}}
{{- define "seaweedfs.tmpDirVolumeMount" -}}
{{- $securityContext := index . 1 -}}
{{- if and $securityContext.enabled $securityContext.readOnlyRootFilesystem
(ne (include "seaweedfs.tmpDirCovered" .) "true") }}
- name: seaweedfs-tmp
mountPath: /tmp
{{- end }}
{{- end -}}
{{/* Whether the mysql filer store is selected; a flag the chart cannot read counts as selected. */}}
{{- define "seaweedfs.filer.mysqlEnabled" -}}
{{- $merged := dict -}}
@@ -105,13 +142,13 @@ true
{{- end -}}
{{- end -}}
{{/* Whether admin authentication is enabled from any supported source:
admin.secret (adminPassword or existingSecret), or WEED_ADMIN_PASSWORD
supplied via extraEnvironmentVars / secretExtraEnvironmentVars (which
weed admin picks up through viper's AutomaticEnv). A secret-backed
entry counts as enabled even though the chart cannot read its value. */}}
{{/* Whether the admin non-loopback bind guard is satisfied: admin.secret
(adminPassword or existingSecret), WEED_ADMIN_PASSWORD via
extraEnvironmentVars / secretExtraEnvironmentVars, or
admin.allowInsecureBind. A secret-backed entry counts as enabled even
though the chart cannot read its value. */}}
{{- define "seaweedfs.admin.authEnabled" -}}
{{- if or .Values.admin.secret.existingSecret .Values.admin.secret.adminPassword -}}
{{- if or .Values.admin.secret.existingSecret .Values.admin.secret.adminPassword .Values.admin.allowInsecureBind -}}
true
{{- else -}}
{{- $merged := dict -}}
@@ -467,8 +504,9 @@ true
{{- $pvcName := printf "%s-%s-%s-%d" $dir.name $seaweedfsName $volumeName $e }}
{{- $currentPVC := (lookup "v1" "PersistentVolumeClaim" $.Release.Namespace $pvcName) }}
{{- if $currentPVC }}
{{- $oldSize := include "seaweedfs.resource-quantity" $currentPVC.spec.resources.requests.storage }}
{{- $newSize := include "seaweedfs.resource-quantity" $desiredSize }}
{{- /* include returns a string such as "6.442450944e+10"; convert back to a number, or gt compares lexically */}}
{{- $oldSize := include "seaweedfs.resource-quantity" $currentPVC.spec.resources.requests.storage | float64 }}
{{- $newSize := include "seaweedfs.resource-quantity" $desiredSize | float64 }}
{{- if gt $newSize $oldSize }}
{{- $commands = append $commands (printf "kubectl patch pvc %s-%s-%s-%d -p '{\"spec\":{\"resources\":{\"requests\":{\"storage\":\"%s\"}}}}'" $dir.name $seaweedfsName $volumeName $e $desiredSize) }}
{{- end }}
@@ -16,6 +16,12 @@
{{- $clusterUpper := upper $clusterAlias }}
{{- $clusterMasterKey := printf "WEED_CLUSTER_%s_MASTER" $clusterUpper }}
{{- $clusterFilerKey := printf "WEED_CLUSTER_%s_FILER" $clusterUpper }}
{{- $podSecurityContext := .Values.filer.podSecurityContext }}
{{- $containerSecurityContext := .Values.filer.containerSecurityContext }}
{{- if .Values.allInOne.enabled }}
{{- $podSecurityContext = .Values.allInOne.podSecurityContext }}
{{- $containerSecurityContext = .Values.allInOne.containerSecurityContext }}
{{- end }}
{{- /* Check allInOne mode first */}}
{{- if .Values.allInOne.enabled }}
@@ -71,8 +77,8 @@ spec:
app.kubernetes.io/component: bucket-hook
spec:
restartPolicy: Never
{{- if .Values.filer.podSecurityContext.enabled }}
securityContext: {{- omit .Values.filer.podSecurityContext "enabled" | toYaml | nindent 8 }}
{{- if $podSecurityContext.enabled }}
securityContext: {{- omit $podSecurityContext "enabled" | toYaml | nindent 8 }}
{{- end }}
{{- include "seaweedfs.imagePullSecrets" $ | nindent 6 }}
containers:
@@ -202,8 +208,9 @@ spec:
/usr/bin/weed shell
{{- end }}
{{- end }}
{{- if or $enableAuth (include "seaweedfs.securityConfigEnabled" .) }}
{{- if or (and $containerSecurityContext.enabled $containerSecurityContext.readOnlyRootFilesystem) $enableAuth (include "seaweedfs.securityConfigEnabled" .) }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . $containerSecurityContext "" "") | nindent 10 }}
{{- if $enableAuth }}
- name: config-users
mountPath: /etc/sw
@@ -215,6 +222,14 @@ spec:
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
- name: client-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/client/
{{- end }}
{{- end }}
ports:
- containerPort: {{ .Values.master.port }}
@@ -229,11 +244,12 @@ spec:
resources:
{{- toYaml . | nindent 10 }}
{{- end }}
{{- if .Values.filer.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.filer.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- if $containerSecurityContext.enabled }}
securityContext: {{- omit $containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- if or $enableAuth (include "seaweedfs.securityConfigEnabled" .) }}
{{- if or (and $containerSecurityContext.enabled $containerSecurityContext.readOnlyRootFilesystem) $enableAuth (include "seaweedfs.securityConfigEnabled" .) }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . $containerSecurityContext "" "" false) | nindent 8 }}
{{- if $enableAuth }}
- name: config-users
secret:
@@ -249,5 +265,13 @@ spec:
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
- name: client-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-client-cert
{{- end }}
{{- end }}
{{- end }}
@@ -32,13 +32,32 @@ spec:
spec:
serviceAccountName: {{ $seaweedfsName }}-volume-resize-hook
restartPolicy: Never
{{- if .Values.volume.podSecurityContext.enabled }}
securityContext: {{- omit .Values.volume.podSecurityContext "enabled" | toYaml | nindent 8 }}
{{- end }}
containers:
- name: resize
image: {{ .Values.volume.resizeHook.image }}
{{- if and .Values.volume.containerSecurityContext.enabled .Values.volume.containerSecurityContext.readOnlyRootFilesystem }}
env:
- name: HOME
value: /tmp
{{- end }}
command: ["sh", "-xec"]
args:
- |
{{ $commands | indent 14 }}
{{- if and .Values.volume.containerSecurityContext.enabled .Values.volume.containerSecurityContext.readOnlyRootFilesystem }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.volume.containerSecurityContext "" "") | nindent 12 }}
{{- end }}
{{- if .Values.volume.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.volume.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- if and .Values.volume.containerSecurityContext.enabled .Values.volume.containerSecurityContext.readOnlyRootFilesystem }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.volume.containerSecurityContext "" "" false) | nindent 8 }}
{{- end }}
---
apiVersion: v1
kind: ServiceAccount
@@ -104,6 +104,7 @@ spec:
fi
done
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list $ $volume.containerSecurityContext "" "") | nindent 12 }}
- name: idx
mountPath: /idx
{{- range $dir := $volume.dataDirs }}
@@ -224,6 +225,7 @@ spec:
{{ . }} \
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list $ $volume.containerSecurityContext (tpl (printf "{{ $volumeName := \"%s\" }}%s" $volumeName ($volume.extraVolumeMounts | default "")) $) (tpl ($volume.extraVolumes | default "") $)) | nindent 12 }}
{{- range $dir := $volume.dataDirs }}
{{- if not ( eq $dir.type "custom" ) }}
- name: {{ $dir.name }}
@@ -306,6 +308,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" (printf "{{ $volumeName := \"%s\" }}%s" $volumeName $volume.sidecars) "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list $ $volume.containerSecurityContext (tpl (printf "{{ $volumeName := \"%s\" }}%s" $volumeName ($volume.extraVolumeMounts | default "")) $) (tpl ($volume.extraVolumes | default "") $) (and $initContainers_exists $volume.idx)) | nindent 8 }}
{{- range $dir := $volume.dataDirs }}
@@ -144,6 +144,7 @@ spec:
{{ $arg }}{{- if lt $index (sub (len $.Values.worker.extraArgs) 1) }} \{{ end }}
{{- end }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.worker.containerSecurityContext (tpl (.Values.worker.extraVolumeMounts | default "") .) (tpl (.Values.worker.extraVolumes | default "") .)) | nindent 12 }}
{{- if or (eq .Values.worker.data.type "hostPath") (eq .Values.worker.data.type "emptyDir") (eq .Values.worker.data.type "existingClaim") }}
- name: worker-data
mountPath: {{ .Values.worker.workingDir }}
@@ -262,15 +263,16 @@ spec:
--metrics-ip=0.0.0.0 \
{{- end }}
--max-concurrency={{ .Values.worker.maxExecute }}
{{- if .Values.global.seaweedfs.enableSecurity }}
volumeMounts:
{{- include "seaweedfs.tmpDirVolumeMount" (list . .Values.worker.containerSecurityContext "" "") | nindent 12 }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
- name: worker-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/worker/
{{- end }}
{{- end }}
{{- if .Values.worker.lanceMetricsPort }}
ports:
- containerPort: {{ .Values.worker.lanceMetricsPort }}
@@ -302,6 +304,7 @@ spec:
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.worker.sidecars "context" $) | nindent 8 }}
{{- end }}
volumes:
{{- include "seaweedfs.tmpDirVolume" (list . .Values.worker.containerSecurityContext (tpl (.Values.worker.extraVolumeMounts | default "") .) (tpl (.Values.worker.extraVolumes | default "") .) (ne (include "seaweedfs.worker.lanceNamespaceUrl" .) "")) | nindent 8 }}
{{- if eq .Values.worker.data.type "hostPath" }}
- name: worker-data
hostPath:
+180 -3
View File
@@ -25,6 +25,9 @@ global:
imagePullPolicy: IfNotPresent
restartPolicy: Always
loggingLevel: 1
# Writable temporary storage used when containers run with a read-only root filesystem.
tmpDir:
sizeLimit: ""
enableSecurity: false
masterServer: null
# filerWrite: true mounts security.toml on filer + admin without needing
@@ -190,6 +193,8 @@ master:
podLabels: {}
# Annotations to be added to the master pods
# The chart sets checksum/master-config on master pods; other checksum/* keys
# can be used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the master resources
@@ -258,6 +263,8 @@ master:
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
@@ -266,7 +273,16 @@ master:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
ingress:
@@ -558,6 +574,8 @@ volume:
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
@@ -566,7 +584,16 @@ volume:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
# used to configure livenessProbe on volume-server containers
@@ -773,6 +800,8 @@ filer:
podLabels: {}
# Annotations to be added to the filer pods
# The chart sets checksum/s3config on filer pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the filer resource
@@ -841,6 +870,8 @@ filer:
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
@@ -849,7 +880,16 @@ filer:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
ingresses:
@@ -1078,6 +1118,8 @@ s3:
podLabels: {}
# Annotations to be added to the s3 pods
# The chart sets checksum/s3config on s3 pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the s3 resources
@@ -1121,6 +1163,8 @@ s3:
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
@@ -1129,7 +1173,16 @@ s3:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
# You can also use emptyDir storage:
@@ -1281,7 +1334,33 @@ sftp:
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
# Configure security context for Pod
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# podSecurityContext:
# enabled: true
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
logs:
@@ -1333,10 +1412,11 @@ admin:
# kubelet's httpGet readiness/liveness probes (which dial the pod IP) to ever
# succeed. "0.0.0.0" restores the pre-4.46 behaviour of listening on all
# interfaces. A non-loopback address requires authentication: set
# admin.secret.adminPassword or admin.secret.existingSecret, or supply
# admin.secret.adminPassword or admin.secret.existingSecret, supply
# WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars /
# admin.secretExtraEnvironmentVars; otherwise the admin container will exit
# with a clear error rather than silently staying unready. The whole
# admin.secretExtraEnvironmentVars, or opt out with admin.allowInsecureBind;
# otherwise the admin container will exit with a clear error rather than
# silently staying unready. The whole
# 127.0.0.0/8 range and ::1 are treated as loopback (matching weed admin).
# Set to a loopback address only if you also replace the httpGet probes.
# Note: the -ip flag requires SeaweedFS 4.46 or newer; pinning
@@ -1344,6 +1424,9 @@ admin:
ip: "0.0.0.0"
loggingOverrideLevel: null
# INSECURE: allow binding a non-loopback ip without authentication.
allowInsecureBind: false
# Admin authentication
secret:
# Name of an existing secret containing admin credentials. If set, adminUser and adminPassword below are ignored.
@@ -1424,7 +1507,33 @@ admin:
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
# Configure security context for Pod
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# podSecurityContext:
# enabled: true
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
extraEnvironmentVars: {}
@@ -1581,7 +1690,33 @@ worker:
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
# Configure security context for Pod
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# podSecurityContext:
# enabled: true
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
extraEnvironmentVars: {}
@@ -1788,6 +1923,8 @@ allInOne:
initContainers: "" # Init containers
sidecars: "" # Sidecar containers
annotations: {} # Annotations for the deployment
# The chart sets checksum/master-config and checksum/s3config on all-in-one
# pods; other checksum/* keys can be used for custom rollouts.
podAnnotations: {} # Annotations for the pods
podLabels: {} # Labels for the pods
@@ -1839,6 +1976,8 @@ allInOne:
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
@@ -1847,7 +1986,16 @@ allInOne:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
# Resource management
@@ -1886,7 +2034,33 @@ cosi:
# should have a secret key called seaweedfs_s3_config with an inline json configure
existingConfigSecret: null
# Configure security context for Pod
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# podSecurityContext:
# enabled: true
# runAsUser: 1000
# runAsGroup: 3000
# fsGroup: 2000
# seccompProfile:
# type: RuntimeDefault
podSecurityContext: {}
# Configure security context for Container
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/security-context/
# Example:
# containerSecurityContext:
# enabled: true
# runAsUser: 2000
# runAsGroup: 3000
# runAsNonRoot: true
# privileged: false
# allowPrivilegeEscalation: false
# readOnlyRootFilesystem: true
# capabilities:
# drop:
# - ALL
# seccompProfile:
# type: RuntimeDefault
containerSecurityContext: {}
# used to assign a custom scheduler to cosi pods
@@ -1922,6 +2096,9 @@ certificates:
# Labels to be added to all the created pods
podLabels: {}
# Annotations to be added to all the created pods
# The chart sets checksum/master-config and checksum/s3config on pods whose
# rendered ConfigMaps or Secrets should trigger rollouts. Other checksum/* keys
# can be used for custom rollout annotations.
podAnnotations: {}
networkPolicy:
+1046 -1055
View File
File diff suppressed because it is too large Load Diff

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 54 KiB

+293
View File
@@ -0,0 +1,293 @@
# This file is automatically @generated by Cargo.
# It is not intended for manual editing.
version = 4
[[package]]
name = "aws-lc-rs"
version = "1.18.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ce2b2dcc879c3bae0d371e77c99f2238400ef24ec001394befa67b6e543add9e"
dependencies = [
"aws-lc-sys",
"zeroize",
]
[[package]]
name = "aws-lc-sys"
version = "0.44.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f09fae7be8bb3174e05c6afdb34199e6dc0c7c04ba9fa237b1967adfbde27483"
dependencies = [
"cc",
"cmake",
"dunce",
"fs_extra",
"pkg-config",
]
[[package]]
name = "cc"
version = "1.4.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "509591b7bcd67f4ef775afad7662703b4935daaa6ec0e5605cfb1090b32a2b6d"
dependencies = [
"find-msvc-tools",
"jobserver",
"libc",
"shlex",
]
[[package]]
name = "cfg-if"
version = "1.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
[[package]]
name = "cmake"
version = "0.1.58"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c0f78a02292a74a88ac736019ab962ece0bc380e3f977bf72e376c5d78ff0678"
dependencies = [
"cc",
]
[[package]]
name = "dunce"
version = "1.0.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "92773504d58c093f6de2459af4af33faa518c13451eb8f2b5698ed3d36e7c813"
[[package]]
name = "find-msvc-tools"
version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d45db016d36b838f563236e9193d0ee6ce38f3f68b6c94e914b4929c96bbb890"
[[package]]
name = "fs_extra"
version = "1.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "42703706b716c37f96a77aea830392ad231f44c9e9a67872fa5548707e11b11c"
[[package]]
name = "getrandom"
version = "0.2.17"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0"
dependencies = [
"cfg-if",
"libc",
"wasi",
]
[[package]]
name = "getrandom"
version = "0.4.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "300e883d756b2e4ec94e02791f39b04b522276138852cfc41d9fb7e904106099"
dependencies = [
"cfg-if",
"libc",
"r-efi",
]
[[package]]
name = "jobserver"
version = "0.1.35"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1c00acbd29eabad4a2392fa0e921c874934dbbf4194312ad20f04a0ed67a3cb3"
dependencies = [
"getrandom 0.4.3",
"libc",
]
[[package]]
name = "libc"
version = "0.2.189"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3eaf3ede3fee6db1a4c2ee091bf8a8b4dccdc6d17f656fb07896ee72867612f2"
[[package]]
name = "log"
version = "0.4.33"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad"
[[package]]
name = "once_cell"
version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "pkg-config"
version = "0.3.34"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f6b464fbc74e149a392436b17d523f769e057cb6877f6a5c4618bc6f11800548"
[[package]]
name = "r-efi"
version = "6.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f8dcc9c7d52a811697d2151c701e0d08956f92b0e24136cf4cf27b57a6a0d9bf"
[[package]]
name = "ring"
version = "0.17.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7"
dependencies = [
"cc",
"cfg-if",
"getrandom 0.2.17",
"libc",
"untrusted",
"windows-sys",
]
[[package]]
name = "rustls"
version = "0.23.43"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
dependencies = [
"aws-lc-rs",
"log",
"once_cell",
"rustls-pki-types",
"rustls-webpki",
"subtle",
"zeroize",
]
[[package]]
name = "rustls-pki-types"
version = "1.15.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96"
dependencies = [
"zeroize",
]
[[package]]
name = "rustls-webpki"
version = "0.103.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0527518605e68109d875e248ea259b6758801cf165e4b2c2733ae3b51f12535a"
dependencies = [
"aws-lc-rs",
"ring",
"rustls-pki-types",
"untrusted",
]
[[package]]
name = "seaweed-common"
version = "0.1.0"
dependencies = [
"rustls",
]
[[package]]
name = "shlex"
version = "2.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f8fadd59c855ef2080decdef8ff161eb6661b86933c9d82e5ba29dc602a55aba"
[[package]]
name = "subtle"
version = "2.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292"
[[package]]
name = "untrusted"
version = "0.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "wasi"
version = "0.11.1+wasi-snapshot-preview1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b"
[[package]]
name = "windows-sys"
version = "0.52.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "282be5f36a8ce781fad8c8ae18fa3f9beff57ec1b52cb3de0789201425d9a33d"
dependencies = [
"windows-targets",
]
[[package]]
name = "windows-targets"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9b724f72796e036ab90c1021d4780d4d3d648aca59e491e6b98e725b84e99973"
dependencies = [
"windows_aarch64_gnullvm",
"windows_aarch64_msvc",
"windows_i686_gnu",
"windows_i686_gnullvm",
"windows_i686_msvc",
"windows_x86_64_gnu",
"windows_x86_64_gnullvm",
"windows_x86_64_msvc",
]
[[package]]
name = "windows_aarch64_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32a4622180e7a0ec044bb555404c800bc9fd9ec262ec147edd5989ccd0c02cd3"
[[package]]
name = "windows_aarch64_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09ec2a7bb152e2252b53fa7803150007879548bc709c039df7627cabbd05d469"
[[package]]
name = "windows_i686_gnu"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8e9b5ad5ab802e97eb8e295ac6720e509ee4c243f69d781394014ebfe8bbfa0b"
[[package]]
name = "windows_i686_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0eee52d38c090b3caa76c563b86c3a4bd71ef1a819287c19d586d7334ae8ed66"
[[package]]
name = "windows_i686_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "240948bc05c5e7c6dabba28bf89d89ffce3e303022809e73deaefe4f6ec56c66"
[[package]]
name = "windows_x86_64_gnu"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "147a5c80aabfbf0c7d901cb5895d1de30ef2907eb21fbbab29ca94c5b08b1a78"
[[package]]
name = "windows_x86_64_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "24d5b23dc417412679681396f2b49f3de8c1473deb516bd34410872eff51ed0d"
[[package]]
name = "windows_x86_64_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec"
[[package]]
name = "zeroize"
version = "1.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e"
+27
View File
@@ -0,0 +1,27 @@
[package]
name = "seaweed-common"
version = "0.1.0"
edition = "2024"
# The lower of the two consumers' floors (seaweed-volume 1.91.1,
# seaweed-worker 1.94.1), so depending on this crate cannot raise either
# tree's MSRV. Verified with `cargo +1.91.1 check --all-targets`.
rust-version = "1.91.1"
description = "Helpers shared by the SeaweedFS Rust volume server and the Rust plugin workers"
# There is no root manifest: seaweed-volume and seaweed-worker are separate
# cargo trees with their own lockfiles, and this crate is a path dependency of
# both rather than a member of either. Keeping the lint policy identical in all
# three manifests is what stops them drifting.
[lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[dependencies]
# The same requirement both consumers already write. Cargo unifies all
# semver-compatible `rustls = "0.23"` requirements into one crate per binary,
# which is what makes `install_default_crypto_provider` write the same
# process-wide static the consuming crate reads. rustls is already in both
# trees (the volume server directly, seaweed-worker-core through tonic's
# `tls-aws-lc`), so this adds no crate to either graph.
rustls = "0.23"
+307
View File
@@ -0,0 +1,307 @@
//! SeaweedFS server addresses, the way the Go tree does them.
//!
//! An operator gives a SeaweedFS process an HTTP address and the gRPC port is
//! derived from it rather than asked for separately: `host:port` means gRPC on
//! `port + 10000`, and the explicit `host:port.grpcPort` form names it outright.
//! Dialling the HTTP port by mistake fails as "frame with invalid size", which
//! reads like a protocol bug rather than a wrong port, so the rule is worth its
//! own module. Mirrors `pb.ServerToGrpcAddress` in
//! `weed/pb/grpc_client_server.go`.
//!
//! The volume server and the workers each had their own copy of this and the
//! copies had drifted: the worker's bracketed IPv6 literals and the volume
//! server's did not, so `::1:19333` produced `::1:29333`, which the HTTP
//! authority parser rejects. One implementation, two thin wrappers.
use std::fmt;
use std::num::ParseIntError;
/// SeaweedFS's HTTP↔gRPC port-offset convention.
pub const GRPC_PORT_OFFSET: u16 = 10000;
/// Why an address could not be turned into a gRPC address.
///
/// The `Display` text is the volume server's original wording, because its
/// `parse_grpc_address` wrapper hands it straight to callers that put it in a
/// `Status` or an `io::Error`.
#[derive(Debug, Clone, PartialEq, Eq)]
#[non_exhaustive]
pub enum AddressError {
/// No `:` at all, so there is no port to translate.
MissingPort(String),
/// The HTTP port of the `host:port.grpcPort` form is not a `u16`. It is
/// validated even though it is then discarded, so that a malformed address
/// is rejected here instead of failing later as an opaque connect error.
InvalidHttpPort { port: String, source: ParseIntError },
/// The gRPC port of the `host:port.grpcPort` form is not a `u16`.
InvalidGrpcPort { port: String, source: ParseIntError },
/// The port of the `host:port` form is not a `u16`.
InvalidPort { port: String, source: ParseIntError },
/// `port + GRPC_PORT_OFFSET` leaves the TCP port range, e.g. `host:60000`.
/// Without the check the cast would wrap silently.
ImplicitGrpcPortOutOfRange(u16),
}
impl fmt::Display for AddressError {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
match self {
Self::MissingPort(address) => write!(f, "cannot parse address: {address}"),
Self::InvalidHttpPort { port, source } => {
write!(f, "invalid http port {port:?}: {source}")
}
Self::InvalidGrpcPort { port, source } => {
write!(f, "invalid grpc port {port:?}: {source}")
}
Self::InvalidPort { port, source } => write!(f, "invalid port {port:?}: {source}"),
Self::ImplicitGrpcPortOutOfRange(port) => write!(
f,
"implicit grpc port out of range: {port} + {GRPC_PORT_OFFSET} = {}",
u32::from(*port) + u32::from(GRPC_PORT_OFFSET)
),
}
}
}
impl std::error::Error for AddressError {
fn source(&self) -> Option<&(dyn std::error::Error + 'static)> {
match self {
Self::InvalidHttpPort { source, .. }
| Self::InvalidGrpcPort { source, .. }
| Self::InvalidPort { source, .. } => Some(source),
Self::MissingPort(_) | Self::ImplicitGrpcPortOutOfRange(_) => None,
}
}
}
/// Turn a SeaweedFS server address (`"host:port.grpcPort"` or `"host:port"`)
/// into the `host:grpcPort` form the endpoint builders expect.
///
/// With the trailing `.grpcPort` segment that segment *is* the gRPC port;
/// without it the gRPC port is `port + GRPC_PORT_OFFSET`. An unbracketed IPv6
/// literal comes back bracketed, because otherwise the port reads as part of
/// the address.
pub fn to_grpc_address(server: &str) -> Result<String, AddressError> {
// rfind, not find: an IPv6 literal is full of colons and the port is after
// the last one.
let colon_idx = server
.rfind(':')
.ok_or_else(|| AddressError::MissingPort(server.to_string()))?;
let host = &server[..colon_idx];
let port_part = &server[colon_idx + 1..];
// rfind again rather than split_once: the host may be an IPv4 address, and
// only the part after the last colon is being split here anyway.
if let Some(dot_idx) = port_part.rfind('.') {
let http_port = &port_part[..dot_idx];
let grpc_port = &port_part[dot_idx + 1..];
http_port
.parse::<u16>()
.map_err(|source| AddressError::InvalidHttpPort {
port: http_port.to_string(),
source,
})?;
let grpc_port =
grpc_port
.parse::<u16>()
.map_err(|source| AddressError::InvalidGrpcPort {
port: grpc_port.to_string(),
source,
})?;
return Ok(join_host_port(host, grpc_port));
}
let port: u16 = port_part
.parse()
.map_err(|source| AddressError::InvalidPort {
port: port_part.to_string(),
source,
})?;
let grpc_port = port
.checked_add(GRPC_PORT_OFFSET)
.ok_or(AddressError::ImplicitGrpcPortOutOfRange(port))?;
Ok(join_host_port(host, grpc_port))
}
/// Join a host and a port, bracketing an IPv6 literal that is not bracketed
/// already. Public because the address rule is not the only place that has to
/// put a host and a port back together.
pub fn join_host_port(host: &str, port: u16) -> String {
// An IPv6 literal has to keep its brackets or the port reads as part of it.
if host.contains(':') && !host.starts_with('[') {
format!("[{host}]:{port}")
} else {
format!("{host}:{port}")
}
}
#[cfg(test)]
mod tests {
use super::{AddressError, GRPC_PORT_OFFSET, join_host_port, to_grpc_address};
// ---- the volume server's cases -------------------------------------
#[test]
fn dotted_form_states_the_grpc_port() {
assert_eq!(
to_grpc_address("127.0.0.1:8080.18080").unwrap(),
"127.0.0.1:18080"
);
assert_eq!(
to_grpc_address("192.168.1.66:8080.18080").unwrap(),
"192.168.1.66:18080"
);
}
#[test]
fn implicit_form_adds_the_offset() {
assert_eq!(
to_grpc_address("127.0.0.1:8080").unwrap(),
"127.0.0.1:18080"
);
assert_eq!(
to_grpc_address("192.168.1.66:8080").unwrap(),
"192.168.1.66:18080"
);
assert_eq!(
to_grpc_address("localhost:9333").unwrap(),
"localhost:19333"
);
}
#[test]
fn the_dotted_grpc_port_comes_back_normalised() {
// The volume server's copy validated this segment as a u16 and then
// emitted the original text, so a padded or signed port produced an
// authority the URI parser rejects. The parsed value is emitted now.
assert_eq!(to_grpc_address("host:8080.018080").unwrap(), "host:18080");
assert_eq!(to_grpc_address("host:8080.+18080").unwrap(), "host:18080");
}
#[test]
fn an_ipv4_host_is_not_confused_with_the_dotted_port() {
// Regression: a naive split on '.' breaks on IP addresses.
assert_eq!(
to_grpc_address("10.0.0.1:8080.18080").unwrap(),
"10.0.0.1:18080"
);
assert_eq!(to_grpc_address("10.0.0.1:8080").unwrap(), "10.0.0.1:18080");
}
#[test]
fn rejects_a_non_numeric_http_port_in_the_dotted_form() {
let err = to_grpc_address("host:abc.18080").unwrap_err();
assert!(
matches!(err, AddressError::InvalidHttpPort { .. }),
"{err:?}"
);
assert!(err.to_string().contains("invalid http port"), "{err}");
}
#[test]
fn rejects_a_non_numeric_grpc_port_in_the_dotted_form() {
let err = to_grpc_address("host:8080.xyz").unwrap_err();
assert!(
matches!(err, AddressError::InvalidGrpcPort { .. }),
"{err:?}"
);
assert!(err.to_string().contains("invalid grpc port"), "{err}");
}
#[test]
fn rejects_an_implicit_port_that_leaves_the_tcp_range() {
let err = to_grpc_address("127.0.0.1:60000").unwrap_err();
assert!(
matches!(err, AddressError::ImplicitGrpcPortOutOfRange(60000)),
"{err:?}"
);
assert!(err.to_string().contains("out of range"), "{err}");
}
#[test]
fn the_messages_are_the_volume_servers_wording_verbatim() {
// parse_grpc_address hands these straight to callers that put them in a
// Status or an io::Error, so the whole string is the contract, not just
// the substring the older tests match on. Only the two variants whose
// text is entirely ours are pinned exactly; the other three end in a
// std ParseIntError message, which is std's to reword.
assert_eq!(
to_grpc_address("127.0.0.1:60000").unwrap_err().to_string(),
"implicit grpc port out of range: 60000 + 10000 = 70000"
);
assert_eq!(
to_grpc_address("hostname").unwrap_err().to_string(),
"cannot parse address: hostname"
);
}
#[test]
fn rejects_an_address_without_a_port() {
for source in ["hostname", "no-colon", "localhost"] {
let err = to_grpc_address(source).unwrap_err();
assert!(matches!(err, AddressError::MissingPort(_)), "{err:?}");
assert!(err.to_string().contains("cannot parse"), "{err}");
}
}
// ---- the worker's cases --------------------------------------------
#[test]
fn derives_the_grpc_port() {
assert_eq!(
to_grpc_address("localhost:23646").unwrap(),
"localhost:33646"
);
assert_eq!(
to_grpc_address("127.0.0.1:9333").unwrap(),
"127.0.0.1:19333"
);
}
#[test]
fn honours_an_explicit_grpc_port() {
assert_eq!(
to_grpc_address("localhost:23646.33999").unwrap(),
"localhost:33999"
);
}
#[test]
fn rejects_what_it_cannot_parse() {
let err = to_grpc_address("localhost:notaport").unwrap_err();
assert!(matches!(err, AddressError::InvalidPort { .. }), "{err:?}");
assert!(err.to_string().contains("invalid port"), "{err}");
}
// ---- IPv6, which only the worker's copy handled --------------------
#[test]
fn brackets_ipv6_literals() {
assert_eq!(to_grpc_address("::1:23646").unwrap(), "[::1]:33646");
assert_eq!(to_grpc_address("::1:9333").unwrap(), "[::1]:19333");
assert_eq!(
to_grpc_address("fe80::1:9333.19333").unwrap(),
"[fe80::1]:19333"
);
}
#[test]
fn leaves_an_already_bracketed_literal_alone() {
assert_eq!(to_grpc_address("[::1]:9333").unwrap(), "[::1]:19333");
assert_eq!(to_grpc_address("[::1]:9333.19333").unwrap(), "[::1]:19333");
}
#[test]
fn join_host_port_brackets_only_unbracketed_literals() {
assert_eq!(join_host_port("127.0.0.1", 19333), "127.0.0.1:19333");
assert_eq!(join_host_port("localhost", 19333), "localhost:19333");
assert_eq!(join_host_port("::1", 19333), "[::1]:19333");
assert_eq!(join_host_port("[::1]", 19333), "[::1]:19333");
}
// ---- the offset itself ---------------------------------------------
#[test]
fn the_offset_is_the_seaweedfs_convention() {
assert_eq!(GRPC_PORT_OFFSET, 10000);
}
}
+11
View File
@@ -0,0 +1,11 @@
//! Helpers the SeaweedFS Rust volume server and the Rust plugin workers both need.
//!
//! `seaweed-volume` and `seaweed-worker` are separate cargo trees with separate
//! lockfiles and no root manifest, so anything both of them need was, until this
//! crate existed, written twice. The two things in here are the ones where a
//! second copy is a correctness risk rather than a typing cost: the HTTP↔gRPC
//! address rule, which two copies had already drifted on, and the process-wide
//! rustls provider, which only works if every binary installs the same one.
pub mod address;
pub mod tls;
+28
View File
@@ -0,0 +1,28 @@
//! The process-wide rustls crypto provider.
//!
//! Both binaries link aws-lc-rs and ring transitively — in the volume server
//! through the AWS SDK and reqwest, in the lance worker through lance's `aws`
//! backend and reqwest — so rustls cannot auto-select a provider and tonic's
//! client TLS panics on first use. Each binary has to pin one, and it has to be
//! the same one, which is why the choice lives here rather than in either tree.
use rustls::crypto::aws_lc_rs;
/// Pin rustls's process-wide default provider to aws-lc-rs, matching the
/// volume server's TLS config. Idempotent: the first call wins and every
/// later one is a no-op, so callers do not have to coordinate.
pub fn install_default_crypto_provider() {
let _ = aws_lc_rs::default_provider().install_default();
}
#[cfg(test)]
mod tests {
use super::install_default_crypto_provider;
#[test]
fn installing_is_idempotent_and_leaves_a_default_behind() {
install_default_crypto_provider();
install_default_crypto_provider();
assert!(rustls::crypto::CryptoProvider::get_default().is_some());
}
}
+121 -127
View File
@@ -503,7 +503,7 @@ dependencies = [
"rustls-pki-types",
"tokio",
"tokio-rustls",
"tower 0.5.3",
"tower",
"tracing",
]
@@ -628,13 +628,13 @@ dependencies = [
[[package]]
name = "axum"
version = "0.7.9"
version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "edca88bc138befd0323b20752846e6587272d3b03b0343c8ea28a6f819e6e71f"
checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90"
dependencies = [
"async-trait",
"axum-core",
"bytes",
"form_urlencoded",
"futures-util",
"http 1.4.0",
"http-body 1.0.1",
@@ -648,14 +648,13 @@ dependencies = [
"multer",
"percent-encoding",
"pin-project-lite",
"rustversion",
"serde",
"serde_core",
"serde_json",
"serde_path_to_error",
"serde_urlencoded",
"sync_wrapper",
"tokio",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -663,19 +662,17 @@ dependencies = [
[[package]]
name = "axum-core"
version = "0.4.5"
version = "0.5.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09f2bd6146b97ae3359fa0cc6d6b376d9539582c7b4220f041a33ec24c226199"
checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1"
dependencies = [
"async-trait",
"bytes",
"futures-util",
"futures-core",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"mime",
"pin-project-lite",
"rustversion",
"sync_wrapper",
"tower-layer",
"tower-service",
@@ -1654,19 +1651,13 @@ dependencies = [
"futures-core",
"futures-sink",
"http 1.4.0",
"indexmap 2.13.1",
"indexmap",
"slab",
"tokio",
"tokio-util",
"tracing",
]
[[package]]
name = "hashbrown"
version = "0.12.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a9ee70c43aaf417c914396645a0fa852624801b24ebb7ae78fe8272889ac888"
[[package]]
name = "hashbrown"
version = "0.14.5"
@@ -1860,7 +1851,7 @@ dependencies = [
"libc",
"percent-encoding",
"pin-project-lite",
"socket2 0.6.3",
"socket2",
"tokio",
"tower-service",
"tracing",
@@ -2027,16 +2018,6 @@ dependencies = [
"quick-error",
]
[[package]]
name = "indexmap"
version = "1.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bd070e393353796e801d209ad339e89596eb4c8d430d18ede6a1cced8fafbd99"
dependencies = [
"autocfg",
"hashbrown 0.12.3",
]
[[package]]
name = "indexmap"
version = "2.13.1"
@@ -2250,9 +2231,9 @@ dependencies = [
[[package]]
name = "matchit"
version = "0.7.3"
version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0e7465ac9959cc2b1404e8e2367b43684a6d13790fe23056cc8c6c5a6b7bcb94"
checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3"
[[package]]
name = "md-5"
@@ -2619,17 +2600,18 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4c5cc86750666a3ed20bdaf5ca2a0344f9c67674cae0515bec2da16fbaa47db"
dependencies = [
"fixedbitset 0.4.2",
"indexmap 2.13.1",
"indexmap",
]
[[package]]
name = "petgraph"
version = "0.7.1"
version = "0.8.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3672b37090dbd86368a4145bc067582552b29c27377cad4e0a306c97f9bd7772"
checksum = "8701b58ea97060d5e5b155d383a69952a60943f0e6dfe30b04c287beb0b27455"
dependencies = [
"fixedbitset 0.5.7",
"indexmap 2.13.1",
"hashbrown 0.15.5",
"indexmap",
]
[[package]]
@@ -2836,12 +2818,12 @@ dependencies = [
[[package]]
name = "prost"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2796faa41db3ec313a31f7624d9286acf277b52de526150b7e69f3debf891ee5"
checksum = "528ac67416ff8646872a3c02cad9cc4ee5dc9f9540c9b10771855c95cb2e5ae1"
dependencies = [
"bytes",
"prost-derive 0.13.5",
"prost-derive 0.14.4",
]
[[package]]
@@ -2867,19 +2849,20 @@ dependencies = [
[[package]]
name = "prost-build"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042"
dependencies = [
"heck",
"itertools 0.14.0",
"log",
"multimap",
"once_cell",
"petgraph 0.7.1",
"petgraph 0.8.3",
"prettyplease",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"pulldown-cmark",
"pulldown-cmark-to-cmark",
"regex",
"syn",
"tempfile",
@@ -2900,9 +2883,9 @@ dependencies = [
[[package]]
name = "prost-derive"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf"
dependencies = [
"anyhow",
"itertools 0.14.0",
@@ -2922,11 +2905,11 @@ dependencies = [
[[package]]
name = "prost-types"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "52c2c1bf36ddb1a1c396b3601a3cec27c2462e45f07c386894ec3ccf5332bd16"
checksum = "f94967dc7688f3054c7fac87473ffae4cc4c3904800e2d9f5b857246d8963b0a"
dependencies = [
"prost 0.13.5",
"prost 0.14.4",
]
[[package]]
@@ -2993,6 +2976,26 @@ version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.11.0",
"memchr",
"unicase",
]
[[package]]
name = "pulldown-cmark-to-cmark"
version = "22.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab1ad36992cead65f02aa399a373a42730922f1525d988172634fdefdecb8a60"
dependencies = [
"pulldown-cmark",
]
[[package]]
name = "pxfm"
version = "0.1.28"
@@ -3018,7 +3021,7 @@ dependencies = [
"quinn-udp",
"rustc-hash",
"rustls",
"socket2 0.6.3",
"socket2",
"thiserror 2.0.18",
"tokio",
"tracing",
@@ -3056,7 +3059,7 @@ dependencies = [
"cfg_aliases",
"libc",
"once_cell",
"socket2 0.6.3",
"socket2",
"tracing",
"windows-sys 0.60.2",
]
@@ -3262,8 +3265,8 @@ dependencies = [
"tokio",
"tokio-rustls",
"tokio-util",
"tower 0.5.3",
"tower-http 0.6.8",
"tower",
"tower-http",
"tower-service",
"url",
"wasm-bindgen",
@@ -3484,6 +3487,13 @@ version = "1.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "94143f37725109f92c262ed2cf5e59bce7498c01bcc1502d7b9afe439a4e9f49"
[[package]]
name = "seaweed-common"
version = "0.1.0"
dependencies = [
"rustls",
]
[[package]]
name = "sec1"
version = "0.3.0"
@@ -3719,16 +3729,6 @@ version = "1.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1b6b67fb9a61334225b5b790716f609cd58395f895b3fe8b328786812a40bc3b"
[[package]]
name = "socket2"
version = "0.5.10"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e22376abed350d73dd1cd119b57ffccad95b4e585a7cda43e286245ce23c0678"
dependencies = [
"libc",
"windows-sys 0.52.0",
]
[[package]]
name = "socket2"
version = "0.6.3"
@@ -3990,7 +3990,7 @@ dependencies = [
"parking_lot 0.12.5",
"pin-project-lite",
"signal-hook-registry",
"socket2 0.6.3",
"socket2",
"tokio-macros",
"windows-sys 0.61.2",
]
@@ -4077,7 +4077,7 @@ version = "0.22.27"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
dependencies = [
"indexmap 2.13.1",
"indexmap",
"serde",
"serde_spanned",
"toml_datetime",
@@ -4093,11 +4093,10 @@ checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
[[package]]
name = "tonic"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "877c5b330756d856ffcc4553ab34a5684481ade925ecc54bcd1bf02b1d0d4d52"
checksum = "ac2a5518c70fa84342385732db33fb3f44bc4cc748936eb5833d2df34d6445ef"
dependencies = [
"async-stream",
"async-trait",
"axum",
"base64",
@@ -4111,13 +4110,12 @@ dependencies = [
"hyper-util",
"percent-encoding",
"pin-project",
"prost 0.13.5",
"rustls-pemfile",
"socket2 0.5.10",
"socket2",
"sync_wrapper",
"tokio",
"tokio-rustls",
"tokio-stream",
"tower 0.4.13",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -4125,49 +4123,55 @@ dependencies = [
[[package]]
name = "tonic-build"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9557ce109ea773b399c9b9e5dca39294110b74f1f342cb347a80d1fce8c26a11"
checksum = "c68f61875ac5293cf72e6c8cf0158086428c82c37229e98c840878f1706b0322"
dependencies = [
"prettyplease",
"proc-macro2",
"prost-build 0.13.5",
"prost-types 0.13.5",
"quote",
"syn",
]
[[package]]
name = "tonic-reflection"
version = "0.12.3"
name = "tonic-prost"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "878d81f52e7fcfd80026b7fdb6a9b578b3c3653ba987f87f0dce4b64043cba27"
checksum = "50849f68853be452acf590cde0b146665b8d507b3b8af17261df47e02c209ea0"
dependencies = [
"prost 0.13.5",
"prost-types 0.13.5",
"tokio",
"tokio-stream",
"bytes",
"prost 0.14.4",
"tonic",
]
[[package]]
name = "tower"
version = "0.4.13"
name = "tonic-prost-build"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8fa9be0de6cf49e536ce1851f987bd21a43b771b09473c3549a6c853db37c1c"
checksum = "654e5643eff75d7f8c99197ce1440ed19a3474eada74c12bbac488b2cafdae27"
dependencies = [
"futures-core",
"futures-util",
"indexmap 1.9.3",
"pin-project",
"pin-project-lite",
"rand 0.8.7",
"slab",
"prettyplease",
"proc-macro2",
"prost-build 0.14.4",
"prost-types 0.14.4",
"quote",
"syn",
"tempfile",
"tonic-build",
]
[[package]]
name = "tonic-reflection"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "acccd136a4bf19810a1fde9c74edc6129b42a66b44d0c1c8aaa67aeb49a146a7"
dependencies = [
"prost 0.14.4",
"prost-types 0.14.4",
"tokio",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
"tokio-stream",
"tonic",
"tonic-prost",
]
[[package]]
@@ -4178,26 +4182,12 @@ checksum = "ebe5ef63511595f1344e2d5cfa636d973292adc0eec1f0ad45fae9f0851ab1d4"
dependencies = [
"futures-core",
"futures-util",
"indexmap",
"pin-project-lite",
"slab",
"sync_wrapper",
"tokio",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
name = "tower-http"
version = "0.5.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1e9cd434a998747dd2c4276bc96ee2e0c7a2eadf3cae88e52be55a05fa9053f5"
dependencies = [
"bitflags 2.11.0",
"bytes",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"pin-project-lite",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
@@ -4216,9 +4206,10 @@ dependencies = [
"http-body 1.0.1",
"iri-string",
"pin-project-lite",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
@@ -4495,7 +4486,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bb0e353e6a2fbdc176932bbaab493762eb1255a7900fe0fea1a2f96c296cc909"
dependencies = [
"anyhow",
"indexmap 2.13.1",
"indexmap",
"wasm-encoder",
"wasmparser",
]
@@ -4521,7 +4512,7 @@ checksum = "47b807c72e1bac69382b3a6fb3dbe8ea4c0ed87ff5629b8685ae6b9a611028fe"
dependencies = [
"bitflags 2.11.0",
"hashbrown 0.15.5",
"indexmap 2.13.1",
"indexmap",
"semver",
]
@@ -4564,6 +4555,7 @@ dependencies = [
"aws-config",
"aws-credential-types",
"aws-sdk-s3",
"aws-smithy-runtime-api",
"aws-types",
"axum",
"base64",
@@ -4582,7 +4574,6 @@ dependencies = [
"image",
"jsonwebtoken",
"kamadak-exif",
"lazy_static",
"libc",
"md-5",
"memmap2",
@@ -4591,8 +4582,8 @@ dependencies = [
"parking_lot 0.12.5",
"pprof",
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"protoc-bin-vendored",
"rand 0.10.2",
"redb",
@@ -4601,6 +4592,7 @@ dependencies = [
"rustls",
"rustls-pemfile",
"rusty-leveldb",
"seaweed-common",
"serde",
"serde_json",
"serde_urlencoded",
@@ -4613,13 +4605,15 @@ dependencies = [
"tokio-stream",
"toml",
"tonic",
"tonic-build",
"tonic-prost",
"tonic-prost-build",
"tonic-reflection",
"tower 0.4.13",
"tower-http 0.5.2",
"tower",
"tower-http",
"tracing",
"tracing-subscriber",
"uuid",
"windows-sys 0.61.2",
"x509-parser",
"xxhash-rust",
]
@@ -4966,7 +4960,7 @@ checksum = "b7c566e0f4b284dd6561c786d9cb0142da491f46a9fbed79ea69cdad5db17f21"
dependencies = [
"anyhow",
"heck",
"indexmap 2.13.1",
"indexmap",
"prettyplease",
"syn",
"wasm-metadata",
@@ -4997,7 +4991,7 @@ checksum = "9d66ea20e9553b30172b5e831994e35fbde2d165325bec84fc43dbf6f4eb9cb2"
dependencies = [
"anyhow",
"bitflags 2.11.0",
"indexmap 2.13.1",
"indexmap",
"log",
"serde",
"serde_derive",
@@ -5016,7 +5010,7 @@ checksum = "ecc8ac4bc1dc3381b7f59c34f00b67e18f910c2c0f50015669dde7def656a736"
dependencies = [
"anyhow",
"id-arena",
"indexmap 2.13.1",
"indexmap",
"log",
"semver",
"serde",
+29 -9
View File
@@ -23,32 +23,44 @@ default = ["5bytes"]
# Pulls redb's experimental_cursor (and therefore experimental-api-5).
redb-experimental-cursor = ["redb/experimental_cursor"]
[lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
# Every `unsafe` block states its precondition, right above the block.
undocumented_unsafe_blocks = "warn"
[dependencies]
# Helpers the Rust plugin workers (seaweed-worker) need as well. A path
# dependency because the two trees are separate cargo workspaces with no
# common root manifest.
seaweed-common = { path = "../seaweed-common" }
# Async runtime
tokio = { version = "1", features = ["full"] }
tokio-stream = { version = "0.1", features = ["net"] }
tokio-io-timeout = "1"
# gRPC + protobuf
tonic = { version = "0.12", features = ["tls"] }
tonic-reflection = "0.12"
prost = "0.13"
prost-types = "0.13"
tonic = { version = "0.14", features = ["tls-aws-lc"] }
tonic-prost = "0.14"
tonic-reflection = "0.14"
prost = "0.14"
prost-types = "0.14"
# HTTP server
axum = { version = "0.7", features = ["multipart"] }
axum = { version = "0.8", features = ["multipart"] }
http-body = "1"
hyper = { version = "1", features = ["full"] }
hyper-util = { version = "0.1", features = ["tokio", "service", "server-auto", "http1", "http2"] }
tower = "0.4"
tower-http = { version = "0.5", features = ["cors", "trace"] }
tower = { version = "0.5", features = ["util"] }
tower-http = { version = "0.6", features = ["cors", "trace"] }
# CLI
clap = { version = "4", features = ["derive"] }
# Metrics
prometheus = { version = "0.13", default-features = false, features = ["process"] }
lazy_static = "1"
# JWT
jsonwebtoken = { version = "10", features = ["rust_crypto"] }
@@ -138,11 +150,19 @@ aws-types = "1"
[target.'cfg(unix)'.dependencies]
pprof = { version = "0.15", features = ["prost-codec"] }
# GetDiskFreeSpaceExW for per-path disk capacity on Windows (0.61.2 already
# in the tree via tempfile/mio, so this unifies rather than adding a version).
[target.'cfg(windows)'.dependencies]
windows-sys = { version = "0.61", features = ["Win32_Storage_FileSystem"] }
[dev-dependencies]
tempfile = "3"
# Already a transitive dependency of aws-sdk-s3 at a single locked version;
# needed directly only for the canned HttpClient in remote_storage::s3 tests.
aws-smithy-runtime-api = "1"
[build-dependencies]
tonic-build = "0.12"
tonic-prost-build = "0.14"
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry.
+1 -1
View File
@@ -12,7 +12,7 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
tonic_build::configure()
tonic_prost_build::configure()
.build_server(true)
.build_client(true)
// filer.proto uses proto3 optional, which protoc rejects without this
+6
View File
@@ -168,6 +168,7 @@ message VacuumVolumeCheckRequest {
}
message VacuumVolumeCheckResponse {
double garbage_ratio = 1;
bool disk_space_low = 4; // the volume is read-only solely because its disk is low on space — a cause compaction itself reclaims
}
message VacuumVolumeCompactRequest {
@@ -259,6 +260,10 @@ message VolumeDeleteRequest {
// when true, do not remove the cloud-tier object backing the volume.
// used for moves where another server is taking over the same .vif.
bool keep_remote_data = 3;
// when true, delete only if every needle is deleted: the volume held
// data once but nothing is live anymore. Passing either check,
// only_empty or this one, is enough to delete.
bool only_garbage = 4;
}
message VolumeDeleteResponse {
}
@@ -471,6 +476,7 @@ message VolumeEcShardsDeleteRequest {
repeated uint32 shard_ids = 3;
bool full_teardown = 4; // pre-encode cleanup: wipe every EC artifact + generation for this volume, not just shard_ids
int64 encode_ts_ns = 5; // full_teardown generation fence: delete only a disk whose .vif generation is strictly OLDER than this; preserve same-or-newer, generation 0, and an unreadable .vif. 0 => wipe-all (shell pre-encode / pre-upgrade)
uint32 delete_generations_older_than = 6; // post-commit cleanup: delete only staged <base>.*.v<N> artifacts with N strictly below this; 0 disables
}
message VolumeEcShardsDeleteResponse {
bool full_teardown_done = 1; // set by a new server that performed full_teardown; absent from an old server lets the caller detect the silent no-op
File diff suppressed because it is too large Load Diff
+41 -35
View File
@@ -6,19 +6,22 @@ use seaweed_volume::config::{self, VolumeServerConfig};
use seaweed_volume::metrics;
use seaweed_volume::pb::volume_server_pb::volume_server_server::VolumeServerServer;
use seaweed_volume::security::tls::{
build_rustls_server_config, build_rustls_server_config_with_grpc_client_auth,
install_default_crypto_provider, GrpcClientAuthPolicy, TlsPolicy,
GrpcClientAuthPolicy, TlsPolicy, build_rustls_server_config,
build_rustls_server_config_with_grpc_client_auth, install_default_crypto_provider,
};
use seaweed_volume::security::{Guard, SigningKey};
#[cfg(unix)]
use seaweed_volume::server::debug::build_debug_router;
use seaweed_volume::server::grpc_client::load_outgoing_grpc_tls;
use seaweed_volume::server::grpc_client::{
GRPC_INITIAL_WINDOW_SIZE, GRPC_KEEPALIVE_INTERVAL, GRPC_KEEPALIVE_TIMEOUT,
GRPC_MAX_MESSAGE_SIZE, load_outgoing_grpc_tls,
};
use seaweed_volume::server::grpc_server::VolumeGrpcService;
#[cfg(unix)]
use seaweed_volume::server::profiling::CpuProfileSession;
use seaweed_volume::server::request_id::GrpcRequestIdLayer;
use seaweed_volume::server::volume_server::{
build_metrics_router, RuntimeMetricsConfig, VolumeServerState,
RuntimeMetricsConfig, VolumeServerState, build_metrics_router,
};
use seaweed_volume::server::write_queue::WriteQueue;
use seaweed_volume::storage::store::Store;
@@ -31,10 +34,10 @@ type CpuProfileParam = Option<CpuProfileSession>;
#[cfg(not(unix))]
type CpuProfileParam = Option<()>;
const GRPC_MAX_MESSAGE_SIZE: usize = 1 << 30;
const GRPC_KEEPALIVE_INTERVAL: std::time::Duration = std::time::Duration::from_secs(60);
const GRPC_KEEPALIVE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(20);
const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
// The two settings that only make sense for the inbound server. The rest of
// this server's HTTP/2 tuning — keepalive, window sizes, message size — is
// imported from `server::grpc_client` above, which is also what the outgoing
// clients dial with, so the two directions cannot drift apart.
const GRPC_MAX_HEADER_LIST_SIZE: u32 = 8 * 1024 * 1024;
const GRPC_MAX_CONCURRENT_STREAMS: u32 = 1000;
@@ -343,9 +346,6 @@ async fn run(
pre_stop_seconds: config.pre_stop_seconds,
volume_state_notify: tokio::sync::Notify::new(),
write_queue: std::sync::OnceLock::new(),
s3_tier_registry: std::sync::RwLock::new(
seaweed_volume::remote_storage::s3_tier::S3TierRegistry::new(),
),
read_mode: config.read_mode,
allow_untrusted_remote_endpoints: config.allow_untrusted_remote_endpoints,
master_url,
@@ -371,6 +371,9 @@ async fn run(
.to_string_lossy()
.into_owned()
},
ec_decodes_in_flight: std::sync::Mutex::new(std::collections::HashSet::new()),
ec_decode_tail: std::sync::Mutex::new(std::collections::HashSet::new()),
ec_decode_tail_notify: tokio::sync::Notify::new(),
});
// Load persisted state from disk if it exists (matches Go's State.Load on startup)
@@ -671,8 +674,7 @@ async fn run(
})
.await
} else {
let incoming =
tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
let incoming = tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
info!("gRPC server listening on {}", grpc_local_addr);
build_grpc_server_builder()
.layer(GrpcRequestIdLayer)
@@ -1058,15 +1060,17 @@ mod tests {
#[test]
fn test_grpc_server_tls_returns_none_when_files_are_missing() {
assert!(build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none()
);
}
#[test]
@@ -1088,19 +1092,21 @@ mod tests {
"-----BEGIN CERTIFICATE-----\nZmFrZQ==\n-----END CERTIFICATE-----\n",
);
assert!(build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none()
);
}
#[test]
+225 -124
View File
@@ -6,7 +6,7 @@ use prometheus::{
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge,
IntGaugeVec, Opts, Registry, TextEncoder,
};
use std::sync::Once;
use std::sync::{LazyLock, Once};
use crate::version;
@@ -16,220 +16,320 @@ pub struct PushGatewayConfig {
pub interval_seconds: u32,
}
lazy_static::lazy_static! {
pub static ref REGISTRY: Registry = Registry::new();
pub static REGISTRY: LazyLock<Registry> = LazyLock::new(Registry::new);
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static ref REQUEST_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_request_total", "Volume server requests"),
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static REQUEST_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_request_total",
"Volume server requests",
),
&["type", "code"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Request duration histogram with label `type` (HTTP method).
pub static ref REQUEST_DURATION: HistogramVec = HistogramVec::new(
/// Request duration histogram with label `type` (HTTP method).
pub static REQUEST_DURATION: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_request_seconds",
"Volume server request duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
/// Handler-level operation counter with label `type`.
pub static ref HANDLER_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_handler_total", "Volume server handler counters"),
/// Handler-level operation counter with label `type`.
pub static HANDLER_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_handler_total",
"Volume server handler counters",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
/// Vacuuming compact counter with label `success` (true/false).
pub static ref VACUUMING_COMPACT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_compact_count", "Counter of volume vacuuming Compact counter"),
/// Vacuuming compact counter with label `success` (true/false).
pub static VACUUMING_COMPACT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_compact_count",
"Counter of volume vacuuming Compact counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming commit counter with label `success` (true/false).
pub static ref VACUUMING_COMMIT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_commit_count", "Counter of volume vacuuming commit counter"),
/// Vacuuming commit counter with label `success` (true/false).
pub static VACUUMING_COMMIT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_commit_count",
"Counter of volume vacuuming commit counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static ref VACUUMING_HISTOGRAM: HistogramVec = HistogramVec::new(
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static VACUUMING_HISTOGRAM: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_vacuuming_seconds",
"Volume vacuuming duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
/// Volumes per collection and type (volume/ec_shards).
pub static ref VOLUME_GAUGE: GaugeVec = GaugeVec::new(
/// Volumes per collection and type (volume/ec_shards).
pub static VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_volumes", "Number of volumes"),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Read-only volumes per collection and type.
pub static ref READ_ONLY_VOLUME_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_read_only_volumes", "Number of read-only volumes."),
/// Read-only volumes per collection and type.
pub static READ_ONLY_VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_read_only_volumes",
"Number of read-only volumes.",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Maximum number of volumes this server can hold.
pub static ref MAX_VOLUMES: IntGauge = IntGauge::new(
/// Maximum number of volumes this server can hold.
pub static MAX_VOLUMES: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_max_volumes",
"Maximum number of volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static ref DISK_SIZE_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_total_disk_size", "Actual disk size used by volumes"),
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static DISK_SIZE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_total_disk_size",
"Actual disk size used by volumes",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
/// Disk resource usage per directory and type (all/used/free/avail).
pub static ref RESOURCE_GAUGE: GaugeVec = GaugeVec::new(
/// Disk resource usage per directory and type (all/used/free/avail).
pub static RESOURCE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_resource", "Server resource usage"),
&["name", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
/// In-flight requests per HTTP method.
pub static ref INFLIGHT_REQUESTS_GAUGE: IntGaugeVec = IntGaugeVec::new(
Opts::new("SeaweedFS_volumeServer_in_flight_requests", "Current number of in-flight requests being handled by volume server."),
/// In-flight requests per HTTP method.
pub static INFLIGHT_REQUESTS_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_in_flight_requests",
"Current number of in-flight requests being handled by volume server.",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent download limit in bytes.
pub static ref CONCURRENT_DOWNLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent download limit in bytes.
pub static CONCURRENT_DOWNLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_download_limit",
"Limit for total concurrent download size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent upload limit in bytes.
pub static ref CONCURRENT_UPLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent upload limit in bytes.
pub static CONCURRENT_UPLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_upload_limit",
"Limit for total concurrent upload size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight download bytes.
pub static ref INFLIGHT_DOWNLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight download bytes.
pub static INFLIGHT_DOWNLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_download_size",
"In flight total download size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight upload bytes.
pub static ref INFLIGHT_UPLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight upload bytes.
pub static INFLIGHT_UPLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_upload_size",
"In flight total upload size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static ref UPLOAD_ERROR_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created");
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static UPLOAD_ERROR_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created")
});
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static ref SCRUB_LAST_TIME_SECONDS: GaugeVec = GaugeVec::new(
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static SCRUB_LAST_TIME_SECONDS: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_last_time_seconds",
"Last scrub execution time, as seconds since UNIX epoch.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_VOLUME_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static SCRUB_VOLUME_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_volume_failures",
"Counter of overall volumes with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_SHARD_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static SCRUB_SHARD_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_shard_failures",
"Counter of overall EC shards with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static ref STORAGE_IO_ERROR_COUNTER: IntCounter = IntCounter::new(
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static STORAGE_IO_ERROR_COUNTER: LazyLock<IntCounter> = LazyLock::new(|| {
IntCounter::new(
"SeaweedFS_volumeServer_storage_io_error_total",
"Counter of storage read/write EIO errors on volumes and EC shards.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static ref IO_QUARANTINE_GAUGE: IntGaugeVec = IntGaugeVec::new(
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static IO_QUARANTINE_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_io_quarantine",
"Number of volumes or EC shards quarantined due to storage IO errors.",
),
&["kind"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Legacy aliases for backward compat with existing code ----
// ---- Legacy aliases for backward compat with existing code ----
/// Total number of volumes on this server (flat gauge).
pub static ref VOLUMES_TOTAL: IntGauge = IntGauge::new(
"volume_server_volumes_total",
"Total number of volumes",
).expect("metric can be created");
/// Total number of volumes on this server (flat gauge).
pub static VOLUMES_TOTAL: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new("volume_server_volumes_total", "Total number of volumes")
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static ref DISK_SIZE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk size in bytes per directory.
pub static DISK_SIZE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_size_bytes", "Disk size in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Disk free bytes per directory.
pub static ref DISK_FREE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk free bytes per directory.
pub static DISK_FREE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_free_bytes", "Disk free space in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current number of in-flight requests (flat gauge).
pub static ref INFLIGHT_REQUESTS: IntGauge = IntGauge::new(
/// Current number of in-flight requests (flat gauge).
pub static INFLIGHT_REQUESTS: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_inflight_requests",
"Current number of in-flight requests",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Total number of files stored across all volumes.
pub static ref VOLUME_FILE_COUNT: IntGauge = IntGauge::new(
/// Total number of files stored across all volumes.
pub static VOLUME_FILE_COUNT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_volume_file_count",
"Total number of files stored across all volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Build info (Go: BuildInfo) ----
// ---- Build info (Go: BuildInfo) ----
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static ref BUILD_INFO: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created");
}
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static BUILD_INFO: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created")
});
/// Generate exponential bucket boundaries for histograms.
fn exponential_buckets(start: f64, factor: f64, count: usize) -> Vec<f64> {
@@ -249,6 +349,7 @@ pub const DOWNLOAD_LIMIT_COND: &str = "downloadLimitCondition";
pub const UPLOAD_LIMIT_COND: &str = "uploadLimitCondition";
pub const READ_PROXY_REQ: &str = "readProxyRequest";
pub const READ_REDIRECT_REQ: &str = "readRedirectRequest";
pub const READ_DELETED_NEEDLE: &str = "readDeletedNeedle";
pub const EMPTY_READ_PROXY_LOC: &str = "emptyReadProxyLocaction";
pub const FAILED_READ_PROXY_REQ: &str = "failedReadProxyRequest";
@@ -377,10 +478,8 @@ fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
type_value = Some(label.get_value().to_string());
}
}
if matches_collection {
if let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
if matches_collection && let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
}
}
@@ -427,7 +526,7 @@ pub async fn push_metrics_once(
#[cfg(test)]
mod tests {
use super::*;
use axum::{routing::put, Router};
use axum::{Router, routing::put};
use std::sync::{Arc, Mutex};
#[test]
@@ -499,7 +598,9 @@ mod tests {
register_metrics();
VOLUME_GAUGE.with_label_values(&["pics", "volume"]).set(2.0);
VOLUME_GAUGE.with_label_values(&["pics", "ec_shards"]).set(3.0);
VOLUME_GAUGE
.with_label_values(&["pics", "ec_shards"])
.set(3.0);
READ_ONLY_VOLUME_GAUGE
.with_label_values(&["pics", "volume"])
.set(1.0);
@@ -119,7 +119,11 @@ pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
/// reachable for callers whose target legitimately sits on an internal network
/// (peer volume servers), while still blocking loopback, link-local (IMDS) and
/// unspecified. Mirrors Go's `checkBlockedIPPolicy`.
pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool) -> Result<(), String> {
pub fn check_blocked_ip_policy(
endpoint: &str,
ip: IpAddr,
allow_private: bool,
) -> Result<(), String> {
// Normalize IPv4-mapped IPv6 (`::ffff:a.b.c.d`) to its IPv4 form so the
// IPv4 deny rules apply. The OS routes these to the embedded IPv4 address,
// so without this `::ffff:127.0.0.1` / `::ffff:169.254.169.254` would slip
@@ -173,10 +177,10 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if let IpAddr::V6(v6) = ip
&& let Some(v4) = embedded_transition_ipv4(v6)
{
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
Ok(())
}
@@ -214,9 +218,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
// Authority is everything up to the first '/', '?', or '#'.
let after = &trimmed[scheme_end + 3..];
let authority_end = after
.find(|c| c == '/' || c == '?' || c == '#')
.unwrap_or(after.len());
let authority_end = after.find(['/', '?', '#']).unwrap_or(after.len());
let authority = &after[..authority_end];
// Strip optional userinfo ("user:pass@").
@@ -233,7 +235,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
return Err(format!(
"remote endpoint {:?} has a malformed IPv6 host",
endpoint
))
));
}
}
} else {
@@ -309,7 +311,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
return Err("replica target is empty".to_string());
}
if trimmed.contains("://") || trimmed.contains(['/', '?', '#', '@', '\\']) {
return Err(format!("replica target {:?} must be a bare host:port", target));
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
// Require an explicit host:port, handling `[IPv6]:port`. A bracketless IPv6
@@ -318,12 +323,22 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let host = if let Some(rest) = trimmed.strip_prefix('[') {
match rest.split_once(']') {
Some((h, port)) if port.starts_with(':') && port.len() > 1 => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
} else {
match trimmed.rsplit_once(':') {
Some((h, port)) if !port.is_empty() && !h.contains(':') => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
};
@@ -342,7 +357,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let addrs = resolve_host(host).await?;
if addrs.is_empty() {
return Err(format!("resolve replica target host {:?}: no addresses", host));
return Err(format!(
"resolve replica target host {:?}: no addresses",
host
));
}
for ip in addrs {
check_blocked_ip_policy(target, ip, true)?;
@@ -350,6 +368,56 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
Ok(())
}
/// Resolve `host`, re-apply the replica deny list (private peers allowed) to
/// every resolved address, and connect to the first one that passes -- the
/// connect-time twin of [`validate_replica_target`], so a hostname whose DNS
/// answer flips to a blocked address after the up-front check is still refused.
/// Mirrors Go's `guardedDialerPolicy` with allowPrivate=true.
pub async fn guarded_tcp_connect(
host: &str,
port: u16,
endpoint: &str,
) -> std::io::Result<tokio::net::TcpStream> {
use std::io::{Error, ErrorKind};
let denied = |e: String| Error::new(ErrorKind::PermissionDenied, e);
if is_blocked_imds_host(&host.to_ascii_lowercase()) {
return Err(denied(format!(
"remote endpoint {:?} targets instance metadata service",
endpoint
)));
}
if let Ok(ip) = host.parse::<IpAddr>() {
check_blocked_ip_policy(endpoint, ip, true).map_err(denied)?;
return tokio::net::TcpStream::connect((ip, port)).await;
}
let lookup = tokio::net::lookup_host((host.to_string(), port));
let addrs = tokio::time::timeout(std::time::Duration::from_secs(2), lookup)
.await
.map_err(|_| {
Error::new(
ErrorKind::TimedOut,
format!("resolve remote endpoint host {:?}: timed out", host),
)
})??;
let mut first_block_err: Option<String> = None;
for addr in addrs {
if let Err(e) = check_blocked_ip_policy(endpoint, addr.ip(), true) {
if first_block_err.is_none() {
first_block_err = Some(e);
}
continue;
}
return tokio::net::TcpStream::connect(addr).await;
}
Err(denied(first_block_err.unwrap_or_else(|| {
format!("resolve remote endpoint host {:?}: no addresses", host)
})))
}
#[cfg(test)]
mod tests {
use super::*;
@@ -380,22 +448,30 @@ mod tests {
#[test]
fn rejects_empty_and_bad_scheme() {
assert!(precheck_endpoint("").unwrap_err().contains("empty"));
assert!(precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https"));
assert!(precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https"));
assert!(
precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https")
);
assert!(
precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https")
);
}
#[test]
fn rejects_imds_hostnames() {
assert!(precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service"));
assert!(precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service"));
assert!(
precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service")
);
assert!(
precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service")
);
}
#[test]
@@ -415,27 +491,41 @@ mod tests {
#[test]
fn check_blocked_ip_matches_resolved_categories() {
// Mirror Go's "host resolves to X" cases at the address level.
assert!(check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local"));
assert!(check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT"));
assert!(check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local")
);
assert!(
check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT")
);
assert!(
check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private")
);
assert!(check_blocked_ip("e", ip("52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("2606:4700:4700::1111")).is_ok());
}
@@ -478,33 +568,45 @@ mod tests {
assert!(check_blocked_ip("e", ip("2001::f7f7:f7f7")).is_ok());
assert!(check_blocked_ip("e", ip("::808:808")).is_ok());
// Bracketed transition literal via the full endpoint path.
assert!(precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata"));
assert!(
precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata")
);
}
#[test]
fn rejects_ipv4_mapped_ipv6() {
// IPv4-mapped IPv6 must be unmapped so the IPv4 rules catch it.
assert!(check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private")
);
// A mapped public address still passes, and genuine IPv6 loopback is
// still caught by the V6 path.
assert!(check_blocked_ip("e", ip("::ffff:52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback"));
assert!(
check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback")
);
// Bracketed mapped literal via the full endpoint path.
assert!(precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback"));
assert!(
precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback")
);
}
#[test]
@@ -514,60 +616,86 @@ mod tests {
assert!(check_blocked_ip_policy("e", ip("192.168.1.5"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("100.64.0.42"), true).is_ok());
// Loopback / IMDS / unspecified stay blocked even when private is allowed.
assert!(check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified"));
assert!(
check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified")
);
}
#[tokio::test]
async fn validate_replica_target_rejects_and_allows() {
// A path plus a trailing ?a= would otherwise swallow ?type=replicate.
assert!(validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("")
.await
.unwrap_err()
.contains("empty"));
assert!(
validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("")
.await
.unwrap_err()
.contains("empty")
);
// Legitimate peer volume servers on private networks pass.
assert!(validate_replica_target("10.0.0.7:8080").await.is_ok());
assert!(validate_replica_target("192.168.1.5:8080").await.is_ok());
+1 -1
View File
@@ -7,7 +7,7 @@ pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::{validate_remote_endpoint, validate_replica_target};
pub use endpoint_guard::{guarded_tcp_connect, validate_remote_endpoint, validate_replica_target};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
+221 -23
View File
@@ -2,9 +2,10 @@
//!
//! Works with AWS S3, MinIO, SeaweedFS S3, and all S3-compatible providers.
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::Client;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::error::{DisplayErrorContext, SdkError};
use aws_sdk_s3::primitives::ByteStream;
use super::{RemoteEntry, RemoteStorageClient, RemoteStorageError};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
@@ -25,6 +26,23 @@ impl S3RemoteStorageClient {
endpoint: &str,
force_path_style: bool,
) -> Self {
let client = Client::from_conf(
Self::config_builder(access_key, secret_key, region, endpoint, force_path_style)
.build(),
);
S3RemoteStorageClient { client, conf }
}
/// Build the SDK config for the given credentials and endpoint. Split out so
/// tests can attach a canned HTTP client before building the [`Client`].
fn config_builder(
access_key: &str,
secret_key: &str,
region: &str,
endpoint: &str,
force_path_style: bool,
) -> aws_sdk_s3::config::Builder {
let region = if region.is_empty() {
"us-east-1"
} else {
@@ -49,9 +67,7 @@ impl S3RemoteStorageClient {
s3_config = s3_config.endpoint_url(endpoint);
}
let client = Client::from_conf(s3_config.build());
S3RemoteStorageClient { client, conf }
s3_config
}
}
@@ -75,13 +91,14 @@ impl RemoteStorageClient for S3RemoteStorageClient {
req = req.range(format!("bytes={}-", offset));
}
let resp = req.send().await.map_err(|e| {
let msg = format!("{}", e);
if msg.contains("NoSuchKey") || msg.contains("404") {
let resp = req.send().await.map_err(|e| match e {
// Go compares `aerr.Code()` to NoSuchKey on GET
// (s3_storage_client.go:436): a bare 404 maps to "NotFound"
// and stays a generic error, as it does here.
SdkError::ServiceError(ref se) if se.err().is_no_such_key() => {
RemoteStorageError::ObjectNotFound(format!("{}/{}", loc.bucket, key))
} else {
RemoteStorageError::Other(format!("s3 get object: {}", e))
}
e => RemoteStorageError::Other(format!("s3 get object: {}", DisplayErrorContext(&e))),
})?;
let data = resp
@@ -108,7 +125,9 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.body(ByteStream::from(data.to_vec()))
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 put object: {}", e)))?;
.map_err(|e| {
RemoteStorageError::Other(format!("s3 put object: {}", DisplayErrorContext(&e)))
})?;
Ok(RemoteEntry {
size: data.len() as i64,
@@ -134,13 +153,18 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.key(key)
.send()
.await
.map_err(|e| {
let msg = format!("{}", e);
if msg.contains("404") || msg.contains("NotFound") {
.map_err(|e| match e {
// Go checks only the raw HTTP status on HEAD
// (s3_storage_client.go:373): a HEAD response carries no
// error body, so a 404 is not-found whatever code the SDK
// assigns, and a non-404 is not.
SdkError::ServiceError(ref se) if se.raw().status().as_u16() == 404 => {
RemoteStorageError::ObjectNotFound(format!("{}/{}", loc.bucket, key))
} else {
RemoteStorageError::Other(format!("s3 head object: {}", e))
}
e => RemoteStorageError::Other(format!(
"s3 head object: {}",
DisplayErrorContext(&e)
)),
})?;
Ok(RemoteEntry {
@@ -160,18 +184,17 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.key(key)
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 delete object: {}", e)))?;
.map_err(|e| {
RemoteStorageError::Other(format!("s3 delete object: {}", DisplayErrorContext(&e)))
})?;
Ok(())
}
async fn list_buckets(&self) -> Result<Vec<String>, RemoteStorageError> {
let resp = self
.client
.list_buckets()
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 list buckets: {}", e)))?;
let resp = self.client.list_buckets().send().await.map_err(|e| {
RemoteStorageError::Other(format!("s3 list buckets: {}", DisplayErrorContext(&e)))
})?;
Ok(resp
.buckets()
@@ -184,3 +207,178 @@ impl RemoteStorageClient for S3RemoteStorageClient {
&self.conf
}
}
#[cfg(test)]
pub(crate) mod tests {
use super::*;
use aws_sdk_s3::config::http::{HttpRequest, HttpResponse};
use aws_sdk_s3::config::retry::RetryConfig;
use aws_sdk_s3::config::{HttpClient, RuntimeComponents};
use aws_sdk_s3::primitives::SdkBody;
use aws_smithy_runtime_api::client::http::{
HttpConnector, HttpConnectorFuture, HttpConnectorSettings, SharedHttpConnector,
};
use aws_smithy_runtime_api::http::StatusCode;
/// An SDK HTTP client that answers every request with one canned response,
/// so the error-mapping paths can be exercised without a network or a
/// running S3 server.
#[derive(Debug, Clone)]
pub(crate) struct CannedResponse {
pub(crate) status: u16,
pub(crate) body: &'static str,
}
impl HttpConnector for CannedResponse {
fn call(&self, _request: HttpRequest) -> HttpConnectorFuture {
let status = StatusCode::try_from(self.status).expect("valid HTTP status");
HttpConnectorFuture::ready(Ok(HttpResponse::new(status, SdkBody::from(self.body))))
}
}
impl HttpClient for CannedResponse {
fn http_connector(
&self,
_settings: &HttpConnectorSettings,
_components: &RuntimeComponents,
) -> SharedHttpConnector {
SharedHttpConnector::new(self.clone())
}
}
fn client_with(status: u16, body: &'static str) -> S3RemoteStorageClient {
let config = S3RemoteStorageClient::config_builder(
"AKIATEST",
"secret",
"us-east-1",
"http://127.0.0.1:1",
true,
)
.http_client(CannedResponse { status, body })
.retry_config(RetryConfig::disabled())
.build();
S3RemoteStorageClient {
client: Client::from_conf(config),
conf: RemoteConf::default(),
}
}
fn location() -> RemoteStorageLocation {
RemoteStorageLocation {
name: "remote".to_string(),
bucket: "bucket".to_string(),
path: "/dir/missing".to_string(),
..Default::default()
}
}
pub(crate) const NO_SUCH_KEY: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>NoSuchKey</Code><Message>The specified key does not exist.</Message><Key>dir/missing</Key></Error>"#;
const NOT_FOUND_BODY: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>NotFound</Code><Message>Not Found</Message></Error>"#;
const ACCESS_DENIED: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>AccessDenied</Code><Message>Access Denied</Message></Error>"#;
#[tokio::test]
async fn get_no_such_key_is_object_not_found() {
let err = client_with(404, NO_SUCH_KEY)
.read_file(&location(), 0, 0)
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::ObjectNotFound(path) if path == "bucket/dir/missing"),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn get_bare_404_is_not_object_not_found() {
// Go compares codes, not statuses, on GET: a body-less 404 stays generic.
let err = client_with(404, "")
.read_file(&location(), 0, 0)
.await
.unwrap_err();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
}
#[tokio::test]
async fn head_404_is_object_not_found() {
let err = client_with(404, "")
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::ObjectNotFound(path) if path == "bucket/dir/missing"),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn head_404_with_foreign_error_body_is_object_not_found() {
// The raw status check makes a 404 not-found whatever body it carries.
let err = client_with(404, NO_SUCH_KEY)
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(err, RemoteStorageError::ObjectNotFound(_)),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn head_not_found_code_on_a_non_404_status_is_not_object_not_found() {
// A NotFound body on a non-404 status stays an error, as in Go.
let err = client_with(400, NOT_FOUND_BODY)
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::Other(msg) if msg.contains("NotFound")),
"expected Other naming the code, got {err:?}"
);
}
#[tokio::test]
async fn get_access_denied_keeps_service_error_code() {
let err = client_with(403, ACCESS_DENIED)
.read_file(&location(), 0, 0)
.await
.unwrap_err();
let msg = err.to_string();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
assert!(
msg.contains("AccessDenied"),
"message should carry the S3 error code, got: {msg}"
);
assert!(
!msg.ends_with("service error"),
"message should not be the bare SdkError Display, got: {msg}"
);
}
#[tokio::test]
async fn head_access_denied_keeps_service_error_code() {
let err = client_with(403, ACCESS_DENIED)
.stat_file(&location())
.await
.unwrap_err();
let msg = err.to_string();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
assert!(
msg.contains("AccessDenied"),
"message should carry the S3 error code, got: {msg}"
);
}
}
+410 -59
View File
@@ -7,15 +7,66 @@ use std::collections::HashMap;
use std::future::Future;
use std::sync::{Arc, OnceLock, RwLock};
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::types::{CompletedMultipartUpload, CompletedPart};
use aws_sdk_s3::Client;
use aws_sdk_s3::config::http::HttpResponse;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::error::{DisplayErrorContext, SdkError};
use aws_sdk_s3::operation::get_object::GetObjectError;
use aws_sdk_s3::operation::head_object::HeadObjectError;
use aws_sdk_s3::types::{CompletedMultipartUpload, CompletedPart};
use tokio::io::{AsyncReadExt, AsyncSeekExt, AsyncWriteExt};
use tokio::sync::Semaphore;
/// Concurrency limit for multipart upload/download (matches Go's s3manager).
const CONCURRENCY: usize = 5;
/// A tier transfer failure. The variant is what callers match on; the
/// message is the operator-facing text.
#[derive(Debug, thiserror::Error)]
pub enum TierError {
/// The remote object does not exist.
#[error("{0}")]
NotFound(String),
/// An S3 request or a local file operation failed.
#[error("{0}")]
Io(String),
/// The tier I/O runtime could not be built or dropped the task.
#[error("{0}")]
RuntimeUnavailable(String),
/// The progress callback asked to stop.
#[error("{0}")]
Aborted(String),
}
// Not-found rules as in remote_storage/s3.rs: HEAD by the raw 404 status,
// GET by the NoSuchKey code only.
fn head_object_error(key: &str, e: SdkError<HeadObjectError, HttpResponse>) -> TierError {
let message = format!("failed to head object {}: {}", key, DisplayErrorContext(&e));
match e {
SdkError::ServiceError(ref se) if se.raw().status().as_u16() == 404 => {
TierError::NotFound(message)
}
_ => TierError::Io(message),
}
}
fn get_object_error(
key: &str,
range: &str,
e: SdkError<GetObjectError, HttpResponse>,
) -> TierError {
let message = format!(
"failed to get object {} range {}: {}",
key,
range,
DisplayErrorContext(&e)
);
match e {
SdkError::ServiceError(ref se) if se.err().is_no_such_key() => TierError::NotFound(message),
_ => TierError::Io(message),
}
}
/// Configuration for an S3 tier backend.
#[derive(Debug, Clone)]
pub struct S3TierConfig {
@@ -89,7 +140,7 @@ impl S3TierBackend {
&self,
file_path: &str,
progress_fn: F,
) -> Result<(String, u64), String>
) -> Result<(String, u64), TierError>
where
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
@@ -97,7 +148,7 @@ impl S3TierBackend {
let metadata = tokio::fs::metadata(file_path)
.await
.map_err(|e| format!("failed to stat file {}: {}", file_path, e))?;
.map_err(|e| TierError::Io(format!("failed to stat file {}: {}", file_path, e)))?;
let file_size = metadata.len();
// Calculate part size: start at 64MB, scale up for very large files (matches Go)
@@ -119,11 +170,16 @@ impl S3TierBackend {
)
.send()
.await
.map_err(|e| format!("failed to create multipart upload: {}", e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to create multipart upload: {}",
DisplayErrorContext(&e)
))
})?;
let upload_id = create_resp
.upload_id()
.ok_or_else(|| "no upload_id in multipart upload response".to_string())?
.ok_or_else(|| TierError::Io("no upload_id in multipart upload response".to_string()))?
.to_string();
// Build list of (part_number, offset, size) for all parts
@@ -159,19 +215,21 @@ impl S3TierBackend {
let _permit = sem
.acquire()
.await
.map_err(|e| format!("semaphore error: {}", e))?;
.map_err(|e| TierError::Io(format!("semaphore error: {}", e)))?;
// Read this part's data from the file at the correct offset
let mut file = tokio::fs::File::open(&fp)
.await
.map_err(|e| format!("failed to open file {}: {}", fp, e))?;
.map_err(|e| TierError::Io(format!("failed to open file {}: {}", fp, e)))?;
file.seek(std::io::SeekFrom::Start(off))
.await
.map_err(|e| format!("failed to seek to offset {}: {}", off, e))?;
.map_err(|e| {
TierError::Io(format!("failed to seek to offset {}: {}", off, e))
})?;
let mut buf = vec![0u8; size];
file.read_exact(&mut buf)
.await
.map_err(|e| format!("failed to read file at offset {}: {}", off, e))?;
file.read_exact(&mut buf).await.map_err(|e| {
TierError::Io(format!("failed to read file at offset {}: {}", off, e))
})?;
let upload_part_resp = client
.upload_part()
@@ -183,7 +241,12 @@ impl S3TierBackend {
.send()
.await
.map_err(|e| {
format!("failed to upload part {} at offset {}: {}", pn, off, e)
TierError::Io(format!(
"failed to upload part {} at offset {}: {}",
pn,
off,
DisplayErrorContext(&e)
))
})?;
let e_tag = upload_part_resp.e_tag().unwrap_or_default().to_string();
@@ -202,9 +265,9 @@ impl S3TierBackend {
};
(guard.1)(uploaded as i64, pct)
};
progress_result?;
progress_result.map_err(TierError::Aborted)?;
Ok::<_, String>(
Ok::<_, TierError>(
CompletedPart::builder()
.e_tag(e_tag)
.part_number(pn)
@@ -219,7 +282,7 @@ impl S3TierBackend {
for handle in handles {
let part = handle
.await
.map_err(|e| format!("upload task panicked: {}", e))??;
.map_err(|e| TierError::Io(format!("upload task panicked: {}", e)))??;
completed_parts.push(part);
}
@@ -236,9 +299,14 @@ impl S3TierBackend {
.multipart_upload(completed_upload)
.send()
.await
.map_err(|e| format!("failed to complete multipart upload: {}", e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to complete multipart upload: {}",
DisplayErrorContext(&e)
))
})?;
Ok::<(), String>(())
Ok::<(), TierError>(())
}
.await;
@@ -281,7 +349,7 @@ impl S3TierBackend {
dest_path: &str,
key: &str,
progress_fn: F,
) -> Result<u64, String>
) -> Result<u64, TierError>
where
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
@@ -293,7 +361,7 @@ impl S3TierBackend {
.key(key)
.send()
.await
.map_err(|e| format!("failed to head object {}: {}", key, e))?;
.map_err(|e| head_object_error(key, e))?;
let file_size = head_resp.content_length().unwrap_or(0) as u64;
@@ -305,10 +373,12 @@ impl S3TierBackend {
.truncate(true)
.open(dest_path)
.await
.map_err(|e| format!("failed to open dest file {}: {}", dest_path, e))?;
.map_err(|e| {
TierError::Io(format!("failed to open dest file {}: {}", dest_path, e))
})?;
file.set_len(file_size)
.await
.map_err(|e| format!("failed to set file length: {}", e))?;
.map_err(|e| TierError::Io(format!("failed to set file length: {}", e)))?;
}
let part_size: u64 = 64 * 1024 * 1024;
@@ -344,7 +414,7 @@ impl S3TierBackend {
let _permit = sem
.acquire()
.await
.map_err(|e| format!("semaphore error: {}", e))?;
.map_err(|e| TierError::Io(format!("semaphore error: {}", e)))?;
let end = off + size - 1;
let range = format!("bytes={}-{}", off, end);
@@ -356,13 +426,13 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(&key, &range, e))?;
let body = get_resp
.body
.collect()
.await
.map_err(|e| format!("failed to read body: {}", e))?;
.map_err(|e| TierError::Io(format!("failed to read body: {}", e)))?;
let bytes = body.into_bytes();
// Write at the correct offset (like Go's WriteAt)
@@ -370,13 +440,17 @@ impl S3TierBackend {
.write(true)
.open(&dp)
.await
.map_err(|e| format!("failed to open dest file {}: {}", dp, e))?;
.map_err(|e| {
TierError::Io(format!("failed to open dest file {}: {}", dp, e))
})?;
file.seek(std::io::SeekFrom::Start(off))
.await
.map_err(|e| format!("failed to seek to offset {}: {}", off, e))?;
.map_err(|e| {
TierError::Io(format!("failed to seek to offset {}: {}", off, e))
})?;
file.write_all(&bytes)
.await
.map_err(|e| format!("failed to write to {}: {}", dp, e))?;
.map_err(|e| TierError::Io(format!("failed to write to {}: {}", dp, e)))?;
// Report progress. The lock is released before the result is
// propagated so an aborting callback cannot poison the mutex
@@ -392,9 +466,9 @@ impl S3TierBackend {
};
(guard.1)(downloaded as i64, pct)
};
progress_result?;
progress_result.map_err(TierError::Aborted)?;
Ok::<_, String>(())
Ok::<_, TierError>(())
}));
}
@@ -402,7 +476,7 @@ impl S3TierBackend {
for handle in handles {
handle
.await
.map_err(|e| format!("download task panicked: {}", e))??;
.map_err(|e| TierError::Io(format!("download task panicked: {}", e)))??;
}
// fsync the file so its content is durable before the caller trims the .vif
@@ -411,16 +485,21 @@ impl S3TierBackend {
.write(true)
.open(dest_path)
.await
.map_err(|e| format!("failed to open {} for fsync: {}", dest_path, e))?;
.map_err(|e| TierError::Io(format!("failed to open {} for fsync: {}", dest_path, e)))?;
synced
.sync_all()
.await
.map_err(|e| format!("failed to fsync {}: {}", dest_path, e))?;
.map_err(|e| TierError::Io(format!("failed to fsync {}: {}", dest_path, e)))?;
Ok(file_size)
}
pub async fn read_range(&self, key: &str, offset: u64, size: usize) -> Result<Vec<u8>, String> {
pub async fn read_range(
&self,
key: &str,
offset: u64,
size: usize,
) -> Result<Vec<u8>, TierError> {
let end = offset + (size as u64).saturating_sub(1);
let range = format!("bytes={}-{}", offset, end);
let resp = self
@@ -431,29 +510,35 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(key, &range, e))?;
let body = resp
.body
.collect()
.await
.map_err(|e| format!("failed to read object {} body: {}", key, e))?;
.map_err(|e| TierError::Io(format!("failed to read object {} body: {}", key, e)))?;
Ok(body.into_bytes().to_vec())
}
/// Delete a file from S3.
pub async fn delete_file(&self, key: &str) -> Result<(), String> {
pub async fn delete_file(&self, key: &str) -> Result<(), TierError> {
self.client
.delete_object()
.bucket(&self.bucket)
.key(key)
.send()
.await
.map_err(|e| format!("failed to delete object {}: {}", key, e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to delete object {}: {}",
key,
DisplayErrorContext(&e)
))
})?;
Ok(())
}
pub fn delete_file_blocking(&self, key: &str) -> Result<(), String> {
pub fn delete_file_blocking(&self, key: &str) -> Result<(), TierError> {
let client = self.client.clone();
let bucket = self.bucket.clone();
let key = key.to_string();
@@ -464,7 +549,13 @@ impl S3TierBackend {
.key(&key)
.send()
.await
.map_err(|e| format!("failed to delete object {}: {}", key, e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to delete object {}: {}",
key,
DisplayErrorContext(&e)
))
})?;
Ok(())
})
}
@@ -474,7 +565,7 @@ impl S3TierBackend {
key: &str,
offset: u64,
size: usize,
) -> Result<Vec<u8>, String> {
) -> Result<Vec<u8>, TierError> {
let client = self.client.clone();
let bucket = self.bucket.clone();
let key = key.to_string();
@@ -488,13 +579,12 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(&key, &range, e))?;
let body = resp
.body
.collect()
.await
.map_err(|e| format!("failed to read object {} body: {}", key, e))?;
let body =
resp.body.collect().await.map_err(|e| {
TierError::Io(format!("failed to read object {} body: {}", key, e))
})?;
Ok(body.into_bytes().to_vec())
})
}
@@ -555,18 +645,279 @@ pub fn global_s3_tier_registry() -> &'static RwLock<S3TierRegistry> {
GLOBAL_S3_TIER_REGISTRY.get_or_init(|| RwLock::new(S3TierRegistry::new()))
}
fn block_on_tier_future<F, T>(future: F) -> Result<T, String>
where
F: Future<Output = Result<T, String>> + Send + 'static,
T: Send + 'static,
{
std::thread::spawn(move || {
let runtime = tokio::runtime::Builder::new_current_thread()
/// The one process-wide runtime for tiered-S3 I/O issued from synchronous
/// storage code. A per-call runtime tore down the SDK's pooled connections
/// after every 64 KiB chunk, re-dialing TLS per read; a long-lived runtime
/// keeps the pool warm.
///
/// Built on first use. A build failure is returned, not cached or panicked:
/// callers sit inside `Volume::destroy` and needle reads, whose own error
/// paths must run, and a later call may succeed.
static TIER_RUNTIME: std::sync::Mutex<Option<tokio::runtime::Runtime>> =
std::sync::Mutex::new(None);
fn tier_handle() -> Result<tokio::runtime::Handle, TierError> {
let mut slot = TIER_RUNTIME
.lock()
.unwrap_or_else(|poisoned| poisoned.into_inner());
if slot.is_none() {
let runtime = tokio::runtime::Builder::new_multi_thread()
.worker_threads(2)
.thread_name("tier-io")
.enable_all()
.build()
.map_err(|e| format!("failed to build tokio runtime: {}", e))?;
runtime.block_on(future)
})
.join()
.map_err(|_| "tier runtime thread panicked".to_string())?
.map_err(|e| {
TierError::RuntimeUnavailable(format!(
"failed to build the tier I/O tokio runtime: {}",
e
))
})?;
*slot = Some(runtime);
}
Ok(slot.as_ref().expect("just initialised").handle().clone())
}
/// Run `future` on the tier runtime and block the calling thread until it
/// finishes. The caller may be a worker of *another* tokio runtime, so this
/// waits on a channel rather than `Handle::block_on`, which panics when
/// called from inside any runtime context.
fn block_on_tier_future<F, T>(future: F) -> Result<T, TierError>
where
F: Future<Output = Result<T, TierError>> + Send + 'static,
T: Send + 'static,
{
let handle = tier_handle()?;
let task = handle.spawn(future);
let (tx, rx) = std::sync::mpsc::sync_channel(1);
handle.spawn(async move {
// The receiver only goes away if the caller was unwound; nothing to
// report then.
let _ = tx.send(task.await);
});
match rx.recv() {
Ok(Ok(result)) => result,
Ok(Err(join_error)) => Err(describe_join_error(join_error)),
Err(_) => Err(TierError::RuntimeUnavailable(
"tier I/O runtime dropped the task before it finished".to_string(),
)),
}
}
/// Turn a `JoinError` into a message that keeps the panic payload, so an
/// SDK panic surfaces as "boom" rather than a fixed "thread panicked".
fn describe_join_error(join_error: tokio::task::JoinError) -> TierError {
if join_error.is_panic() {
let payload = join_error.into_panic();
let message = if let Some(s) = payload.downcast_ref::<&str>() {
(*s).to_string()
} else if let Some(s) = payload.downcast_ref::<String>() {
s.clone()
} else {
"non-string panic payload".to_string()
};
TierError::Io(format!("tier I/O task panicked: {}", message))
} else {
TierError::RuntimeUnavailable(format!("tier I/O task failed: {}", join_error))
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::remote_storage::s3::tests::{CannedResponse, NO_SUCH_KEY};
use std::collections::HashSet;
use tokio::runtime::Handle;
fn probe() -> Result<(tokio::runtime::Id, Option<String>), TierError> {
block_on_tier_future(async {
Ok((
Handle::current().id(),
std::thread::current().name().map(str::to_string),
))
})
}
#[test]
fn block_on_tier_future_reuses_one_runtime() {
let (first_runtime, first_thread) = probe().expect("first call");
let (second_runtime, second_thread) = probe().expect("second call");
assert_eq!(
first_runtime, second_runtime,
"each call must run on the same long-lived tier runtime"
);
assert_eq!(first_thread.as_deref(), Some("tier-io"));
assert_eq!(second_thread.as_deref(), Some("tier-io"));
let mut runtimes = HashSet::new();
for _ in 0..20 {
let (id, _) = probe().expect("probe");
runtimes.insert(id);
}
assert_eq!(runtimes.len(), 1);
}
#[test]
fn block_on_tier_future_returns_the_value_and_the_error() {
assert_eq!(block_on_tier_future(async { Ok(7u32) }).unwrap(), 7);
let err = block_on_tier_future::<_, u32>(async { Err(TierError::NotFound("nope".into())) })
.unwrap_err();
assert!(
matches!(&err, TierError::NotFound(m) if m == "nope"),
"{err:?}"
);
}
#[test]
fn block_on_tier_future_works_from_a_std_thread() {
let (id, _) = std::thread::spawn(probe)
.join()
.expect("probe thread")
.expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn block_on_tier_future_works_from_spawn_blocking() {
let (id, _) = tokio::task::spawn_blocking(probe)
.await
.expect("spawn_blocking")
.expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
// Called straight from another runtime's async context: the case that
// would panic with `Handle::block_on` ("Cannot start a runtime from
// within a runtime").
#[tokio::test]
async fn block_on_tier_future_works_from_a_current_thread_runtime() {
let (id, _) = probe().expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn block_on_tier_future_works_from_a_multi_thread_runtime_worker() {
let (id, _) = probe().expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
#[test]
fn block_on_tier_future_reports_the_panic_payload() {
let err = block_on_tier_future::<_, ()>(async {
if std::hint::black_box(true) {
panic!("boom {}", 42);
}
Ok(())
})
.expect_err("a panicking future must be an error");
assert!(matches!(err, TierError::Io(_)), "got: {err:?}");
assert!(err.to_string().contains("boom 42"), "got: {err}");
assert!(err.to_string().contains("panicked"), "got: {err}");
}
#[test]
fn block_on_tier_future_reports_a_str_panic_payload() {
let err = block_on_tier_future::<_, ()>(async {
if std::hint::black_box(true) {
panic!("static boom");
}
Ok(())
})
.expect_err("a panicking future must be an error");
assert!(err.to_string().contains("static boom"), "got: {err}");
}
fn backend_answering(status: u16, body: &'static str) -> S3TierBackend {
let config = aws_sdk_s3::Config::builder()
.behavior_version(BehaviorVersion::latest())
.region(Region::new("us-east-1"))
.credentials_provider(Credentials::new("AKIATEST", "secret", None, None, "test"))
.endpoint_url("http://127.0.0.1:1")
.force_path_style(true)
.http_client(CannedResponse { status, body })
.retry_config(aws_sdk_s3::config::retry::RetryConfig::disabled())
.build();
S3TierBackend {
client: Client::from_conf(config),
bucket: "bucket".to_string(),
storage_class: "STANDARD".to_string(),
}
}
#[tokio::test]
async fn download_head_404_is_not_found() {
let tmp = tempfile::tempdir().unwrap();
let dest = tmp.path().join("1.dat");
let err = backend_answering(404, "")
.download_file(dest.to_str().unwrap(), "missing", |_, _| Ok(()))
.await
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
assert!(
err.to_string()
.starts_with("failed to head object missing: "),
"{err}"
);
}
#[tokio::test]
async fn download_head_403_is_io() {
let tmp = tempfile::tempdir().unwrap();
let dest = tmp.path().join("1.dat");
let err = backend_answering(403, "")
.download_file(dest.to_str().unwrap(), "denied", |_, _| Ok(()))
.await
.unwrap_err();
assert!(matches!(err, TierError::Io(_)), "{err:?}");
}
#[tokio::test]
async fn read_range_no_such_key_is_not_found() {
let err = backend_answering(404, NO_SUCH_KEY)
.read_range("missing", 0, 8)
.await
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
assert!(
err.to_string()
.starts_with("failed to get object missing range bytes=0-7: "),
"{err}"
);
}
#[tokio::test]
async fn read_range_bare_404_is_io() {
// As in Go, GET is not-found by the NoSuchKey code, not the status.
let err = backend_answering(404, "")
.read_range("missing", 0, 8)
.await
.unwrap_err();
assert!(matches!(err, TierError::Io(_)), "{err:?}");
}
#[test]
fn read_range_blocking_no_such_key_is_not_found() {
let err = backend_answering(404, NO_SUCH_KEY)
.read_range_blocking("missing", 0, 8)
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
}
#[test]
fn backend_name_to_type_id_splits_on_dot() {
assert_eq!(
backend_name_to_type_id("s3"),
("s3".to_string(), "default".to_string())
);
assert_eq!(
backend_name_to_type_id("s3.eu"),
("s3".to_string(), "eu".to_string())
);
assert_eq!(
backend_name_to_type_id("s3.a.b"),
(String::new(), String::new())
);
}
}
+10 -8
View File
@@ -10,7 +10,7 @@ use std::collections::HashSet;
use std::net::IpAddr;
use std::time::{SystemTime, UNIX_EPOCH};
use jsonwebtoken::{decode, encode, Algorithm, DecodingKey, EncodingKey, Header, Validation};
use jsonwebtoken::{Algorithm, DecodingKey, EncodingKey, Header, Validation, decode, encode};
use serde::{Deserialize, Serialize};
// ============================================================================
@@ -297,10 +297,10 @@ impl Guard {
/// Extract host from "host:port" or "[::1]:port" format.
fn extract_host(addr: &str) -> String {
// Handle IPv6 with brackets
if addr.starts_with('[') {
if let Some(end) = addr.find(']') {
return addr[1..end].to_string();
}
if addr.starts_with('[')
&& let Some(end) = addr.find(']')
{
return addr[1..end].to_string();
}
// Handle host:port
if let Some(pos) = addr.rfind(':') {
@@ -481,9 +481,11 @@ mod tests {
let token = gen_jwt(&key, 3600, "3,01637037d6").unwrap();
// Correct file ID
assert!(guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok());
assert!(
guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok()
);
// Wrong file ID
let err = guard.check_jwt_for_file(Some(&token), "4,deadbeef", true);
+8 -7
View File
@@ -3,12 +3,12 @@ use std::fmt;
use std::sync::Arc;
use rustls::client::danger::HandshakeSignatureValid;
use rustls::crypto::aws_lc_rs;
use rustls::crypto::CryptoProvider;
use rustls::crypto::aws_lc_rs;
use rustls::pki_types::UnixTime;
use rustls::pki_types::{CertificateDer, PrivateKeyDer};
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::server::WebPkiClientVerifier;
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::{
CipherSuite, DigitallySignedStruct, DistinguishedName, RootCertStore, ServerConfig,
SignatureScheme, SupportedCipherSuite, SupportedProtocolVersion,
@@ -120,10 +120,11 @@ impl ClientCertVerifier for CommonNameVerifier {
// aws-lc-rs and ring both get linked transitively, so rustls can't auto-select
// a provider and tonic's client TLS panics on first use. Pin the default to
// aws-lc-rs, matching the server config. Idempotent.
pub fn install_default_crypto_provider() {
let _ = aws_lc_rs::default_provider().install_default();
}
// aws-lc-rs, matching the server config. Idempotent. The body lives in
// seaweed-common so this binary and the Rust plugin workers cannot end up
// installing different providers; re-exported here so callers keep their
// import path.
pub use seaweed_common::tls::install_default_crypto_provider;
pub fn build_rustls_server_config(
cert_path: &str,
@@ -376,7 +377,7 @@ fn go_tls_version_for_supported(version: &SupportedProtocolVersion) -> GoTlsVers
#[cfg(test)]
mod tests {
use super::{build_supported_versions, common_name_is_allowed, parse_cipher_suites, TlsPolicy};
use super::{TlsPolicy, build_supported_versions, common_name_is_allowed, parse_cipher_suites};
use rustls::crypto::aws_lc_rs;
use std::collections::HashSet;
+2 -2
View File
@@ -1,9 +1,9 @@
use axum::Router;
use axum::body::Body;
use axum::extract::Query;
use axum::http::{header, StatusCode};
use axum::http::{StatusCode, header};
use axum::response::{IntoResponse, Response};
use axum::routing::{any, get};
use axum::Router;
use pprof::protos::Message;
use serde::Deserialize;
+352 -45
View File
@@ -1,16 +1,42 @@
//! Construction of the volume server's *outgoing* gRPC clients: TLS material,
//! endpoint tuning, dial bounds, and the three client constructors every call
//! site goes through.
//!
//! The keepalive, window-size and message-size constants below are shared with
//! the *inbound* server built in `main.rs`, which imports them from here rather
//! than declaring its own. Changing one therefore changes both directions at
//! once, which is deliberate: a volume server talks to its peers with the same
//! HTTP/2 settings it offers them.
use std::error::Error;
use std::fmt;
use std::time::Duration;
use hyper::http::Uri;
use tonic::service::interceptor::InterceptedService;
use tonic::transport::{Certificate, Channel, ClientTlsConfig, Endpoint, Identity};
use tonic::{Request, Status};
use crate::config::VolumeServerConfig;
use crate::pb::filer_pb::seaweed_filer_client::SeaweedFilerClient;
use crate::pb::master_pb::seaweed_client::SeaweedClient;
use crate::pb::volume_server_pb::volume_server_client::VolumeServerClient;
use crate::server::request_id::outgoing_request_id_interceptor;
pub const GRPC_MAX_MESSAGE_SIZE: usize = 1 << 30;
const GRPC_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(60);
const GRPC_KEEPALIVE_TIMEOUT: Duration = Duration::from_secs(20);
const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
pub const GRPC_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(60);
pub const GRPC_KEEPALIVE_TIMEOUT: Duration = Duration::from_secs(20);
pub const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
/// Bound on the TCP connect of every outgoing dial. `build_grpc_endpoint` is
/// private and `connect_channel` is the only way out of this module, so every
/// call site picks this up whether it thinks about timeouts or not.
///
/// It bounds the TCP handshake only — tonic hands it to
/// `HttpConnector::set_connect_timeout`. A peer that completes the handshake
/// and then stalls in the TLS or HTTP/2 exchange is not covered; callers that
/// need that bound wrap the whole dial (see `connect_ping_target`).
const GRPC_CONNECT_TIMEOUT: Duration = Duration::from_secs(5);
#[derive(Clone, Debug)]
pub struct OutgoingGrpcTlsConfig {
@@ -40,7 +66,9 @@ pub fn load_outgoing_grpc_tls(
(&config.grpc_client_cert_file, &config.grpc_client_key_file)
} else {
if !config.grpc_client_cert_file.is_empty() || !config.grpc_client_key_file.is_empty() {
tracing::warn!("grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key");
tracing::warn!(
"grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key"
);
}
(&config.grpc_cert_file, &config.grpc_key_file)
};
@@ -79,7 +107,7 @@ pub fn grpc_endpoint_uri(grpc_host_port: &str, tls: Option<&OutgoingGrpcTlsConfi
format!("{}://{}", scheme, grpc_host_port)
}
pub fn build_grpc_endpoint(
fn build_grpc_endpoint(
grpc_host_port: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
) -> Result<Endpoint, GrpcClientError> {
@@ -115,6 +143,184 @@ pub fn build_grpc_endpoint(
Ok(endpoint)
}
/// Connect `endpoint` through a connector that re-validates every resolved
/// address at connect time (Go's `guardedDialerPolicy` mirror), pinning a
/// validated copy/tail source against DNS rebinding. `allow_untrusted`
/// preserves the plain connect for operators that opted out.
pub async fn connect_guarded(
endpoint: Endpoint,
target: &str,
allow_untrusted: bool,
) -> Result<Channel, GrpcClientError> {
if allow_untrusted {
return endpoint
.connect()
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)));
}
let target_owned = target.to_string();
let connector = tower::service_fn(move |uri: Uri| {
let target = target_owned.clone();
async move {
let host = uri.host().unwrap_or_default().to_string();
let port = uri.port_u16().unwrap_or(80);
crate::remote_storage::guarded_tcp_connect(&host, port, &target)
.await
.map(hyper_util::rt::TokioIo::new)
}
});
endpoint
.connect_with_connector(connector)
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)))
}
/// How a dial is bounded.
///
/// `connect_timeout` is handed to the TCP connector. `request_timeout` becomes
/// [`Endpoint::timeout`], which tonic installs as a `GrpcTimeout` layer in
/// front of *every* request the resulting channel carries — it is not a
/// property of one call.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct GrpcDialOptions {
/// Bound on establishing the connection to the peer.
pub connect_timeout: Duration,
/// Deadline applied to each RPC on the channel, or `None` to leave them
/// unbounded.
pub request_timeout: Option<Duration>,
}
impl GrpcDialOptions {
/// A short request/response call: connect within 5 s, answer within 10 s.
pub fn unary() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: Some(Duration::from_secs(10)),
}
}
/// A call the peer may take a while to answer: connect within 5 s, answer
/// within 30 s.
pub fn long() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: Some(Duration::from_secs(30)),
}
}
/// A bounded connect with no deadline on the RPCs themselves.
///
/// `request_timeout` must stay `None` here. [`Endpoint::timeout`] is not a
/// transfer budget: tonic layers it as a `GrpcTimeout` around the
/// response future, which resolves when the server's *first response
/// headers* arrive, so it bounds how long the peer may take to start
/// answering — per request, for every request the channel carries. A 10 s
/// value picked to suit one short call would therefore also be the header
/// deadline for the `VolumeCopy` that shares the dial, and a busy source
/// that takes longer than that to open its file would lose the whole copy.
/// `VolumeCopy`, `VolumeTailSender` and `VolumeEcShardsCopy` have never
/// carried one.
pub fn stream() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: None,
}
}
}
/// Dial a peer and return a connected channel.
///
/// The error carries only the transport failure: every caller already wraps it
/// with the address and the operation it was attempting.
pub async fn connect_channel(
grpc_host_port: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
opts: GrpcDialOptions,
) -> Result<Channel, GrpcClientError> {
let mut endpoint =
build_grpc_endpoint(grpc_host_port, tls)?.connect_timeout(opts.connect_timeout);
if let Some(request_timeout) = opts.request_timeout {
endpoint = endpoint.timeout(request_timeout);
}
endpoint
.connect()
.await
.map_err(|e| GrpcClientError(e.to_string()))
}
/// Dial a copy/tail source and return a connected channel, re-validating every
/// resolved address at connect time.
///
/// The guarded equivalent of [`connect_channel`]: same `opts` bounds, but the
/// dial goes through [`connect_guarded`] so a source address that passed
/// validation cannot be re-pointed by DNS between the check and the connect.
/// The bounds are applied to the endpoint *before* delegating, so the
/// `allow_untrusted` opt-out is timed too.
///
/// `target` is the caller-facing source address (the unparsed
/// `"ip:port.grpcPort"` form), which is what the guard pins against; the error
/// carries only the transport failure, as every caller already wraps it with
/// the address and the operation it was attempting.
pub async fn connect_channel_guarded(
grpc_host_port: &str,
target: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
opts: GrpcDialOptions,
allow_untrusted: bool,
) -> Result<Channel, GrpcClientError> {
let mut endpoint =
build_grpc_endpoint(grpc_host_port, tls)?.connect_timeout(opts.connect_timeout);
if let Some(request_timeout) = opts.request_timeout {
endpoint = endpoint.timeout(request_timeout);
}
connect_guarded(endpoint, target, allow_untrusted).await
}
/// The outgoing request-id interceptor as a concrete type, so the client
/// aliases below can name it.
pub type RequestIdInterceptor = fn(Request<()>) -> Result<Request<()>, Status>;
/// A volume-server client with the request-id interceptor attached.
pub type VolumeServerGrpcClient =
VolumeServerClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// A master client with the request-id interceptor attached.
pub type MasterGrpcClient = SeaweedClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// A filer client with the request-id interceptor attached.
pub type FilerGrpcClient = SeaweedFilerClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// Wrap a connected channel in a volume-server client that forwards the
/// current request id and lifts both message-size limits.
pub fn volume_server_client(channel: Channel) -> VolumeServerGrpcClient {
VolumeServerClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Wrap a connected channel in a master client that forwards the current
/// request id and lifts both message-size limits.
pub fn master_client(channel: Channel) -> MasterGrpcClient {
SeaweedClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Wrap a connected channel in a filer client that forwards the current
/// request id and lifts both message-size limits.
pub fn filer_client(channel: Channel) -> FilerGrpcClient {
SeaweedFilerClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Parse a SeaweedFS server address (`"ip:port.grpcPort"` or
/// `"ip:port"`) into the `host:grpcPort` form `build_grpc_endpoint`
/// expects. With the trailing `.grpcPort` segment, that segment IS
@@ -124,53 +330,27 @@ pub fn build_grpc_endpoint(
/// Shared between `grpc_server.rs` and the distributed-EC-read path
/// in `store_ec.rs` — keep this as the single source of truth so the
/// HTTP↔gRPC port translation can't drift between callers.
///
/// The rule itself lives in `seaweed_common::address`, which the Rust
/// plugin workers share; this wrapper only flattens the typed error
/// back to the `String` its callers already handle. Unbracketed IPv6
/// literals come back bracketed, which this copy used to get wrong.
pub fn parse_grpc_address(source: &str) -> Result<String, String> {
let colon_idx = source
.rfind(':')
.ok_or_else(|| format!("cannot parse address: {}", source))?;
let host = &source[..colon_idx];
let port_part = &source[colon_idx + 1..];
if let Some(dot_idx) = port_part.rfind('.') {
// Format: "ip:port.grpcPort". Validate BOTH ports as u16
// so a malformed HTTP port (e.g. `host:abc.18080`) is
// rejected here rather than tripping a downstream
// `build_grpc_endpoint` URI parse failure with a less
// useful error.
let http_port = &port_part[..dot_idx];
let grpc_port = &port_part[dot_idx + 1..];
http_port
.parse::<u16>()
.map_err(|e| format!("invalid http port {:?}: {}", http_port, e))?;
grpc_port
.parse::<u16>()
.map_err(|e| format!("invalid grpc port {:?}: {}", grpc_port, e))?;
return Ok(format!("{}:{}", host, grpc_port));
}
// Format: "ip:port" → grpc = port + 10000. Reject inputs whose
// implicit grpc port would overflow the TCP port range (e.g.
// `host:60000` produces 70000 — invalid). Without this check
// the cast silently wraps and the endpoint call later fails
// with an opaque connection error.
let port: u16 = port_part
.parse()
.map_err(|e| format!("invalid port {:?}: {}", port_part, e))?;
let grpc_port = port as u32 + 10000;
if grpc_port > u16::MAX as u32 {
return Err(format!(
"implicit grpc port out of range: {} + 10000 = {}",
port, grpc_port
));
}
Ok(format!("{}:{}", host, grpc_port))
seaweed_common::address::to_grpc_address(source).map_err(|e| e.to_string())
}
#[cfg(test)]
mod tests {
use super::{build_grpc_endpoint, grpc_endpoint_uri, load_outgoing_grpc_tls};
use super::{
GrpcDialOptions, build_grpc_endpoint, connect_channel, grpc_endpoint_uri,
load_outgoing_grpc_tls, volume_server_client,
};
use crate::config::{NeedleMapKind, ReadMode, VolumeServerConfig};
use crate::pb::volume_server_pb;
use crate::security::tls::TlsPolicy;
use crate::server::request_id::scope_request_id;
use std::sync::{Arc, Mutex};
use std::time::Duration;
const TEST_CERT_PEM: &str = "-----BEGIN CERTIFICATE-----\nMIIBPDCB76ADAgECAhRuRPQgeAu43BT/M7EfAWSdapVdYDAFBgMrZXAwFDESMBAG\nA1UEAwwJbG9jYWxob3N0MB4XDTI2MDcwNTE2MTUyOVoXDTM2MDcwMjE2MTUyOVow\nFDESMBAGA1UEAwwJbG9jYWxob3N0MCowBQYDK2VwAyEAr/3bNIFI+8V32oCiY6y+\nXRFmZpdNQ2g//VtRkT+nQg+jUzBRMB0GA1UdDgQWBBTsy9tLf1zPiXCQfgci6zNi\ndEzRSjAfBgNVHSMEGDAWgBTsy9tLf1zPiXCQfgci6zNidEzRSjAPBgNVHRMBAf8E\nBTADAQH/MAUGAytlcANBAIvsdw0IbvOBBkb9cd7BfMJfIP9pQQrAL03pCRWJFnFh\nSysaLVgFXI4T078IiaM874oO+iB+5vNbWEpc7CkGow4=\n-----END CERTIFICATE-----\n";
const TEST_KEY_PEM: &str = "-----BEGIN PRIVATE KEY-----\nMC4CAQAwBQYDK2VwBCIEIHbyn71Kk+Y7KT3sBctit7uZpErpoH6qDbFj6P8qGaZH\n-----END PRIVATE KEY-----\n";
@@ -365,4 +545,131 @@ mod tests {
let err = parse_grpc_address("hostname").unwrap_err();
assert!(err.contains("cannot parse"), "{}", err);
}
#[test]
fn test_parse_grpc_address_brackets_ipv6_literals() {
use super::parse_grpc_address;
// This used to come back as `::1:29333`, which is not a valid
// authority: `build_grpc_endpoint` reads the last colon as the port
// separator and rejects the rest.
assert_eq!(parse_grpc_address("::1:19333").unwrap(), "[::1]:29333");
assert_eq!(parse_grpc_address("::1:9333.19333").unwrap(), "[::1]:19333");
// Already bracketed, so it is left alone.
assert_eq!(parse_grpc_address("[::1]:9333").unwrap(), "[::1]:19333");
}
#[test]
fn test_build_grpc_endpoint_accepts_an_ipv6_master_address() {
use super::parse_grpc_address;
let endpoint = build_grpc_endpoint(&parse_grpc_address("::1:9333").unwrap(), None).unwrap();
assert_eq!(endpoint.uri().port_u16(), Some(19333));
}
/// A minimal HTTP/2 server that records the gRPC request headers it is
/// sent and answers every call with a trailers-only `unimplemented`. It is
/// enough to prove what a helper-built client puts on the wire, without
/// standing up the whole `VolumeServer` service behind a tonic server.
async fn serve_header_capture() -> (u16, Arc<Mutex<Option<String>>>) {
use hyper::service::service_fn;
use hyper_util::rt::{TokioExecutor, TokioIo};
let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let seen: Arc<Mutex<Option<String>>> = Arc::new(Mutex::new(None));
let captured = Arc::clone(&seen);
tokio::spawn(async move {
while let Ok((stream, _)) = listener.accept().await {
let captured = Arc::clone(&captured);
tokio::spawn(async move {
let _ = hyper::server::conn::http2::Builder::new(TokioExecutor::new())
.serve_connection(
TokioIo::new(stream),
service_fn(move |req: hyper::Request<hyper::body::Incoming>| {
let captured = Arc::clone(&captured);
async move {
let value = req
.headers()
.get("x-amz-request-id")
.and_then(|v| v.to_str().ok())
.map(str::to_string);
*captured.lock().unwrap() = value;
Ok::<_, std::convert::Infallible>(
hyper::http::Response::builder()
.status(200)
.header("content-type", "application/grpc")
.header("grpc-status", "12")
.body(tonic::body::Body::empty())
.unwrap(),
)
}
}),
)
.await;
});
}
});
(port, seen)
}
#[tokio::test]
async fn test_helper_built_client_sends_the_scoped_request_id() {
let (port, seen) = serve_header_capture().await;
let channel = connect_channel(
&format!("127.0.0.1:{}", port),
None,
GrpcDialOptions::unary(),
)
.await
.expect("dial the header-capturing server");
let mut client = volume_server_client(channel);
// The interceptor has a request id to forward only inside a scope, so
// the call has to run inside one for this to test anything.
let _ = scope_request_id("REQUEST-ID-ON-THE-WIRE".to_string(), async move {
client
.ping(volume_server_pb::PingRequest {
target: String::new(),
target_type: String::new(),
})
.await
})
.await;
assert_eq!(
seen.lock().unwrap().as_deref(),
Some("REQUEST-ID-ON-THE-WIRE"),
"a client built by volume_server_client must carry the outgoing request id"
);
}
#[test]
fn test_dial_presets_match_the_call_sites_they_replace() {
assert_eq!(
GrpcDialOptions::unary().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::unary().request_timeout,
Some(Duration::from_secs(10))
);
assert_eq!(
GrpcDialOptions::long().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::long().request_timeout,
Some(Duration::from_secs(30))
);
assert_eq!(
GrpcDialOptions::stream().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::stream().request_timeout,
None,
"a streaming dial must not put a per-request deadline on the channel"
);
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+13 -9
View File
@@ -36,16 +36,20 @@ pub fn collect_mem_status() -> volume_server_pb::MemStatus {
#[cfg(target_os = "linux")]
fn get_system_memory_linux() -> Option<(u64, u64)> {
unsafe {
let mut info: libc::sysinfo = std::mem::zeroed();
if libc::sysinfo(&mut info) == 0 {
let unit = info.mem_unit as u64;
let total = info.totalram as u64 * unit;
let free = info.freeram as u64 * unit;
return Some((total, free));
}
// SAFETY: `libc::sysinfo` is plain data — integers and trailing padding,
// no pointers and no restricted niches — so the all-zero value is a valid
// one for the kernel to overwrite.
let mut info: libc::sysinfo = unsafe { std::mem::zeroed() };
// SAFETY: `&mut info` is a live, aligned, exclusive pointer to a
// `sysinfo` that the kernel only writes through, and its fields are read
// below only after the call reports success.
if unsafe { libc::sysinfo(&mut info) } != 0 {
return None;
}
None
let unit = info.mem_unit as u64;
let total = info.totalram as u64 * unit;
let free = info.freeram as u64 * unit;
Some((total, free))
}
#[cfg(target_os = "linux")]
+113
View File
@@ -1,3 +1,8 @@
use tonic::Status;
use crate::remote_storage::s3_tier::TierError;
use crate::storage::volume::VolumeError;
#[cfg(unix)]
pub mod debug;
pub mod grpc_client;
@@ -13,3 +18,111 @@ pub mod store_ec;
pub mod ui;
pub mod volume_server;
pub mod write_queue;
/// Map a storage error onto the gRPC code that describes it.
impl From<VolumeError> for Status {
fn from(err: VolumeError) -> Self {
let message = err.to_string();
match err {
VolumeError::NotFound
| VolumeError::VolumeNotFound(_)
| VolumeError::Tier(TierError::NotFound(_)) => Status::not_found(message),
VolumeError::ReadOnly(_) | VolumeError::NotEmpty => {
Status::failed_precondition(message)
}
VolumeError::InsufficientSpace { .. } => Status::resource_exhausted(message),
VolumeError::AlreadyExists => Status::already_exists(message),
_ => Status::internal(message),
}
}
}
/// Same mapping, with the RPC's own context prefixed (`compact volume 7: ...`).
pub fn status_with_context(context: &str, err: VolumeError) -> Status {
let status = Status::from(err);
Status::new(status.code(), format!("{context}: {}", status.message()))
}
/// Render a configured disk directory as an absolute path for display, so the
/// status JSON and the UI show the same thing for a relative `-dir`. Falls
/// back to the configured spelling when the current directory cannot be read.
pub(crate) fn absolute_display_path(path: &str) -> String {
let p = std::path::Path::new(path);
if p.is_absolute() {
return path.to_string();
}
std::env::current_dir()
.map(|cwd| cwd.join(p).to_string_lossy().to_string())
.unwrap_or_else(|_| path.to_string())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::storage::types::VolumeId;
#[test]
fn test_volume_error_maps_to_grpc_code() {
use tonic::Code;
let code = |e: VolumeError| Status::from(e).code();
assert_eq!(
code(VolumeError::VolumeNotFound(VolumeId(7))),
Code::NotFound
);
assert_eq!(code(VolumeError::NotFound), Code::NotFound);
assert_eq!(
code(VolumeError::ReadOnly(VolumeId(7))),
Code::FailedPrecondition
);
assert_eq!(
VolumeError::ReadOnly(VolumeId(7)).to_string(),
"volume 7 is read only"
);
assert_eq!(
code(VolumeError::InsufficientSpace {
vid: VolumeId(7),
required: 2,
free: 1,
}),
Code::ResourceExhausted
);
assert_eq!(code(VolumeError::AlreadyExists), Code::AlreadyExists);
assert_eq!(code(VolumeError::NotInitialized), Code::Internal);
assert_eq!(
code(TierError::NotFound("gone".into()).into()),
Code::NotFound
);
for tier in [
TierError::Io("io".into()),
TierError::RuntimeUnavailable("rt".into()),
TierError::Aborted("bye".into()),
] {
assert_eq!(code(tier.into()), Code::Internal);
}
let status = status_with_context(
"backend s3.default copy file /data/1.dat",
TierError::NotFound("failed to head object k: NotFound".into()).into(),
);
assert_eq!(status.code(), Code::NotFound);
assert_eq!(
status.message(),
"backend s3.default copy file /data/1.dat: failed to head object k: NotFound"
);
let status = status_with_context(
"compact volume 7",
VolumeError::InsufficientSpace {
vid: VolumeId(7),
required: 2,
free: 1,
},
);
assert_eq!(status.code(), Code::ResourceExhausted);
assert_eq!(
status.message(),
"compact volume 7: not enough free space: required 2, free 1"
);
}
}
+3 -3
View File
@@ -29,11 +29,11 @@ impl<S> Layer<S> for GrpcRequestIdLayer {
impl<S, B> Service<http::Request<B>> for GrpcRequestIdService<S>
where
S: Service<http::Request<B>, Response = http::Response<tonic::body::BoxBody>> + Send + 'static,
S: Service<http::Request<B>, Response = http::Response<tonic::body::Body>> + Send + 'static,
S::Future: Send + 'static,
B: Send + 'static,
{
type Response = http::Response<tonic::body::BoxBody>;
type Response = http::Response<tonic::body::Body>;
type Error = S::Error;
type Future = Pin<Box<dyn Future<Output = Result<Self::Response, Self::Error>> + Send>>;
@@ -57,7 +57,7 @@ where
let future = self.inner.call(request);
Box::pin(async move {
let mut response: http::Response<tonic::body::BoxBody> =
let mut response: http::Response<tonic::body::Body> =
scope_request_id(request_id.clone(), future).await?;
if let Ok(value) = HeaderValue::from_str(&request_id) {
response.headers_mut().insert("x-amz-request-id", value);
File diff suppressed because it is too large Load Diff
+7 -17
View File
@@ -1,5 +1,6 @@
use std::fmt::Write as _;
use crate::server::absolute_display_path;
use crate::server::server_stats;
use crate::server::volume_server::VolumeServerState;
use crate::storage::store::Store;
@@ -423,13 +424,12 @@ fn collect_ui_data(
shard_id: shard.shard_id,
size: shard_size,
});
if created_at == "-" {
if let Ok(metadata) = std::fs::metadata(shard.file_name()) {
if let Ok(modified) = metadata.modified() {
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
if created_at == "-"
&& let Ok(metadata) = std::fs::metadata(shard.file_name())
&& let Ok(modified) = metadata.modified()
{
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
let preferred_size = ec_volume.dat_file_size.max(0) as u64;
@@ -451,16 +451,6 @@ fn collect_ui_data(
(disk_rows, volumes, remote_volumes, ec_volumes)
}
fn absolute_display_path(path: &str) -> String {
let p = std::path::Path::new(path);
if p.is_absolute() {
return path.to_string();
}
std::env::current_dir()
.map(|cwd| cwd.join(p).to_string_lossy().to_string())
.unwrap_or_else(|_| path.to_string())
}
fn join_i64(values: &[i64]) -> String {
values
.iter()
+39 -27
View File
@@ -14,12 +14,12 @@ use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32, Ordering};
use std::sync::{Arc, RwLock};
use axum::{
extract::{connect_info::ConnectInfo, Request, State},
http::{header, HeaderValue, Method, StatusCode},
Router,
extract::{Request, State, connect_info::ConnectInfo},
http::{HeaderValue, Method, StatusCode, header},
middleware::{self, Next},
response::{IntoResponse, Response},
routing::{any, get},
Router,
};
use crate::config::ReadMode;
@@ -73,8 +73,6 @@ pub struct VolumeServerState {
pub volume_state_notify: tokio::sync::Notify,
/// Optional batched write queue for improved throughput under load.
pub write_queue: std::sync::OnceLock<WriteQueue>,
/// Registry of S3 tier backends for tiered storage operations.
pub s3_tier_registry: std::sync::RwLock<crate::remote_storage::s3_tier::S3TierRegistry>,
/// Read mode: local, proxy, or redirect for non-local volumes.
pub read_mode: ReadMode,
/// If true, FetchAndWriteNeedle skips remote S3 endpoint validation,
@@ -116,6 +114,21 @@ pub struct VolumeServerState {
pub cli_white_list: Vec<String>,
/// Path to state.pb file for persisting VolumeServerState across restarts.
pub state_file_path: String,
/// Volumes with an EC decode in flight. A dropped request leaves the
/// blocking job running; this keeps a retry from racing it on the
/// same volume files.
pub ec_decodes_in_flight:
std::sync::Mutex<std::collections::HashSet<crate::storage::types::VolumeId>>,
/// Volumes whose EC decode is in its publishing tail (journal catch-up,
/// .idx write, compaction). Local .ecj appenders wait on
/// `ec_decode_tail_notify` while their vid is listed, so no committed
/// delete falls between the last catch_up and the .cpd/.cpx swap —
/// the per-volume slice of Go's EcVolume.ecjFileAccessLock.
pub ec_decode_tail:
std::sync::Mutex<std::collections::HashSet<crate::storage::types::VolumeId>>,
/// Wakes .ecj appenders waiting on `ec_decode_tail` when a decode's
/// publishing tail ends.
pub ec_decode_tail_notify: tokio::sync::Notify,
}
impl VolumeServerState {
@@ -200,9 +213,7 @@ pub fn to_http_address(addr: &str) -> std::borrow::Cow<'_, str> {
// rather than being silently rewritten. Mirrors the validation already
// done in `to_grpc_address` for the inverse direction.
if let (Ok(_), Ok(_)) = (http_port.parse::<u16>(), grpc_port.parse::<u16>()) {
return std::borrow::Cow::Owned(
addr[..ports_sep_index + 1 + dot_idx].to_string(),
);
return std::borrow::Cow::Owned(addr[..ports_sep_index + 1 + dot_idx].to_string());
}
}
std::borrow::Cow::Borrowed(addr)
@@ -312,16 +323,15 @@ async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Requ
)
}
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -358,16 +368,15 @@ async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Req
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -432,13 +441,13 @@ pub fn build_admin_router_with_ui(state: Arc<VolumeServerState>, ui_enabled: boo
.route("/healthz", get(handlers::healthz_handler))
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(admin_store_handler))
.route("/:path", any(admin_store_handler))
.route("/:vid/:fid", any(admin_store_handler))
.route("/:vid/:fid/:filename", any(admin_store_handler))
.route("/{path}", any(admin_store_handler))
.route("/{vid}/{fid}", any(admin_store_handler))
.route("/{vid}/{fid}/{filename}", any(admin_store_handler))
.fallback(admin_store_handler);
if ui_enabled {
// Note: /stats/* endpoints are commented out in Go's volume_server.go (L130-134).
@@ -455,13 +464,13 @@ pub fn build_public_router(state: Arc<VolumeServerState>) -> Router {
Router::new()
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(public_store_handler))
.route("/:path", any(public_store_handler))
.route("/:vid/:fid", any(public_store_handler))
.route("/:vid/:fid/:filename", any(public_store_handler))
.route("/{path}", any(public_store_handler))
.route("/{vid}/{fid}", any(public_store_handler))
.route("/{vid}/{fid}/{filename}", any(public_store_handler))
.fallback(public_store_handler)
.layer(middleware::from_fn(common_headers_middleware))
.with_state(state)
@@ -516,7 +525,10 @@ mod tests {
// "host:abc.def"), and silently rewriting it would just hide the bug.
assert_eq!(to_http_address("host:abc.def"), "host:abc.def");
assert_eq!(to_http_address("host:9333.notaport"), "host:9333.notaport");
assert_eq!(to_http_address("host:notaport.19333"), "host:notaport.19333");
assert_eq!(
to_http_address("host:notaport.19333"),
"host:notaport.19333"
);
// Out-of-range ports must not be silently truncated either.
assert_eq!(to_http_address("host:99999.19333"), "host:99999.19333");
}
+69 -8
View File
@@ -3,8 +3,8 @@
//! Instead of each upload handler directly calling `write_needle`, writes are
//! submitted to a queue. A background worker drains the queue in batches (up to
//! 128 entries), groups them by volume ID, and processes them together under a
//! single store lock. Requests that asked for `fsync` are flushed by
//! `write_needle` itself, one flush per durable write.
//! single store lock. Durable writes to a volume share their .dat and .idx
//! flushes (see `Volume::write_needles_grouped`).
use std::sync::Arc;
@@ -159,8 +159,12 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
let mut store = state.store.write().unwrap();
for (vid, entries) in groups {
for (mut needle, fsync, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle, fsync);
let (mut writes, senders): (Vec<_>, Vec<_>) = entries
.into_iter()
.map(|(needle, fsync, response_tx)| ((needle, fsync), response_tx))
.unzip();
let results = store.write_volume_needles(vid, &mut writes);
for (response_tx, result) in senders.into_iter().zip(results) {
// Send result back; ignore error if receiver dropped.
let _ = response_tx.send(result);
}
@@ -178,8 +182,8 @@ mod tests {
use crate::server::volume_server::RuntimeMetricsConfig;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::store::Store;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
use std::sync::RwLock;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
let store = Store::new(NeedleMapKind::InMemory);
let guard = Guard::new(&[], SigningKey(vec![]), 0, SigningKey(vec![]), 0);
@@ -207,9 +211,6 @@ mod tests {
pre_stop_seconds: 0,
volume_state_notify: tokio::sync::Notify::new(),
write_queue: std::sync::OnceLock::new(),
s3_tier_registry: std::sync::RwLock::new(
crate::remote_storage::s3_tier::S3TierRegistry::new(),
),
read_mode: crate::config::ReadMode::Local,
allow_untrusted_remote_endpoints: false,
master_url: String::new(),
@@ -228,6 +229,9 @@ mod tests {
security_file: String::new(),
cli_white_list: vec![],
state_file_path: String::new(),
ec_decodes_in_flight: std::sync::Mutex::new(std::collections::HashSet::new()),
ec_decode_tail: std::sync::Mutex::new(std::collections::HashSet::new()),
ec_decode_tail_notify: tokio::sync::Notify::new(),
})
}
@@ -315,6 +319,63 @@ mod tests {
}
}
/// The queue hands a volume's batch to the grouped path, so ten durable
/// writes cost one .dat sync and one .idx sync, not ten of each.
#[test]
fn test_process_batch_group_commits_fsync_writes() {
use crate::config::MinFreeSpace;
use crate::storage::types::{DiskType, NeedleId};
use crate::storage::volume::VolumeSpec;
let tmp = tempfile::TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let state = make_test_state();
{
let mut store = state.store.write().unwrap();
store
.add_location(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
store
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
}
let mut receivers = Vec::new();
let batch = (1..=10u64)
.map(|id| {
let (response_tx, response_rx) = oneshot::channel();
receivers.push(response_rx);
WriteRequest {
volume_id: VolumeId(1),
needle: Needle {
id: NeedleId(id),
cookie: 0x1111.into(),
data: vec![id as u8; 8],
data_size: 8,
..Needle::default()
},
fsync: true,
response_tx,
}
})
.collect();
process_batch(state.clone(), batch);
for mut rx in receivers {
assert!(matches!(rx.try_recv().unwrap(), Ok((_, _, false))));
}
let store = state.store.read().unwrap();
let (_, vol) = store.find_volume(VolumeId(1)).unwrap();
assert_eq!(vol.sync_counts_for_test(), (1, 1));
}
#[tokio::test]
async fn test_write_queue_dropped_sender() {
// When the queue is dropped, subsequent submits should fail gracefully.
+336 -197
View File
@@ -15,15 +15,17 @@ use tracing::warn;
use crate::config::MinFreeSpace;
use crate::storage::erasure_coding::ec_bitrot::remove_bitrot_sidecars;
use crate::storage::erasure_coding::ec_shard::{
EcVolumeShard, DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE, ERASURE_CODING_SMALL_BLOCK_SIZE,
EcVolumeShard, ShardId,
};
use crate::storage::erasure_coding::ec_volume::{
ECJ_COMPACT_TMP_EXT, EcVolume, is_usable_ecx_file,
};
use crate::storage::erasure_coding::ec_volume::EcVolume;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::super_block::{ReplicaPlacement, SUPER_BLOCK_SIZE};
use crate::storage::super_block::SUPER_BLOCK_SIZE;
use crate::storage::types::*;
use crate::storage::volume::{
remove_volume_files, volume_file_name, VifVolumeInfo, Volume, VolumeError,
VifVolumeInfo, Volume, VolumeError, VolumeSpec, remove_volume_files, volume_file_name,
};
/// A single disk location managing volumes in one directory.
@@ -131,10 +133,10 @@ impl DiskLocation {
for entry in entries {
let entry = entry?;
let name = entry.file_name().into_string().unwrap_or_default();
if let Some((collection, vid)) = parse_volume_filename(&name) {
if seen.insert((collection.clone(), vid)) {
dat_files.push((collection, vid));
}
if let Some((collection, vid)) = parse_volume_filename(&name)
&& seen.insert((collection.clone(), vid))
{
dat_files.push((collection, vid));
}
}
@@ -205,7 +207,6 @@ impl DiskLocation {
continue;
}
// Load existing data only; never create a phantom `.dat`. A lone
// `.vif`/`.idx` (e.g. an EC sidecar whose `.ecx` is on a sibling
// disk) would otherwise have Volume::new write an 8-byte stub that
@@ -280,30 +281,33 @@ impl DiskLocation {
let opened = Mutex::new(Vec::with_capacity(to_load.len()));
std::thread::scope(|scope| {
for _ in 0..workers {
scope.spawn(|| loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
match Volume::new(
&self.directory,
&self.idx_directory,
collection,
*vid,
needle_map_kind,
None, // replica placement read from superblock
None, // TTL read from superblock
0, // no preallocate on load
Version::current(),
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
scope.spawn(|| {
loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
// Replica placement and TTL are read back from the
// superblock, and a load never preallocates.
match Volume::new(
&self.directory,
&self.idx_directory,
*vid,
needle_map_kind,
&VolumeSpec {
collection,
..Default::default()
},
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
}
}
}
}
@@ -327,10 +331,10 @@ impl DiskLocation {
.strip_suffix(".cpc")
.or_else(|| name.strip_suffix(".cpd"))
.or_else(|| name.strip_suffix(".cpx"));
if let Some(stem) = stem {
if let Some(key) = parse_collection_volume_id(stem) {
pending.insert(key);
}
if let Some(stem) = stem
&& let Some(key) = parse_collection_volume_id(stem)
{
pending.insert(key);
}
}
}
@@ -374,8 +378,10 @@ impl DiskLocation {
let mut expected_shard_size: Option<i64> = None;
let dat_exists = match fs::metadata(&dat_path) {
Ok(meta) if meta.len() > SUPER_BLOCK_SIZE as u64 => {
expected_shard_size =
Some(calculate_expected_shard_size(meta.len() as i64, data_shards));
expected_shard_size = Some(calculate_expected_shard_size(
meta.len() as i64,
data_shards,
));
true
}
Ok(_) => false,
@@ -399,7 +405,13 @@ impl DiskLocation {
if size != prev {
// Inconsistent sizes signal corruption or mixed
// generations; not trusted for deletion -> keep.
warn!(volume_id = vid.0, shard = i, size, expected = prev, "EC shard size mismatch; keeping shards");
warn!(
volume_id = vid.0,
shard = i,
size,
expected = prev,
"EC shard size mismatch; keeping shards"
);
return true;
}
} else {
@@ -426,11 +438,16 @@ impl DiskLocation {
if shard_count == 0 {
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size) {
if actual < expected {
warn!(volume_id = vid.0, actual, expected, "shards smaller than the .dat's full encode; reclaiming the complete .dat");
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size)
&& actual < expected
{
warn!(
volume_id = vid.0,
actual,
expected,
"shards smaller than the .dat's full encode; reclaiming the complete .dat"
);
return false;
}
true
}
@@ -445,13 +462,16 @@ impl DiskLocation {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
const MAX_SHARD_COUNT: usize = 32;
// Remove index files from idx directory (.ecx, .ecj)
// Remove index files from idx directory (.ecx, .ecj, and a compaction
// tmp a crash may have left beside the .ecj)
rm_if_present(format!("{}.ecx", idx_base))?;
rm_if_present(format!("{}.ecj", idx_base))?;
rm_if_present(format!("{}{}", idx_base, ECJ_COMPACT_TMP_EXT))?;
// Also try data directory in case .ecx/.ecj were created before -dir.idx was configured
if self.idx_directory != self.directory {
rm_if_present(format!("{}.ecx", base))?;
rm_if_present(format!("{}.ecj", base))?;
rm_if_present(format!("{}{}", base, ECJ_COMPACT_TMP_EXT))?;
}
// Remove all EC shard files (.ec00 ~ .ec31)
@@ -464,6 +484,13 @@ impl DiskLocation {
if self.idx_directory != self.directory {
remove_bitrot_sidecars(&idx_base)?;
}
// Staged 2PC generations (<base>.ecNN.v<N>, versioned .ecx/.ecj/.vif)
// belong to this volume's EC state too; leaving them orphans the files.
crate::storage::erasure_coding::ec_shard::remove_ec_generation_files(&base, 0)?;
if self.idx_directory != self.directory {
crate::storage::erasure_coding::ec_shard::remove_ec_generation_files(&idx_base, 0)?;
}
Ok(())
}
@@ -510,10 +537,10 @@ impl DiskLocation {
pub(crate) fn ec_generation_ts_ns(&self, collection: &str, vid: VolumeId) -> Option<i64> {
for dir in [&self.directory, &self.idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Ok(s) = fs::read_to_string(&vif) {
if let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s) {
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if let Ok(s) = fs::read_to_string(&vif)
&& let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s)
{
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if self.directory == self.idx_directory {
break;
@@ -545,27 +572,19 @@ impl DiskLocation {
pub fn create_volume(
&mut self,
vid: VolumeId,
collection: &str,
needle_map_kind: NeedleMapKind,
replica_placement: Option<ReplicaPlacement>,
ttl: Option<crate::storage::needle::ttl::TTL>,
preallocate: u64,
version: Version,
spec: &VolumeSpec<'_>,
) -> Result<(), VolumeError> {
let mut v = Volume::new(
&self.directory,
&self.idx_directory,
collection,
vid,
needle_map_kind,
replica_placement,
ttl,
preallocate,
version,
spec,
)?;
v.location_disk_space_low = self.is_disk_space_low.clone();
crate::metrics::VOLUME_GAUGE
.with_label_values(&[collection, "volume"])
.with_label_values(&[spec.collection, "volume"])
.inc();
self.volumes.insert(vid, v);
Ok(())
@@ -591,13 +610,20 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
only_empty: bool,
only_garbage: bool,
keep_remote_data: bool,
) -> Result<(), VolumeError> {
// Refuse before removing: a refused destroy must leave it mounted.
if let Some(v) = self.volumes.get(&vid)
&& v.is_compacting()
{
return Err(v.compacting_error());
}
if let Some(mut v) = self.volumes.remove(&vid) {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&v.collection, "volume"])
.dec();
v.destroy(only_empty, keep_remote_data)?;
v.destroy(only_empty, only_garbage, keep_remote_data)?;
Ok(())
} else {
Err(VolumeError::NotFound)
@@ -618,7 +644,7 @@ impl DiskLocation {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&v.collection, "volume"])
.dec();
if let Err(e) = v.destroy(false, false) {
if let Err(e) = v.destroy(false, false, false) {
warn!(volume_id = vid.0, error = %e, "delete collection: failed to destroy volume");
}
}
@@ -669,8 +695,7 @@ impl DiskLocation {
pub fn free_volume_count(&self) -> i32 {
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let max = self.max_volume_count.load(Ordering::Relaxed);
let free_count = (max as i64 - self.volumes.len() as i64)
* DATA_SHARDS_COUNT as i64
let free_count = (max as i64 - self.volumes.len() as i64) * DATA_SHARDS_COUNT as i64
- self.ec_shard_count() as i64;
let effective_free = free_count / DATA_SHARDS_COUNT as i64;
if effective_free > 0 {
@@ -773,22 +798,18 @@ impl DiskLocation {
/// Mirrors `DiskLocation.HasEcxFileOnDisk` in
/// `weed/storage/disk_location_ec.go`. Skips entries that are
/// directories so a stray dir named `<collection>_<vid>.ecx` doesn't
/// register as a present index file.
/// register as a present index file. A 0-byte `.ecx` is a corrupt stub
/// left by a failed EC distribute copy; it must not steer placement
/// toward this disk, so it counts as absent (Go requires `Size() > 0`).
pub fn has_ecx_file_on_disk(&self, collection: &str, vid: VolumeId) -> bool {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
let idx_path = format!("{}.ecx", idx_base);
if let Ok(meta) = fs::metadata(&idx_path) {
if !meta.is_dir() {
return true;
}
if is_usable_ecx_file(&format!("{}.ecx", idx_base)) {
return true;
}
if self.idx_directory != self.directory {
let data_base = volume_file_name(&self.directory, collection, vid);
let data_path = format!("{}.ecx", data_base);
if let Ok(meta) = fs::metadata(&data_path) {
if !meta.is_dir() {
return true;
}
if is_usable_ecx_file(&format!("{}.ecx", data_base)) {
return true;
}
}
false
@@ -799,6 +820,20 @@ impl DiskLocation {
self.ec_volumes.remove(&vid)
}
/// Drop the in-memory EC volume for vid and close its descriptors without
/// deleting files, so a following unlink frees the inodes instead of
/// leaving open fds serving the old bytes. Mirrors Go's unloadEcVolume.
pub fn unload_ec_volume(&mut self, vid: VolumeId) {
if let Some(mut ec_vol) = self.ec_volumes.remove(&vid) {
for _ in 0..ec_vol.shard_count() {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&ec_vol.collection, "ec_shards"])
.dec();
}
ec_vol.close();
}
}
/// Mount EC shards for a volume on this location.
///
/// `source_disk_type` is the source volume's disk type carried on the
@@ -811,7 +846,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
source_disk_type: &str,
) -> Result<(), VolumeError> {
let idx_dir = self.idx_directory.clone();
@@ -833,7 +868,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
idx_dir: &str,
source_disk_type: &str,
) -> Result<(), VolumeError> {
@@ -845,14 +880,10 @@ impl DiskLocation {
// propagate the error to the caller.
let created = !self.ec_volumes.contains_key(&vid);
if created {
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid)
.map_err(VolumeError::Io)?;
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid).map_err(VolumeError::Io)?;
self.ec_volumes.insert(vid, ec_vol);
}
let ec_vol = self
.ec_volumes
.get_mut(&vid)
.expect("just inserted above");
let ec_vol = self.ec_volumes.get_mut(&vid).expect("just inserted above");
// When the orchestrator supplied a source disk type on the Mount
// RPC, override the EC volume's disk type so heartbeats report
// under the source volume's disk type (#9423). When the caller
@@ -871,10 +902,10 @@ impl DiskLocation {
// keep the existing registration (mirrors Go's AddEcVolumeShard
// added=false) — re-adding would replace a serving fd and bump
// the ec_shards gauge without growing the mounted count.
if ec_vol.has_shard(shard_id as u8) {
if ec_vol.has_shard(shard_id) {
continue;
}
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id as u8);
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id);
shard.disk_type = ec_vol.disk_type.clone();
if let Err(e) = ec_vol.add_shard(shard) {
// The shard was dropped (its descriptors closed) inside the
@@ -902,14 +933,14 @@ impl DiskLocation {
/// caller passes a shard that lives on a sibling disk
/// (cross-disk reconcile makes that the common case for the same
/// `vid` after reconciliation).
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[u32]) {
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[ShardId]) {
if let Some(ec_vol) = self.ec_volumes.get_mut(&vid) {
let collection = ec_vol.collection.clone();
for &shard_id in shard_ids {
if !ec_vol.has_shard(shard_id as u8) {
if !ec_vol.has_shard(shard_id) {
continue;
}
ec_vol.remove_shard(shard_id as u8);
let _ = ec_vol.remove_shard(shard_id);
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&collection, "ec_shards"])
.dec();
@@ -969,7 +1000,7 @@ impl DiskLocation {
}
entries.sort();
let mut same_volume_shards: Vec<(String, u32)> = Vec::new(); // (filename, shard_id)
let mut same_volume_shards: Vec<(String, ShardId)> = Vec::new(); // (filename, shard_id)
let mut prev_vid: Option<VolumeId> = None;
let mut prev_collection: String = String::new();
@@ -1034,7 +1065,12 @@ impl DiskLocation {
/// Validate + mount a (collection, vid) group when its `.ecx` is
/// found. Mirrors `handleFoundEcxFile` in
/// `weed/storage/disk_location_ec.go`.
fn handle_found_ecx_file(&mut self, shards: &[(String, u32)], collection: &str, vid: VolumeId) {
fn handle_found_ecx_file(
&mut self,
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) {
let base = volume_file_name(&self.directory, collection, vid);
let dat_path = format!("{}.dat", base);
let dat_exists = check_dat_file_exists(&dat_path);
@@ -1048,7 +1084,7 @@ impl DiskLocation {
return;
}
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if let Err(e) = self.mount_ec_shards(vid, collection, &shard_ids, "") {
// A mount failure (corrupt/locked .ecx, EMFILE, transient I/O) is
// not proof the shards are disposable -- validate_ec_volume already
@@ -1057,8 +1093,7 @@ impl DiskLocation {
// delete on a load error.
warn!(
volume_id = vid.0,
"Failed to load EC shards: {}; keeping files for retry",
e,
"Failed to load EC shards: {}; keeping files for retry", e,
);
self.unmount_ec_shards(vid, &shard_ids);
}
@@ -1071,7 +1106,7 @@ impl DiskLocation {
/// distributed-EC shards waiting for cross-disk reconciliation.
fn check_orphaned_shards(
&self,
shards: &[(String, u32)],
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) -> bool {
@@ -1107,7 +1142,7 @@ impl DiskLocation {
/// Close all volumes.
pub fn close(&mut self) {
for (_, v) in self.volumes.iter_mut() {
for v in self.volumes.values_mut() {
v.close();
}
self.volumes.clear();
@@ -1127,20 +1162,60 @@ pub fn get_disk_stats(path: &str) -> (u64, u64) {
Ok(p) => p,
Err(_) => return (0, 0),
};
unsafe {
let mut stat: libc::statvfs = std::mem::zeroed();
if libc::statvfs(c_path.as_ptr(), &mut stat) == 0 {
let all = stat.f_blocks as u64 * stat.f_frsize as u64;
let free = stat.f_bavail as u64 * stat.f_frsize as u64;
return (all, free);
}
// SAFETY: `libc::statvfs` is plain data — integers and reserved
// padding, no pointers and no restricted niches — so the all-zero
// value is a valid one for the call to overwrite.
let mut stat: libc::statvfs = unsafe { std::mem::zeroed() };
// SAFETY: `c_path` is a live NUL-terminated `CString` that outlives
// the call, and `&mut stat` is a live, aligned, exclusive pointer the
// kernel only writes through; the fields are read below only after
// the call reports success.
if unsafe { libc::statvfs(c_path.as_ptr(), &mut stat) } != 0 {
return (0, 0);
}
(0, 0)
let all = stat.f_blocks as u64 * stat.f_frsize as u64;
let free = stat.f_bavail as u64 * stat.f_frsize as u64;
(all, free)
}
#[cfg(not(unix))]
#[cfg(windows)]
{
let _ = path;
(0, 0)
use std::os::windows::ffi::OsStrExt;
// Canonicalize so symlinks, `.`/`..` segments, and relative paths
// resolve to the real location before querying. `\\?\`-prefixed
// extended-length paths and UNC (`\\?\UNC\...`) are passed through
// untouched: GetDiskFreeSpaceExW accepts them as-is.
let canonical = match std::fs::canonicalize(path) {
Ok(p) => p,
Err(_) => return (0, 0),
};
// UTF-16 with trailing NUL for the Win32 wide-string call.
let mut wide: Vec<u16> = canonical.as_os_str().encode_wide().collect();
// UNC directory names must end in a backslash for GetDiskFreeSpaceExW.
if !wide.ends_with(&[0x5C]) {
wide.push(0x5C);
}
wide.push(0);
// SAFETY: `wide` is NUL-terminated; the out-params are valid u64
// writes; the call has no other preconditions.
unsafe {
let mut free_available: u64 = 0;
let mut total: u64 = 0;
let ok = windows_sys::Win32::Storage::FileSystem::GetDiskFreeSpaceExW(
wide.as_ptr(),
&mut free_available,
&mut total,
std::ptr::null_mut(),
);
if ok == 0 {
return (0, 0);
}
return (total, free_available);
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("get_disk_stats is implemented for unix and windows only");
}
}
@@ -1176,7 +1251,12 @@ fn rm_if_present(path: String) -> io::Result<()> {
}
}
fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &str, vid: VolumeId) -> usize {
fn ec_data_shards_from_vif(
directory: &str,
idx_directory: &str,
collection: &str,
vid: VolumeId,
) -> usize {
for dir in [directory, idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Some(ds) = fs::read_to_string(&vif)
@@ -1184,10 +1264,9 @@ fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &st
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
.and_then(|vi| vi.ec_shard_config)
.map(|c| c.data_shards as usize)
&& ds > 0
{
if ds > 0 {
return ds;
}
return ds;
}
if directory == idx_directory {
break;
@@ -1223,7 +1302,7 @@ fn parse_collection_volume_id(base: &str) -> Option<(String, VolumeId)> {
/// `pub(crate)` re-export of [`parse_ec_shard_extension`] for the
/// cross-disk reconcile in `store_ec_reconcile.rs`.
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<ShardId> {
parse_ec_shard_extension(ext)
}
@@ -1237,7 +1316,7 @@ pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
/// shardId > 255` guard. The 3-digit form (`.ec100`–`.ec255`) is
/// retained so the parser can still recognise shards from custom
/// 32+ ratios that fit in a u8 even though OSS only ships 10+4.
fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
fn parse_ec_shard_extension(ext: &str) -> Option<ShardId> {
let rest = ext.strip_prefix(".ec")?;
if rest.len() < 2 || rest.len() > 3 {
return None;
@@ -1246,7 +1325,7 @@ fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
if id > 255 {
return None;
}
Some(id)
ShardId::try_from(id).ok()
}
/// Robust check that a `.dat` with actual data exists. An empty `.dat`
@@ -1308,7 +1387,10 @@ fn remove_empty_ec_dat_stub(volume_name: &str, idx_name: &str, vid: VolumeId) ->
return false;
}
warn!(volume_id = vid.0, "removing leftover empty .dat stub for EC volume");
warn!(
volume_id = vid.0,
"removing leftover empty .dat stub for EC volume"
);
let _ = fs::remove_file(&dat_path);
let _ = fs::remove_file(format!("{}.idx", idx_name));
true
@@ -1331,6 +1413,17 @@ mod tests {
use super::*;
use tempfile::TempDir;
/// get_disk_stats must report real capacity for a real path on every
/// platform (Windows included) — consumers treat total==0 as "unknown"
/// and leave available_space at 0, which breaks volume assignment.
#[test]
fn test_get_disk_stats_reports_capacity_for_real_path() {
let tmp = TempDir::new().unwrap();
let (total, free) = get_disk_stats(tmp.path().to_str().unwrap());
assert!(total > 0, "expected total>0, got {total}");
assert!(free > 0, "expected free>0, got {free}");
}
/// When `-dir.idx` is configured the EC `.vif` may live in the idx
/// directory; the sweep must look there too, not only the data dir.
#[test]
@@ -1353,7 +1446,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&vif).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&vif).unwrap(),
)
.unwrap();
assert!(
remove_empty_ec_dat_stub(&vbase, &ibase, VolumeId(42)),
@@ -1369,16 +1466,30 @@ mod tests {
fn test_validate_ec_volume_partial_dat_next_to_full_shards_keeps() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(70));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let full = calculate_expected_shard_size(30 * 1024 * 1024, ds);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(full as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(full as u64)
.unwrap();
}
// Partial .dat: bigger than a superblock so it is not swept as a stub,
// but smaller than what these shards encode.
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(5 * 1024 * 1024).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(5 * 1024 * 1024)
.unwrap();
assert!(
loc.validate_ec_volume("", VolumeId(70)),
"full-size shards beside a smaller (stale/partial) .dat must be kept",
@@ -1392,15 +1503,29 @@ mod tests {
fn test_validate_ec_volume_interrupted_encode_reclaims() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(71));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let dat_size = 30 * 1024 * 1024i64;
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(dat_size as u64).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(dat_size as u64)
.unwrap();
let partial = calculate_expected_shard_size(dat_size, ds) / 3;
assert!(partial > 0);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(partial as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(partial as u64)
.unwrap();
}
assert!(
!loc.validate_ec_volume("", VolumeId(71)),
@@ -1445,7 +1570,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
// A .vif with no EC config reads as generation 0 (recovered/pre-upgrade live volume).
@@ -1454,12 +1583,20 @@ mod tests {
version: 3,
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&no_cfg).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&no_cfg).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(0));
// idx-dir fallback: only the idx dir holds the .vif.
std::fs::remove_file(format!("{}.vif", dbase)).unwrap();
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
}
@@ -1499,16 +1636,8 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 1);
assert!(loc.find_volume(VolumeId(1)).is_some());
@@ -1532,24 +1661,15 @@ mod tests {
Vec::new(),
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(
VolumeId(2),
"test",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "test",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1592,12 +1712,11 @@ mod tests {
.unwrap();
loc.create_volume(
VolumeId(9),
"good",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "good",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1641,29 +1760,13 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(
VolumeId(2),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(VolumeId(2), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 2);
loc.delete_volume(VolumeId(1), false, false).unwrap();
loc.delete_volume(VolumeId(1), false, false, false).unwrap();
assert_eq!(loc.volumes_len(), 1);
assert!(loc.find_volume(VolumeId(1)).is_none());
}
@@ -1684,32 +1787,29 @@ mod tests {
loc.create_volume(
VolumeId(1),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(2),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(3),
"docs",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "docs",
..Default::default()
},
)
.unwrap();
assert_eq!(loc.volumes_len(), 3);
@@ -1719,6 +1819,34 @@ mod tests {
assert!(loc.find_volume(VolumeId(3)).is_some());
}
/// A 0-byte `.ecx` is the stub a failed EC distribute copy leaves behind.
/// Go's HasEcxFileOnDisk requires Size() > 0 so the stub cannot pin
/// placement to a disk that has no usable index.
#[test]
fn test_has_ecx_file_on_disk_ignores_zero_byte_stub() {
let tmp = TempDir::new().unwrap();
let data = tmp.path().join("data");
let idx = tmp.path().join("idx");
fs::create_dir_all(&data).unwrap();
fs::create_dir_all(&idx).unwrap();
let loc = DiskLocation::new(
data.to_str().unwrap(),
idx.to_str().unwrap(),
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
fs::write(idx.join("pics_7.ecx"), b"").unwrap();
assert!(!loc.has_ecx_file_on_disk("pics", VolumeId(7)));
// A real index in the data dir still counts, stub or no stub.
fs::write(data.join("pics_7.ecx"), [0u8; 16]).unwrap();
assert!(loc.has_ecx_file_on_disk("pics", VolumeId(7)));
}
#[test]
fn test_disk_location_delete_collection_removes_ec_volumes() {
let tmp = TempDir::new().unwrap();
@@ -1735,6 +1863,8 @@ mod tests {
let shard_path = format!("{}/pics_7.ec00", dir);
std::fs::write(&shard_path, b"ec-shard").unwrap();
// An EC volume needs its .ecx to mount.
std::fs::write(format!("{}/pics_7.ecx", dir), [0u8; 16]).unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "").unwrap();
assert!(loc.has_ec_volume(VolumeId(7)));
@@ -1772,7 +1902,9 @@ mod tests {
// mount_ec_shards with source_disk_type="ssd" — simulating the
// VolumeEcShardsMount RPC path.
std::fs::write(format!("{}/pics_7.ec00", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd").unwrap();
std::fs::write(format!("{}/pics_7.ecx", dir), [0u8; 16]).unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd")
.unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume mounted");
assert_eq!(
@@ -1789,7 +1921,9 @@ mod tests {
std::fs::write(format!("{}/pics_7.ec01", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[1], "").unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume still mounted");
let ec_vol = loc
.find_ec_volume(VolumeId(7))
.expect("ec volume still mounted");
assert_eq!(
ec_vol.disk_type,
DiskType::Ssd,
@@ -1868,10 +2002,12 @@ mod tests {
// A collection name unique to this test: the gauge is process-global
// and sibling tests running in parallel touch other labels.
std::fs::write(format!("{}/dupmount_11.ec00", dir), b"shard bytes").unwrap();
std::fs::write(format!("{}/dupmount_11.ecx", dir), [0u8; 16]).unwrap();
let gauge = crate::metrics::VOLUME_GAUGE.with_label_values(&["dupmount", "ec_shards"]);
let before = gauge.get();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "").unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.expect("a duplicate mount must succeed as a no-op");
@@ -1947,8 +2083,11 @@ mod tests {
let path = format!("{}/{}_{}.ec{:02}", dir, collection, vid.0, sid);
std::fs::write(&path, b"shard data nonempty").unwrap();
}
std::fs::write(format!("{}/{}_{}.ecx", dir, collection, vid.0), vec![0u8; 20])
.unwrap();
std::fs::write(
format!("{}/{}_{}.ecx", dir, collection, vid.0),
vec![0u8; 20],
)
.unwrap();
std::fs::write(format!("{}/{}_{}.ecj", dir, collection, vid.0), b"").unwrap();
std::fs::write(
format!("{}/{}_{}.vif", dir, collection, vid.0),
@@ -30,6 +30,7 @@ use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums, EcShardConfig,
};
use crate::storage::erasure_coding::ec_shard::MAX_SHARD_COUNT;
use crate::storage::io::read_exact_at;
use crate::storage::needle::crc::CRC;
/// Canonical extension for the checksum sidecar. Generation 0 (legacy/fresh
@@ -164,16 +165,16 @@ pub fn remove_bitrot_sidecars(base: &str) -> io::Result<()> {
};
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res {
if first_err.is_none() {
first_err = Some(e);
}
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
record(rm(format!("{}{}", base, BITROT_SIDECAR_EXT).into()));
let path = Path::new(base);
if let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) {
let prefix = format!("{}{}.v", fname.to_string_lossy(), BITROT_SIDECAR_EXT);
let prefix = format!("{}{}.v", fname.display(), BITROT_SIDECAR_EXT);
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries.flatten() {
@@ -203,7 +204,7 @@ pub fn new_encode_uuid() -> Vec<u8> {
/// Reports whether `block_size` is a power of two in [1 MiB, MAX_BITROT_BLOCK_SIZE].
pub fn is_pow2_multiple_of_1mib(block_size: u32) -> bool {
block_size >= (1 << 20) && block_size <= MAX_BITROT_BLOCK_SIZE && block_size.count_ones() == 1
((1 << 20)..=MAX_BITROT_BLOCK_SIZE).contains(&block_size) && block_size.count_ones() == 1
}
/// Returns ceil(covered_size / block_size).
@@ -402,7 +403,7 @@ pub fn validate_manifest(
total
));
}
let mut seen = vec![false; MAX_SHARD_COUNT];
let mut seen = [false; MAX_SHARD_COUNT];
for s in &prot.shards {
if s.shard_id >= total as u32 {
return Err(format!(
@@ -537,7 +538,7 @@ pub fn verify_shard_blocks(
break;
}
let to_read = to_read as usize;
read_full_at(f, &mut buf[..to_read], offset as u64)?;
read_exact_at(f, &mut buf[..to_read], offset as u64)?;
if CRC::new(&buf[..to_read]).0 != *want_crc {
mismatched.push(i);
}
@@ -546,33 +547,6 @@ pub fn verify_shard_blocks(
Ok(mismatched)
}
/// Reads exactly `buf.len()` bytes from `f` at `offset`, erroring on early EOF.
fn read_full_at(f: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
let mut total = 0usize;
while total < buf.len() {
#[cfg(unix)]
let n = {
use std::os::unix::fs::FileExt;
f.read_at(&mut buf[total..], offset + total as u64)?
};
#[cfg(not(unix))]
let n = {
use std::io::{Read, Seek, SeekFrom};
let mut fc = f.try_clone()?;
fc.seek(SeekFrom::Start(offset + total as u64))?;
fc.read(&mut buf[total..])?
};
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"short read on shard block",
));
}
total += n;
}
Ok(())
}
/// Builds the `EcShardConfig` proto for the given layout. The bitrot sidecar
/// carries its own top-level encode_uuid, so the nested config leaves it empty.
pub fn ec_shard_config(data_shards: u32, parity_shards: u32, block_size: i64) -> EcShardConfig {
@@ -627,7 +601,10 @@ mod tests {
save_bitrot_sidecar(path, &prot).unwrap();
let bytes = std::fs::read(path).unwrap();
let hex: String = bytes.iter().map(|b| format!("{:02x}", b)).collect();
assert_eq!(hex, CANONICAL_HEX, "Rust .ecsum bytes drifted from the Go canonical form");
assert_eq!(
hex, CANONICAL_HEX,
"Rust .ecsum bytes drifted from the Go canonical form"
);
let _ = std::fs::remove_file(path);
}
@@ -662,7 +639,11 @@ mod tests {
format!("{}.ecsum.v1", base),
format!("{}.ecsum.v7", base),
] {
assert!(!std::path::Path::new(&p).exists(), "{} should be removed", p);
assert!(
!std::path::Path::new(&p).exists(),
"{} should be removed",
p
);
}
assert!(std::path::Path::new(&keep_shard).exists());
assert!(std::path::Path::new(&keep_other_vid).exists());
@@ -681,7 +662,9 @@ mod tests {
assert!(!is_pow2_multiple_of_1mib(1 << 19)); // 512 KiB, too small
assert!(!is_pow2_multiple_of_1mib(3 << 20)); // 3 MiB, not pow2
assert!(!is_pow2_multiple_of_1mib(128 * 1024 * 1024)); // pow2 but > MAX_BITROT_BLOCK_SIZE
assert!(!is_pow2_multiple_of_1mib(DEFAULT_BITROT_BLOCK_SIZE as u32 + 1));
assert!(!is_pow2_multiple_of_1mib(
DEFAULT_BITROT_BLOCK_SIZE as u32 + 1
));
}
#[test]
@@ -735,12 +718,7 @@ mod tests {
#[test]
fn test_save_load_roundtrip() {
let tmp = tempfile::TempDir::new().unwrap();
let path = tmp
.path()
.join("vol.ecsum")
.to_str()
.unwrap()
.to_string();
let path = tmp.path().join("vol.ecsum").to_str().unwrap().to_string();
let mut builder = ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64);
builder.write(b"hello world");
@@ -901,8 +879,7 @@ mod tests {
assert_eq!(resolve_status(&notfound, 0, 10, 4), BitrotStatus::Off);
// Integrity failure => Invalid.
let bad: Result<EcBitrotProtection, BitrotLoadError> =
Err(BitrotLoadError::BadMagic(0));
let bad: Result<EcBitrotProtection, BitrotLoadError> = Err(BitrotLoadError::BadMagic(0));
assert_eq!(resolve_status(&bad, 0, 10, 4), BitrotStatus::Invalid);
// Generation mismatch => Off.
@@ -3,10 +3,12 @@
//! Rebuilds the original .dat + .idx files from data shards (.ec00-.ec09)
//! and the sorted index (.ecx) + deletion journal (.ecj).
use std::collections::HashSet;
use std::fs::File;
use std::io::{self, Read, Write};
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::erasure_coding::ec_volume::read_ecj_ids;
use crate::storage::idx;
use crate::storage::needle::needle::get_actual_size;
use crate::storage::super_block::SUPER_BLOCK_SIZE;
@@ -20,18 +22,21 @@ use crate::storage::volume::{fsync_dir, volume_file_name};
/// `dir` is used both for reading `.ec00` and `.ecx`. For split-disk
/// reconciled volumes call [`find_dat_file_size_with_dirs`] instead.
pub fn find_dat_file_size(dir: &str, collection: &str, volume_id: VolumeId) -> io::Result<i64> {
find_dat_file_size_with_dirs(dir, dir, collection, volume_id)
let deleted = read_ecj_deletions(&[dir], collection, volume_id)?;
find_dat_file_size_with_dirs(dir, dir, collection, volume_id, &deleted)
}
/// Like [`find_dat_file_size`] but lets the caller pass separate dirs
/// for `.ec00` (the data shard) and `.ecx` (the sealed index). This
/// is the form needed when shards are split across data dirs and the
/// `.ecx` lives on a sibling disk's idx dir (#9252).
/// `.ecx` lives on a sibling disk's idx dir (#9252). Needles in `deleted`
/// count as deleted.
pub fn find_dat_file_size_with_dirs(
ec00_dir: &str,
ecx_dir: &str,
collection: &str,
volume_id: VolumeId,
deleted: &HashSet<NeedleId>,
) -> io::Result<i64> {
let ec00_base = volume_file_name(ec00_dir, collection, volume_id);
let ecx_base = volume_file_name(ecx_dir, collection, volume_id);
@@ -53,9 +58,9 @@ pub fn find_dat_file_size_with_dirs(
for i in 0..entry_count {
let start = i * NEEDLE_MAP_ENTRY_SIZE;
let (_, offset, size) =
let (key, offset, size) =
idx_entry_from_bytes(&ecx_data[start..start + NEEDLE_MAP_ENTRY_SIZE]);
if size.is_deleted() {
if size.is_deleted() || deleted.contains(&key) {
continue;
}
let entry_stop = offset.to_actual_offset() + get_actual_size(size, version);
@@ -67,72 +72,163 @@ pub fn find_dat_file_size_with_dirs(
Ok(dat_size)
}
/// Reconstruct a .dat file from EC data shards.
///
/// Reads from .ec00-.ec09 and writes a new .dat file. All data shards
/// must live in `dir`. For the cross-disk reconciled layout where
/// shards are split across multiple data dirs of the same node, use
/// [`write_dat_file_from_shards_with_dirs`] instead.
pub fn write_dat_file_from_shards(
dir: &str,
/// Whether the `.ecx` in `ecx_dir` indexes a needle deleted neither there nor
/// in `deleted`.
pub fn has_live_needles(
ecx_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let dirs: Vec<String> = (0..data_shards).map(|_| dir.to_string()).collect();
write_dat_file_from_shards_with_dirs(
dir,
collection,
volume_id,
dat_file_size,
encoded_dat_file_size,
data_shards,
&dirs,
large_block_size,
small_block_size,
)
deleted: &HashSet<NeedleId>,
) -> io::Result<bool> {
let ecx_base = volume_file_name(ecx_dir, collection, volume_id);
let ecx_data = std::fs::read(format!("{}.ecx", ecx_base))?;
let (entries, _) = ecx_data.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>();
Ok(entries.iter().any(|entry| {
let (key, _, size) = idx_entry_from_bytes(entry);
!size.is_deleted() && !deleted.contains(&key)
}))
}
/// Reconstruct a .dat file from EC data shards, taking the source
/// directory for each shard separately.
///
/// `dat_dir` is where the produced `.dat` is written. `shard_dirs[i]`
/// is the directory holding shard `i`. For the simple "all shards in
/// one dir" case both can be the same value.
/// Distinct needle ids journaled in the `.ecj` of any of `dirs`. Go folds the
/// journal into the `.ecx` (RebuildEcxFile) before a decode; reading it leaves
/// the sealed index untouched. Only NotFound means "no journal".
pub fn read_ecj_deletions(
dirs: &[&str],
collection: &str,
volume_id: VolumeId,
) -> io::Result<HashSet<NeedleId>> {
Ok(EcjDeletions::read(dirs, collection, volume_id)?.ids)
}
/// The ids [`read_ecj_deletions`] returns, plus how far each journal was read
/// so that ids journaled later can be added.
pub struct EcjDeletions {
pub ids: HashSet<NeedleId>,
/// Each distinct journal path and the whole-record length read so far.
journals: Vec<(String, u64)>,
}
impl EcjDeletions {
pub fn read(dirs: &[&str], collection: &str, volume_id: VolumeId) -> io::Result<Self> {
let mut journals: Vec<(String, u64)> = Vec::new();
for dir in dirs {
let path = format!("{}.ecj", volume_file_name(dir, collection, volume_id));
if !journals.iter().any(|(p, _)| *p == path) {
journals.push((path, 0));
}
}
let mut deletions = EcjDeletions {
ids: HashSet::new(),
journals,
};
deletions.catch_up()?;
Ok(deletions)
}
/// Adds the ids appended to each journal since the last read. A journal
/// that shrank is read again from the start — and the whole set rebuilt,
/// since ids already folded in from the truncated tail may have been a
/// rolled-back append. Non-regular journals (the FIFOs the tests stand
/// in for a blocking disk) stat empty and cannot be rolled back, so they
/// never count as shrunk.
pub fn catch_up(&mut self) -> io::Result<()> {
let mut shrank = false;
for (path, read_to) in &self.journals {
if *read_to == 0 {
continue;
}
match std::fs::metadata(path) {
Ok(m) => {
if m.is_file() && m.len() < *read_to {
shrank = true;
break;
}
}
// A journal that was read before and is now gone shrank to
// nothing (e.g. the volume was destroyed mid-scan) — the ids
// read from it no longer reflect committed content.
Err(e) if e.kind() == io::ErrorKind::NotFound => {
shrank = true;
break;
}
Err(e) => return Err(e),
}
}
if shrank {
self.ids.clear();
for (_, read_to) in &mut self.journals {
*read_to = 0;
}
}
for (path, read_to) in &mut self.journals {
let file = match File::open(&*path) {
Ok(file) => file,
Err(e) if e.kind() == io::ErrorKind::NotFound => continue,
Err(e) => return Err(e),
};
let len = file.metadata()?.len();
if len < *read_to {
*read_to = 0;
}
read_ecj_ids(&file, *read_to, len, &mut self.ids)?;
*read_to = len - len % NEEDLE_ID_SIZE as u64;
}
Ok(())
}
/// Clears the set and re-reads every journal from the start. Run under
/// the caller's store read lock (no append in flight): the result is
/// then exactly the committed content — an earlier unlocked read may
/// have folded in bytes a rolled-back append later truncated, or missed
/// a record re-appended to the very offset a rollback freed.
pub fn rescan(&mut self) -> io::Result<()> {
self.ids.clear();
for (_, read_to) in &mut self.journals {
*read_to = 0;
}
self.catch_up()
}
}
/// What it takes to rebuild a volume's .dat from its EC data shards.
///
/// Mirrors Go's `WriteDatFile(baseFileName, datFileSize,
/// encodedDatFileSize, shardFileNames)` shape — Go passes per-shard
/// paths so a reconciled volume with shards split across disks of the
/// same volume server can still be decoded back to a regular .dat
/// (seaweedfs/seaweedfs#9252).
#[derive(Clone, Copy, Debug)]
pub struct DatRebuild<'a> {
/// Where the produced `.dat` is written.
pub dat_dir: &'a str,
pub collection: &'a str,
pub volume_id: VolumeId,
/// The number of bytes to write, i.e. the live data extent from
/// [`find_dat_file_size`].
pub dat_file_size: i64,
/// The .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Zero when the .vif does not
/// record the encode-time size; the layout is then inferred from the
/// shard size.
pub encoded_dat_file_size: i64,
pub data_shards: usize,
/// `shard_dirs[i]` is the directory holding shard `i`. `None` means every
/// data shard sits in `dat_dir`.
pub shard_dirs: Option<&'a [String]>,
/// The volume's shard block layout, e.g. `EcVolume::large_block_size()`
/// / `small_block_size()` from its .vif EC config.
pub large_block_size: usize,
pub small_block_size: usize,
}
/// Reconstruct a .dat file from EC data shards.
///
/// `dat_file_size` is the number of bytes to write, i.e. the live data
/// extent from [`find_dat_file_size`]. `encoded_dat_file_size` is the
/// .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Pass zero when the .vif does
/// not record the encode-time size to infer the layout from the shard
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
write_dat_file(
/// Reads from .ec00-.ec09 and writes a new .dat file, from one directory or
/// from the per-shard directories of a cross-disk reconciled volume.
pub fn write_dat_file_from_shards(spec: &DatRebuild<'_>) -> io::Result<()> {
let DatRebuild {
dat_dir,
collection,
volume_id,
@@ -142,21 +238,15 @@ pub fn write_dat_file_from_shards_with_dirs(
shard_dirs,
large_block_size,
small_block_size,
)
}
#[allow(clippy::too_many_arguments)]
fn write_dat_file(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
} = *spec;
let same_dir: Vec<String>;
let shard_dirs: &[String] = match shard_dirs {
Some(dirs) => dirs,
None => {
same_dir = vec![dat_dir.to_string(); data_shards];
&same_dir
}
};
if data_shards == 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
@@ -233,10 +323,10 @@ fn write_dat_file(
// Read large blocks
while encoded_remaining >= large_row_size && remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = large_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -255,10 +345,10 @@ fn write_dat_file(
// Read small blocks
while remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = small_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -298,56 +388,87 @@ fn write_dat_file(
write_result
}
/// Fails when the decoded `.dat` in `dat_dir` is shorter than the
/// `dat_file_size` bytes its EC index references: the caller deletes the
/// shards next, and they are the only other copy of the needles past the cut.
/// A longer file passes.
pub fn verify_decoded_dat_file(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
) -> io::Result<()> {
let dat_path = format!("{}.dat", volume_file_name(dat_dir, collection, volume_id));
let size = std::fs::metadata(&dat_path)?.len();
if (size as i64) < dat_file_size {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"decoded {} is {} bytes, short of the {} its ec index references",
dat_path, size, dat_file_size
),
));
}
Ok(())
}
/// Write .idx file from .ecx index + .ecj deletion journal.
///
/// Copies sorted .ecx entries to .idx, then appends tombstones for
/// deleted needles from .ecj.
/// See [`write_idx_file_from_ec_index_with_dirs`]; everything lives in `dir`.
pub fn write_idx_file_from_ec_index(
dir: &str,
collection: &str,
volume_id: VolumeId,
) -> io::Result<()> {
let base = volume_file_name(dir, collection, volume_id);
let ecx_path = format!("{}.ecx", base);
let ecj_path = format!("{}.ecj", base);
let idx_path = format!("{}.idx", base);
let deleted = read_ecj_deletions(&[dir], collection, volume_id)?;
let dat_file_size = find_dat_file_size_with_dirs(dir, dir, collection, volume_id, &deleted)?;
write_idx_file_from_ec_index_with_dirs(dir, dir, collection, volume_id, &deleted, dat_file_size)
}
/// Write the `.idx` for a `.dat` decoded to `dat_file_size` bytes, from the
/// `.ecx` in `ecx_dir`, into `idx_dir`.
///
/// Copies the `.ecx` rows, then appends one tombstone per row whose needle is
/// in `deleted`. A deleted needle at or past `dat_file_size` was cut from the
/// `.dat`, so its row is dropped: a row pointing past the end of the `.dat`
/// makes the volume load read-only.
pub fn write_idx_file_from_ec_index_with_dirs(
ecx_dir: &str,
idx_dir: &str,
collection: &str,
volume_id: VolumeId,
deleted: &HashSet<NeedleId>,
dat_file_size: i64,
) -> io::Result<()> {
let ecx_path = format!("{}.ecx", volume_file_name(ecx_dir, collection, volume_id));
let idx_path = format!("{}.idx", volume_file_name(idx_dir, collection, volume_id));
// Write to a temp file and atomically rename into place, so a crash
// mid-write never leaves a partial .idx at the final name beside the
// source shards.
let tmp_path = format!("{}.tmp", idx_path);
let write_result = (|| -> io::Result<()> {
// Copy .ecx to the temp .idx
std::fs::copy(&ecx_path, &tmp_path)?;
// Append deletions from .ecj as tombstones. Read the journal directly
// and treat only NotFound as "no journal": Path::exists would also
// swallow a permission/IO error and silently skip deletions, which
// would resurrect deleted needles as live.
let mut idx_file = std::fs::OpenOptions::new()
.write(true)
.append(true)
.open(&tmp_path)?;
match std::fs::read(&ecj_path) {
Ok(ecj_data) => {
let count = ecj_data.len() / NEEDLE_ID_SIZE;
for i in 0..count {
let start = i * NEEDLE_ID_SIZE;
let needle_id = NeedleId::from_bytes(&ecj_data[start..start + NEEDLE_ID_SIZE]);
idx::write_index_entry(
&mut idx_file,
needle_id,
Offset::default(),
TOMBSTONE_FILE_SIZE,
)?;
}
let mut ecx_file = File::open(&ecx_path)?;
let mut idx_file = io::BufWriter::new(File::create(&tmp_path)?);
let mut tombstoned = Vec::new();
idx::walk_index_file(&mut ecx_file, 0, |key, offset, size| {
let is_deleted = size.is_deleted() || deleted.contains(&key);
if is_deleted && offset.to_actual_offset() >= dat_file_size {
return Ok(());
}
Err(e) if e.kind() == io::ErrorKind::NotFound => {}
Err(e) => return Err(e),
idx::write_index_entry(&mut idx_file, key, offset, size)?;
if !size.is_deleted() && deleted.contains(&key) {
tombstoned.push(key);
}
Ok(())
})?;
for key in tombstoned {
idx::write_index_entry(&mut idx_file, key, Offset::default(), TOMBSTONE_FILE_SIZE)?;
}
// fsync, rename, then fsync the dir so the decoded .idx is durable and
// atomically published before the caller deletes the source shards.
let idx_file = idx_file.into_inner().map_err(|e| e.into_error())?;
idx_file.sync_all()?;
drop(idx_file);
// Windows rename does not replace an existing file on every version;
@@ -372,7 +493,7 @@ mod tests {
use crate::storage::erasure_coding::ec_encoder;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -384,13 +505,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -430,16 +547,17 @@ mod tests {
std::fs::remove_file(format!("{}/1.idx", dir)).unwrap();
// Reconstruct from EC shards
write_dat_file_from_shards(
dir,
"",
VolumeId(1),
original_dat_size as i64,
original_dat_size as i64,
write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(1),
dat_file_size: original_dat_size as i64,
encoded_dat_file_size: original_dat_size as i64,
data_shards,
block_size as usize,
block_size as usize,
)
shard_dirs: None,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
})
.unwrap();
write_idx_file_from_ec_index(dir, "", VolumeId(1)).unwrap();
@@ -459,13 +577,9 @@ mod tests {
let v2 = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -485,29 +599,29 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
// No shard files exist, so de-striping must fail and publish nothing:
// neither the final .dat nor a partial .dat.tmp may remain.
let res = write_dat_file_from_shards(
dir,
"",
VolumeId(7),
100,
100,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
);
let res = write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(7),
dat_file_size: 100,
encoded_dat_file_size: 100,
data_shards: 10,
shard_dirs: None,
large_block_size: ERASURE_CODING_LARGE_BLOCK_SIZE,
small_block_size: ERASURE_CODING_SMALL_BLOCK_SIZE,
});
assert!(res.is_err());
assert!(!std::path::Path::new(&format!("{}/7.dat", dir)).exists());
assert!(!std::path::Path::new(&format!("{}/7.dat.tmp", dir)).exists());
}
// Decoding when .vif does not record the encode-time size: the layout is
// inferred from the shard size, except when that is an exact large-block
// multiple and the live extent reaches the ambiguous region.
#[test]
fn test_write_dat_file_fallback_layout() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -545,11 +659,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -567,7 +683,17 @@ mod tests {
-> io::Result<Vec<u8>> {
let out = format!("{}/{}", dir, sub);
std::fs::create_dir_all(&out).unwrap();
write_dat_file(&out, "", VolumeId(1), live, encoded, 10, shard_dirs, LARGE, SMALL)?;
write_dat_file_from_shards(&DatRebuild {
dat_dir: &out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live,
encoded_dat_file_size: encoded,
data_shards: 10,
shard_dirs: Some(shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})?;
Ok(std::fs::read(format!("{}/1.dat", out)).unwrap())
};
@@ -581,14 +707,20 @@ mod tests {
// each shard exactly one large block, indistinguishable from one large row
let (dir, shard_dirs, _) = encode("ambig1", large_row_size - 1);
let err = decode_to(&dir, "out", large_row_size / 2, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
// two-row equivalent: decoding within the agreed prefix still works
let (dir, shard_dirs, original) = encode("ambig2", 2 * large_row_size - 1);
let decoded = decode_to(&dir, "outa", large_row_size, 0, &shard_dirs).unwrap();
assert_eq!(&original[..large_row_size as usize], &decoded[..]);
let err = decode_to(&dir, "outb", large_row_size + 1, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
}
// Decoding after deletions moved the live extent below the large-block row
@@ -597,7 +729,7 @@ mod tests {
#[test]
fn test_write_dat_file_after_tail_deletion() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -637,11 +769,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -657,17 +791,17 @@ mod tests {
std::fs::create_dir(&out_dir).unwrap();
let out = out_dir.to_str().unwrap();
let decode = |live_size: i64, encoded_size: i64| -> Vec<u8> {
write_dat_file(
out,
"",
VolumeId(1),
live_size,
encoded_size,
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live_size,
encoded_dat_file_size: encoded_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.unwrap();
let path = format!("{}/1.dat", out);
let decoded = std::fs::read(&path).unwrap();
@@ -702,17 +836,131 @@ mod tests {
assert_ne!(&original[..(large_row_size / 2) as usize], &control[..]);
// the live extent can never exceed the encode-time size
assert!(write_dat_file(
out,
"",
VolumeId(1),
dat_size + 1,
dat_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
.is_err());
assert!(
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: dat_size + 1,
encoded_dat_file_size: dat_size,
data_shards,
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.is_err()
);
}
/// A journal many chunks long that repeats a few ids reads back as those
/// ids, from each dir once, whatever its length.
#[test]
fn test_read_ecj_deletions_collects_distinct_ids_across_dirs() {
let tmp = TempDir::new().unwrap();
let data = tmp.path().join("data");
let idx = tmp.path().join("idx");
let missing = tmp.path().join("missing");
std::fs::create_dir_all(&data).unwrap();
std::fs::create_dir_all(&idx).unwrap();
let (data, idx, missing) = (
data.to_str().unwrap(),
idx.to_str().unwrap(),
missing.to_str().unwrap(),
);
let entry = |id: u64| {
let mut buf = [0u8; NEEDLE_ID_SIZE];
NeedleId(id).to_bytes(&mut buf);
buf
};
// Past two load chunks of three repeating ids, then an id only in the
// last chunk and a torn trailing record.
let mut ecj = Vec::new();
while ecj.len() <= 2 * (1 << 20) {
for id in [1, 2, 3] {
ecj.extend_from_slice(&entry(id));
}
}
ecj.extend_from_slice(&entry(7));
ecj.extend_from_slice(&entry(8)[..3]);
std::fs::write(format!("{idx}/1.ecj"), &ecj).unwrap();
std::fs::write(format!("{data}/1.ecj"), entry(9)).unwrap();
let ids = read_ecj_deletions(&[data, idx, idx, missing], "", VolumeId(1)).unwrap();
let expected: HashSet<NeedleId> = [1, 2, 3, 7, 9].into_iter().map(NeedleId).collect();
assert_eq!(ids, expected);
}
#[test]
fn test_verify_decoded_dat_file_rejects_a_short_dat() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let dat_path = format!("{dir}/1.dat");
let err = verify_decoded_dat_file(dir, "", VolumeId(1), 100).unwrap_err();
assert_eq!(err.kind(), io::ErrorKind::NotFound);
std::fs::write(&dat_path, vec![0u8; 99]).unwrap();
let err = verify_decoded_dat_file(dir, "", VolumeId(1), 100).unwrap_err();
assert!(err.to_string().contains("short of the 100"), "{err}");
std::fs::write(&dat_path, vec![0u8; 100]).unwrap();
verify_decoded_dat_file(dir, "", VolumeId(1), 100).unwrap();
std::fs::write(&dat_path, vec![0u8; 101]).unwrap();
verify_decoded_dat_file(dir, "", VolumeId(1), 100).unwrap();
}
/// Ids appended after the first read, including the rest of a record torn
/// at that point, are picked up; a journal that shrank is read again.
#[test]
fn test_ecj_deletions_catch_up_reads_appended_ids() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let ecj_path = format!("{dir}/1.ecj");
let entry = |id: u64| {
let mut buf = [0u8; NEEDLE_ID_SIZE];
NeedleId(id).to_bytes(&mut buf);
buf
};
let append = |bytes: &[u8]| {
let mut f = std::fs::OpenOptions::new()
.create(true)
.append(true)
.open(&ecj_path)
.unwrap();
f.write_all(bytes).unwrap();
};
let ids = |d: &EcjDeletions| {
let mut ids: Vec<u64> = d.ids.iter().map(|id| id.0).collect();
ids.sort();
ids
};
// No journal yet.
let mut deletions = EcjDeletions::read(&[dir, dir], "", VolumeId(1)).unwrap();
assert!(deletions.ids.is_empty());
append(&entry(1));
append(&entry(2)[..3]);
deletions.catch_up().unwrap();
assert_eq!(ids(&deletions), [1]);
append(&entry(2)[3..]);
append(&entry(3));
deletions.catch_up().unwrap();
assert_eq!(ids(&deletions), [1, 2, 3]);
// A shrunk journal is rebuilt from its surviving content: ids folded
// in from the truncated tail may have been rolled back and must not
// linger as phantom tombstones.
std::fs::write(&ecj_path, entry(9)).unwrap();
deletions.catch_up().unwrap();
assert_eq!(ids(&deletions), [9]);
// A journal removed since it was read is the extreme shrink: its
// earlier ids must go with it, not linger.
std::fs::remove_file(&ecj_path).unwrap();
deletions.catch_up().unwrap();
assert!(deletions.ids.is_empty());
}
}
@@ -5,16 +5,12 @@
use std::fs::File;
use std::io;
#[cfg(not(unix))]
use std::io::{Read, Seek, SeekFrom};
use reed_solomon_erasure::galois_8::ReedSolomon;
use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums,
};
use crate::pb::volume_server_pb::{ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums};
use crate::storage::erasure_coding::ec_bitrot::{
self, ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
self, DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::idx;
@@ -50,7 +46,7 @@ pub fn write_ec_files(
let dat_size = dat_file.metadata()?.len() as i64;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Create shard files
let total_shards = data_shards + parity_shards;
@@ -77,11 +73,13 @@ pub fn write_ec_files(
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
ENCODE_BUFFER_SIZE,
block_size as usize,
block_size as usize,
EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: ENCODE_BUFFER_SIZE,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
},
)?;
// Close all shards
@@ -162,7 +160,7 @@ pub fn rebuild_ec_files(
}
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
@@ -175,7 +173,7 @@ pub fn rebuild_ec_files(
let mut shard_size = 0;
for (i, shard) in shards.iter_mut().enumerate() {
if !missing_shard_ids.contains(&(i as u32)) {
if let Ok(_) = shard.open() {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
@@ -185,7 +183,7 @@ pub fn rebuild_ec_files(
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
if let Ok(_) = alt.open() {
if alt.open().is_ok() {
let size = alt.file_size();
if size > shard_size {
shard_size = size;
@@ -251,12 +249,8 @@ pub fn rebuild_ec_files(
}
// Reconstruct missing shards
rs.reconstruct(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon reconstruct: {:?}", e),
)
})?;
rs.reconstruct(&mut buffers)
.map_err(|e| io::Error::other(format!("reed-solomon reconstruct: {:?}", e)))?;
// Write recovered data into the missing shards
for i in missing_shard_ids {
@@ -296,7 +290,7 @@ pub fn verify_ec_shards(
parity_shards: usize,
) -> io::Result<(Vec<u32>, Vec<String>)> {
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<Option<EcVolumeShard>> = (0..total_shards)
@@ -378,27 +372,27 @@ pub fn verify_ec_shards(
if !read_failed {
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
if !is_valid {
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
if let Ok(is_valid) = rs.verify(&slice_ptrs)
&& !is_valid
{
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for i in data_shards..total_shards {
verify_buffers[i].fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for buf in &mut verify_buffers[data_shards..total_shards] {
buf.fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
}
}
@@ -431,22 +425,23 @@ pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::R
// Read all idx entries
let mut idx_file = File::open(idx_path)?;
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
let mut last: std::collections::HashMap<NeedleId, (Offset, Size)> =
std::collections::HashMap::new();
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
entries.push((key, offset, size));
last.insert(key, (offset, size));
Ok(())
})?;
// Sort by NeedleId, then by actual offset so later entries come last
entries.sort_by_key(|&(key, offset, _)| (key, offset.to_actual_offset()));
// Remove duplicates (keep last/latest entry for each key).
// dedup_by_key keeps the first in each run, so we reverse first,
// dedup, then reverse back.
entries.reverse();
entries.dedup_by_key(|entry| entry.0);
entries.reverse();
let mut entries: Vec<(NeedleId, Offset, Size)> = last
.into_iter()
.filter_map(|(key, (offset, size))| {
if size.is_deleted() || offset.is_zero() {
None
} else {
Some((key, offset, size))
}
})
.collect();
entries.sort_by_key(|&(key, _o, _s)| key);
// Write sorted entries to .ecx
let mut ecx_file = File::create(ecx_path)?;
@@ -490,7 +485,7 @@ pub fn rebuild_ecx_file(
.collect();
for (i, shard) in shards.iter_mut().enumerate() {
if let Err(_) = shard.open() {
if shard.open().is_err() {
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
@@ -507,7 +502,7 @@ pub fn rebuild_ecx_file(
}
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("cannot open data shard for ecx rebuild"),
"cannot open data shard for ecx rebuild".to_string(),
));
}
}
@@ -515,7 +510,7 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
let total_data_size = shard_size * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
@@ -538,7 +533,7 @@ pub fn rebuild_ecx_file(
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
(shard_size - 1).max(0)
};
// Read version from superblock (first byte of logical data)
@@ -587,7 +582,8 @@ pub fn rebuild_ecx_file(
}
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
let needle_id = NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let needle_id =
NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let size = Size::from_bytes(&header_buf[COOKIE_SIZE + NEEDLE_ID_SIZE..header_size]);
// Validate: stop if we hit zero cookie+id (end of data)
@@ -640,7 +636,6 @@ pub fn rebuild_ecx_file(
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
@@ -707,30 +702,50 @@ fn read_from_data_shards(
/// the uniform block is.
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// Shape of one encode run: the Reed-Solomon split and the block sizes that
/// fix where every byte of the .dat lands in the shards. Mirrors Go's
/// `ECContext`. `buffer_size` must divide both block sizes.
#[derive(Clone, Copy, Debug)]
pub(crate) struct EcEncodeLayout {
pub(crate) data_shards: usize,
pub(crate) parity_shards: usize,
/// Bytes of each shard's block handled per sub-batch; bounds memory at
/// `total_shards * buffer_size` however large the blocks are.
pub(crate) buffer_size: usize,
pub(crate) large_block_size: usize,
pub(crate) small_block_size: usize,
}
/// Encode the .dat file data into shard files.
///
/// Uses a two-phase approach matching Go's ec_encoder.go:
/// 1. Process as many large blocks as possible
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
dat_size: i64,
rs: &ReedSolomon,
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
buffer_size: usize,
large_block_size: usize,
small_block_size: usize,
layout: EcEncodeLayout,
) -> io::Result<()> {
let EcEncodeLayout {
data_shards,
parity_shards,
buffer_size,
large_block_size,
small_block_size,
} = layout;
let total_shards = data_shards + parity_shards;
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
.map(|_| vec![0u8; buffer_size])
.collect();
let mut buffers: Vec<Vec<u8>> = (0..total_shards).map(|_| vec![0u8; buffer_size]).collect();
let mut run = EncodeRun {
dat_file,
rs,
buffers: &mut buffers,
shards,
builders,
data_shards,
};
let mut remaining = dat_size;
let mut offset: u64 = 0;
@@ -739,16 +754,7 @@ pub(crate) fn encode_dat_file(
let large_row_size = large_block_size * data_shards;
while remaining >= large_row_size as i64 {
encode_data(
dat_file,
offset,
large_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, large_block_size)?;
offset += large_row_size as u64;
remaining -= large_row_size as i64;
}
@@ -758,16 +764,7 @@ pub(crate) fn encode_dat_file(
while remaining > 0 {
let to_process = remaining.min(small_row_size as i64);
encode_data(
dat_file,
offset,
small_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, small_block_size)?;
offset += to_process as u64;
remaining -= to_process;
}
@@ -775,108 +772,66 @@ pub(crate) fn encode_dat_file(
Ok(())
}
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
/// Everything one encode run streams through: the source .dat, the codec, a
/// buffer per shard, and the per-shard file and checksum sinks.
struct EncodeRun<'a> {
dat_file: &'a File,
rs: &'a ReedSolomon,
buffers: &'a mut [Vec<u8>],
shards: &'a mut [EcVolumeShard],
builders: &'a mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
encode_one_batch(
dat_file,
row_offset + (b * buffer_size) as u64,
block_size,
rs,
buffers,
shards,
builders,
data_shards,
)?;
}
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
impl EncodeRun<'_> {
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches
/// so arbitrarily large blocks never require block-sized allocations.
/// Mirrors Go's encodeData.
fn encode_row(&mut self, row_offset: u64, block_size: usize) -> io::Result<()> {
let buffer_size = self.buffers[0].len();
if !block_size.is_multiple_of(buffer_size) {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
)
})?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
for (i, buf) in buffers.iter().enumerate() {
shards[i].write_all(buf)?;
builders[i].write(buf);
}
Ok(())
}
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let mut n = 0;
while n < buf.len() {
#[cfg(unix)]
let r = {
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buf[n..], offset + n as u64)?
};
#[cfg(not(unix))]
let r = {
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(offset + n as u64))?;
f.read(&mut buf[n..])?
};
if r == 0 {
break;
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
self.encode_one_batch(row_offset + (b * buffer_size) as u64, block_size)?;
}
n += r;
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block
/// in this row. Mirrors Go's encodeDataOneBatch.
fn encode_one_batch(&mut self, offset: u64, block_size: usize) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the
// buffers are reused across batches, so the tail must be cleared
// explicitly.
for (i, buf) in self.buffers[..self.data_shards].iter_mut().enumerate() {
let read_offset = offset + (i * block_size) as u64;
let n = crate::storage::io::read_full_at(self.dat_file, buf, read_offset)?;
buf[n..].fill(0);
}
// Encode parity shards
self.rs
.encode(&mut *self.buffers)
.map_err(|e| io::Error::other(format!("reed-solomon encode: {:?}", e)))?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk
// length.
for (i, buf) in self.buffers.iter().enumerate() {
self.shards[i].write_all(buf)?;
self.builders[i].write(buf);
}
Ok(())
}
Ok(n)
}
#[cfg(test)]
@@ -884,7 +839,7 @@ mod tests {
use super::*;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -896,13 +851,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -947,13 +898,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=n {
@@ -1026,13 +973,9 @@ mod tests {
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=20 {
@@ -1064,7 +1007,10 @@ mod tests {
let victim = format!("{}/1.ec03", dir);
let full = std::fs::metadata(&victim).unwrap().len();
assert!(full > 0, "encoded shard should be non-empty");
let f = std::fs::OpenOptions::new().write(true).open(&victim).unwrap();
let f = std::fs::OpenOptions::new()
.write(true)
.open(&victim)
.unwrap();
f.set_len(full / 2).unwrap();
drop(f);
@@ -1218,19 +1164,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_uniform_layout() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1260,7 +1202,10 @@ mod tests {
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
let rebuilt = std::fs::read(&ecx_path).unwrap();
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
assert_eq!(
canonical, rebuilt,
"rebuilt .ecx must match the encode-time .ecx"
);
}
// A truncated data shard must FAIL the .ecx rebuild, not publish the
@@ -1268,19 +1213,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(3),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1378,13 +1319,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1462,13 +1399,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1500,13 +1433,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
vid,
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
@@ -1597,4 +1526,86 @@ mod tests {
details
);
}
#[test]
fn test_encode_drops_tombstone_last_wins() {
use crate::storage::idx;
use crate::storage::types::{NeedleId, Offset, Size, TOMBSTONE_FILE_SIZE};
let tmp = tempfile::TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let idx_path = format!("{}/t.idx", dir);
let ecx_path = format!("{}/t.ecx", dir);
let key = NeedleId(12345);
{
let mut f = std::fs::File::create(&idx_path).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
}
super::write_sorted_ecx_from_idx(&idx_path, &ecx_path).unwrap();
let mut found = false;
{
let mut f = std::fs::File::open(&ecx_path).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found = true;
}
Ok(())
})
.unwrap();
}
assert!(!found, "tombstoned key must not appear in .ecx");
let idx2 = format!("{}/t2.idx", dir);
let ecx2 = format!("{}/t2.ecx", dir);
{
let mut f = std::fs::File::create(&idx2).unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(2048), Size(200))
.unwrap();
}
super::write_sorted_ecx_from_idx(&idx2, &ecx2).unwrap();
let mut found2 = false;
{
let mut f = std::fs::File::open(&ecx2).unwrap();
idx::walk_index_file(&mut f, 0, |k, o, s| {
if k == key {
found2 = true;
assert_eq!(o.to_actual_offset(), 2048);
assert_eq!(s, Size(200));
}
Ok(())
})
.unwrap();
}
assert!(found2, "re-created key must appear live");
// Zero offset with non-negative size is also a deletion: Go
// readNeedleMap (`if !offset.IsZero() && !size.IsDeleted() { Set }
// else { Delete }`) and CompactNeedleMap::load_from_idx both treat
// it as deleted. Encode must drop it too, or the .ecx live-map
// mismatches replay.
let idx3 = format!("{}/t3.idx", dir);
let ecx3 = format!("{}/t3.ecx", dir);
{
let mut f = std::fs::File::create(&idx3).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), Size(0)).unwrap();
}
super::write_sorted_ecx_from_idx(&idx3, &ecx3).unwrap();
let mut found3 = false;
{
let mut f = std::fs::File::open(&ecx3).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found3 = true;
}
Ok(())
})
.unwrap();
}
assert!(
!found3,
"zero-offset row must not appear in .ecx even with non-negative size"
);
}
}
@@ -16,6 +16,20 @@ pub const ERASURE_CODING_SMALL_BLOCK_SIZE: usize = 1024 * 1024; // 1MB
pub type ShardId = u8;
/// Validate a wire shard id. `ShardId` is `u8` but only 0..MAX_SHARD_COUNT are valid.
/// Rejects 256 (would truncate to 0 and delete .ec00) and 270 (would alias 14).
pub fn shard_id_try_from(v: u32) -> Result<ShardId, String> {
if v < MAX_SHARD_COUNT as u32 {
Ok(v as ShardId)
} else {
Err(format!(
"invalid shard id {} (max {})",
v,
MAX_SHARD_COUNT - 1
))
}
}
/// A single erasure-coded shard file.
pub struct EcVolumeShard {
pub volume_id: VolumeId,
@@ -73,28 +87,14 @@ impl EcVolumeShard {
Ok(())
}
/// Read data at a specific offset.
/// Read data at a specific offset, filling `buf` unless the shard ends first.
pub fn read_at(&self, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let file = self
.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(not(unix))]
{
use std::io::{Read, Seek, SeekFrom};
// File::read_at is unix-only; fall back to seek + read.
// We need a mutable reference for seek/read, so clone the handle.
let mut f = file.try_clone()?;
f.seek(SeekFrom::Start(offset))?;
f.read(buf)
}
crate::storage::io::read_full_at(file, buf, offset)
}
/// Write data to the shard file (appends).
@@ -102,7 +102,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
file.write_all(data)?;
self.ecd_file_size += data.len() as i64;
Ok(())
@@ -123,7 +123,7 @@ impl EcVolumeShard {
pub fn try_clone_file(&self) -> io::Result<File> {
self.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?
.ok_or_else(|| io::Error::other("shard file not open"))?
.try_clone()
}
@@ -195,6 +195,104 @@ impl ShardBits {
}
}
/// Parses the generation of a 2PC-staged `<base>.v<N>` file: `None` means the
/// name is not a generation file of `base`.
pub fn ec_file_generation(name: &str, base: &str) -> Option<u32> {
let suffix = name.strip_prefix(&format!("{}.v", base))?;
match suffix.parse::<u32>() {
Ok(g) if g > 0 => Some(g),
_ => None,
}
}
/// Removes 2PC generation files staged under `base`:
/// `<base>.ecNN.v<N>`, `<base>.ecx.v<N>`, `<base>.ecj.v<N>`, `<base>.ecsum.v<N>`
/// and `<base>.vif.v<N>`. `generations_older_than == 0` removes every
/// generation; otherwise only generations strictly below it. Returns the
/// first real removal failure. Mirrors Go's `RemoveEcGenerationFiles`.
pub fn remove_ec_generation_files(base: &str, generations_older_than: u32) -> io::Result<()> {
let path = std::path::Path::new(base);
let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) else {
return Ok(());
};
let ec_prefix = format!("{}.ec", fname.to_string_lossy());
let vif_name = format!("{}.vif", fname.to_string_lossy());
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries {
let entry = match entry {
Ok(entry) => entry,
Err(e) => {
// A skipped entry means an incomplete sweep; report it
// instead of pretending the cleanup finished.
record(Err(e));
continue;
}
};
let name = entry.file_name().to_string_lossy().into_owned();
let Some((artifact, _)) = name.rsplit_once(".v") else {
continue;
};
if artifact != vif_name && !artifact.starts_with(&ec_prefix) {
continue;
}
let Some(generation) = ec_file_generation(&name, artifact) else {
continue;
};
if generations_older_than > 0 && generation >= generations_older_than {
continue;
}
record(match fs::remove_file(entry.path()) {
Err(e) if e.kind() != io::ErrorKind::NotFound => Err(e),
_ => Ok(()),
});
}
}
Err(e) if e.kind() != io::ErrorKind::NotFound => record(Err(e)),
Err(_) => {}
}
match first_err {
Some(e) => Err(e),
None => Ok(()),
}
}
/// Removes every staged generation `<shard_file>.v<N>` of one shard file.
/// Returns true when at least one generation file was removed.
pub fn remove_ec_shard_generations(shard_file: &str) -> io::Result<bool> {
let path = std::path::Path::new(shard_file);
let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) else {
return Ok(false);
};
let fname = fname.to_string_lossy().into_owned();
let mut removed = false;
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries {
let entry = entry?;
let name = entry.file_name().to_string_lossy().into_owned();
if ec_file_generation(&name, &fname).is_some() {
match fs::remove_file(entry.path()) {
Err(e) if e.kind() != io::ErrorKind::NotFound => return Err(e),
_ => removed = true,
}
}
}
}
Err(e) if e.kind() != io::ErrorKind::NotFound => return Err(e),
Err(_) => {}
}
Ok(removed)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -251,4 +349,42 @@ mod tests {
let shard = EcVolumeShard::new("/data", "", VolumeId(7), 13);
assert_eq!(shard.file_name(), "/data/7.ec13");
}
#[test]
fn test_shard_id_try_from_u32_rejects_overflow() {
use super::{MAX_SHARD_COUNT, shard_id_try_from};
assert_eq!(shard_id_try_from(0).unwrap(), 0u8);
assert_eq!(shard_id_try_from(14).unwrap(), 14u8);
assert_eq!(shard_id_try_from(31).unwrap(), 31u8);
assert!(shard_id_try_from(32).is_err());
assert!(shard_id_try_from(256).is_err());
assert!(shard_id_try_from(270).is_err());
assert!(shard_id_try_from(u32::MAX).is_err());
assert_eq!(MAX_SHARD_COUNT, 32);
}
#[test]
fn test_shard_batch_validation_is_atomic_rejects_without_partial_prefix() {
use super::shard_id_try_from;
// The mount/unmount handlers pre-validate the ENTIRE req.shard_ids into
// a Vec<ShardId> BEFORE acquiring the write lock or mutating any EC
// state. This test pins the validation half of that contract at the
// unit level: a batch like [0, 32] must fail as a whole, so by
// construction no validated prefix (e.g. shard 0) is ever applied.
// The handler-level tests below assert the no-state-change half.
let batch = vec![0u32, 32u32];
let validated: Result<Vec<_>, _> =
batch.iter().map(|&sid| shard_id_try_from(sid)).collect();
assert!(
validated.is_err(),
"batch {:?} must be rejected as a whole",
batch
);
// A fully-valid batch still validates cleanly.
let ok: Result<Vec<_>, _> = [0u32, 1u32, 13u32]
.iter()
.map(|&sid| shard_id_try_from(sid))
.collect();
assert_eq!(ok.unwrap(), vec![0u8, 1u8, 13u8]);
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,290 @@
//! Set-union merge of `.ecj` deletion journals for EC shard copy / index
//! recovery. Mirrors Go's `weed/storage/erasure_coding/ecj_merge.go`.
//!
//! An EC volume's deletion journal (`<vid>.ecj`) is a *set* of deleted needle
//! ids stored as 8-byte big-endian records. Shard copy and index recovery fold
//! a peer's journal into the local one; they must append only the ids the
//! local journal lacks, or every `ec_balance` round trip doubles the file.
//!
//! The journal is only ever appended to, never replaced: a mounted `EcVolume`
//! holds it open, and a rename would leave that handle writing to an unlinked
//! inode, losing every later delete at the next mount. A mounted volume merges
//! through [`EcVolume::merge_journal`](super::ec_volume::EcVolume::merge_journal);
//! [`append_ecj_ids`] is for a journal no volume has open.
use std::collections::HashSet;
use std::fs::{self, OpenOptions};
use std::io::{self, Read, Seek, SeekFrom, Write};
use std::path::Path;
use crate::storage::types::{NEEDLE_ID_SIZE, NeedleId};
use crate::storage::volume::fsync_dir;
use crate::storage::volume_open::open_volume_file;
/// Bytes per read when scanning a journal; a multiple of `NEEDLE_ID_SIZE`.
const ECJ_READ_CHUNK_BYTES: usize = 1 << 20;
/// Decodes `.ecj` records from a byte stream split at arbitrary boundaries (a
/// CopyFile stream, chunked reads), collecting the distinct ids. Memory
/// follows the number of distinct ids, not the journal's length. A trailing
/// partial record is never decoded.
#[derive(Default)]
pub(crate) struct EcjIdDecoder {
ids: HashSet<NeedleId>,
partial: [u8; NEEDLE_ID_SIZE],
pending: usize,
}
impl EcjIdDecoder {
pub(crate) fn push(&mut self, mut bytes: &[u8]) {
if self.pending > 0 {
let n = (NEEDLE_ID_SIZE - self.pending).min(bytes.len());
self.partial[self.pending..self.pending + n].copy_from_slice(&bytes[..n]);
self.pending += n;
bytes = &bytes[n..];
if self.pending < NEEDLE_ID_SIZE {
return;
}
self.ids.insert(NeedleId::from_bytes(&self.partial));
self.pending = 0;
}
let mut records = bytes.chunks_exact(NEEDLE_ID_SIZE);
for record in &mut records {
self.ids.insert(NeedleId::from_bytes(record));
}
let rest = records.remainder();
self.partial[..rest.len()].copy_from_slice(rest);
self.pending = rest.len();
}
pub(crate) fn into_ids(self) -> HashSet<NeedleId> {
self.ids
}
}
/// Read the distinct ids of the journal at `path` in bounded chunks. A missing
/// file reads as empty. Also returns the whole-record length read; a torn
/// trailing partial record is excluded from it.
pub(crate) fn read_ecj_ids(path: &str) -> io::Result<(HashSet<NeedleId>, u64)> {
let mut file = match fs::File::open(path) {
Ok(f) => f,
Err(e) if e.kind() == io::ErrorKind::NotFound => return Ok((HashSet::new(), 0)),
Err(e) => return Err(e),
};
let len = file.metadata()?.len();
let size = len - len % NEEDLE_ID_SIZE as u64;
let mut decoder = EcjIdDecoder::default();
let mut buf = vec![0u8; (ECJ_READ_CHUNK_BYTES as u64).min(size) as usize];
let mut off = 0u64;
while off < size {
let want = (buf.len() as u64).min(size - off) as usize;
file.read_exact(&mut buf[..want])?;
decoder.push(&buf[..want]);
off += want as u64;
}
Ok((decoder.into_ids(), size))
}
/// The ids of `incoming` that `has` does not report, ascending, so a merge
/// appends deterministic output.
pub(crate) fn ecj_delta(
incoming: &HashSet<NeedleId>,
has: impl Fn(&NeedleId) -> bool,
) -> Vec<NeedleId> {
let mut delta: Vec<NeedleId> = incoming.iter().copied().filter(|id| !has(id)).collect();
delta.sort_unstable();
delta
}
pub(crate) fn encode_ecj_ids(ids: &[NeedleId]) -> Vec<u8> {
let mut buf = vec![0u8; ids.len() * NEEDLE_ID_SIZE];
for (record, id) in buf.chunks_exact_mut(NEEDLE_ID_SIZE).zip(ids) {
id.to_bytes(record);
}
buf
}
/// Append to the journal at `path` the ids of `incoming` that `local` lacks,
/// in one write and one fsync, returning how many were added. `local` and
/// `size` come from [`read_ecj_ids`] on the same path: if the journal's
/// whole-record length is no longer `size`, returns `Ok(None)` so the caller
/// re-reads. A torn tail past `size` is truncated first so the new records
/// stay aligned.
pub(crate) fn append_ecj_ids(
path: &str,
local: &HashSet<NeedleId>,
incoming: &HashSet<NeedleId>,
size: u64,
) -> io::Result<Option<usize>> {
let delta = ecj_delta(incoming, |id| local.contains(id));
if delta.is_empty() {
return Ok(Some(0));
}
let created = !Path::new(path).exists();
let mut file = open_volume_file(OpenOptions::new().read(true).write(true).create(true), path)?;
let len = file.metadata()?.len();
if len - len % NEEDLE_ID_SIZE as u64 != size {
return Ok(None);
}
if len != size {
file.set_len(size)?;
}
let appended = file
.seek(SeekFrom::Start(size))
.and_then(|_| file.write_all(&encode_ecj_ids(&delta)))
.and_then(|_| file.sync_all());
if let Err(e) = appended {
let _ = file.set_len(size);
return Err(e);
}
if created {
fsync_dir(path)?;
}
Ok(Some(delta.len()))
}
#[cfg(test)]
mod tests {
use super::*;
fn ids(v: &[u64]) -> HashSet<NeedleId> {
v.iter().map(|&id| NeedleId(id)).collect()
}
fn bytes(v: &[u64]) -> Vec<u8> {
encode_ecj_ids(&v.iter().map(|&id| NeedleId(id)).collect::<Vec<_>>())
}
fn records(path: &str) -> Vec<u64> {
let data = fs::read(path).expect("read ecj");
assert_eq!(data.len() % NEEDLE_ID_SIZE, 0, "journal must stay aligned");
data.chunks_exact(NEEDLE_ID_SIZE)
.map(|c| NeedleId::from_bytes(c).0)
.collect()
}
/// Run the unmounted merge the way the server does: read, then append.
fn merge_file(path: &str, incoming: &HashSet<NeedleId>) -> usize {
let (local, size) = read_ecj_ids(path).expect("read");
append_ecj_ids(path, &local, incoming, size)
.expect("append")
.expect("journal unchanged")
}
#[test]
fn decoder_handles_records_split_across_chunks() {
let mut stream = bytes(&[1, 2, 3, 2, 0x0102030405060708]);
stream.extend_from_slice(&[9, 9, 9]);
for chunk in 1..=stream.len() {
let mut d = EcjIdDecoder::default();
for piece in stream.chunks(chunk) {
d.push(piece);
}
assert_eq!(
d.into_ids(),
ids(&[1, 2, 3, 0x0102030405060708]),
"chunk {}",
chunk
);
}
}
#[test]
fn read_ecj_ids_dedups_and_ignores_torn_tail() {
let dir = tempfile::tempdir().expect("tempdir");
let missing = dir.path().join("missing.ecj");
let (got, size) = read_ecj_ids(missing.to_str().unwrap()).expect("read");
assert!(got.is_empty());
assert_eq!(size, 0);
let torn = dir.path().join("torn.ecj");
let mut data = bytes(&[1, 2, 1]);
data.extend_from_slice(&[7, 7, 7]);
fs::write(&torn, data).unwrap();
let (got, size) = read_ecj_ids(torn.to_str().unwrap()).expect("read");
assert_eq!(got, ids(&[1, 2]));
assert_eq!(size, 3 * NEEDLE_ID_SIZE as u64);
// A bloated journal repeating a few ids across several read chunks
// keeps only the distinct ids.
let bloated = dir.path().join("bloated.ecj");
let data: Vec<u8> = (0..300_000u64).flat_map(|i| bytes(&[i % 3])).collect();
fs::write(&bloated, &data).unwrap();
let (got, size) = read_ecj_ids(bloated.to_str().unwrap()).expect("read");
assert_eq!(got, ids(&[0, 1, 2]));
assert_eq!(size, data.len() as u64);
}
#[test]
fn appends_only_missing_ids() {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join("vol.ecj");
let path = path.to_str().unwrap();
fs::write(path, bytes(&[1, 2, 3])).unwrap();
assert_eq!(merge_file(path, &ids(&[3, 4])), 1);
assert_eq!(records(path), vec![1, 2, 3, 4]);
}
#[test]
fn round_trip_stays_constant() {
// A->B->A->B 20 times: the old append path doubled the journal each trip.
let dir = tempfile::tempdir().expect("tempdir");
let a = dir.path().join("a.ecj");
let b = dir.path().join("b.ecj");
let (a, b) = (a.to_str().unwrap(), b.to_str().unwrap());
fs::write(a, bytes(&[1, 2])).unwrap();
fs::write(b, bytes(&[2, 3])).unwrap();
for i in 0..20 {
let (src, dst) = if i % 2 == 0 { (a, b) } else { (b, a) };
let (incoming, _) = read_ecj_ids(src).expect("read");
merge_file(dst, &incoming);
}
for path in [a, b] {
let mut got = records(path);
got.sort_unstable();
assert_eq!(got, vec![1, 2, 3]);
}
}
#[test]
fn nothing_new_leaves_journal_alone() {
let dir = tempfile::tempdir().expect("tempdir");
let missing = dir.path().join("missing.ecj");
assert_eq!(merge_file(missing.to_str().unwrap(), &ids(&[])), 0);
assert!(
!missing.exists(),
"an empty merge must not create a journal"
);
let path = dir.path().join("vol.ecj");
let path = path.to_str().unwrap();
fs::write(path, bytes(&[1, 2])).unwrap();
assert_eq!(merge_file(path, &ids(&[2, 1])), 0);
assert_eq!(records(path), vec![1, 2]);
}
#[test]
fn repairs_torn_tail() {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join("vol.ecj");
let path = path.to_str().unwrap();
let mut data = bytes(&[1, 2]);
data.extend_from_slice(&[9, 9, 9]);
fs::write(path, data).unwrap();
assert_eq!(merge_file(path, &ids(&[3])), 1);
assert_eq!(records(path), vec![1, 2, 3]);
}
#[test]
fn rejects_changed_journal() {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join("vol.ecj");
let path = path.to_str().unwrap();
fs::write(path, bytes(&[1])).unwrap();
let (local, size) = read_ecj_ids(path).expect("read");
fs::write(path, bytes(&[1, 5])).unwrap();
let outcome = append_ecj_ids(path, &local, &ids(&[2]), size).expect("append");
assert_eq!(outcome, None);
assert_eq!(records(path), vec![1, 5]);
}
}
@@ -0,0 +1,326 @@
//! Process-wide coordination of everything that touches one `.ecj` path.
//!
//! Mount-time compaction replaces a deletion journal with a new inode. That is
//! only safe while nothing else in this process can write the old one:
//!
//! - **Holders** are mounted `EcVolume`s with an append handle on the path. A
//! store holds one `EcVolume` per disk location, and a shard mount or a
//! cross-disk reconcile can point one disk's volume at another disk's
//! `.ecj`, so several holders of one path are normal. A holder that keeps
//! appending to a replaced inode acknowledges deletes that are gone at the
//! next mount.
//! - **Writers** append to or replace the path by name without holding it
//! open across calls: `ReceiveFile` of an EC `.ecj`, and the unmounted
//! append in `merge_ec_journal`, which `VolumeEcShardsCopy` and EC index
//! recovery funnel a peer's journal through. Bytes they write after the
//! compactor sized the journal would be dropped by the rename.
//!
//! Compaction therefore runs only while its caller is the sole holder and no
//! writer is active, and while it runs no holder may open the path and no
//! writer may start. Both wait instead; a compaction rewrites only the distinct
//! id set, so the wait is short.
//!
//! No writer active at the reservation is not enough: one that ran while the
//! holder loaded the journal, or after, and has finished may have rewritten it
//! in place to the same length (`ReceiveFile` truncates and refills), which
//! the inode-and-size re-check cannot see. So each writer bumps the path's
//! write generation as it starts, and a holder may compact only if no writer
//! was active when it registered and the generation has not moved since.
//!
//! Paths are keyed by their canonical parent directory, so two disk locations
//! that spell one directory differently still meet here.
use std::collections::HashMap;
use std::path::{Path, PathBuf};
use std::sync::{Condvar, LazyLock, Mutex, MutexGuard};
#[derive(Default)]
struct PathState {
holders: usize,
writers: usize,
compacting: bool,
/// Writers that have started on the path. Lives as long as the entry,
/// which a registered holder keeps.
write_gen: u64,
}
impl PathState {
fn idle(&self) -> bool {
self.holders == 0 && self.writers == 0 && !self.compacting
}
}
struct Registry {
paths: Mutex<HashMap<PathBuf, PathState>>,
changed: Condvar,
}
static REGISTRY: LazyLock<Registry> = LazyLock::new(|| Registry {
paths: Mutex::new(HashMap::new()),
changed: Condvar::new(),
});
fn lock() -> MutexGuard<'static, HashMap<PathBuf, PathState>> {
// The critical sections only adjust counters and cannot panic midway, so
// a poisoned lock still guards consistent state.
REGISTRY.paths.lock().unwrap_or_else(|e| e.into_inner())
}
/// Canonical key for `path`: its resolved parent directory joined with the
/// file name. The file itself may not exist yet (a copy creates it), so only
/// the directory is resolved.
fn key_for(path: &str) -> PathBuf {
let p = Path::new(path);
let (Some(parent), Some(name)) = (p.parent(), p.file_name()) else {
return std::path::absolute(p).unwrap_or_else(|_| p.to_path_buf());
};
let parent = if parent.as_os_str().is_empty() {
Path::new(".")
} else {
parent
};
let dir = std::fs::canonicalize(parent)
.or_else(|_| std::path::absolute(parent))
.unwrap_or_else(|_| parent.to_path_buf());
dir.join(name)
}
/// Block until no compaction is running on `key`, then apply `f` to its state.
fn update_when_not_compacting<R>(key: &Path, f: impl FnOnce(&mut PathState) -> R) -> R {
let mut paths = lock();
while paths.get(key).is_some_and(|s| s.compacting) {
paths = REGISTRY
.changed
.wait(paths)
.unwrap_or_else(|e| e.into_inner());
}
f(paths.entry(key.to_path_buf()).or_default())
}
fn release(key: &Path, f: impl FnOnce(&mut PathState)) {
let mut paths = lock();
if let Some(state) = paths.get_mut(key) {
f(state);
if state.idle() {
paths.remove(key);
}
}
drop(paths);
REGISTRY.changed.notify_all();
}
/// A mounted `EcVolume`'s registration as a holder of its `.ecj`. Taken before
/// the journal is opened and released when dropped.
pub(crate) struct EcjHold {
key: PathBuf,
/// The path's write generation when the hold was taken, and whether a
/// writer was active then. Taken before the journal is opened and loaded,
/// so they cover every write the load might have missed.
write_gen: u64,
writer_at_start: bool,
}
impl EcjHold {
/// Register as a holder of `ecj_path`, first waiting out any compaction in
/// progress so the handle opened afterwards is on the final inode.
pub(crate) fn acquire(ecj_path: &str) -> Self {
let key = key_for(ecj_path);
let (write_gen, writer_at_start) = update_when_not_compacting(&key, |s| {
s.holders += 1;
(s.write_gen, s.writers > 0)
});
EcjHold {
key,
write_gen,
writer_at_start,
}
}
/// Reserve the path for a compaction, or `None` when another holder or an
/// active writer could still reach the current inode, or when a writer
/// has run on the path since the hold was taken, so the journal may no
/// longer be what the holder loaded.
pub(crate) fn try_begin_compaction(&self) -> Option<EcjCompaction> {
let mut paths = lock();
let state = paths.get_mut(&self.key)?;
if state.holders != 1 || state.writers != 0 || state.compacting {
return None;
}
if self.writer_at_start || state.write_gen != self.write_gen {
return None;
}
state.compacting = true;
Some(EcjCompaction {
key: self.key.clone(),
})
}
}
impl Drop for EcjHold {
fn drop(&mut self) {
release(&self.key, |s| s.holders = s.holders.saturating_sub(1));
}
}
/// An exclusive reservation of a `.ecj` path for compaction. Holders and
/// writers wait until it is dropped.
pub(crate) struct EcjCompaction {
key: PathBuf,
}
impl Drop for EcjCompaction {
fn drop(&mut self) {
release(&self.key, |s| s.compacting = false);
}
}
/// An out-of-band writer (shard copy, index recovery, `ReceiveFile`) on a
/// `.ecj` path.
/// Compaction does not start while one is alive.
pub(crate) struct EcjWrite {
key: PathBuf,
}
impl Drop for EcjWrite {
fn drop(&mut self) {
release(&self.key, |s| s.writers = s.writers.saturating_sub(1));
}
}
/// Register as a writer of `ecj_path`, waiting out any compaction in progress.
/// Blocks; async callers use [`begin_ecj_write_async`].
pub(crate) fn begin_ecj_write(ecj_path: &str) -> EcjWrite {
let key = key_for(ecj_path);
update_when_not_compacting(&key, |s| {
s.writers += 1;
s.write_gen += 1;
});
EcjWrite { key }
}
/// [`begin_ecj_write`] for async handlers: the wait runs on the blocking pool
/// so a compaction in progress never stalls a runtime worker.
pub(crate) async fn begin_ecj_write_async(ecj_path: &str) -> EcjWrite {
let path = ecj_path.to_string();
match tokio::task::spawn_blocking(move || begin_ecj_write(&path)).await {
Ok(write) => write,
// Only a panic inside the registry lands here, and it leaves no count
// behind; registering inline is still correct, merely blocking.
Err(_) => begin_ecj_write(ecj_path),
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::sync::mpsc;
use std::time::Duration;
use tempfile::TempDir;
fn ecj(dir: &TempDir) -> String {
dir.path().join("1.ecj").to_str().unwrap().to_string()
}
#[test]
fn sole_holder_may_compact() {
let dir = TempDir::new().unwrap();
let hold = EcjHold::acquire(&ecj(&dir));
assert!(hold.try_begin_compaction().is_some());
}
#[test]
fn second_holder_blocks_compaction() {
let dir = TempDir::new().unwrap();
let a = EcjHold::acquire(&ecj(&dir));
let b = EcjHold::acquire(&ecj(&dir));
assert!(a.try_begin_compaction().is_none());
drop(b);
assert!(a.try_begin_compaction().is_some());
}
#[test]
fn active_writer_blocks_compaction() {
let dir = TempDir::new().unwrap();
let hold = EcjHold::acquire(&ecj(&dir));
let w = begin_ecj_write(&ecj(&dir));
assert!(hold.try_begin_compaction().is_none());
drop(w);
// This hold loaded before the write; a later one may compact.
drop(hold);
let hold = EcjHold::acquire(&ecj(&dir));
assert!(hold.try_begin_compaction().is_some());
}
/// A writer that ran after the hold was taken, or was already running
/// then, may have changed the journal the holder loaded, even though it
/// has finished by the time compaction asks.
#[test]
fn finished_writer_since_hold_blocks_compaction() {
let dir = TempDir::new().unwrap();
let path = ecj(&dir);
let hold = EcjHold::acquire(&path);
drop(begin_ecj_write(&path));
assert!(
hold.try_begin_compaction().is_none(),
"a writer that started after the hold",
);
drop(hold);
let w = begin_ecj_write(&path);
let hold = EcjHold::acquire(&path);
drop(w);
assert!(
hold.try_begin_compaction().is_none(),
"a writer active when the hold was taken",
);
drop(hold);
let hold = EcjHold::acquire(&path);
assert!(
hold.try_begin_compaction().is_some(),
"no writer since the hold",
);
}
#[test]
fn differently_spelled_paths_share_one_key() {
let dir = TempDir::new().unwrap();
std::fs::create_dir(dir.path().join("sub")).unwrap();
let plain = ecj(&dir);
let dotted = dir
.path()
.join("sub")
.join("..")
.join("1.ecj")
.to_str()
.unwrap()
.to_string();
let a = EcjHold::acquire(&plain);
let _b = EcjHold::acquire(&dotted);
assert!(a.try_begin_compaction().is_none());
}
#[test]
fn writer_waits_for_compaction_to_finish() {
let dir = TempDir::new().unwrap();
let path = ecj(&dir);
let hold = EcjHold::acquire(&path);
let compaction = hold.try_begin_compaction().unwrap();
let (tx, rx) = mpsc::channel();
let p = path.clone();
let t = std::thread::spawn(move || {
let _w = begin_ecj_write(&p);
tx.send(()).unwrap();
});
assert!(
rx.recv_timeout(Duration::from_millis(100)).is_err(),
"a writer must not start while a compaction holds the path",
);
drop(compaction);
rx.recv_timeout(Duration::from_secs(5))
.expect("writer must proceed once the compaction ends");
t.join().unwrap();
}
}
@@ -9,9 +9,11 @@ pub mod ec_encoder;
pub mod ec_locate;
pub mod ec_shard;
pub mod ec_volume;
pub mod ecj_merge;
pub(crate) mod ecj_registry;
pub use ec_shard::{
EcVolumeShard, ShardId, DATA_SHARDS_COUNT, MAX_SHARD_COUNT, MIN_TOTAL_DISKS,
PARITY_SHARDS_COUNT, TOTAL_SHARDS_COUNT,
DATA_SHARDS_COUNT, EcVolumeShard, MAX_SHARD_COUNT, MIN_TOTAL_DISKS, PARITY_SHARDS_COUNT,
ShardId, TOTAL_SHARDS_COUNT,
};
pub use ec_volume::EcVolume;
+136 -8
View File
@@ -21,12 +21,26 @@ where
let mut buf = vec![0u8; NEEDLE_MAP_ENTRY_SIZE * ROWS_TO_READ];
loop {
let count = match reader.read(&mut buf) {
Ok(0) => return Ok(()),
Ok(n) => n,
Err(ref e) if e.kind() == io::ErrorKind::UnexpectedEof => return Ok(()),
Err(e) => return Err(e),
};
// Fill the batch before decoding: `read` may return a count that is
// not a multiple of the entry size, and a split entry would misalign
// every later row. Go is immune: `ReadAt` fills or errors.
let mut count = 0;
let mut eof = false;
while count < buf.len() {
match reader.read(&mut buf[count..]) {
Ok(0) => {
eof = true;
break;
}
Ok(n) => count += n,
Err(ref e) if e.kind() == io::ErrorKind::Interrupted => continue,
Err(ref e) if e.kind() == io::ErrorKind::UnexpectedEof => {
eof = true;
break;
}
Err(e) => return Err(e),
}
}
let mut i = 0;
while i + NEEDLE_MAP_ENTRY_SIZE <= count {
@@ -34,6 +48,11 @@ where
f(key, offset, size)?;
i += NEEDLE_MAP_ENTRY_SIZE;
}
// A trailing partial entry at EOF is ignored, as Go does on `io.EOF`.
if eof {
return Ok(());
}
}
}
@@ -57,7 +76,7 @@ pub fn check_index_file<R: Read + Seek>(
errs.push(format!("walk index file: {}", e));
}
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3 .0.cmp(&b.3 .0)));
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3.0.cmp(&b.3.0)));
// Offset-0 logical tombstones (remote-tier deletes) occupy no physical extent,
// so they cannot overlap anything — exclude them from the overlap check. They
@@ -177,6 +196,111 @@ mod tests {
data
}
/// Reader that hands back at most `chunk` bytes per `read`. 7 is coprime
/// with the 17-byte entry size, so nearly every read ends mid-entry. With
/// `interrupts`, every other call fails with `ErrorKind::Interrupted`.
struct ShortReader {
inner: Cursor<Vec<u8>>,
chunk: usize,
interrupts: bool,
interrupt_next: bool,
}
impl ShortReader {
fn new(data: Vec<u8>, interrupts: bool) -> Self {
ShortReader {
inner: Cursor::new(data),
chunk: 7,
interrupts,
interrupt_next: false,
}
}
}
impl Read for ShortReader {
fn read(&mut self, buf: &mut [u8]) -> io::Result<usize> {
if self.interrupt_next {
self.interrupt_next = false;
return Err(io::Error::from(io::ErrorKind::Interrupted));
}
self.interrupt_next = self.interrupts;
let n = buf.len().min(self.chunk);
self.inner.read(&mut buf[..n])
}
}
impl Seek for ShortReader {
fn seek(&mut self, pos: SeekFrom) -> io::Result<u64> {
self.inner.seek(pos)
}
}
fn walk_all<R: Read + Seek>(reader: &mut R, start_from: u64) -> Vec<(NeedleId, i64, Size)> {
let mut collected = Vec::new();
walk_index_file(reader, start_from, |key, offset, size| {
collected.push((key, offset.to_actual_offset(), size));
Ok(())
})
.unwrap();
collected
}
/// More than one ROWS_TO_READ batch, so the walk crosses a buffer refill.
fn many_entries() -> Vec<(NeedleId, Offset, Size)> {
(0..(ROWS_TO_READ as u64 * 2 + 37))
.map(|i| {
(
NeedleId(i * 7 + 1),
Offset::from_actual_offset(i as i64 * 128),
Size(i as i32 + 1),
)
})
.collect()
}
#[test]
fn test_walk_index_file_short_reads_keep_alignment() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
assert_eq!(expected.len(), ROWS_TO_READ * 2 + 37);
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_walk_index_file_retries_interrupted_reads() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
let mut short = ShortReader::new(data, true);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_walk_index_file_short_reads_start_from() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
let start = ROWS_TO_READ as u64 + 5;
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, start), expected[start as usize..]);
}
#[test]
fn test_walk_index_file_ignores_trailing_partial_entry() {
// A torn final entry is dropped without an error, as Go does on io.EOF.
let entries = many_entries();
let mut data = idx_bytes(&entries);
data.extend_from_slice(&[0xAB; NEEDLE_MAP_ENTRY_SIZE - 1]);
let expected = walk_all(&mut Cursor::new(idx_bytes(&entries)), 0);
assert_eq!(walk_all(&mut Cursor::new(data.clone()), 0), expected);
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_check_index_file_clean() {
let data = idx_bytes(&[
@@ -213,7 +337,11 @@ mod tests {
let size = data.len() as i64;
let (count, errs) = check_index_file(&mut Cursor::new(data), size, Version(3));
assert_eq!(count, 2, "tombstone row is still counted: {:?}", errs);
assert!(errs.is_empty(), "offset-0 tombstone must not overlap: {:?}", errs);
assert!(
errs.is_empty(),
"offset-0 tombstone must not overlap: {:?}",
errs
);
}
#[test]
+208
View File
@@ -0,0 +1,208 @@
//! Positional file reads.
//!
//! Every read here is "these bytes at this offset", never "the next bytes".
//! The handles are shared — `.dat` and `.idx` descriptors are borrowed from
//! [`file_pool`](super::needle_map::file_pool), a mounted EC shard's handle is
//! duplicated into a scrub plan — so no caller may rely on a file position.
//!
//! On unix that is `pread(2)` through `std::os::unix::fs::FileExt`. On Windows
//! it is `seek_read`, which passes the offset through `OVERLAPPED`, so the read
//! itself is independent of the current cursor.
//!
//! What these helpers replace is `try_clone()` + `seek()` + `read()`. A
//! duplicated handle shares one kernel file offset with the original, so that
//! sequence is two syscalls against state another thread can move in between:
//! the seek positions the offset, a concurrent reader or an append moves it,
//! and the read returns bytes from somewhere else entirely. `seek_read` carries
//! its own offset in a single call, so there is no window.
//!
//! `seek_read` does still advance the cursor as a side effect — Windows updates
//! the file pointer even for an `OVERLAPPED` read — which nothing here relies
//! on. A caller that genuinely needs a private position must open the file
//! again rather than duplicate a handle; see `Volume::dat_scan_plan` in
//! [`storage::volume`](super::volume).
use std::fs::File;
use std::io;
/// Reads exactly `buf.len()` bytes from `file` starting at `offset`.
///
/// Fails with [`io::ErrorKind::UnexpectedEof`] if the file ends first.
pub(crate) fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)?;
}
#[cfg(windows)]
{
if read_full_at(file, buf, offset)? < buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
Ok(())
}
/// Reads up to `buf.len()` bytes from `file` starting at `offset`, returning
/// how many were read.
///
/// A short read — including `0` at or past end of file — is not an error; use
/// [`read_exact_at`] when the whole buffer must be filled.
pub(crate) fn read_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
file.seek_read(buf, offset)
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
}
/// Reads into `buf` at `offset` until it is full or the file ends, retrying
/// interrupted reads; returns how many bytes were read.
///
/// Unlike [`read_at`], a count below `buf.len()` always means end of file.
pub(crate) fn read_full_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
fill_at(|b, at| read_at(file, b, at), buf, offset)
}
fn fill_at(
mut read: impl FnMut(&mut [u8], u64) -> io::Result<usize>,
buf: &mut [u8],
offset: u64,
) -> io::Result<usize> {
let mut filled = 0;
while filled < buf.len() {
match read(&mut buf[filled..], offset + filled as u64) {
Ok(0) => break,
Ok(n) => filled += n,
Err(err) if err.kind() == io::ErrorKind::Interrupted => {}
Err(err) => return Err(err),
}
}
Ok(filled)
}
#[cfg(test)]
mod tests {
use super::{fill_at, read_at, read_exact_at, read_full_at};
use std::io::{ErrorKind, Write};
fn temp_file(bytes: &[u8]) -> tempfile::NamedTempFile {
let mut f = tempfile::NamedTempFile::new().expect("temp file");
f.write_all(bytes).expect("write");
f.flush().expect("flush");
f
}
#[test]
fn read_exact_at_fills_the_whole_buffer() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 10];
read_exact_at(f.as_file(), &mut buf, 0).expect("read");
assert_eq!(&buf, b"0123456789");
}
#[test]
fn read_exact_at_reads_from_the_offset() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 4];
read_exact_at(f.as_file(), &mut buf, 3).expect("read");
assert_eq!(&buf, b"3456");
// The helper is positional: a second read at a lower offset sees the
// bytes at that offset, not wherever the first read left a cursor.
let mut again = [0u8; 4];
read_exact_at(f.as_file(), &mut again, 1).expect("read");
assert_eq!(&again, b"1234");
}
#[test]
fn read_exact_at_short_file_is_unexpected_eof() {
let f = temp_file(b"0123");
let mut buf = [0u8; 8];
let err = read_exact_at(f.as_file(), &mut buf, 0).expect_err("short file");
assert_eq!(err.kind(), ErrorKind::UnexpectedEof);
}
#[test]
fn read_at_allows_a_short_read_at_eof() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 8];
let n = read_at(f.as_file(), &mut buf, 6).expect("read");
assert_eq!(n, 4);
assert_eq!(&buf[..n], b"6789");
// Entirely past the end is zero bytes, not an error.
let n = read_at(f.as_file(), &mut buf, 10).expect("read");
assert_eq!(n, 0);
}
/// A source that returns at most `chunk` bytes per call and fails with
/// `Interrupted` on its first call, like a network mount under a signal.
fn chunked(src: &[u8], chunk: usize) -> impl FnMut(&mut [u8], u64) -> std::io::Result<usize> {
let mut interrupted = false;
move |buf, at| {
if !interrupted {
interrupted = true;
return Err(ErrorKind::Interrupted.into());
}
let at = (at as usize).min(src.len());
let n = buf.len().min(chunk).min(src.len() - at);
buf[..n].copy_from_slice(&src[at..at + n]);
Ok(n)
}
}
#[test]
fn fill_at_fills_across_short_and_interrupted_reads() {
let src: Vec<u8> = (0..=255).collect();
let mut buf = [0u8; 100];
let n = fill_at(chunked(&src, 7), &mut buf, 50).expect("read");
assert_eq!(n, buf.len());
assert_eq!(&buf[..], &src[50..150]);
}
#[test]
fn fill_at_stops_at_end_of_source() {
let src: Vec<u8> = (0..=255).collect();
let mut buf = [0u8; 100];
let n = fill_at(chunked(&src, 7), &mut buf, 200).expect("read");
assert_eq!(n, 56);
assert_eq!(&buf[..n], &src[200..]);
}
#[test]
fn fill_at_propagates_other_errors() {
let mut buf = [0u8; 8];
let err = fill_at(|_, _| Err(ErrorKind::PermissionDenied.into()), &mut buf, 0)
.expect_err("error");
assert_eq!(err.kind(), ErrorKind::PermissionDenied);
}
#[test]
fn read_full_at_returns_the_short_count_only_at_eof() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 8];
assert_eq!(read_full_at(f.as_file(), &mut buf, 0).expect("read"), 8);
assert_eq!(&buf, b"01234567");
assert_eq!(read_full_at(f.as_file(), &mut buf, 6).expect("read"), 4);
assert_eq!(&buf[..4], b"6789");
assert_eq!(read_full_at(f.as_file(), &mut buf, 10).expect("read"), 0);
}
}
+354
View File
@@ -0,0 +1,354 @@
//! Consecutive storage-media error tracking shared by `Volume` and
//! `EcVolume`. Mirrors Go's `weed/storage/io_error.go`.
use std::io;
use std::sync::Mutex;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
/// Consecutive storage-media errors allowed before the volume is quarantined.
pub(crate) const IO_ERROR_TOLERANCE: i32 = 3;
/// Returns true for I/O errors that indicate faulty storage media, not
/// transient/network failures. On Unix this is EIO; on Windows it covers
/// ERROR_CRC and ERROR_IO_DEVICE, which the kernel returns for failing disks.
pub(crate) fn is_storage_io_error(e: &io::Error) -> bool {
#[cfg(unix)]
{
e.raw_os_error() == Some(libc::EIO)
}
#[cfg(windows)]
{
const ERROR_CRC: i32 = 23;
const ERROR_IO_DEVICE: i32 = 1117;
return e.raw_os_error() == Some(ERROR_CRC) || e.raw_os_error() == Some(ERROR_IO_DEVICE);
}
#[cfg(not(any(unix, windows)))]
{
false
}
}
/// Consecutive storage-media error state for one volume. `quarantined` is
/// sticky: once set it survives later successful I/O and is lifted only by
/// `reset_io_error_state`.
#[derive(Default)]
pub(crate) struct IoErrorTracker {
last: Mutex<Option<String>>,
/// The consecutive error count in the low 32 bits and, in the high 32,
/// how many times it has been cleared. They share one word so that
/// `record_success_at` updates both in one step: reads record their
/// outcomes here without the volume's write lock.
streak: AtomicU64,
quarantined: AtomicBool,
}
const STREAK_COUNT_BITS: u64 = 0xffff_ffff;
fn streak_count(streak: u64) -> i32 {
(streak & STREAK_COUNT_BITS) as i32
}
/// `streak` with its count cleared and one more clear on record.
fn streak_cleared(streak: u64) -> u64 {
(streak >> 32).wrapping_add(1) << 32
}
/// A point in the error streak, taken where a write landed whose success
/// is only recorded later. See `IoErrorTracker::record_success_at`.
#[derive(Clone, Copy)]
pub(crate) struct StreakMark(u64);
impl IoErrorTracker {
/// `Some(e)` records a failure, `None` a success. Only storage-media
/// failures count; every other outcome clears the count and last error.
pub(crate) fn check_read_write_error(&self, err: Option<&io::Error>) {
if let Some(e) = err
&& is_storage_io_error(e)
{
self.streak.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
}
self.clear_count();
self.clear_last();
}
fn clear_count(&self) {
self.update_streak(|streak| Some(streak_cleared(streak)));
}
fn clear_last(&self) {
if let Ok(mut guard) = self.last.lock()
&& guard.is_some()
{
*guard = None;
}
}
/// Apply `f` to the streak atomically; `None` leaves it as it is.
/// Returns the streak `f` produced, if any.
fn update_streak(&self, mut f: impl FnMut(u64) -> Option<u64>) -> Option<u64> {
let mut updated = None;
let _ = self
.streak
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |streak| {
updated = f(streak);
updated
});
updated
}
pub(crate) fn mark(&self) -> StreakMark {
StreakMark(self.streak.load(Ordering::Relaxed))
}
/// Record a success as if it had come at `mark`: the errors counted
/// before the mark are cleared and the ones counted since still stand,
/// as they would had each outcome been recorded in order. A streak
/// cleared since the mark is left as it is.
pub(crate) fn record_success_at(&self, mark: StreakMark) {
let before = streak_count(mark.0);
let updated = self.update_streak(|streak| {
if streak >> 32 != mark.0 >> 32 {
return None;
}
if streak_count(streak) <= before {
return Some(streak_cleared(streak));
}
Some(streak - before as u64)
});
// The last error stays when one counted since the mark is left.
if updated.is_some_and(|streak| streak_count(streak) == 0) {
self.clear_last();
}
}
/// The last recorded error, the consecutive count, and the quarantine flag.
pub(crate) fn get_io_error_state(&self) -> (Option<String>, i32, bool) {
let err = self.last.lock().ok().and_then(|g| g.clone());
let count = self.count();
let quarantined = self.quarantined.load(Ordering::Relaxed);
(err, count, quarantined)
}
fn count(&self) -> i32 {
streak_count(self.streak.load(Ordering::Relaxed))
}
pub(crate) fn should_quarantine(&self) -> bool {
self.quarantined.load(Ordering::Relaxed) || self.count() >= IO_ERROR_TOLERANCE
}
pub(crate) fn mark_io_quarantined(&self) {
self.quarantined.store(true, Ordering::Relaxed);
}
pub(crate) fn reset_io_error_state(&self) {
self.clear_count();
self.quarantined.store(false, Ordering::Relaxed);
if let Ok(mut guard) = self.last.lock() {
*guard = None;
}
}
#[cfg(test)]
pub(crate) fn set_last_io_error_for_test(&self, err: Option<&str>) {
if let Ok(mut guard) = self.last.lock() {
*guard = err.map(|value| value.to_string());
}
if err.is_some() {
self.update_streak(|streak| {
Some((streak & !STREAK_COUNT_BITS) | IO_ERROR_TOLERANCE as u64)
});
} else {
self.clear_count();
}
}
}
// The tracker only reacts to errors `is_storage_io_error` recognises, which is
// nothing at all on a platform that is neither Unix nor Windows.
#[cfg(all(test, any(unix, windows)))]
mod tests {
use super::*;
/// An OS error the platform reports for failing storage media.
#[cfg(unix)]
fn media_error() -> io::Error {
io::Error::from_raw_os_error(libc::EIO)
}
/// An OS error the platform reports for failing storage media.
#[cfg(windows)]
fn media_error() -> io::Error {
const ERROR_IO_DEVICE: i32 = 1117;
io::Error::from_raw_os_error(ERROR_IO_DEVICE)
}
#[test]
fn check_read_write_error_counts_consecutive_media_errors() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let (last, count, quarantined) = tracker.get_io_error_state();
assert_eq!(last, Some(media_error().to_string()));
assert_eq!(count, 2);
assert!(!quarantined);
}
#[test]
fn success_clears_the_count_and_the_last_error() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn non_media_error_clears_the_count() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&io::Error::new(
io::ErrorKind::NotFound,
"no such file",
)));
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn should_quarantine_only_once_the_tolerance_is_reached() {
let tracker = IoErrorTracker::default();
for _ in 1..IO_ERROR_TOLERANCE {
tracker.check_read_write_error(Some(&media_error()));
assert!(!tracker.should_quarantine());
}
tracker.check_read_write_error(Some(&media_error()));
assert!(tracker.should_quarantine());
}
#[test]
fn quarantine_survives_later_successful_io() {
let tracker = IoErrorTracker::default();
tracker.mark_io_quarantined();
tracker.check_read_write_error(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, true));
assert!(tracker.should_quarantine());
}
#[test]
fn success_at_a_mark_keeps_only_the_errors_after_it() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.check_read_write_error(Some(&media_error()));
tracker.record_success_at(mark);
assert_eq!(
tracker.get_io_error_state(),
(Some(media_error().to_string()), 1, false)
);
}
#[test]
fn success_at_a_mark_with_nothing_after_it_clears_the_streak() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.record_success_at(mark);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn success_at_a_mark_leaves_a_streak_cleared_since() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.check_read_write_error(None);
tracker.check_read_write_error(Some(&media_error()));
tracker.record_success_at(mark);
assert_eq!(
tracker.get_io_error_state(),
(Some(media_error().to_string()), 1, false)
);
}
/// Reads update the tracker without the volume's write lock, so a
/// success replayed at a mark must not lose the errors they record
/// while it runs.
#[test]
fn success_at_a_mark_keeps_concurrent_errors() {
use std::sync::{Arc, Barrier};
const READERS: i32 = 4;
const ERRORS: i32 = 200;
for _ in 0..500 {
let tracker = Arc::new(IoErrorTracker::default());
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
let start = Arc::new(Barrier::new(READERS as usize + 1));
let readers: Vec<_> = (0..READERS)
.map(|_| {
let (tracker, start) = (tracker.clone(), start.clone());
std::thread::spawn(move || {
start.wait();
for _ in 0..ERRORS {
tracker.check_read_write_error(Some(&media_error()));
}
})
})
.collect();
start.wait();
while tracker.get_io_error_state().1 < 2 + READERS * ERRORS / 2 {
std::hint::spin_loop();
}
tracker.record_success_at(mark);
for reader in readers {
reader.join().unwrap();
}
// The two errors before the mark are cleared; every error the
// readers recorded after it stands.
assert_eq!(tracker.get_io_error_state().1, READERS * ERRORS);
}
}
#[test]
fn reset_io_error_state_lifts_the_quarantine() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.mark_io_quarantined();
tracker.reset_io_error_state();
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
assert!(!tracker.should_quarantine());
}
#[test]
fn test_helper_arms_a_sustained_error() {
let tracker = IoErrorTracker::default();
tracker.set_last_io_error_for_test(Some("input/output error"));
assert!(tracker.should_quarantine());
assert_eq!(
tracker.get_io_error_state(),
(
Some("input/output error".to_string()),
IO_ERROR_TOLERANCE,
false
)
);
tracker.set_last_io_error_for_test(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
}
+3
View File
@@ -1,9 +1,12 @@
pub mod disk_location;
pub mod erasure_coding;
pub mod idx;
pub(crate) mod io;
pub(crate) mod io_error;
pub mod needle;
pub mod needle_map;
pub mod store;
pub mod store_ec_journal;
pub mod store_ec_mirror;
pub mod store_ec_reconcile;
pub mod super_block;
+4 -2
View File
@@ -21,7 +21,7 @@ impl CRC {
/// Legacy `.Value()` function — deprecated in Go but needed for backward compat check.
/// Formula: (crc >> 15 | crc << 17) + 0xa282ead8
pub fn legacy_value(&self) -> u32 {
(self.0 >> 15 | self.0 << 17).wrapping_add(0xa282ead8)
self.0.rotate_right(15).wrapping_add(0xa282ead8)
}
}
@@ -67,7 +67,9 @@ mod tests {
fn test_crc_legacy_value() {
let crc = CRC(0x12345678);
let v = crc.legacy_value();
let expected = (0x12345678u32 >> 15 | 0x12345678u32 << 17).wrapping_add(0xa282ead8);
// (0x12345678 >> 15 | 0x12345678 << 17) + 0xa282ead8, worked out by hand so
// the test checks the rotate rather than restating it.
let expected = 0x4f730f40_u32;
assert_eq!(v, expected);
}
}
+4
View File
@@ -1,4 +1,8 @@
pub mod crc;
#[expect(
clippy::module_inception,
reason = "needle/needle.rs mirrors the Go package layout"
)]
pub mod needle;
pub mod ttl;
+120 -33
View File
@@ -198,8 +198,8 @@ impl Needle {
/// the data payload from disk at all, matching Go's `ReadNeedleMeta`.
pub fn read_paged_meta(
&mut self,
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
offset: i64,
expected_size: Size,
version: Version,
@@ -560,7 +560,7 @@ impl Needle {
// Padding to 8-byte alignment
let padding = padding_length(self.size, version).0 as usize;
buf.extend(std::iter::repeat(0u8).take(padding));
buf.extend(std::iter::repeat_n(0u8, padding));
buf
}
@@ -581,23 +581,19 @@ impl Needle {
// ============================================================================
/// Compute padding to align needle to NEEDLE_PADDING_SIZE (8 bytes).
///
/// The sum is formed in i64: a size read from a corrupt header can sit near
/// `i32::MAX`, and adding the header, checksum and timestamp widths to it in
/// i32 would overflow (a panic with overflow checks, a wrapped padding
/// without). The result is at most NEEDLE_PADDING_SIZE, so it fits `Size`.
pub fn padding_length(needle_size: Size, version: Version) -> Size {
if version == VERSION_3 {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32
+ needle_size.0
+ NEEDLE_CHECKSUM_SIZE as i32
+ TIMESTAMP_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
let fixed = if version == VERSION_3 {
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE + TIMESTAMP_SIZE
} else {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32 + needle_size.0 + NEEDLE_CHECKSUM_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
}
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE
};
let unpadded = fixed as i64 + needle_size.0 as i64;
Size((NEEDLE_PADDING_SIZE as i64 - unpadded % NEEDLE_PADDING_SIZE as i64) as i32)
}
/// Body length = Size + Checksum + [Timestamp] + Padding.
@@ -619,6 +615,30 @@ pub fn get_actual_size(size: Size, version: Version) -> i64 {
NEEDLE_HEADER_SIZE as i64 + needle_body_length(size, version)
}
/// Validate a wire-supplied needle body size before any `as usize` cast.
/// Rejects negative/deleted sizes and bodies larger than the gRPC max message.
/// Size(0) is allowed: empty/anomalous entries and tombstones read as size 0
/// (actual_size = header+checksum+pad > 0, safe alloc, no wrap).
/// Transport cap only: storage paths must NOT use this cap — see volume.rs
/// guards (a >1GiB stored needle from a high-limit cluster must remain
/// readable/compaction-safe). Keep `get_actual_size` unchanged (it
/// intentionally returns negative for deleted index entries).
pub fn validate_wire_size(size: Size) -> Result<(), String> {
if size.0 < 0 {
return Err(format!("invalid needle size {}", size.0));
}
// Keep in sync with canonical `GRPC_MAX_MESSAGE_SIZE` in server/grpc_client.rs:10
// (duplicated here to avoid a storage->server import and prevent drift).
const WIRE_MAX_NEEDLE_SIZE: i32 = 1 << 30;
if size.0 > WIRE_MAX_NEEDLE_SIZE {
return Err(format!(
"needle size {} exceeds max {}",
size.0, WIRE_MAX_NEEDLE_SIZE
));
}
Ok(())
}
/// Read 5 bytes as a u64 (big-endian, zero-padded high bytes).
fn bytes_to_u64_5(bytes: &[u8]) -> u64 {
assert!(bytes.len() >= 5);
@@ -729,6 +749,14 @@ pub fn parse_needle_id_cookie(s: &str) -> Result<(NeedleId, Cookie), String> {
(s, None)
};
// Every length check and the split below are in BYTES, so a multi-byte
// character would let `split` land inside one and panic the slice. Hex is
// ASCII by definition; reject anything else up front, as Go's ParseUint
// does a step later.
if !hex_part.is_ascii() {
return Err("KeyHash must be ASCII hex.".to_string());
}
// Go: len(key_hash_string) <= CookieSize*2 => error (must be > 8 hex chars)
if hex_part.len() <= COOKIE_SIZE * 2 {
return Err("KeyHash is too short.".to_string());
@@ -770,7 +798,9 @@ pub fn parse_needle_id_cookie(s: &str) -> Result<(NeedleId, Cookie), String> {
#[derive(Debug, thiserror::Error)]
pub enum NeedleError {
#[error("size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}")]
#[error(
"size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}"
)]
SizeMismatch {
offset: i64,
id: NeedleId,
@@ -806,6 +836,30 @@ pub enum NeedleError {
mod tests {
use super::*;
/// A fid whose hex part carries multi-byte UTF-8 must be rejected, not
/// panic. `split` is a byte offset into `hex_part`; before the ASCII guard
/// `&hex_part[..split]` could land inside a character. `GET /3,ééééa` is
/// nine bytes, so it passes the length checks and splits at byte 1 —
/// halfway through the first `é`. Go's `ParseUint` just errors.
#[test]
fn parse_needle_id_cookie_rejects_non_ascii_instead_of_panicking() {
for s in ["ééééa", "ééééaaaaa", "0123456é9abc", "ééééa_1"] {
assert!(
parse_needle_id_cookie(s).is_err(),
"non-ASCII fid {:?} must be an error",
s
);
}
}
/// The ASCII guard must not change any accepted input.
#[test]
fn parse_needle_id_cookie_still_accepts_ascii_hex() {
let (id, cookie) = parse_needle_id_cookie("01637037d6").unwrap();
assert_eq!(id, NeedleId(0x01));
assert_eq!(cookie, Cookie(0x637037d6));
}
#[test]
fn test_parse_header() {
let mut buf = [0u8; NEEDLE_HEADER_SIZE];
@@ -824,11 +878,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v3() {
let mut n = Needle::default();
n.cookie = Cookie(42);
n.id = NeedleId(100);
n.data = b"hello world".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(42),
id: NeedleId(100),
data: b"hello world".to_vec(),
flags: 0,
..Needle::default()
};
n.set_has_name();
n.name = b"test.txt".to_vec();
n.name_size = 8;
@@ -867,11 +923,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v2() {
let mut n = Needle::default();
n.cookie = Cookie(77);
n.id = NeedleId(200);
n.data = b"data v2".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(77),
id: NeedleId(200),
data: b"data v2".to_vec(),
flags: 0,
..Needle::default()
};
let bytes = n.write_bytes(VERSION_2);
let expected_size = get_actual_size(n.size, VERSION_2);
@@ -886,10 +944,12 @@ mod tests {
#[test]
fn test_read_bytes_meta_only_handles_tombstone_v3() {
let mut tombstone = Needle::default();
tombstone.cookie = Cookie(0x1234abcd);
tombstone.id = NeedleId(300);
tombstone.append_at_ns = 999_999;
let mut tombstone = Needle {
cookie: Cookie(0x1234abcd),
id: NeedleId(300),
append_at_ns: 999_999,
..Needle::default()
};
let bytes = tombstone.write_bytes(VERSION_3);
@@ -917,6 +977,21 @@ mod tests {
}
}
#[test]
fn padding_length_does_not_overflow_on_a_corrupt_size() {
// A header read from a corrupt or truncated file can carry any i32
// size. The scanners bound it against the bytes left before sizing a
// buffer, but on a volume with more than 2 GiB left a size near
// i32::MAX passes that bound, so the padding arithmetic itself must
// not overflow. Overflow checks are on in test builds, so an i32 sum
// here would panic rather than wrap.
for version in [VERSION_2, VERSION_3] {
let padding = padding_length(Size(i32::MAX), version).0 as i64;
assert!((1..=NEEDLE_PADDING_SIZE as i64).contains(&padding));
assert_eq!(get_actual_size(Size(i32::MAX), version) % 8, 0);
}
}
#[test]
fn test_file_id_parse() {
let fid = FileId::parse("3,01637037d6").unwrap();
@@ -961,4 +1036,16 @@ mod tests {
assert_eq!(fid.key, NeedleId(0x123));
assert_eq!(fid.cookie, Cookie(0));
}
#[test]
fn test_validate_wire_size_boundaries() {
assert!(validate_wire_size(Size(-100)).is_err());
assert!(validate_wire_size(Size(-1)).is_err());
assert!(validate_wire_size(Size(0)).is_ok());
assert!(validate_wire_size(Size(1024)).is_ok());
assert!(validate_wire_size(Size(1)).is_ok());
assert!(validate_wire_size(Size(1 << 30)).is_ok());
assert!(validate_wire_size(Size((1 << 30) + 1)).is_err());
assert!(validate_wire_size(Size(i32::MAX)).is_err());
}
}
+122 -25
View File
@@ -80,8 +80,14 @@ impl TTL {
if s.is_empty() {
return Ok(TTL::EMPTY);
}
// The unit is read as the last BYTE and the count as everything before
// it, so a trailing multi-byte character would split inside itself and
// panic. A TTL is digits plus a one-letter unit; reject the rest.
if !s.is_ascii() {
return Err(format!("invalid TTL {:?}: must be ASCII", s));
}
let last_byte = s.as_bytes()[s.len() - 1];
let (num_str, unit_byte) = if last_byte >= b'0' && last_byte <= b'9' {
let (num_str, unit_byte) = if last_byte.is_ascii_digit() {
// All digits — default to minutes (matching Go)
(s, b'm')
} else {
@@ -144,40 +150,73 @@ fn fit_ttl_count(count: u32, unit: u8) -> TTL {
const MINUTE_SECS: u64 = 60;
// First pass: try exact fits from largest to smallest
if seconds % YEAR_SECS == 0 && seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
if seconds.is_multiple_of(YEAR_SECS) && seconds / YEAR_SECS < 256 {
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
if seconds % MONTH_SECS == 0 && seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
if seconds.is_multiple_of(MONTH_SECS) && seconds / MONTH_SECS < 256 {
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds % WEEK_SECS == 0 && seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
if seconds.is_multiple_of(WEEK_SECS) && seconds / WEEK_SECS < 256 {
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds % DAY_SECS == 0 && seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
if seconds.is_multiple_of(DAY_SECS) && seconds / DAY_SECS < 256 {
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds % HOUR_SECS == 0 && seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
if seconds.is_multiple_of(HOUR_SECS) && seconds / HOUR_SECS < 256 {
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
// Minutes: truncating division
if seconds / MINUTE_SECS < 256 {
return TTL { count: (seconds / MINUTE_SECS) as u8, unit: TTL_UNIT_MINUTE };
return TTL {
count: (seconds / MINUTE_SECS) as u8,
unit: TTL_UNIT_MINUTE,
};
}
// Second pass: truncating division from smallest to largest
if seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
if seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
TTL::EMPTY
}
@@ -207,6 +246,16 @@ impl fmt::Display for TTL {
mod tests {
use super::*;
/// `?ttl=5%C3%A9` must be an error, not a panic. The unit is taken as the
/// last *byte*, so a trailing multi-byte character made `&s[..s.len()-1]`
/// split inside it.
#[test]
fn ttl_read_rejects_non_ascii_instead_of_panicking() {
for s in ["5é", "é", "3🦀", "12é"] {
assert!(TTL::read(s).is_err(), "non-ASCII TTL {:?} must error", s);
}
}
#[test]
fn test_ttl_parse() {
let ttl = TTL::read("3m").unwrap();
@@ -225,7 +274,13 @@ mod tests {
// 24h normalizes to 1d via fitTtlCount
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl.to_seconds(), 86400);
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -271,12 +326,24 @@ mod tests {
fn test_ttl_overflow_normalizes() {
// Go's ReadTTL calls fitTtlCount: 300m = 18000s = 5h (exact fit)
let ttl = TTL::read("300m").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_HOUR
}
);
// 256h = 921600s. Doesn't fit in hours (256 >= 256), doesn't fit exact in days.
// Second pass: 921600/86400 = 10 (truncated) < 256 -> 10d
let ttl = TTL::read("256h").unwrap();
assert_eq!(ttl, TTL { count: 10, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 10,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -284,19 +351,49 @@ mod tests {
// Go's ReadTTL calls fitTtlCount which normalizes to coarsest unit.
// 120m -> 2h, 7d -> 1w, 24h -> 1d.
let ttl = TTL::read("120m").unwrap();
assert_eq!(ttl, TTL { count: 2, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 2,
unit: TTL_UNIT_HOUR
}
);
let ttl = TTL::read("7d").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_WEEK });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_WEEK
}
);
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
// Values that don't simplify stay as-is
let ttl = TTL::read("5d").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_DAY
}
);
let ttl = TTL::read("3m").unwrap();
assert_eq!(ttl, TTL { count: 3, unit: TTL_UNIT_MINUTE });
assert_eq!(
ttl,
TTL {
count: 3,
unit: TTL_UNIT_MINUTE
}
);
}
}
+337 -225
View File
@@ -97,12 +97,13 @@ impl NeedleMapMetric {
self.file_byte_count
.fetch_add(new_size.0 as u64, Ordering::Relaxed);
// Go: if oldSize > 0 && oldSize.IsValid() { LogDeletionCounter(oldSize) }
if let Some(old_val) = old {
if old_val.size.0 > 0 && old_val.size.is_valid() {
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
if let Some(old_val) = old
&& old_val.size.0 > 0
&& old_val.size.is_valid()
{
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
}
@@ -194,7 +195,12 @@ impl NeedleMapKind {
// ============================================================================
/// Trait for appending to an index file.
pub trait IdxFileWriter: Write + Send + Sync {
///
/// The file is opened without append mode and each row is written at the
/// current `idx_file_offset` — the same positioned-write model the Go
/// server uses — because an append-mode handle cannot truncate on Windows,
/// where the std library keeps it strictly append-only.
pub trait IdxFileWriter: Write + Seek + Send + Sync {
fn sync_all(&self) -> io::Result<()>;
/// Truncate the file to `len` bytes. Used to remove an orphan .idx row
/// left by a failed redb commit so `idx_file_offset` stays a contiguous
@@ -223,6 +229,15 @@ pub struct CompactNeedleMap {
metric: NeedleMapMetric,
idx_file: Option<Box<dyn IdxFileWriter>>,
idx_file_offset: u64,
/// The file holds bytes past `idx_file_offset` that must be trimmed
/// before another row can land aligned.
idx_torn: bool,
}
impl Default for CompactNeedleMap {
fn default() -> Self {
Self::new()
}
}
impl CompactNeedleMap {
@@ -233,6 +248,7 @@ impl CompactNeedleMap {
metric: NeedleMapMetric::default(),
idx_file: None,
idx_file_offset: 0,
idx_torn: false,
}
}
@@ -240,6 +256,8 @@ impl CompactNeedleMap {
pub fn load_from_idx<R: Read + Seek>(reader: &mut R, version: Version) -> io::Result<Self> {
let mut nm = CompactNeedleMap::new();
idx::walk_index_file(reader, 0, |key, offset, size| {
// A read-only load attaches no writer, so this is its only size.
nm.idx_file_offset += NEEDLE_MAP_ENTRY_SIZE as u64;
nm.metric.maybe_set_max_needle_end(offset, size, version);
if offset.is_zero() || size.is_deleted() {
nm.delete_from_map(key);
@@ -255,6 +273,7 @@ impl CompactNeedleMap {
pub fn set_idx_file(&mut self, file: Box<dyn IdxFileWriter>, offset: u64) {
self.idx_file = Some(file);
self.idx_file_offset = offset;
self.idx_torn = false;
}
/// True when an .idx file writer is attached. A read-only load leaves
@@ -269,8 +288,8 @@ impl CompactNeedleMap {
/// Insert or update an entry. Appends to .idx file if present.
pub fn put(&mut self, key: NeedleId, offset: Offset, size: Size) -> io::Result<()> {
// Persist to idx file BEFORE mutating in-memory state for crash consistency
if let Some(ref mut idx_file) = self.idx_file {
idx::write_index_entry(idx_file, key, offset, size)?;
self.append_to_index_file(key, offset, size)?;
if self.idx_file.is_some() {
self.idx_file_offset += NEEDLE_MAP_ENTRY_SIZE as u64;
}
@@ -280,6 +299,41 @@ impl CompactNeedleMap {
Ok(())
}
/// Write one row to the .idx file at `idx_file_offset`. A row left
/// half-written by a failed write is trimmed back to the offset so the
/// next row still lands aligned; while the trim keeps failing no row is
/// written at all, or it would sit off alignment and parse as garbage on
/// load. The offset itself is advanced by the caller once the row counts.
fn append_to_index_file(
&mut self,
key: NeedleId,
offset: Offset,
size: Size,
) -> io::Result<()> {
let Some(idx_file) = self.idx_file.as_mut() else {
return Ok(());
};
if self.idx_torn {
match idx_file.truncate_to(self.idx_file_offset) {
Ok(()) => self.idx_torn = false,
Err(e) => {
return Err(io::Error::other(format!(
"index file still holds a torn row: {e}"
)));
}
}
}
idx_file.seek(io::SeekFrom::Start(self.idx_file_offset))?;
if let Err(e) = idx::write_index_entry(idx_file, key, offset, size) {
if let Err(te) = idx_file.truncate_to(self.idx_file_offset) {
self.idx_torn = true;
tracing::warn!("failed to trim torn .idx row: {}", te);
}
return Err(e);
}
Ok(())
}
/// Look up a needle.
pub fn get(&self, key: NeedleId) -> Option<NeedleValue> {
self.map.get(key)
@@ -304,8 +358,8 @@ impl CompactNeedleMap {
}
// Always write tombstone to idx file (matching Go)
if let Some(ref mut idx_file) = self.idx_file {
idx::write_index_entry(idx_file, key, offset, TOMBSTONE_FILE_SIZE)?;
self.append_to_index_file(key, offset, TOMBSTONE_FILE_SIZE)?;
if self.idx_file.is_some() {
self.idx_file_offset += NEEDLE_MAP_ENTRY_SIZE as u64;
}
@@ -410,9 +464,9 @@ impl CompactNeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
{
self.map.ascending_visit(f)
}
@@ -455,6 +509,9 @@ pub struct RedbNeedleMap {
metric: NeedleMapMetric,
idx_file: Option<Box<dyn IdxFileWriter>>,
idx_file_offset: u64,
/// The file holds bytes past `idx_file_offset` that must be trimmed
/// before another row can land aligned.
idx_torn: bool,
/// Puts/deletes since the last durable checkpoint.
writes_since_checkpoint: u32,
}
@@ -465,9 +522,9 @@ impl RedbNeedleMap {
/// loses at most the writes since the last checkpoint from redb, and
/// the next load replays them from .idx.
fn begin_write_no_fsync(db: &Database) -> io::Result<redb::WriteTransaction> {
let mut txn = db.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = db
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
let _ = txn.set_durability(Durability::None);
Ok(txn)
}
@@ -501,7 +558,7 @@ impl RedbNeedleMap {
pub fn checkpoint(&mut self, sync_idx: bool) -> io::Result<()> {
let txn = self.begin_checkpoint(sync_idx)?;
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
self.writes_since_checkpoint = 0;
Ok(())
}
@@ -516,17 +573,17 @@ impl RedbNeedleMap {
if sync_idx {
self.sync()?;
}
let mut txn = self.db_or_err()?.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = self
.db_or_err()?
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
txn.set_quick_repair(true);
if self.idx_file.is_some() {
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, self.idx_file_offset).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, self.idx_file_offset)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
Ok(txn)
}
@@ -538,22 +595,20 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.create(db_path)
.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb create error: {}", e))
})?;
.map_err(|e| io::Error::other(format!("redb create error: {}", e)))?;
// Ensure tables exist
let txn = Self::begin_write_no_fsync(&db)?;
{
let _table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let _meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table meta: {}", e))
})?;
let _table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let _meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table meta: {}", e)))?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
Ok(RedbNeedleMap {
db: Some(db),
@@ -563,6 +618,7 @@ impl RedbNeedleMap {
metric: NeedleMapMetric::default(),
idx_file: None,
idx_file_offset: 0,
idx_torn: false,
writes_since_checkpoint: 0,
})
}
@@ -572,16 +628,14 @@ impl RedbNeedleMap {
fn save_idx_size_meta(&self, idx_size: u64) -> io::Result<()> {
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
{
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, idx_size).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, idx_size)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
txn.commit().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb commit meta: {}", e))
})?;
txn.commit()
.map_err(|e| io::Error::other(format!("redb commit meta: {}", e)))?;
Ok(())
}
@@ -590,22 +644,26 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `meta`; bind the match so the
// temporary Result is dropped before `meta`.
let result = match meta.get(META_IDX_SIZE) {
// ReadableTable::get guard borrows `meta`; edition 2024 drops the tail
// expression's temporaries before `meta`, so no extra binding is needed.
match meta.get(META_IDX_SIZE) {
Ok(Some(guard)) => Ok(Some(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get meta: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get meta: {}", e))),
}
}
/// Test-only read of META `idx_size` through the live handle. See
/// [`test_support::live_meta_idx_size`] for why durability tests use
/// this instead of copying the open `.rdb`.
#[cfg(test)]
pub(crate) fn live_meta_idx_size(&self) -> Option<u64> {
self.read_idx_size_meta().unwrap()
}
/// Load from an .idx file, reusing an existing .rdb if it is consistent.
@@ -648,7 +706,7 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.open(db_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open: {}", e)))?;
let mut nm = RedbNeedleMap {
db: Some(db),
@@ -658,19 +716,17 @@ impl RedbNeedleMap {
metric: NeedleMapMetric::default(),
idx_file: None,
idx_file_offset: 0,
idx_torn: false,
writes_since_checkpoint: 0,
};
let stored_idx_size = nm
.read_idx_size_meta()?
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "no idx_size in redb meta"))?;
.ok_or_else(|| io::Error::other("no idx_size in redb meta"))?;
if stored_idx_size > idx_size {
// .idx shrank — corrupted or truncated, need full rebuild
return Err(io::Error::new(
io::ErrorKind::Other,
"idx file smaller than stored size",
));
return Err(io::Error::other("idx file smaller than stored size"));
}
// Counters come from the whole .idx history, never from the table,
@@ -683,40 +739,37 @@ impl RedbNeedleMap {
let start_entry = stored_idx_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
idx::walk_index_file(reader, start_entry, |key, offset, size| {
let key_u64: u64 = key.into();
if offset.is_zero() || size.is_deleted() {
// Delete: store a tombstone (negative size, original
// offset) over a live value; already deleted is a no-op.
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64) {
if old.size.is_valid() {
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
)
})?;
}
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64)
&& old.size.is_valid()
{
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
} else {
let packed = pack_needle_value(&NeedleValue { offset, size });
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
Ok(())
})?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
}
@@ -734,10 +787,7 @@ impl RedbNeedleMap {
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
@@ -790,13 +840,13 @@ impl RedbNeedleMap {
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
if !unlinked {
table.retain(|_, _| false).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb retain: {}", e))
})?;
table
.retain(|_, _| false)
.map_err(|e| io::Error::other(format!("redb retain: {}", e)))?;
}
#[cfg(not(feature = "redb-experimental-cursor"))]
@@ -804,30 +854,33 @@ impl RedbNeedleMap {
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
}
#[cfg(feature = "redb-experimental-cursor")]
{
let mut cursor = table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
let mut cursor =
table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
cursor.insert_before(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
cursor
.insert_before(key_u64, packed.as_slice())
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
}
cursor.close().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb cursor close: {}", e))
@@ -835,7 +888,7 @@ impl RedbNeedleMap {
}
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
Ok(())
@@ -855,6 +908,7 @@ impl RedbNeedleMap {
pub fn set_idx_file(&mut self, file: Box<dyn IdxFileWriter>, offset: u64) {
self.idx_file = Some(file);
self.idx_file_offset = offset;
self.idx_torn = false;
}
/// True when an .idx file writer is attached. See CompactNeedleMap.
@@ -871,9 +925,7 @@ impl RedbNeedleMap {
// commit leaves an orphan row in .idx that redb doesn't reflect, and
// advancing the offset here would let a later checkpoint record it as
// reflected, making the reload skip it permanently.
if let Some(ref mut idx_file) = self.idx_file {
idx::write_index_entry(idx_file, key, offset, size)?;
}
self.append_to_index_file(key, offset, size)?;
let key_u64: u64 = key.into();
let packed = pack_needle_value(&NeedleValue { offset, size });
@@ -901,23 +953,16 @@ impl RedbNeedleMap {
Ok(t) => t,
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb open_table: {}", e),
));
return Err(io::Error::other(format!("redb open_table: {}", e)));
}
};
let result = match table.insert(key_u64, packed.as_slice()) {
match table.insert(key_u64, packed.as_slice()) {
Ok(prev) => prev.and_then(|g| packed_to_needle_value(g.value())),
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
};
result
}
};
match txn.commit() {
Ok(()) => old,
@@ -925,8 +970,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -935,12 +979,9 @@ impl RedbNeedleMap {
// visible and redb refuses further writes. Keep
// the .idx row (do NOT truncate) and reopen from
// .idx to repair redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after put commit error failed: {}",
reopen_err
);
tracing::warn!("redb reopen after put commit error failed: {}", reopen_err);
}
return Err(err);
}
@@ -955,6 +996,41 @@ impl RedbNeedleMap {
Ok(())
}
/// Write one row to the .idx file at `idx_file_offset`. A row left
/// half-written by a failed write is trimmed back to the offset so the
/// next row still lands aligned; while the trim keeps failing no row is
/// written at all, or it would sit off alignment and parse as garbage on
/// load. The offset itself is advanced by the caller once the row counts.
fn append_to_index_file(
&mut self,
key: NeedleId,
offset: Offset,
size: Size,
) -> io::Result<()> {
let Some(idx_file) = self.idx_file.as_mut() else {
return Ok(());
};
if self.idx_torn {
match idx_file.truncate_to(self.idx_file_offset) {
Ok(()) => self.idx_torn = false,
Err(e) => {
return Err(io::Error::other(format!(
"index file still holds a torn row: {e}"
)));
}
}
}
idx_file.seek(io::SeekFrom::Start(self.idx_file_offset))?;
if let Err(e) = idx::write_index_entry(idx_file, key, offset, size) {
if let Err(te) = idx_file.truncate_to(self.idx_file_offset) {
self.idx_torn = true;
tracing::warn!("failed to trim torn .idx row: {}", te);
}
return Err(e);
}
Ok(())
}
/// Look up a needle. A redb failure is an ERROR, not an absent needle:
/// answering "not found" would turn a database problem into a read miss
/// and let a delete report success without recording a tombstone.
@@ -968,39 +1044,32 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `table`; bind the match so the
// temporary Result is dropped before `table`.
let result = match table.get(key_u64) {
// ReadableTable::get guard borrows `table`; edition 2024 drops the tail
// expression's temporaries before `table`, so no extra binding is needed.
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
/// Mark a needle as deleted. Appends tombstone to .idx file, negates size in redb.
pub fn delete(&mut self, key: NeedleId, offset: Offset) -> io::Result<Option<Size>> {
let key_u64: u64 = key.into();
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let old = match table.get(key_u64) {
Ok(Some(guard)) => packed_to_needle_value(guard.value()),
Ok(None) => None,
Err(e) => {
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
));
return Err(io::Error::other(format!("redb get: {}", e)));
}
};
let Some(old) = old.filter(|nv| nv.size.is_valid()) else {
@@ -1008,9 +1077,7 @@ impl RedbNeedleMap {
return Ok(None);
};
if let Some(ref mut idx_file) = self.idx_file {
idx::write_index_entry(idx_file, key, offset, TOMBSTONE_FILE_SIZE)?;
}
self.append_to_index_file(key, offset, TOMBSTONE_FILE_SIZE)?;
let deleted_nv = NeedleValue {
offset: old.offset,
@@ -1021,10 +1088,7 @@ impl RedbNeedleMap {
drop(table);
if let Err(e) = insert_res {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
match txn.commit() {
Ok(()) => {}
@@ -1032,8 +1096,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -1042,7 +1105,7 @@ impl RedbNeedleMap {
// and redb refuses further writes. Keep the .idx row
// (do NOT truncate) and reopen from .idx to repair
// redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after delete commit error failed: {}",
@@ -1104,11 +1167,14 @@ impl RedbNeedleMap {
/// a failed redb commit. Without this the next successful write appends
/// after the orphan, `idx_file_offset` advances past it, and a later
/// checkpoint records an offset that makes the reload skip the orphan.
/// When the trim fails the file is latched torn so no later row lands
/// after bytes the map does not reflect.
fn truncate_idx_to_offset(&mut self) {
if let Some(ref mut idx_file) = self.idx_file {
if let Err(e) = idx_file.truncate_to(self.idx_file_offset) {
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
if let Some(ref mut idx_file) = self.idx_file
&& let Err(e) = idx_file.truncate_to(self.idx_file_offset)
{
self.idx_torn = true;
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
}
@@ -1146,24 +1212,19 @@ impl RedbNeedleMap {
let read_file = std::fs::OpenOptions::new()
.read(true)
.open(&idx_path)
.map_err(|e| {
io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e))
})?;
.map_err(|e| io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e)))?;
let actual_idx_size = read_file.metadata()?.len();
let mut reader = io::BufReader::new(read_file);
let reopened = Self::load_from_idx(
&self.rdb_path,
&mut reader,
self.version,
self.cache_bytes,
)?;
let reopened =
Self::load_from_idx(&self.rdb_path, &mut reader, self.version, self.cache_bytes)?;
// Preserve the append writer and the paths/version/cache; adopt the
// repaired database, metrics, and idx_file_offset from the reload.
self.db = reopened.db;
self.metric = reopened.metric;
self.idx_file_offset = actual_idx_size;
self.idx_torn = reopened.idx_torn;
// The reopen replayed all rows since the last durable checkpoint
// non-durably; start the counter fresh.
self.writes_since_checkpoint = 0;
@@ -1198,10 +1259,10 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let mut file = std::fs::OpenOptions::new()
.write(true)
@@ -1212,18 +1273,17 @@ impl RedbNeedleMap {
// redb iterates in key order (u64 ascending)
let iter = table
.iter()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb iter: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb iter: {}", e)))?;
for entry in iter {
let (key_guard, val_guard) = entry.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb iter next: {}", e))
})?;
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb iter next: {}", e)))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if let Some(nv) = packed_to_needle_value(bytes) {
if nv.size.is_valid() {
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
if let Some(nv) = packed_to_needle_value(bytes)
&& nv.size.is_valid()
{
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
}
file.sync_all()?;
@@ -1231,9 +1291,10 @@ impl RedbNeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, mut f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, mut f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
let txn = self
.db_or_err()
@@ -1413,6 +1474,18 @@ impl NeedleMap {
}
}
/// Skew the live file count away from what the `.idx` holds, so tests
/// can build a volume whose reported count disagrees with a reload.
#[cfg(test)]
pub(crate) fn add_file_count_for_test(&self, delta: i64) {
let metric = match self {
NeedleMap::InMemory(nm) => &nm.metric,
NeedleMap::Redb(nm) => &nm.metric,
NeedleMap::SortedFile(_) => panic!("sorted-file needle maps are read-only"),
};
metric.file_count.fetch_add(delta, Ordering::Relaxed);
}
/// Largest (offset + actual size) seen during the load walk; 0 if the
/// map is empty. Used at volume load to detect .idx entries that
/// reference past the end of .dat (issue #8928) without a second scan.
@@ -1471,9 +1544,10 @@ impl NeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
match self {
NeedleMap::InMemory(nm) => nm.ascending_visit(f),
@@ -1494,7 +1568,7 @@ impl NeedleMap {
// The visitor never fails, so neither can this.
let _ = nm.ascending_visit(|id, nv| {
entries.push((id, *nv));
Ok(())
Ok::<(), std::convert::Infallible>(())
});
Ok(entries)
}
@@ -1512,22 +1586,30 @@ impl NeedleMap {
pub(crate) mod test_support {
use super::*;
/// The `.idx` size recorded in the durable state of the `.rdb` at
/// `rdb_path`, read from a copy taken while the map may still be open:
/// exactly what a crash would leave behind. `None` when nothing durable
/// has been recorded yet.
pub(crate) fn durable_idx_size(rdb_path: &Path) -> Option<u64> {
let copy = rdb_path.with_extension("crash-copy.rdb");
std::fs::copy(rdb_path, &copy).unwrap();
let db = Database::open(&copy).unwrap();
let txn = db.begin_read().unwrap();
let meta = txn.open_table(META_TABLE).ok()?;
let size = meta.get(META_IDX_SIZE).unwrap().map(|g| g.value());
drop(meta);
drop(txn);
drop(db);
let _ = std::fs::remove_file(&copy);
size
/// The `.idx` size in the map's META table, read through the live
/// handle.
///
/// The load path records the `.idx` size with `Durability::None`, and
/// every `put`/`delete` also commits non-durably, so before the first
/// checkpoint this is the load-time value (`Some(0)` for a fresh map) —
/// NOT the crash-durable `None` a copy of the open `.rdb` would show.
/// A live read is the only portable observation: redb 4.2.0 takes an
/// exclusive whole-file lock, which is advisory on Unix but mandatory
/// on Windows, so copying the open `.rdb` fails there with OS error 33.
///
/// It still pins the property under test: the only *durable* META
/// writer is `checkpoint`, so any value other than the load-time one
/// proves a checkpoint recorded progress — and the post-checkpoint
/// value equals the durable one, because checkpoints commit with
/// `Durability::Immediate`. What is lost vs the old copy: strict crash
/// fidelity — a hard crash pre-checkpoint would leave META absent
/// rather than `Some(0)` (loader-equivalent outcomes: full rebuild vs
/// replay-from-0, both correct). A clean close+reopen cannot recover
/// that distinction either: dropping the `Database` flushes pending
/// non-durable commits, so a reopened handle reads `Some(0)` just like
/// the live one.
pub(crate) fn live_meta_idx_size(nm: &RedbNeedleMap) -> Option<u64> {
nm.live_meta_idx_size()
}
}
@@ -1682,6 +1764,7 @@ mod tests {
.read(true)
.write(true)
.create(true)
.truncate(false)
.open(&idx_path)
.unwrap();
let idx_size = idx_file.metadata().unwrap().len();
@@ -1694,7 +1777,7 @@ mod tests {
)
.unwrap();
let writer = std::fs::OpenOptions::new()
.append(true)
.write(true)
.open(&idx_path)
.unwrap();
nm.set_idx_file(Box::new(writer), idx_size);
@@ -1857,7 +1940,7 @@ mod tests {
let mut live = 0u64;
nm.ascending_visit(|_, _| {
live += 1;
Ok(())
Ok::<(), String>(())
})
.unwrap();
assert_eq!(live, N - 1);
@@ -2070,7 +2153,7 @@ mod tests {
let mut visited = Vec::new();
nm.ascending_visit(|id, nv| {
visited.push((id, nv.size));
Ok(())
Ok::<(), String>(())
})
.unwrap();
@@ -2168,8 +2251,14 @@ mod tests {
// server opens one redb database per volume, so the process-wide
// ceiling is roughly (volumes x budget).
assert_eq!(NeedleMapKind::Redb.redb_cache_bytes(), 4 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbMedium.redb_cache_bytes(), 8 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbLarge.redb_cache_bytes(), 16 * 1024 * 1024);
assert_eq!(
NeedleMapKind::RedbMedium.redb_cache_bytes(),
8 * 1024 * 1024
);
assert_eq!(
NeedleMapKind::RedbLarge.redb_cache_bytes(),
16 * 1024 * 1024
);
}
#[test]
@@ -2198,7 +2287,7 @@ mod tests {
#[test]
fn test_redb_checkpoint_is_explicit_and_due_every_interval() {
use test_support::durable_idx_size;
use test_support::live_meta_idx_size;
// Every non-durable redb commit leaves bookkeeping behind until a
// durable one clears it, so a writable map asks for a checkpoint on
@@ -2208,8 +2297,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, _idx_path) = open_writable_redb(dir.path());
for i in 1..EXPECTED_INTERVAL {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
assert!(!nm.checkpoint_due(), "due after only {i} writes");
}
nm.put(
@@ -2219,20 +2312,29 @@ mod tests {
)
.unwrap();
assert!(nm.checkpoint_due());
assert_eq!(durable_idx_size(&db_path), None, "put() must not commit durably");
// No checkpoint taken yet: META still holds the load-time .idx size.
// put() only commits non-durably, so the live value is unchanged.
assert_eq!(
live_meta_idx_size(&nm),
Some(0),
"put() must not record checkpoint progress"
);
nm.checkpoint(true).unwrap();
assert!(!nm.checkpoint_due());
assert_eq!(
durable_idx_size(&db_path),
live_meta_idx_size(&nm),
Some(EXPECTED_INTERVAL * NEEDLE_MAP_ENTRY_SIZE as u64),
"checkpoint records how much of the .idx the table reflects"
);
// Snapshot the .rdb while the map is still open: what a crash leaves.
// Everything is durable after the checkpoint, so the map is closed
// first and the snapshot sees the same bytes on every platform.
// (Copying while open fails on Windows, where redb's file lock is
// mandatory: what a crash leaves.)
drop(nm);
let crash_copy = dir.path().join("crash.rdb");
std::fs::copy(&db_path, &crash_copy).unwrap();
drop(nm);
let db = Database::open(&crash_copy).unwrap();
let txn = db.begin_read().unwrap();
let table = txn.open_table(NEEDLE_TABLE).unwrap();
@@ -2247,8 +2349,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
nm.close();
drop(nm);
@@ -2272,8 +2378,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
// Drop without close(): redb makes the table durable on drop, but the
// recorded .idx size stays at its load-time value (0), so the reload
@@ -2344,12 +2454,14 @@ mod tests {
reloaded.deleted_count(),
reloaded.deleted_size(),
);
assert_eq!(
after, live,
"close_first={close_first} rebuild={rebuild}"
);
assert_eq!(after, live, "close_first={close_first} rebuild={rebuild}");
assert_eq!(reloaded.get(NeedleId(1)).unwrap().unwrap().size, Size(200));
assert!(reloaded.get(NeedleId(2)).unwrap().map_or(true, |v| v.size.is_deleted()));
assert!(
reloaded
.get(NeedleId(2))
.unwrap()
.is_none_or(|v| v.size.is_deleted())
);
}
}
}
@@ -31,7 +31,7 @@ struct CompactEntry {
}
impl CompactEntry {
fn to_needle_value(&self) -> NeedleValue {
fn to_needle_value(self) -> NeedleValue {
NeedleValue {
offset: Offset::from_bytes(&self.offset),
size: self.size,
@@ -205,6 +205,18 @@ mod tests {
use std::os::unix::fs::FileExt;
borrowed.read_exact_at(&mut buf, 0).unwrap();
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = 0;
while filled < buf.len() {
let n = borrowed.seek_read(&mut buf[filled..], at).unwrap();
assert!(n != 0, "unexpected EOF in seek_read");
filled += n;
at += n as u64;
}
}
assert_eq!(&buf, b"first");
}
@@ -132,9 +132,7 @@ mod tests {
let mut seen = SeenKeys::new(10_000, FALSE_POSITIVE_RATE);
// Fresh keys may occasionally collide (that is the false-positive
// rate), but only rarely.
let fresh_reported_seen = (0..10_000u64)
.filter(|&key| seen.test_and_add(key))
.count();
let fresh_reported_seen = (0..10_000u64).filter(|&key| seen.test_and_add(key)).count();
assert!(
fresh_reported_seen < 50,
"fresh keys reported seen: {fresh_reported_seen}"
@@ -18,6 +18,7 @@ use std::sync::{Mutex, RwLock};
use super::file_pool::pooled_index_files;
use crate::storage::idx;
use crate::storage::io::read_exact_at;
use crate::storage::needle_map::{CompactNeedleMap, NeedleMapMetric, NeedleValue};
use crate::storage::types::*;
@@ -133,9 +134,7 @@ impl SortedFileNeedleMap {
}
let file = pooled_index_files()
.borrow(&self.db_file_name, false)
.map_err(|e| {
io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e))
})?;
.map_err(|e| io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e)))?;
match search_sorted_index(&file, self.db_file_size, key)? {
Some((_, offset, size)) => Ok(Some(NeedleValue { offset, size })),
None => Ok(None),
@@ -226,10 +225,7 @@ impl SortedFileNeedleMap {
.fail_sdx_mark
.load(std::sync::atomic::Ordering::Relaxed)
{
return Err(io::Error::new(
io::ErrorKind::Other,
"injected .sdx mark failure",
));
return Err(io::Error::other("injected .sdx mark failure"));
}
let mut buf = [0u8; SIZE_SIZE];
TOMBSTONE_FILE_SIZE.to_bytes(&mut buf);
@@ -309,7 +305,7 @@ impl SortedFileNeedleMap {
let rows = rows_per_read.min(entry_count - done) as usize;
let bytes = &mut block[..rows * NEEDLE_MAP_ENTRY_SIZE];
read_exact_at(&file, bytes, done * NEEDLE_MAP_ENTRY_SIZE as u64)?;
for entry in bytes.chunks_exact(NEEDLE_MAP_ENTRY_SIZE) {
for entry in bytes.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0 {
let (key, offset, size) = idx_entry_from_bytes(entry);
if !size.is_valid() || pending.contains_key(&key) {
continue; // deleted in place, or still awaiting that mark
@@ -321,9 +317,11 @@ impl SortedFileNeedleMap {
Ok(())
}
pub fn ascending_visit<F>(&self, mut f: F) -> Result<(), String>
/// Visit all live entries in ascending order by needle ID.
pub fn ascending_visit<F, E>(&self, mut f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
let mut visit_error = None;
self.visit_live_entries(|id, nv| {
@@ -333,7 +331,7 @@ impl SortedFileNeedleMap {
}
Ok(())
})
.map_err(|e| visit_error.take().unwrap_or_else(|| e.to_string()))
.map_err(|e| visit_error.take().unwrap_or_else(|| E::from(e.to_string())))
}
pub fn iter_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
@@ -525,32 +523,6 @@ fn search_sorted_index(
Ok(None)
}
fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = offset;
while filled < buf.len() {
let n = file.seek_read(&mut buf[filled..], at)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
filled += n;
at += n as u64;
}
Ok(())
}
}
fn write_at(file: &File, buf: &[u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
@@ -703,10 +675,11 @@ mod tests {
// without a reload — the same contract Go's Get has, where callers
// check size.is_deleted().
assert!(m.get(NeedleId(2)).unwrap().unwrap().size.is_deleted());
assert!(m
.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none());
assert!(
m.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none()
);
assert!(!m.get(NeedleId(1)).unwrap().unwrap().size.is_deleted());
}
@@ -1002,7 +975,8 @@ mod tests {
// The retry is a no-op: no second tombstone, no double counting.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
assert_eq!(m.deleted_count(), deleted_before + 2);
@@ -1032,7 +1006,8 @@ mod tests {
);
// And a retry must not append a second tombstone for it.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
}
@@ -1062,7 +1037,7 @@ mod tests {
let mut visited = Vec::new();
m.ascending_visit(|id, _| {
visited.push(id);
Ok(())
Ok::<(), String>(())
})
.unwrap();
assert_eq!(visited, vec![NeedleId(2)]);
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,374 @@
//! Merging a peer's `.ecj` deletion ids into a local EC journal. Mirrors Go's
//! `Store.MergeEcJournal` (`weed/storage/store_ec_journal.go`).
use std::collections::HashSet;
use std::io;
use std::path::Path;
use std::sync::RwLock;
use crate::storage::erasure_coding::ecj_merge::{append_ecj_ids, read_ecj_ids};
use crate::storage::store::Store;
use crate::storage::types::{NeedleId, VolumeId};
/// How often an unmounted merge re-reads a journal that changed under it. Only
/// a mount-delete-unmount or a concurrent merge between the read and the
/// append changes it, so one retry is nearly always enough.
const ECJ_MERGE_ATTEMPTS: usize = 5;
/// Fold a peer's deletion `ids` into the local journal of EC volume `vid` on
/// the receiving disk, the one whose data directory is `data_dir`; `ecj_path`
/// is that journal's path in the disk's index directory. Appends only the ids
/// the journal lacks and returns how many it added. Blocking: call it from
/// `spawn_blocking`.
///
/// A mounted volume owns its journal: the merge goes through its open handle
/// and in-memory set. That is the receiving disk's own runtime for `vid`,
/// wherever its journal lives (it may sit in the data dir rather than
/// `ecj_path`'s index dir), else a sibling runtime journaling into `ecj_path`
/// itself: disks sharing one index directory, or reconciliation mounting `vid`
/// on a disk that journals into another's (#9212). Otherwise `ecj_path` is
/// appended to under the store write lock, which mounts take, so no mount can
/// open it mid-append. The read that computes the delta runs outside the lock,
/// and a journal that changed in between — including by a concurrent merge —
/// is re-read.
pub fn merge_ec_journal(
store: &RwLock<Store>,
vid: VolumeId,
data_dir: &str,
ecj_path: &str,
ids: &HashSet<NeedleId>,
) -> io::Result<usize> {
merge_ec_journal_with(store, vid, data_dir, ecj_path, ids, read_ecj_ids)
}
/// `merge_ec_journal` with the unlocked journal read injected, so a test can
/// mount the volume between that read and the append.
fn merge_ec_journal_with(
store: &RwLock<Store>,
vid: VolumeId,
data_dir: &str,
ecj_path: &str,
ids: &HashSet<NeedleId>,
mut read: impl FnMut(&str) -> io::Result<(HashSet<NeedleId>, u64)>,
) -> io::Result<usize> {
for _ in 0..ECJ_MERGE_ATTEMPTS {
{
let mut store = store
.write()
.map_err(|_| io::Error::other("store lock poisoned"))?;
if let Some(primary) = mounted_ec_journal(&store, vid, data_dir, ecj_path)? {
let ecv = store.locations[primary]
.find_ec_volume_mut(vid)
.expect("mounted journal runtime");
let journal_path = ecv.ecj_file_name();
let added = ecv.merge_journal(ids)?;
// Publish to the holders of the file the merge wrote to — the
// picked runtime's journal may live outside ecj_path, and a
// holder of a different file must not claim ids it lacks.
publish_to_journal_siblings(&store, primary, vid, &journal_path, ids);
return Ok(added);
}
}
let (local, size) = read(ecj_path)?;
let mut store = store
.write()
.map_err(|_| io::Error::other("store lock poisoned"))?;
if let Some(primary) = mounted_ec_journal(&store, vid, data_dir, ecj_path)? {
// Mounted since the read: its handle owns the journal now.
let ecv = store.locations[primary]
.find_ec_volume_mut(vid)
.expect("mounted journal runtime");
let journal_path = ecv.ecj_file_name();
let added = ecv.merge_journal(ids)?;
publish_to_journal_siblings(&store, primary, vid, &journal_path, ids);
return Ok(added);
}
// The path append registers as a writer so a mount compacting this
// journal cannot swap its inode underneath it (the write itself is
// already serialized with mounts by the store lock).
let _ecj_write =
crate::storage::erasure_coding::ecj_registry::begin_ecj_write(ecj_path);
if let Some(added) = append_ecj_ids(ecj_path, &local, ids, size)? {
return Ok(added);
}
}
Err(io::Error::other(format!(
"ec volume {}: journal {} kept changing during merge",
vid.0, ecj_path
)))
}
/// The disk index of the runtime holding `ecj_path` open, if any: the disk at
/// `data_dir`'s own, else the first sibling journaling into it.
fn mounted_ec_journal(
store: &Store,
vid: VolumeId,
data_dir: &str,
ecj_path: &str,
) -> io::Result<Option<usize>> {
let owner = store
.locations
.iter()
.position(|loc| Path::new(&loc.directory) == Path::new(data_dir))
.ok_or_else(|| {
io::Error::other(format!(
"ec volume {}: no disk at {} owns journal {}",
vid.0, data_dir, ecj_path
))
})?;
let runtime = if store.locations[owner].has_ec_volume(vid) {
Some(owner)
} else {
store.locations.iter().position(|loc| {
loc.find_ec_volume(vid)
.is_some_and(|ecv| Path::new(&ecv.ecj_file_name()) == Path::new(ecj_path))
})
};
Ok(runtime)
}
/// Publishes merged ids into every other runtime journaling into `ecj_path`.
fn publish_to_journal_siblings(
store: &Store,
primary: usize,
vid: VolumeId,
ecj_path: &str,
ids: &HashSet<NeedleId>,
) {
for (i, loc) in store.locations.iter().enumerate() {
if i == primary {
continue;
}
if let Some(ecv) = loc.find_ec_volume(vid) {
if Path::new(&ecv.ecj_file_name()) == Path::new(ecj_path) {
ecv.publish_merged_ids(ids);
}
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::config::MinFreeSpace;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::DiskType;
use crate::storage::volume::{VifEcShardConfig, VifVolumeInfo};
use tempfile::TempDir;
const COLLECTION: &str = "c";
const VID: VolumeId = VolumeId(9);
/// A store with one disk per entry of `data`, all sharing `idx` when given,
/// else each indexing into its own data dir.
fn make_store(tmp: &TempDir, data: &[&str], idx: Option<&str>) -> RwLock<Store> {
let mut store = Store::new(NeedleMapKind::InMemory);
for d in data {
let dir = tmp.path().join(d).to_string_lossy().into_owned();
let idx_dir = idx
.map(|i| tmp.path().join(i).to_string_lossy().into_owned())
.unwrap_or_else(|| dir.clone());
std::fs::create_dir_all(&dir).unwrap();
std::fs::create_dir_all(&idx_dir).unwrap();
store
.add_location(
&dir,
&idx_dir,
100,
DiskType::HardDrive,
MinFreeSpace::Percent(0.0),
Vec::new(),
)
.unwrap();
}
RwLock::new(store)
}
fn dir(tmp: &TempDir, d: &str) -> String {
tmp.path().join(d).to_string_lossy().into_owned()
}
fn records(ids: &[u64]) -> Vec<u8> {
let ids: Vec<NeedleId> = ids.iter().copied().map(NeedleId).collect();
crate::storage::erasure_coding::ecj_merge::encode_ecj_ids(&ids)
}
fn id_set(ids: &[u64]) -> HashSet<NeedleId> {
ids.iter().copied().map(NeedleId).collect()
}
/// Shard 0 of `VID` and its `.vif` in `data_dir`.
fn write_shard0(data_dir: &str) {
let base = format!("{}/{}_{}", data_dir, COLLECTION, VID.0);
std::fs::write(format!("{}.ec00", base), b"shard data nonempty").unwrap();
let vif = VifVolumeInfo {
version: 3,
ec_shard_config: Some(VifEcShardConfig {
data_shards: 10,
parity_shards: 4,
..Default::default()
}),
..Default::default()
};
std::fs::write(
format!("{}.vif", base),
serde_json::to_string(&vif).unwrap(),
)
.unwrap();
}
/// `VID`'s `.ecx` and a `.ecj` holding `deleted` in `dir`; returns the
/// journal path.
fn write_index(dir: &str, deleted: &[u64]) -> String {
let base = format!("{}/{}_{}", dir, COLLECTION, VID.0);
std::fs::write(format!("{}.ecx", base), vec![0u8; 16]).unwrap();
let ecj = format!("{}.ecj", base);
std::fs::write(&ecj, records(deleted)).unwrap();
ecj
}
fn deleted_on(store: &RwLock<Store>, disk: usize, id: u64) -> bool {
store.read().unwrap().locations[disk]
.find_ec_volume(VID)
.expect("mounted")
.is_needle_deleted(NeedleId(id))
}
/// Disks sharing one index directory all hold the same journal path.
/// Copying shards onto a disk that has not mounted `vid` must still reach
/// the sibling runtime holding that journal open.
#[test]
fn shared_index_dir_reaches_sibling_mount() {
let tmp = TempDir::new().unwrap();
let store = make_store(&tmp, &["d0", "d1"], Some("idx"));
write_shard0(&dir(&tmp, "d0"));
let ecj = write_index(&dir(&tmp, "idx"), &[1]);
store.write().unwrap().locations[0]
.mount_ec_shards(VID, COLLECTION, &[0], "")
.unwrap();
assert_eq!(
store.read().unwrap().locations[0]
.find_ec_volume(VID)
.unwrap()
.ecj_file_name(),
ecj
);
let added =
merge_ec_journal(&store, VID, &dir(&tmp, "d1"), &ecj, &id_set(&[1, 2])).unwrap();
assert_eq!(added, 1);
assert!(
deleted_on(&store, 0, 2),
"the mounted sibling must see id 2"
);
assert_eq!(std::fs::read(&ecj).unwrap(), records(&[1, 2]));
}
/// Every runtime holding the journal open must see merged ids in memory.
#[test]
fn shared_journal_reaches_every_holder() {
let tmp = TempDir::new().unwrap();
let store = make_store(&tmp, &["d0", "d1"], Some("idx"));
write_shard0(&dir(&tmp, "d0"));
write_shard0(&dir(&tmp, "d1"));
let ecj = write_index(&dir(&tmp, "idx"), &[1]);
for i in 0..2 {
store.write().unwrap().locations[i]
.mount_ec_shards(VID, COLLECTION, &[0], "")
.unwrap();
}
let added =
merge_ec_journal(&store, VID, &dir(&tmp, "d1"), &ecj, &id_set(&[1, 2])).unwrap();
assert_eq!(added, 1);
assert!(
deleted_on(&store, 0, 2) && deleted_on(&store, 1, 2),
"every journal holder must see the merged id"
);
assert_eq!(std::fs::read(&ecj).unwrap(), records(&[1, 2]));
}
/// The picked runtime may journal to a different file than the copied
/// one — its index lives in its data directory while a sibling's lives
/// in the index directory. The ids must be published only to holders of
/// the file they were written to.
#[test]
fn publishes_to_actual_journal_holders() {
let tmp = TempDir::new().unwrap();
let store = make_store(&tmp, &["d0", "d1"], Some("idx"));
write_shard0(&dir(&tmp, "d0"));
write_shard0(&dir(&tmp, "d1"));
let data_ecj = write_index(&dir(&tmp, "d0"), &[1]);
let idx_ecj = write_index(&dir(&tmp, "idx"), &[1]);
for i in 0..2 {
store.write().unwrap().locations[i]
.mount_ec_shards(VID, COLLECTION, &[0], "")
.unwrap();
}
assert_eq!(
store.read().unwrap().locations[0]
.find_ec_volume(VID)
.unwrap()
.ecj_file_name(),
data_ecj
);
let added =
merge_ec_journal(&store, VID, &dir(&tmp, "d0"), &idx_ecj, &id_set(&[1, 2])).unwrap();
assert_eq!(added, 1);
assert_eq!(std::fs::read(&data_ecj).unwrap(), records(&[1, 2]));
assert_eq!(std::fs::read(&idx_ecj).unwrap(), records(&[1]));
assert!(deleted_on(&store, 0, 2));
assert!(
!deleted_on(&store, 1, 2),
"a different journal's holder must not claim the merged id"
);
}
/// A sibling disk can mount `vid` from the receiving disk's index (#9212)
/// while the merge reads the journal unlocked. The merge must go through
/// that mount and report what it added.
#[test]
fn mount_during_read_is_merged_through_and_counted() {
let tmp = TempDir::new().unwrap();
let store = make_store(&tmp, &["d0", "d1"], None);
let owner = dir(&tmp, "d0");
let ecj = write_index(&owner, &[1]);
write_shard0(&dir(&tmp, "d1"));
let mut mounted = false;
let read = |path: &str| {
let read = read_ecj_ids(path);
if !mounted {
store.write().unwrap().locations[1]
.mount_ec_shards_with_idx_dir(VID, COLLECTION, &[0], &owner, "")
.unwrap();
mounted = true;
}
read
};
let added =
merge_ec_journal_with(&store, VID, &owner, &ecj, &id_set(&[1, 2, 3]), read).unwrap();
assert_eq!(added, 2, "the ids merged through the new mount are counted");
assert!(deleted_on(&store, 1, 2) && deleted_on(&store, 1, 3));
assert_eq!(std::fs::read(&ecj).unwrap(), records(&[1, 2, 3]));
}
/// An unmounted journal is merged on disk and a repeat adds nothing; a
/// data dir that is no disk is refused.
#[test]
fn unmounted_journal_is_idempotent() {
let tmp = TempDir::new().unwrap();
let store = make_store(&tmp, &["d0"], Some("idx"));
let ecj = write_index(&dir(&tmp, "idx"), &[1, 2]);
let d0 = dir(&tmp, "d0");
for want in [2, 0, 0] {
let added = merge_ec_journal(&store, VID, &d0, &ecj, &id_set(&[2, 3, 4])).unwrap();
assert_eq!(added, want);
}
assert_eq!(std::fs::read(&ecj).unwrap(), records(&[1, 2, 3, 4]));
assert!(
merge_ec_journal(&store, VID, &dir(&tmp, "elsewhere"), &ecj, &id_set(&[1])).is_err()
);
}
}
+37 -4
View File
@@ -8,7 +8,7 @@ use std::path::Path;
use tracing::{info, warn};
use crate::storage::disk_location::{parse_collection_volume_id_pub, DiskLocation};
use crate::storage::disk_location::{DiskLocation, parse_collection_volume_id_pub};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -131,6 +131,12 @@ impl Store {
let Some(base) = name.strip_suffix(".ecx") else {
continue;
};
// A 0-byte .ecx is a corrupt stub from a failed copy, not a
// credible owner — skip it so the scan keeps looking for a
// real index on a sibling disk (Go's indexEcxOwners).
if !ent.metadata().is_ok_and(|m| m.len() > 0) {
continue;
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
@@ -286,9 +292,7 @@ fn collect_shard_disk_volumes(loc: &DiskLocation) -> HashMap<EcKey, Vec<String>>
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
out.entry(EcKey { collection, vid })
.or_default()
.push(name);
out.entry(EcKey { collection, vid }).or_default().push(name);
}
out
}
@@ -419,4 +423,33 @@ mod tests {
let post = fs::read(dir0.join(format!("{}_{}.ecx", collection, vid))).unwrap();
assert_eq!(post, ecx_local, "mirror overwrote dir0's existing .ecx");
}
/// The mirror shares Go's indexEcxOwners, which skips a 0-byte `.ecx`:
/// a stub must not be chosen as the source to mirror from.
#[test]
fn mirror_owner_index_skips_zero_byte_ecx() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
fs::create_dir_all(&dir0).unwrap();
fs::create_dir_all(&dir1).unwrap();
let collection = "video-recordings";
let vid = 4123u32;
plant_ecx(&dir0, collection, vid, b"");
plant_ecx(&dir1, collection, vid, &[0xA1u8; 20]);
let mut store = Store::new(NeedleMapKind::InMemory);
add_loc(&mut store, &dir0);
add_loc(&mut store, &dir1);
let owners = store.index_ecx_owners_for_mirror();
let owner = owners
.get(&EcKey {
collection: collection.to_string(),
vid: VolumeId(vid),
})
.expect("the valid .ecx on disk 1 must be indexed");
assert_eq!(owner.location, 1);
}
}
+218 -94
View File
@@ -21,7 +21,7 @@ use std::fs;
use tracing::{error, info, warn};
use crate::storage::disk_location::{is_ec_shard_extension, parse_collection_volume_id_pub};
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
use crate::storage::erasure_coding::ec_shard::{DATA_SHARDS_COUNT, ShardId};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -37,6 +37,7 @@ pub(crate) struct EcVolumeMissingIndex {
pub data_dir: String,
}
#[cfg(test)]
pub(crate) fn ec_local_ecx_path(dir: &str, collection: &str, vid: VolumeId) -> String {
if collection.is_empty() {
format!("{}/{}.ecx", dir, vid.0)
@@ -80,6 +81,11 @@ struct EcxOwnerInfo {
idx_dir: String,
}
/// One unit of reconcile work: the disk holding orphan shards, the volume
/// they belong to, the shard files, the `.ecx` owner, and whether the
/// mirror already installed sidecars locally (`use_local_idx`).
type OrphanShardLoad = (usize, EcKey, Vec<(String, ShardId)>, EcxOwnerInfo, bool);
impl Store {
/// Run cross-disk orphan-shard reconciliation. Should be called
/// after every DiskLocation has finished its per-disk EC scan.
@@ -98,7 +104,7 @@ impl Store {
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
let mut to_load: Vec<OrphanShardLoad> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -112,14 +118,11 @@ impl Store {
);
continue;
};
let local_ecx = ec_local_ecx_path(&loc.idx_directory, &key.collection, key.vid);
let local_ecx_in_data = ec_local_ecx_path(&loc.directory, &key.collection, key.vid);
let use_local_idx = std::path::Path::new(&local_ecx).exists()
|| std::path::Path::new(&local_ecx_in_data).exists();
// A 0-byte local stub is not a mirrored index (Go gates this fast
// path on HasEcxFileOnDisk); mount against the owner instead.
let use_local_idx = loc.has_ecx_file_on_disk(&key.collection, key.vid);
if !use_local_idx
&& owner.location == loc_idx
&& owner.idx_dir == loc.idx_directory
if !use_local_idx && owner.location == loc_idx && owner.idx_dir == loc.idx_directory
{
// Same-disk no-op: load_all_ec_shards already
// tried and logged the failure.
@@ -132,7 +135,7 @@ impl Store {
for (loc_idx, key, shards, owner, use_local_idx) in to_load {
let shard_names: Vec<&str> = shards.iter().map(|(n, _)| n.as_str()).collect();
let loc_dir = self.locations[loc_idx].directory.clone();
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if use_local_idx {
info!(
@@ -293,10 +296,10 @@ impl Store {
// may be sole copies of a distributed volume.
let mut node_wide_bits = ev.shard_bits().0;
for other in &self.locations {
if let Some(other_ev) = other.find_ec_volume(*vid) {
if other_ev.collection == ev.collection {
node_wide_bits |= other_ev.shard_bits().0;
}
if let Some(other_ev) = other.find_ec_volume(*vid)
&& other_ev.collection == ev.collection
{
node_wide_bits |= other_ev.shard_bits().0;
}
}
let node_wide = node_wide_bits.count_ones() as usize;
@@ -419,6 +422,12 @@ impl Store {
let Some(base) = name.strip_suffix(".ecx") else {
continue;
};
// A 0-byte .ecx is a corrupt stub from a failed copy, not a
// credible owner — skip it so the scan keeps looking for a
// real index on a sibling disk (Go's indexEcxOwners).
if !ent.metadata().is_ok_and(|m| m.len() > 0) {
continue;
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
@@ -474,13 +483,13 @@ impl Store {
/// Unlike `reconcile_ec_shards_across_disks` it needs no sibling disk, so a
/// single-disk store recovers once its index has been fetched from a peer.
fn load_orphan_ec_shards_with_local_index(&mut self) {
let mut work: Vec<(usize, EcKey, Vec<u32>)> = Vec::new();
let mut work: Vec<(usize, EcKey, Vec<ShardId>)> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
for (key, shards) in collect_orphan_ec_shards(loc, loc_idx) {
if !loc.has_ecx_file_on_disk(&key.collection, key.vid) {
continue;
}
let ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
work.push((loc_idx, key, ids));
}
}
@@ -499,6 +508,53 @@ impl Store {
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, ShardId)>> {
let mut orphans: HashMap<EcKey, Vec<(String, ShardId)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id)
{
continue;
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
#[cfg(test)]
mod tests {
use super::*;
@@ -541,7 +597,13 @@ mod tests {
std::fs::write(&p, b"shard data nonempty").unwrap();
}
fn write_index_files(idx_dir: &str, collection: &str, vid: u32, data_shards: u32, parity_shards: u32) {
fn write_index_files(
idx_dir: &str,
collection: &str,
vid: u32,
data_shards: u32,
parity_shards: u32,
) {
// Minimal sealed .ecx (the loader only opens the file; it
// doesn't parse it during placement).
std::fs::write(
@@ -583,6 +645,34 @@ mod tests {
.unwrap();
}
/// A 0-byte `.ecx` is not a credible owner (Go's indexEcxOwners skips
/// it): picking the stub would hide the valid index on the sibling disk.
#[test]
fn test_index_ecx_owners_skips_zero_byte_stub() {
let (store, _tmp) = make_test_store(2, None);
let d0 = store.locations[0].directory.clone();
let d1 = store.locations[1].directory.clone();
std::fs::write(ec_local_ecx_path(&d0, "pics", VolumeId(7)), b"").unwrap();
write_index_files(&d1, "pics", 7, 10, 4);
let owners = store.index_ecx_owners();
let owner = owners
.get(&EcKey {
collection: "pics".to_string(),
vid: VolumeId(7),
})
.expect("the valid .ecx on disk 1 must be indexed");
assert_eq!(owner.location, 1);
assert_eq!(owner.idx_dir, d1);
// A stub with no real index anywhere owns nothing.
std::fs::write(ec_local_ecx_path(&d0, "pics", VolumeId(8)), b"").unwrap();
assert!(!store.index_ecx_owners().contains_key(&EcKey {
collection: "pics".to_string(),
vid: VolumeId(8),
}));
}
/// An empty `.dat` (<= a superblock, i.e. zero needles) for an EC volume
/// is a leftover stub from the pre-fix loader. It must be swept on startup,
/// not loaded as a phantom empty volume. With the same vid's stub on two
@@ -895,15 +985,74 @@ mod tests {
// dir1 owns the .ecx and so already has shard 1 mounted via
// its own load_all_ec_shards.
let ev1 = store.locations[1].find_ec_volume(VolumeId(vid));
assert!(ev1.is_some(), "baseline broken: dir1 should have mounted shard 1");
assert!(
ev1.is_some(),
"baseline broken: dir1 should have mounted shard 1"
);
// dir0's shards must be reconciled across to its own
// ec_volumes map, pointing at dir1's idx dir.
let ev0 = store.locations[0]
.find_ec_volume(VolumeId(vid))
.expect("dir0 should now have an EcVolume after reconcile");
assert!(ev0.has_shard(0), "shard 0 missing from dir0 after reconcile");
assert!(ev0.has_shard(12), "shard 12 missing from dir0 after reconcile");
assert!(
ev0.has_shard(0),
"shard 0 missing from dir0 after reconcile"
);
assert!(
ev0.has_shard(12),
"shard 12 missing from dir0 after reconcile"
);
}
/// dir0 holds orphan shards next to a 0-byte `.ecx` stub from a failed
/// copy; the real index is on dir1. The stub must not count as a
/// locally-mirrored index (Go gates that fast path on HasEcxFileOnDisk),
/// or the shards get registered against an empty index.
#[test]
fn test_reconcile_ignores_zero_byte_local_ecx_stub() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
std::fs::create_dir_all(&dir0).unwrap();
std::fs::create_dir_all(&dir1).unwrap();
let collection = "grafana-loki";
let vid = 1094u32;
write_shard(dir0.to_str().unwrap(), collection, vid, 0);
write_shard(dir1.to_str().unwrap(), collection, vid, 1);
write_index_files(dir1.to_str().unwrap(), collection, vid, 10, 4);
let mut store = Store::new(NeedleMapKind::InMemory);
for dir in [&dir0, &dir1] {
store
.add_location(
dir.to_str().unwrap(),
dir.to_str().unwrap(),
100,
DiskType::HardDrive,
MinFreeSpace::Percent(0.0),
Vec::new(),
)
.unwrap();
}
// Plant the stub after the startup scan so only the reconcile decision
// is under test, then drop dir0's mount and reconcile again.
store.locations[0].remove_ec_volume(VolumeId(vid));
std::fs::write(
ec_local_ecx_path(dir0.to_str().unwrap(), collection, VolumeId(vid)),
b"",
)
.unwrap();
store.reconcile_ec_shards_across_disks();
let ev0 = store.locations[0]
.find_ec_volume(VolumeId(vid))
.expect("dir0's shard must be mounted against the owner's index");
assert!(ev0.has_shard(0));
assert_eq!(ev0.ecx_actual_dir(), dir1.to_str().unwrap());
}
/// PR 9244 review case: idx_directory is configured but the
@@ -1012,7 +1161,13 @@ mod tests {
assert!(store.locations[0].find_ec_volume(VolumeId(vid)).is_none());
// Shard files must still exist on disk for operator recovery.
for sid in [0u8, 12u8] {
let p = format!("{}/{}_{}.ec{:02}", dir0.to_str().unwrap(), collection, vid, sid);
let p = format!(
"{}/{}_{}.ec{:02}",
dir0.to_str().unwrap(),
collection,
vid,
sid
);
assert!(
std::path::Path::new(&p).exists(),
"orphan shard {} was destroyed",
@@ -1077,10 +1232,12 @@ mod tests {
assert!(ev1.has_shard(6), "dir1 shard missing");
// Nothing left to recover.
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
#[test]
@@ -1112,10 +1269,12 @@ mod tests {
.unwrap();
}
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
/// Helper: build a 2-disk store where reconcile produces the
@@ -1193,7 +1352,11 @@ mod tests {
let vid = VolumeId(7010);
let all = store.find_all_ec_volumes(vid);
assert_eq!(all.len(), 2, "expected one EcVolume per disk holding the vid");
assert_eq!(
all.len(),
2,
"expected one EcVolume per disk holding the vid"
);
// Disk 0 carries shards 0 and 12; disk 1 carries shard 1.
assert!(all[0].has_shard(0));
@@ -1214,7 +1377,7 @@ mod tests {
#[test]
fn test_scrub_plans_reach_every_disk_through_the_store() {
use crate::storage::erasure_coding::ec_volume::{
merge_ec_runtimes, EcChecksumScrubPlan, EcLocalScrubPlan,
EcChecksumScrubPlan, EcLocalScrubPlan, merge_ec_runtimes,
};
let (store, _tmp) = build_split_disk_store(7030);
@@ -1229,8 +1392,15 @@ mod tests {
let merged = merge_ec_runtimes(&runtimes).expect("two runtimes merge");
assert!(merged.slots[0].is_some(), "disk 0's shard 0 unreachable");
assert!(merged.slots[12].is_some(), "disk 0's shard 12 unreachable");
assert!(merged.slots[1].is_some(), "disk 1's shard 1 unreachable — the bug");
assert!(merged.skipped.is_empty(), "same generation: {:?}", merged.skipped);
assert!(
merged.slots[1].is_some(),
"disk 1's shard 1 unreachable — the bug"
);
assert!(
merged.skipped.is_empty(),
"same generation: {:?}",
merged.skipped
);
// Shard 1 is owned by the sibling runtime, not the anchor.
let (owner, _) = merged.slots[1].unwrap();
@@ -1306,7 +1476,7 @@ mod tests {
let vid = VolumeId(7004);
let collection = "grafana-loki";
store.delete_ec_shards(vid, collection, &[1]);
store.delete_ec_shards(vid, collection, &[1]).unwrap();
// Shard 1 file is gone on disk 1.
let p1 = format!(
@@ -1337,17 +1507,22 @@ mod tests {
let (_ev, dirs) = store.collect_ec_shard_dirs(vid, max_shards).unwrap();
// Shards 0 and 12 → disk 0's directory.
assert_eq!(dirs[0].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(dirs[12].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(
dirs[0].as_deref(),
Some(store.locations[0].directory.as_str())
);
assert_eq!(
dirs[12].as_deref(),
Some(store.locations[0].directory.as_str())
);
// Shard 1 → disk 1's directory.
assert_eq!(dirs[1].as_deref(), Some(store.locations[1].directory.as_str()));
assert_eq!(
dirs[1].as_deref(),
Some(store.locations[1].directory.as_str())
);
// Unmounted shards → None.
for sid in [2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13] {
assert_eq!(
dirs[sid], None,
"shard {} unexpectedly reported a dir",
sid,
);
assert_eq!(dirs[sid], None, "shard {} unexpectedly reported a dir", sid,);
}
}
@@ -1656,11 +1831,7 @@ mod tests {
vec![0u8; 20],
)
.unwrap();
std::fs::write(
ec_dir.join(format!("{}_{}.ecj", collection, vid)),
b"",
)
.unwrap();
std::fs::write(ec_dir.join(format!("{}_{}.ecj", collection, vid)), b"").unwrap();
let mut store = Store::new(NeedleMapKind::InMemory);
store
@@ -1759,50 +1930,3 @@ mod tests {
assert!(std::path::Path::new(&format!("{}.ecx", ec_base)).exists());
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
continue;
}
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
+29
View File
@@ -221,6 +221,35 @@ mod tests {
use super::*;
use crate::storage::types::*;
/// Multi-byte input must be an error, not a panic: `to_digit` on the
/// leading characters rejects it before `chars[2]` is ever indexed.
#[test]
fn replica_placement_rejects_non_ascii_instead_of_panicking() {
for s in ["é", "0é", "é0", "🦀", "ééé"] {
assert!(
ReplicaPlacement::from_string(s).is_err(),
"non-ASCII replication {:?} must error",
s
);
}
}
/// The ASCII guard must not change any accepted input, including the
/// zero-padding shorthands.
#[test]
fn replica_placement_still_accepts_ascii_shorthands() {
assert_eq!(
ReplicaPlacement::from_string("1").unwrap(),
ReplicaPlacement::from_string("001").unwrap()
);
assert_eq!(
ReplicaPlacement::from_string("01").unwrap(),
ReplicaPlacement::from_string("001").unwrap()
);
let rp = ReplicaPlacement::from_string("010").unwrap();
assert_eq!(rp.diff_rack_count, 1);
}
#[test]
fn test_super_block_round_trip() {
let sb = SuperBlock {

Some files were not shown because too many files have changed in this diff Show More