Compare commits

..
Author SHA1 Message Date
Chris LuandDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com> d6335c5e94 volume: fix ReadOnly match arm missed by the tuple-variant change
#11544 made VolumeError::ReadOnly carry the volume id; the batch-write
fallback in write_volume_needles still matched and built it as a unit
variant, breaking the build (E0532/E0308).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-01 06:24:02 -07:00
9d3907e36c volume: say "volume N is read only" like Go, so filer retries deletes (#11544)
VolumeError::ReadOnly displayed "volume is read-only". Go's store and
volume say "volume %d is read only", and the filer's deletion classifier
requeues a failed delete only when the error contains "is read only".
Against a Rust volume server a BatchDelete on a read-only volume (tier
move, maintenance) was booked as a permanent failure and the chunk was
never deleted.

ReadOnly now carries the volume id and displays Go's text. The text
reaches clients through BatchDelete results, the HTTP write and delete
error bodies, and gRPC statuses; the gRPC code (FailedPrecondition) and
the HTTP/BatchDelete status codes are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 21:00:04 +08:00
bd953b0f84 volume server: group-commit fsync writes in the write queue (#11543)
* volume: split the write path into reusable steps

do_write_request ran its pre-append checks, the append, the sync
rollback, the index publish and the post-write bookkeeping inline, so a
batched write could only reuse it one needle at a time. Pull the steps
out (check_writable, prepare_write, undo_unsynced_append, publish_write,
finish_write) and the store's volume lookup plus disk-space check
(writable_volume_mut). do_write_request composes them in the same order
with the same early returns; no behaviour change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: group-commit fsync writes in the write queue

The write queue holds one store lock for a batch of up to 128 needles
but wrote them one at a time, so every fsync needle paid its own .dat
sync and its own .idx sync: 2N syncs per batch.

Add Volume::write_needles_grouped, after Go's processBatch. A volume's
entries are split into runs of distinct needle ids (a repeated id starts
a new run, so its dedup and cookie checks see the earlier write). A run
with a durable entry appends everything with append_at_ns chained
through a local, syncs the .dat once, and only then publishes the
entries and syncs the .idx once. A failed .dat sync truncates the .dat
back to the run start (marking the volume unavailable if that fails),
leaves last_append_at_ns and last_modified untouched, and fails every
entry of the run. Runs with no durable entry go through the unchanged
per-needle path. Store::write_volume_needles is the queue's entry point;
the handlers' non-queue path is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: fail closed when a failed append's rollback fails

append_needle discarded the truncate-back result, so a partial write
that could not be rolled back left unindexed bytes on the .dat while
the volume stayed writable; the next append would bury them mid-file,
past the load-time tail check. Route the rollback through
undo_unsynced_append, which marks the volume unavailable when the
truncate fails, so nothing more is appended over an unverified tail.

* volume server: keep a grouped run's I/O error streak from later appends

A synced run stages every append before any entry finishes, so the
success reset in finish_write ran after the failed appends queued
behind the last write to land and erased their media-error streak.
Sent one at a time, those errors would have counted and quarantined
the volume. Skip the reset when an append after the last landed write
added to the streak.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: replay a grouped run's I/O error streak in queue order

Skipping the run's success reset whenever an append after the last
landed write failed kept the errors from before that write as well,
so a run like [EIO, EIO, landed, EIO] reached the quarantine count
that the same writes one at a time (one error) do not. Mark the
streak where each entry is staged and record the run's success at
the last landed write's mark: errors before it are cleared, the ones
after it still count.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: replay a run's I/O error streak in one atomic step

Reads record their outcomes on the tracker without the volume's write
lock, so record_success_at's separate load and store could drop an
error a read counted in between, or restore errors a read had just
cleared. Keep the count and the clear counter in one atomic word and
apply the replay with a single fetch_update.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 20:59:40 +08:00
7944cb4ba2 volume server: keep repeated response headers when proxying a read, like Go (#11542)
In readMode=proxy, proxy_request copied the target's response headers
with HeaderMap::insert, so a header the target sent more than once
(several Set-Cookie, Vary, Link, ...) reached the client with only its
last value. Go's proxyReqToTargetServer adds every value with
w.Header().Add. Append instead of insert; the Server header is still
dropped and status and body handling are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:58:50 +08:00
15d3c65e0c volume server: take the heartbeat's remaining store reads off the runtime (#11540)
* volume server: collect EC heartbeats and adjust volume max off the runtime

The volume pass moved to the blocking pool, but the heartbeat task still
called collect_ec_heartbeat and the following EC shard snapshot, and
Store::maybe_adjust_volume_max, directly on its tokio worker.
maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat
of every writable volume under the store read lock. All of them block the
worker on the node-wide RwLock<Store> whenever a writer holds it or is
queued, and every task sharing that worker stalls with it.

Run the adjustment, on the pulse and after the master changes volume
options, and the EC tick's heartbeat plus shard snapshot through
off_runtime, like the volume pass. apply_master_volume_options now only
reports whether the options changed; the loop adjusts off the runtime.
What is collected and sent, and in what order, is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: keep EC shard deltas out of the volume heartbeat's snapshot

EC shard deltas are the diff between the store's shards and last_ec_shards,
taken when volume_state_notify fires. But the volume tick and the
options-changed heartbeat re-took last_ec_shards from the store too, and a
volume heartbeat carries no shard list: a mount or unmount that landed while
the notify was pending or the volume pass was collecting was absorbed into
the baseline and never sent. The master only learned of it at the next EC
tick, 17 pulses later. The EC tick likewise built its heartbeat and its
baseline under two separate store guards, so a mount between them was lost
the same way.

A volume heartbeat now only takes out of the baseline the expired EC shards
it reports deleted itself, so the next delta does not repeat them. The EC
tick, and the initial EC heartbeat, build the full list and the baseline
under one read guard, still on the blocking pool.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: take the heartbeat's remaining store reads off the runtime

The heartbeat's volume and EC passes and the volume max adjustment already
run on the blocking pool, but several reads of the node-wide RwLock<Store>
were still taken directly on the heartbeat's tokio worker: the digest
report reset before the first heartbeat, the duplicate-UUID directory
lookup and the volume options a master response carries, the EC shard list
a state notification is diffed against, and the deregistration heartbeat
sent on stop and shutdown. The lock is writer-preferring, so with a writer
holding or queued for it each of these parks the worker, and every task
sharing that worker stalls with it.

Run each through off_runtime, which now takes a closure so a pass can carry
what it needs from the master's response. The notify branch's volume
snapshot and EC read become one blocking pass, still under two guards in
the same order. What is collected and sent, and in what order, is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-10-01 20:58:01 +08:00
fea14c01a7 volume server: refuse a tier move while compacting, and a commit once tiered (#11539)
* volume server: refuse a tier move while compacting, and a commit once tiered

A tier move to remote and a vacuum compaction of the same volume could
interleave and leave the volume unreadable:

- A compaction committing while the upload ran swapped .dat/.idx under
  the transfer, which reopens the .dat by path per part. The move then
  published an object holding the old (or a mixed) layout against the
  compacted .idx, and with keep_local_dat_file=false deleted the only
  compacted .dat.
- A tier move finishing while the compaction copy ran (or between the
  copy and the commit) let the commit swap in the compacted .idx while
  the reload served the pre-compaction remote object through it.

The tier move now refuses to start while the volume is compacting, and
re-checks the compaction revision under the store write lock before it
records the remote file; on a mismatch it deletes the uploaded object
and fails with FailedPrecondition, leaving the volume local. Committing
a compaction on a volume that has a remote file is refused and its
.cpd/.cpx removed, since the reload would read the remote object
through the compacted index.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: abort a tier move whose volume was replaced or removed

The tier-up bookkeeping looked the volume up by id only and compared the
compaction revision. A delete and re-create of the same id during the upload
yields a fresh volume at the same revision, so the move recorded the old
volume's object on the new one and, without keep_local_dat_file, removed the
new .dat. An unmounted volume was skipped and the move reported success,
leaving the uploaded object referenced by nothing.

Capture the volume instance (its data-file access control Arc, as the scan
and read plans do) with the revision, and require both under the store write
lock. A replaced volume fails with FailedPrecondition, a missing one with
NotFound; either way nothing is recorded and the object is deleted after the
lock is released. Go fails in both cases because deleting or unmounting closes
the descriptor its copy reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:20:16 +08:00
f2498e122a volume: read EC shards fully, like Go's ReadAt (#11537)
EcVolumeShard::read_at and the scrub plan's EcLocalShard::read_at were a
single pread/seek_read. That may legally return fewer bytes than asked
mid-file (FUSE/NFS/CIFS mounts, a signal, very large requests), and an
Interrupted error was not retried. Callers treat a short count as end of
file or corruption: verify_ec_shards compared a zero tail and reported a
parity mismatch, local scrub reported a broken shard, VolumeEcShardRead
ended the stream early, and decode/rebuild/local needle reads failed.

Add storage::io::read_full_at, which loops until the buffer is full or a
read returns 0 and retries Interrupted, so a short count means EOF. Route
both shard read_at methods through it, replace the encoder's private
read_at_most with it, and reuse it for the Windows read_exact_at loop.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:19:43 +08:00
62481f1673 volume server: ec.decode reads the .ecx from the index dir it was copied to (#11536)
* volume server: ec.decode reads the .ecx from the index dir it was copied to

VolumeEcShardsCopy writes the .ecx/.ecj into the receiver's -dir.idx, so
with a split data/index dir the decode target has no .ecx beside its
shards. VolumeEcShardsToVolume sized the .dat from the right .ecx but
built the .idx from the data dir, failing with NotFound after the .dat
was already published. It now reads .ecx/.ecj from where the EC volume
opened them and writes the .idx beside the .dat, where Go leaves it.

The live-entry check and the .dat size also ignored deletions recorded
only in the .ecj, which Go folds into the .ecx (RebuildEcxFile) first:
a fully deleted volume was decoded instead of reported as having no live
entries, and deleted tail needles were copied into the .dat. Both now
treat journaled ids as deleted, without rewriting the sealed .ecx.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: ec.decode keeps the decoded volume writable and reads every .ecj

The rebuilt .idx copied a journaled tail needle's .ecx row verbatim after
the .dat was cut short before it, so the mount saw a row past EOF and
marked the decoded volume read-only. Rows of deleted needles the .dat no
longer holds are now dropped, and each journaled needle still in the .dat
gets one tombstone instead of one per journal entry.

VolumeEcShardsCopy appends journals collected from other holders into
the idx dir, but the decode read only the .ecj beside the .ecx, which
sits in the data dir when this server generated the shards. It now
reads both, once, in bounded chunks via the loader EcVolume uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: test ec.decode drops a sealed .ecx tail tombstone

Covers the other half of the rule added in the previous commit: a tail
needle tombstoned in the .ecx itself (Go's RebuildEcxFile) is cut from
the .dat, and its row must not reach the rebuilt .idx either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:19:17 +08:00
b8f074b7d3 volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go (#11535)
* volume server: VolumeNeedleStatus reads remote EC shards and reports deleted needles like Go

For an EC volume the handler read only locally mounted shards, so a node
that did not hold the shard with the needle's bytes answered Internal
"ec shard N not available locally". Go's ReadEcShardNeedle fetches the
interval from a peer or reconstructs it. It also mapped every regular
volume read error, including a tombstone, to NotFound "needle not found",
which fs.verify treats as a missing needle; Go returns ErrorDeleted as a
plain error ("already deleted"), which fs.verify skips.

The EC branch now drops the store guard and uses the distributed EC read
the HTTP GET path uses. Errors map like Go: needle absent -> NotFound
"needle not found <decimal id>", tombstoned (regular or EC .ecx/.ecj) ->
Unknown "already deleted", anything else -> Unknown with the error text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: tell EC deletions and vanished volumes apart in VolumeNeedleStatus

The distributed EC reader returned Ok(None) for an absent needle, a needle
a peer reported deleted, and a volume unmounted after the handler's own
existence check. VolumeNeedleStatus answered all three NotFound "needle not
found", which fs.verify -pruneEntries counts as lost data. A reported
deletion was also lost when an earlier interval failed.

The reader now says why it has no needle (EcMiss: NotFound, Deleted,
VolumeNotFound), classifying the local tombstone itself and letting a
reported deletion outrank other interval errors, as Go's ReadEcShardNeedle
does. VolumeNeedleStatus maps Deleted to Unknown "already deleted" and
VolumeNotFound to "volume not found", and drops its separate EC pre-check.
read_ec_shard_needle_distributed keeps its Ok(None) for every miss, so the
other callers are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:18:34 +08:00
cc1ec48151 volume server: collect EC heartbeats and adjust volume max off the runtime (#11532)
The volume pass moved to the blocking pool, but the heartbeat task still
called collect_ec_heartbeat and the following EC shard snapshot, and
Store::maybe_adjust_volume_max, directly on its tokio worker.
maybe_adjust_volume_max runs statvfs on every auto-sized disk and stats the .dat
of every writable volume under the store read lock. All of them block the
worker on the node-wide RwLock<Store> whenever a writer holds it or is
queued, and every task sharing that worker stalls with it.

Run the adjustment, on the pulse and after the master changes volume
options, and the EC tick's heartbeat plus shard snapshot through
off_runtime, like the volume pass. apply_master_volume_options now only
reports whether the options changed; the loop adjusts off the runtime.
What is collected and sent, and in what order, is unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:14:15 +08:00
5ece8dd63c volume server: drop the unused unmount result in VolumeCopy validation, and test-only EC helpers (#11531)
mount_and_reply ignored the Result of store.unmount_volume when a copied
replica failed record count validation, tripping unused_must_use. The Err
branch is unreachable there: the volume was mounted under the same store
write guard, mount_volume refuses an already loaded vid so it is a fresh
Volume with is_compacting false, and a compaction claim needs &mut Volume,
i.e. the store lock. Ignore the result explicitly with a one-line reason.

Store::delete_expired_ec_volumes and Store::remove_ec_volume are called
only from test modules (the heartbeat uses the split
find_expired_ec_volumes / remove_expired_ec_volumes halves), so mark them
#[cfg(test)].

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 20:13:45 +08:00
github-actions[bot] 67f7fad987 docs: regenerate star history chart 2026-10-01 00:56:45 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2b057eccbf build(deps): bump com.fasterxml.jackson.core:jackson-databind from 2.22.2 to 2.22.3 in /test/java/spark (#11548)
build(deps): bump com.fasterxml.jackson.core:jackson-databind

Bumps [com.fasterxml.jackson.core:jackson-databind](https://github.com/FasterXML/jackson-databind) from 2.22.2 to 2.22.3.
- [Commits](https://github.com/FasterXML/jackson-databind/compare/jackson-databind-2.22.2...jackson-databind-2.22.3)

---
updated-dependencies:
- dependency-name: com.fasterxml.jackson.core:jackson-databind
  dependency-version: 2.22.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 01:43:45 +08:00
Javier GarciaandGitHub 8fdcf69eb0 s3api: report the stored checksum in GetObjectAttributes (#11529)
GetObjectAttributes accepted the Checksum attribute but never filled it
in, as its comment said SeaweedFS did not store S3 checksums. PutObject
and CompleteMultipartUpload store them now, and HeadObject returns them.
Fill in Checksum from the same entry fields, with the ChecksumType and
ChecksumCRC64NVME members the response did not have.

Also run ceph/s3-tests' test_get_checksum_object_attributes in CI.
2026-10-01 01:11:19 +08:00
Javier GarciaandGitHub 988fc4f7ba s3api: do not store aws-chunked in an object's Content-Encoding (#11528)
* s3api: do not store aws-chunked in an object's Content-Encoding

aws-chunked in Content-Encoding names the SigV4 streaming framing of
the request body, which the gateway decodes on upload. PutObject and
CreateMultipartUpload stored the header as sent, so an object uploaded
with "gzip, aws-chunked" was served with that Content-Encoding, and one
uploaded with "aws-chunked" alone was served as aws-chunked. S3 drops
aws-chunked and keeps the other encodings.

Also run ceph/s3-tests' test_object_content_encoding_aws_chunked in CI.

* s3api: read every Content-Encoding field, and drop aws-chunked on copy

A client can send aws-chunked and the object's own encoding as separate
Content-Encoding fields. Only the first was read, so "aws-chunked"
followed by "gzip" left the object without its gzip. Combine all the
fields before dropping aws-chunked. CopyObject with the REPLACE
directive stored the requested Content-Encoding as sent: drop
aws-chunked there too.
2026-10-01 01:10:38 +08:00
ihnokimandGitHub 11e8c4c288 master: follow heartbeat read-only changes in the layout's replica flag (#11527)
A replica's read-only flag in the volume layout only moved on registration
and on volume.mark. A change that arrived in the regular heartbeat updated
the node's record, which the writable list follows, but not the layout
flag, which the vacuum sweep reads. So the sweep kept trying volumes on a
disk that had gone read-only while the server ran, and after a restart it
skipped volumes that had since become writable again until the next
restart (issue #11516).

Apply the reported state to the flag for every changed volume. Only the
flag: the writable list stays with EnsureCorrectWritables and its
capacity guards.
2026-10-01 01:09:09 +08:00
ihnokimandGitHub d3cd061c22 shell: say which read-only volumes volume.vacuum leaves alone (#11525)
* shell: say which read-only volumes volume.vacuum leaves alone

volume.vacuum without -volumeId runs the same sweep as the automatic
vacuum, which skips read-only volumes, and the master's response carries
no result. An operator whose disk filled up runs the command, sees it
return, and watches nothing change (issue #11516).

Before issuing the request, list the read-only volumes whose garbage is at
or above the threshold and point at -volumeId, which is the explicit path
PR #9861 opened for them. The help text says the same.

* shell: volume.vacuum hint survives a failed listing and looks at every replica

Review follow-ups: a failed topology listing no longer stops a sweep
without -volumeId, it only drops the hint; a volume counts as read-only
when any replica is, with the garbage ratio taken from the replica that
reports the most, which is what the sweep itself does; a converted index
that reports deletes without sizes is listed rather than hidden; and the
threshold is printed as given instead of rounded to two decimals.

* shell: do not guess a garbage ratio for a converted index

The master cannot compute one for a volume that reports deletes without
their sizes, and a guess of 1 would send the operator to -volumeId for a
volume the server may decline at that threshold. Leave it out and say so.
2026-10-01 01:08:48 +08:00
895d49b55b s3/iam: manage roles through the IAM API, with an opt-in persistent role store (#11522)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store

Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.

Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
  IAM config file the default stays memory; with none it is the filer, as
  for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
  which outlives the file and may be shared by S3 servers with different
  files. They are served from memory beneath the store, as OIDC providers
  are: a stored role of the same name takes precedence, and deleting it
  restores the file's. A config-file role cannot be changed or deleted
  through the API (UnmodifiableEntity), and removing one from the file
  removes it at the next start. An in-memory store holds them as records,
  as before. They have no creation time, so CreateDate is omitted rather
  than reporting when this server started. SetRoleStore installs a store
  the same way, so a store set after startup keeps the config-file roles,
  as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
  cached filer store otherwise serves a peer's stale role for up to its 5m
  TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
  used to report any failed lookup as "role not found". CreateRole proceeds
  only on a confirmed absence, so an unreadable store cannot let it write
  over an existing role.

IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
  AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
  are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
  changing or deleting it through the API is refused (UnmodifiableEntity)
  rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
  <name>.json in the filer, so this also keeps a name from leaving the role
  store's directory. At most 10 managed policies per role (AWS's default
  quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
  policy, as it already is for users and groups: roles attach policies by
  name, so a policy created later under the deleted one's name would
  otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
  refused rather than dropped.

Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
  kept across updates; a config-file role gets a stable ID derived from its
  name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
  and AssumeRole carry the role's ID (claim "rid"), and a request under a role
  whose current ID differs is denied. Resolving a session's policies by role
  name let a session outlive its role: once a role was deleted, a role later
  created under the same name — with a different trust policy and different
  policies — revived every unexpired session of the old one with the new
  role's permissions. Sessions issued before this change carry no ID and are
  unaffected until they expire.

Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: bind every role session to its role, and change roles atomically

Review follow-ups.

Session binding
- The role-ID check ran only when a session carried no policy names, and
  AssumeRole embeds the role's attached policies, so those sessions kept
  their permissions after the role was deleted or recreated. The check
  now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
  instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
  the name alone: a different role put in the file under the same name
  gets a new ID, while an unchanged role keeps its sessions across restarts.

Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
  role is unchanged since the read, and otherwise re-reads and retries. The
  filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
  role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
  Attach/DetachRolePolicy all go through it, so two gateways no longer
  overwrite each other's changes, a change racing a delete no longer
  writes the role back, and of two concurrent creates one gets
  EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
  broken stream instead of returning what arrived, so DeletePolicy's
  attachment check sees every role. ListRoles skips a role deleted between
  listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
  InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
  comment now says so. Role and OIDC provider actions write their own stores.

Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path

Review follow-ups.

- Authorization evaluates the policies of the role definition the session's
  binding was checked against, instead of reading the role again: a role
  replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
  definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
  and take its ID, duration cap and embedded policies from that same
  definition. A role replaced after the caller's trust check by one that does
  not trust the caller now yields AccessDenied, not a session bound to the
  replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
  write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
  routed and locked like the conditional CreateEntry. DeleteRole decides
  against the role it deletes, so a policy attached meanwhile on another
  server is a DeleteConflict, and a delete never removes a role written
  after its check.
- S3 servers watch the role store's configured basePath, not only
  /etc/iam/roles, so a custom path also drops peers' cached roles on change.

Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: state which roles DeletePolicy's attachment check can see

RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* iam: note that a role store set after startup is not watched for peer changes

S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:45:47 +08:00
Khris RichardsonandGitHub 95e0b74fb6 s3/iam: retry a failed OIDC provider refresh until the store answers (#11521)
* s3/iam: retry a failed OIDC provider refresh until the store answers

RefreshOIDCProvidersFromStore reports a failure and nothing retries it. Its
callers can't: a metadata-subscription event reports each change once, so a
refresh that found the filer unreachable on it (the filer restarting, say)
left a peer's new provider untrusted, or a deleted one trusted, until some
unrelated later change. The refresh after a local IAM API mutation has the
same shape. Only the startup load retried.

A failed refresh now retries in the background with the startup load's
backoff until the store answers. At most one retry runs, however many
refreshes fail meanwhile, and installing another store cancels it. The
startup load uses the same path instead of its own.

Seen on a SeaweedFS operator cluster whose filer restarted while an
S3OIDCProvider was created: the gateway logged "OIDC provider refresh after
/etc/iam/oidc-providers change failed: ... fail to dial". The operator's
periodic re-apply happened to recover it; an IAM API client would not.

* s3/iam: never retry or apply a superseded OIDC provider store, and never drop a failure during a retry

Review of the retry (#11521) found two ways to lose the state it protects.

A refresh of store A that failed as store B was installed could start a
retry for A after B's install had cancelled retries. Nothing cancelled it,
and when A answered it replaced B's providers in STS. The installed store
now changes under the retry lock, a store that is no longer current gets no
retry, and a snapshot of a replaced store is never handed to STS, even when
the refresh listed it just before the swap.

A refresh that failed while a retry ran was dropped by the at-most-one
guard, though the retry might already have listed an older snapshot, so the
change the failed refresh would have loaded stayed unloaded. The retry now
runs once more after its success when a failure arrived meanwhile.

Each has a test that fails without its guard.
2026-09-30 17:33:40 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
0978e7f833 vacuum: keep disk-full read-only volumes reclaimable (#11519)
* storage/topology: keep disk-full read-only volumes vacuumable

The vacuum sweep skipped every read-only replica, so a volume that went
read-only because its disk filled could never reclaim its garbage — the
exact situation compaction exists for. The volume server now reports
disk_space_low in VacuumVolumeCheckResponse, and the sweep skips a
read-only replica only when the flag is clear. An explicit volumeId
vacuum is unaffected: it already bypassed the read-only rule.

The field takes number 4: 2 and 3 are downstream-allocated for tombstone
retention, keeping the wire merge clean.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: measure vacuum free space against live bytes

The pre-compaction space check required the current .dat + .idx size
free, which includes the garbage being reclaimed — on a nearly full disk
that estimate can never fit, so the volume stayed garbage-bound forever.
Measure against the estimated compacted output instead: superblock plus
live index entries plus live content bytes, with the existing ten
percent buffer unchanged. Mirrors the same check in the Rust volume
server.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: count per-needle framing in the compacted-size estimate

The live-bytes estimate covered each live needle's content and index
entry but not its .dat framing (header, checksum, timestamp, padding —
~32 bytes on version 3). For small-needle volumes that is more than the
10% headroom, so a disk with space between the estimate and the real
output still ran out mid-compaction. Rust side mirrors the same formula.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* storage: report disk_space_low only when it is the sole read-only cause

Review feedback (ihnokim, greptile, devin): a volume read-only for low
disk space AND an operator mark or I/O quarantine was still eligible for
the automatic sweep, rewriting a copy meant to stay protected. The flag
now reports only the benign sole-cause case in both servers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* topology: fail closed when the read-only lookup misses in the sweep

A heartbeat can drop the volume from the DataNode cache between the
location-list copy and VacuumVolumeCheck; a lookup error previously
skipped the read-only check entirely. Review feedback (coderabbit).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-30 17:32:41 +08:00
Alex HuandGitHub 38ce95d960 s3api: always write XML timestamps with three fractional digits (#11520)
CopyObject responses carried LastModified values such as
"2026-09-29T20:30:04.56Z": trailing zeros of the fractional seconds were
trimmed, and a whole-second value had no fraction at all. AWS S3 always
writes exactly three digits ("...04.560Z"), and clients that parse with a
fixed-width pattern reject anything else. minio-java 8.6.0
(yyyy-MM-dd'T'HH:mm:ss.SSS'Z') throws DateTimeParseException, so roughly
one CopyObject in ten fails on the client even though the copy succeeded.

Two causes:

- xsdDateTime marshalled with "2006-01-02T15:04:05.999999999", which
  drops trailing zeros. It now writes UTC with ".000Z".
- CopyObjectResult.MarshalXML had a pointer receiver, but the handlers
  pass the result by value, so encoding/xml never called it and fell back
  to time.Time's RFC 3339 encoding. It now has a value receiver.
  CopyPartResult had no custom marshaller at all; it now uses xsdDateTime.

Follow-up to #8394 / #8398, which truncated these timestamps to
milliseconds but kept the trimmed format.
2026-09-30 16:22:31 +08:00
github-actions[bot] 9a454bbe2b docs: regenerate star history chart 2026-09-30 00:53:03 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
757917f564 filer: evict remote-cached objects under storage pressure (#11515)
* filer: identify remote-mounted entries safe to drop under disk pressure

ListEvictableRemoteEntries walks every mounted directory directly on the
filer store (no lazy remote listing) and returns entries that hold local
chunks fully synchronized with remote, ordered oldest-cached first.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: evict remote-cached chunks oldest-first and vacuum the garbage

uncacheRemoteEntry applies the same transition remote.uncache does -
cleared chunks plus a reset LastLocalSyncTsNs under the entry path lock -
and evictRemoteCachedEntries serializes passes over all mounts until a
byte target is met. Aged victims are preferred; a second pass accepts any
synchronized cached entry when aged ones cannot cover the request, since
a failed read is worse than a dropped hot object.

Cleared chunks only become disk space after compaction, so
reclaimRemoteCacheSpace pairs each pass with a rate-limited VacuumVolume
call that also picks up orphaned partial fills.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: trigger remote cache eviction under storage pressure

A periodic check (30s) reads disk usage from master topology and evicts
remote-mounted cached chunks once any disk crosses
-filer.remoteCacheEvictThreshold (default 0.9; 0 disables), with a vacuum
pass to reclaim the tombstoned needles.

The cold-read cache path also kicks the same reclaim when a fill fails on
exhausted volumes - the request still falls back to streaming from the
remote, but the cache stops being permanently wedged full.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: flush deletion queue before remote cache vacuum

Vacuum ran immediately after eviction while evicted file IDs still sat
in the asynchronous deletion queue, so compaction saw no garbage and the
cache stayed wedged. Flush the queue synchronously first and shorten the
vacuum cooldown so sustained pressure does not wait five minutes between
reclaim passes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: cover remote cache eviction under capacity pressure

Unit tests pin the eligibility filter and oldest-first ordering; the
integration test runs a constrained two-node setup that saturates the
cache, verifies the oldest synced entry is evicted and vacuumed, and
that a later read re-caches it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: coalesce remote cache reclaim passes

A failed cache fill used to queue behind any in-flight eviction,
stacking full mount traversals during a write-failure storm. Skip the
pass when one is already running; the caller falls back to streaming
from remote regardless.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: stop the remote cache janitor on shutdown

The eviction ticker kept running after Shutdown closed the metadata
store and could traverse a closed store. Give the janitor a context
cancelled from Shutdown and propagate it into its master RPCs and
traversals.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: vacuum only tombstoned volumes and retry deferred passes

VacuumVolume with no volume id swept every collection, compacting
volumes unrelated to the cache fill that failed. Now the reclaim path
collects the vids of file ids actually flushed from the deletion queue
and compacts only those. Vids that land inside the vacuum cooldown stay
in a pending set the janitor retries on each tick, so chunks evicted
just after a sweep are not stranded until the next pressure event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: count only pressured disks when evicting remote cache

The janitor measured the largest excess on one disk but let bytes on
healthy disks satisfy the reclaim target. Split the topology disk view
per physical disk and count only chunk bytes whose volumes sit on an
over-threshold disk; entries contributing nothing there are skipped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: compare remote cache sync time at nanosecond precision

Second-precision mtime comparisons let a local write in the same second
as the last sync still qualify as evictable, discarding unsynced
changes. Compare LastLocalSyncTsNs against full-precision mtime
(mtime_ns round-trips through the entry codec), and apply the same fix
to remote.uncache's inline check.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: invalidate remote sync stamp on local content change

A local overwrite that keeps the remote entry's LastLocalSyncTsNs looks
evictable even though the remote copy no longer matches, and some write
paths stamp mtime at second precision so a timestamp comparison cannot
catch it. UpdateEntry now clears the stamp when chunks change without a
fresh stamp, leaving replicated updates authoritative.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound remote cache master rpcs and vacuum all evicted garbage

VolumeList and VacuumVolume now run under a 30s context so a stalled
master cannot wedge the eviction janitor. The targeted vacuum drops the
garbage threshold so volumes with under 10% deleted bytes still compact.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tolerate straggler fills in remote cache eviction test

Detached fills from the concurrent wave keep racing the final checks:
live chunks legitimately fill both volumes, and a re-cached object can
be evicted again before its commit is observed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: start remote cache eviction loop after filer init

The janitor's first tick dereferences fs.filer; starting the goroutine
before NewFiler assigns it could panic when startup exceeds an interval.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep remote cache vacuum intent across retries

Evicted entries now record their chunk volumes for vacuum directly, so
the intent survives whoever consumes the shared deletion queue first.
A pending volume keeps several vacuum attempts so tombstones that land
late are still compacted, and the janitor retries pending volumes under
the reclaim mutex instead of flushing unrelated deletes every tick.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound each remote cache vacuum request independently

A shared 30s deadline across pending volumes let one slow compaction
cancel the rest. Each VacuumVolume now gets its own context, and pending
volumes keep more attempts since the master reports request acceptance
rather than compaction.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: tighten remote cache reclamation bound

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: treat chunk timestamp changes as content changes

chunksEqual now also compares ModifiedTsNs so an update that rewrites a
chunk record still invalidates the remote sync stamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: run remote cache queue flush under the reclaim context

BatchDelete for flushed file ids now uses the caller's context instead of
context.Background(), so a reclaim pass bounded by shutdown or timeout
stops its deletes too. Other callers keep their existing behavior.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote cache vacuum to evicted volumes

The flush no longer feeds the shared deletion queue's ids into the
pending set — only evicted chunks' volumes are tracked, so ordinary
deletions no longer pick up repeated vacuum attempts. The flush also
runs under a shutdown-immune bounded context and is skipped when no
volume is pending.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retry remote cache vacuums even after unmount

Pending volumes were only retried while a remote mount existed; removing
the last mount skipped every later pass and left evicted bytes allocated.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: reclaim partial cache fills that run out of capacity

A fill that fails midway queues its written chunks for deletion, but
when no entries remain evictable the reclaim pass found no pending
volumes and skipped the flush and vacuum entirely, leaving the partial
garbage to the slow periodic vacuum while the disk stayed full. Mark
the failed fill's chunk volumes pending so the pass tombstones and
compacts them even when nothing was evicted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 22:05:42 +08:00
Sascha KühndelandGitHub a033ab76e5 helm: add base dns names to admin cert (#11518) 2026-09-29 21:05:24 +08:00
Ishwar KanseandGitHub 5fcf12c080 docker: publish release and latest images for linux/ppc64le and linux/s390x (#11517)
Add ppc64le and s390x to the normal and large_disk release image
matrices, next to arm/v7 and 386, and to the manual container_latest
fallback so that rebuilding `latest` does not drop them from the
manifest.

The Go binary is cross-compiled without cgo on the build host, and
Dockerfile.go_build already falls back to empty Rust volume server and
worker placeholders on platforms without pre-built Rust binaries. Only
the small final-stage RUN steps run emulated, so these two platforms use
QEMU on ubuntu-latest like arm/v7.

Refs #5558
2026-09-29 21:04:19 +08:00
Khris RichardsonGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Opus 5.5Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
5da137233d s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart (#11510)
* s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart

The S3 server's IAM config loader never read the documented
`oidcProviderStore` key, so the OIDC provider store was always in memory:
a provider created with CreateOpenIDConnectProvider lived in one gateway's
process, was lost on restart, and was never seen by peers. The
/etc/iam/oidc-providers metadata subscription refreshed from that empty
in-memory store.

- Read `oidcProviderStore` and pass it to the IAM manager. With an IAM
  config file the default stays memory. With no config file (zero-config
  IAM, as `weed filer -s3` and operator-managed clusters run) it defaults
  to the filer: there is nothing static to shadow, and providers created at
  runtime otherwise vanish on restart.
- With a store that outlives the process, load the STS runtime view from it
  at startup, so providers created on an earlier boot or on a peer are
  trusted without waiting for the next mutation.
- If the store cannot be read at startup (a filer not up yet), the load is
  retried in the background with backoff until it succeeds: the metadata
  subscription reports only later changes, so providers already stored would
  otherwise stay unknown to STS until one of them changed.
- Mark records mirrored from STS.Providers as `source: static-config`, and
  at startup delete such records whose provider has left the config, so
  removing a provider from the config file still revokes it. Records created
  through the IAM API are never pruned.
- The filer store reported every failed lookup, an unreachable filer
  included, as ErrOIDCProviderNotFound, which CreateOIDCProvider reads as
  "free to create". Only a confirmed absence is now not-found.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: keep config-file OIDC providers out of a persistent store

Review of the previous commit found that mirroring the IAM config file's
providers into a persistent store, and pruning them when they leave the
file, breaks as soon as S3 servers share a filer:

- a server prunes stored config-file providers its own file does not list,
  including ones a peer's file still defines (a zero-config server prunes
  them all);
- mirroring overwrites an API-created provider with the same ARN and marks
  it config-owned, so a later prune deletes it;
- a failed mirror write or a failed prune leaves a stale record trusted;
- a mirrored record is loaded into STS at startup as an IAM-managed provider
  and shadows the config-file provider, dropping the settings a record does
  not carry (jwksUri, roleMapping, policyClaim, ...).

A persistent store now never receives the config file's providers. STS keeps
serving them from its static configuration, as it always has; the IAM API
lists and returns them from memory, refuses to change or delete them
(UnmodifiableEntity; change them in the file) and to create another provider
with their ARN (EntityAlreadyExists). The store holds only providers created
through the IAM API, and those are what startup loads into STS. There is
nothing to prune, so the source marker is gone. An in-memory store keeps its
behaviour: the config file's providers are records in it, as before.

buildOIDCProviderFromRecord also carries PolicyClaim and
AllowedPrincipalTagKeys now; they were dropped whenever an API-created
provider was loaded into STS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: send UnmodifiableEntity as a 400, not an internal error

The IAM API's error writer had no case for UnmodifiableEntity, which the
previous commit returns for a change to a config-file provider, so it went
out as a 500 ServiceFailure that clients retry. AWS sends it as a 400.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: document stored-over-config precedence, drop invented CreateDate, cancel superseded retries

Follow-ups from review of b881982d2:

- A provider stored under the same ARN as a config-file provider takes
  precedence in the IAM API, matching STS, which already prefers
  IAM-managed providers so that an API call can shadow a bootstrap entry.
  Deleting the stored provider brings the config-file one back. This was
  already the behaviour; it is now documented and tested.
- A config-file provider no longer reports its server's start time as
  CreateDate, which changed on every restart; GetOpenIDConnectProvider now
  omits the date for it. An in-memory store still stamps its copies at load,
  as before.
- The startup retry runs under a cancellable context, is cancelled when
  another store is installed, and retries the store it was started for
  rather than reading the manager's field, so replacing the store neither
  leaves the old retry running nor races with it (go test -race).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* s3/iam: serialize OIDC provider refreshes so an older snapshot cannot restore a deleted provider

Refreshes run concurrently: after an IAM API change, on a peer's change
and in the startup retry. Each lists the store and then hands STS the
result, so a refresh that listed before a DeleteOIDCProvider could finish
after that call's own refresh and keep the deleted provider trusted until
the next change. Refreshes now hold a lock from the read to the hand-off,
and a startup retry cancelled by installing another store drops its
snapshot instead of applying it.

The retry-cancellation test waits for the retry by polling instead of a
fixed sleep.

* s3/iam: route SetOIDCProviderStore through installOIDCProviderStore

A store installed after Initialize skipped the static-provider overlay
and startup hydration: config-file providers disappeared from the IAM
API, ErrOIDCProviderStatic no longer protected them, and stored
providers were never trusted until the next mutation or peer event.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:51:39 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
62d4f9152a iam: evaluate trust policies deny-by-default (#11513)
* iam: evaluate trust policies deny-by-default

EvaluateTrustPolicy seeded its result with the engine's DefaultEffect,
so a non-matching trust-policy statement set still resolved to Allow
when the IAM config sets policy.defaultEffect=Allow. A caller holding
a validly signed token from a registered provider could then assume a
role its trust policy does not admit.

Trust policies now start from implicit deny, matching AWS semantics and
the pre-d751623 behavior of evaluateTrustPolicy; DefaultEffect still
governs identity-policy evaluation.

Upgrade note: deployments on defaultEffect=Allow whose trust policies do
not match their callers will see those assumptions refused.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: cover trust policy implicit deny under DefaultEffect=Allow

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:36:01 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a901c1a5e2 filer: compare IF_ENTRY_EQUAL chunks by fid, not file_id (#11514)
The stored entry came through FindEntry, which restores chunk file ids
from their fid form, while an expected entry built from a metadata-log
event still carries the serialized form (file_id moved into fid). The
proto.Equal saw file_id "" against the restored id and refused every
stamp, so remote.sync re-uploaded each entry and the RemoteEntry stamp
never landed.

Clone both sides and run BeforeEntrySerialization before comparing, so
chunks match on their fid and the file_id spelling is ignored; the stored
entry and the request's ExpectedEntry are left untouched.

Generated with [Devin](https://devin.ai)

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-29 11:34:07 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
67b19b2db5 build(deps): bump com.fasterxml.jackson.core:jackson-databind from 2.22.1 to 2.22.2 in /test/java/spark (#11511)
build(deps): bump com.fasterxml.jackson.core:jackson-databind

Bumps [com.fasterxml.jackson.core:jackson-databind](https://github.com/FasterXML/jackson-databind) from 2.22.1 to 2.22.2.
- [Commits](https://github.com/FasterXML/jackson-databind/compare/jackson-databind-2.22.1...jackson-databind-2.22.2)

---
updated-dependencies:
- dependency-name: com.fasterxml.jackson.core:jackson-databind
  dependency-version: 2.22.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:12:55 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c54f1fe69c build(deps): bump github/codeql-action from 4.38.1 to 4.38.2 (#11509)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.1 to 4.38.2.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.38.1...v4.38.2)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:10:07 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
818ff69270 build(deps): bump cloud.google.com/go/kms from 1.34.0 to 1.35.0 (#11508)
---
updated-dependencies:
- dependency-name: cloud.google.com/go/kms
  dependency-version: 1.35.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:55 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
db89ed198b build(deps): bump github.com/twmb/avro from 1.8.0 to 1.9.0 (#11507)
Bumps [github.com/twmb/avro](https://github.com/twmb/avro) from 1.8.0 to 1.9.0.
- [Commits](https://github.com/twmb/avro/compare/v1.8.0...v1.9.0)

---
updated-dependencies:
- dependency-name: github.com/twmb/avro
  dependency-version: 1.9.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:39 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f7b5ca826b build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity from 1.14.0 to 1.14.1 (#11506)
build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity

Bumps [github.com/Azure/azure-sdk-for-go/sdk/azidentity](https://github.com/Azure/azure-sdk-for-go) from 1.14.0 to 1.14.1.
- [Release notes](https://github.com/Azure/azure-sdk-for-go/releases)
- [Commits](https://github.com/Azure/azure-sdk-for-go/compare/sdk/azcore/v1.14.0...sdk/azidentity/v1.14.1)

---
updated-dependencies:
- dependency-name: github.com/Azure/azure-sdk-for-go/sdk/azidentity
  dependency-version: 1.14.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:24 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
878d23b95a build(deps): bump github.com/dustin/go-humanize from 1.0.1 to 1.1.0 (#11505)
Bumps [github.com/dustin/go-humanize](https://github.com/dustin/go-humanize) from 1.0.1 to 1.1.0.
- [Commits](https://github.com/dustin/go-humanize/compare/v1.0.1...v1.1.0)

---
updated-dependencies:
- dependency-name: github.com/dustin/go-humanize
  dependency-version: 1.1.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:09:01 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1e0b243aed build(deps): bump golang.org/x/oauth2 from 0.36.0 to 0.37.0 (#11504)
Bumps [golang.org/x/oauth2](https://github.com/golang/oauth2) from 0.36.0 to 0.37.0.
- [Commits](https://github.com/golang/oauth2/compare/v0.36.0...v0.37.0)

---
updated-dependencies:
- dependency-name: golang.org/x/oauth2
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-29 10:08:43 +08:00
github-actions[bot] 79e2abdfbf docs: regenerate star history chart 2026-09-29 00:50:14 +00:00
github-actions[bot] 530be3e373 4.48 2026-09-28 15:53:40 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
9b3b12c607 filer: pin-aware reader cache eviction and stream release (#11503)
* filer: synchronize stream pins and release them on transitions

Guard chunkStream.cacher with the ReaderCache lock everywhere: mount
sections share one ChunkReadAt across concurrent reads, and unsynchronized
release could double-unpin. Reads served from the chunk cache now detach
the stream's pin instead of retaining the previous chunk. Eviction prefers
unpinned downloaders so a pinned buffer is not dropped mid-stream. A new
ReleaseStream lets callers drop their pin without destroying the shared
cache; S3 and WebDAV readers use it. lastChunkFid becomes atomic since
concurrent mount reads can update it.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: keep eviction bounded when every downloader is pinned

Both eviction paths still fall back to a pinned victim when no unpinned
one exists, so abandoned stream pins cannot bypass the downloader limit
or stall the memory budget. Budget eviction also rechecks the pin under
the ReaderCache lock at removal time: a stream that pinned the selected
victim in between keeps it mapped and the selection retries.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: restore budget bookkeeping when a victim gets pinned mid-eviction

removeUnpinned losing the pin race left the victim out of the idle list
while still holding its reservation, making it unevictable even as the
pinned fallback. Push it back when the reservation is still live.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 22:37:04 +08:00
Tobias GurtzickandGitHub 3abdef3202 filer: keep shared chunk buffers pinned while another stream reads them (#11502)
The ReaderCache is shared by all streams of a process (every S3 GET, for
instance), but a ChunkReadAt released chunks as if it owned them:

- moving on to the next chunk called UnCache on the previous one,
  destroying the buffer even when other streams were still inside it;
- since #11384 a buffer is dropped once any reader has consumed it to the
  end and no read call is in flight. Streams copy out in slices (256 KiB
  in the S3 gateway), so between two calls a slower stream is not
  attached and loses the buffer to a faster one.

Either way the slower stream refetches the whole chunk from the volume
servers. With many clients downloading the same popular object at once,
each chunk is fetched over and over; in production we saw the S3 gateway
pull ~10 Gbit/s from volume servers while serving ~1 Gbit/s to clients.

A ChunkReadAt now pins the chunk it is positioned in. The pin is taken
and released only under the ReaderCache lock, since concurrent ReadAt
calls on one ChunkReadAt (as in mount) share it. It is released when the
stream reads the chunk to its end, moves to another chunk (including one
served from the chunk cache), or falls back to random reads. A buffer is
dropped once no stream pins it and no read is in progress, if it was
consumed or its last stream left it; a read still in flight when the
stream leaves drops it on detach, as UnCache did via destroy. Eviction by
slot limit and memory budget is unchanged.

lastChunkFid is now guarded as well: concurrent ReadAt calls raced on it.

Tests: two ChunkReadAt instances streaming one object in interleaved
slices fetch each chunk exactly once (2-3 times before); leaving a chunk
for a chunk-cache hit or while another read is in flight releases it;
concurrent ReadAt calls on one ChunkReadAt leave no pins behind under
-race.
2026-09-28 21:56:04 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
43fd5b8d82 volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete

The 2PC generation switch stages each run as <base>.ecNN.v<N> plus
versioned .ecx/.ecj/.vif files. Nothing on the volume server removes
them: isEcDataShardFile only recognises the exact .ecNN name, so the
staged files are invisible to every bookkeeping pass, and even
full_teardown's wipe-all path left them behind. Each re-encode therefore
leaks a full shard set per shard-holding disk.

RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>,
optionally keeping generations at or above a threshold; teardown and the
reconcile wipe remove every generation, and a per-shard delete removes
that shard's staged generations too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: delete staged EC generations older than N via VolumeEcShardsDelete

After a 2PC generation switch commits, the superseded generation's
<base>.*.v<N> files sit on disk with no cleanup path: teardown removes
everything, and a per-shard delete only touches the named shards, so the
executor had no RPC that reclaims just the staged leftovers.

delete_generations_older_than removes staged generation files strictly
below the threshold on every disk. Versioned files are never mounted, so
nothing is unloaded first; the committed generation and the canonical
files are preserved.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: mirror staged EC generation cleanup

Parity with the Go volume server: remove_ec_generation_files sweeps
<base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by
remove_ec_volume_files (which covers both teardown paths) and the new
delete_generations_older_than request field; delete_ec_shards removes a
shard's staged generations along with the canonical file.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: match staged generation filenames literally

filepath.Glob interprets metacharacters in the collection part of the
base name, so a collection like a[bc] could match another volume's
staged files (or miss its own). Scan the directory and compare names
literally instead, mirroring the Rust read_dir implementation.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: report generation-sweep errors and drop the store lock first

- snapshot the location base names under the read lock and run the
  filesystem sweep after dropping it, so a slow disk cannot stall the
  store;
- record per-entry read_dir errors in remove_ec_generation_files and
  propagate them from remove_ec_shard_generations instead of flatten()
  skipping them;
- warn when a staged-shard generation fails to delete rather than
  reporting success with files left behind.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume: fail shard delete when the staged-generation listing fails

A transient ReadDir failure fell back to removing canonical shard names
only: staged .v<N> files survived while the RPC still reported success,
leaving the leak invisible to retrying callers. ENOENT still means the
disk simply has no such directory; other listing errors now propagate.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: propagate staged-generation removal failures

delete_ec_shards logged remove_ec_shard_generations errors and the RPC
returned success while staged .v<N> files remained, diverging from the
Go handler which surfaces the failure. The sweep keeps processing the
remaining shards, retains the first error, and volume_ec_shards_delete
maps it to Status::internal so callers can retry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: notify state change even when the shard sweep errors

delete_ec_shards already deletes and unmounts the shards before
returning a staged-generation failure, so returning early skipped
volume_state_notify and the master kept routing to them until the next
heartbeat. Notify before propagating the error.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 21:55:25 +08:00
150a69fe11 master: make volume capacity reservation timeout configurable (#11426) (#11497)
* master: make volume capacity reservation timeout configurable (#11426)

* master: expire reservations on reads, fix int timeout units

- AvailableSpaceForReservation now expires reservations too: a node that
  is full of reservations is filtered out before TryReserveCapacity can
  clean them, which stranded expired capacity indefinitely.
- Drop TryReserveCapacityWithTimeout: a per-call timeout lets one caller
  expire another's live reservations, and the Node interface stays
  stable for implementations outside this tree.
- parseReservationTimeout no longer routes integer values through
  GetDuration, which read them as nanoseconds; bare numbers are
  seconds. The 5m fallback is now the shared DefaultReservationTimeout.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 18:40:11 +08:00
4fec65d949 filer: demote client-cancelled directory listing log from error (#11495) (#11496)
* filer: demote client-cancelled directory listing log from error (#11495)

* filer: quote path in canceled listing log

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 14:35:41 +08:00
github-actions[bot] f564918685 docs: regenerate star history chart 2026-09-28 00:52:53 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f9289f0570 s3: do not promote ?prefix into the object for non-List actions (#11494)
* s3: do not promote ?prefix into the object for non-List actions

authRequestWithAuthType mapped an empty object to the prefix parameter for
every action, so PUT /bucket?versioning&prefix=x authorized as Write:bucket/x.
An object-scoped grant (Write:bucket/*) could then change bucket versioning,
lifecycle, cors, and object-lock configuration, and the promoted object also
made ResolveS3Action report s3:PutObject to attached IAM policies.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object

resolveS3AuthTarget fed the promoted prefix to ResolveS3Action, so a
bucket-level ?uploads request resolved as s3:GetObject on the prefix ARN
in the admin explicit-deny check. Resolve both action and resource
against the object the bucket listing actually scopes.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: resolve the listing action through the bucket-level object in AuthorizeAction

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: drop the unreachable object-level uploads case from the resolver test

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-28 07:17:58 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
4303b3aa4c s3: keep a listing's start position inside the requested prefix (#11493)
* s3: a list marker that sorts past the prefix leaves nothing to list

AWS scopes a listing to keys under Prefix; StartAfter, Marker and
continuation tokens only reposition inside that range. A marker that
diverges from the prefix at a larger byte is after every key the prefix
can match, so the page is empty. normalizePrefixMarker used to keep such
a marker as the walk cutoff at the bucket root, where the walk descends
into the marker's own directory and returns keys the prefix never names.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the listing variant's action when a prefix is promoted to object

authRequestWithAuthType promotes ?prefix= into the object argument for
the legacy CanDo path. ResolveS3Action treats a non-empty object as
object-level, so a bucket-level ?versions or ?uploads request carrying a
prefix missed its specific action and fell back to the base List action:
an s3:ListBucket grant then covered s3:ListBucketVersions, and an
explicit Deny on the specific action was skipped on the same path.

Resolve the action against the same bucket-level object the resource
ARN already uses.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: treat GET ?uploads as a bucket listing for authorization

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* Update weed/s3api/auth_credentials.go

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
2026-09-28 06:58:44 +08:00
a0ee7ba314 s3: ignore empty intermediate directories in bucketHasUserObjects (#11490) (#11491)
* s3: ignore empty intermediate directories in bucketHasUserObjects (#11490)

* s3: keep nested reserved-named dirs from hiding user objects

Reserved folders (.uploads, *.versions) are internal only at the bucket
root; deeper entries with those names are user key prefixes and must be
walked. Also treat a missing subdirectory as empty via isFilerNotFound
(list errors cross gRPC as status errors, not the sentinel), let names
containing backslashes count as objects, and walk iteratively so empty
chains deeper than the old scan depth no longer report non-empty.

* s3: treat reserved-named directories as internal at every level

Object listing interprets .uploads and *.versions directories as
internal storage wherever they appear, so walking them during the
emptiness check would report invisible version remnants as user objects
and block deletion. A reserved name on a file still counts, matching
listing which only special-cases directories.

* s3: count explicit directory objects under reserved names

A directory object created by PutObject (MIME or prefix-object marker
set) is user data even when named .uploads or *.versions; only a plain
directory with a reserved name is internal storage.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-28 06:56:46 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a976b21010 s3: require dedicated object-lock permissions for x-amz-object-lock-* headers (#11492)
* s3: require dedicated object-lock permissions for x-amz-object-lock-* headers

PutObject, CreateMultipartUpload, and PostPolicy honor the retention and
legal-hold headers after only the route's s3:PutObject check, so a
write-only principal could pin a version under COMPLIANCE retention that
nobody can remove before its retain-until date. On AWS these headers
require s3:PutObjectRetention / s3:PutObjectLegalHold. validateObjectLockHeaders
is the shared funnel for all four call sites; it now authorizes the
corresponding dedicated action when each header is present.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: record the verified POST-policy signer as the request identity

The handler authenticated the form policy signature but stored only the
signer's name, so downstream authorization (the object-lock header check)
re-authenticated the form-signed request as anonymous and evaluated the
wrong principal.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 20:17:12 +08:00
e57f8c4d87 volume server: read GET/HEAD needles off the store lock, and only once (#11487)
The GET/HEAD handler read the needle synchronously on the tokio worker
while holding store.read(): first a stream-info read that loaded the
whole record just to parse its meta, then, for every needle that was not
streamed (small, compressed, chunk manifest, image ops), a second full
read. For a tiered volume each read is an S3 GET under the store lock,
and a writer queued behind it parks every other store reader.

The regular-volume read now runs in spawn_blocking. Under the store guard
it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat
handle or the remote backend, offset, size); the guard is dropped before
any needle data I/O. No data-file lease is held across the read either,
since a writer waits for one while holding the store write lock. The
index size decides the read, as in Go's readNeedle: a HEAD, a ranged read
or a needle above the stream threshold reads only its header and meta
tail (ReadNeedleMeta) and hands off to StreamingBody or the range path;
everything else is read in full once, with its checksum verified. A
compressed or manifest needle found by the meta read is then read in
full once. The range-from-source read also moves to spawn_blocking.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-27 20:13:15 +08:00
67691a1eea volume server: split volume_copy into phases and type the delete-after-status gate (#11485)
volume_copy was one ~400-line handler, and the rule that an existing
local replica is deleted only after the source's ReadVolumeFileStatus
succeeded was held by statement order alone. The keep_remote_data=true
that the pre-copy delete and the failed-copy rollback must share was
kept in sync by a comment pointing from one to the other.

The handler is now a ~60-line orchestrator over connect_to_copy_source,
SourceVolumeStatus::fetch, delete_existing_replica,
plan_copy_destination and a VolumeCopyJob whose run() drives
preallocate_dat, transfer_files, finish_copied_files and
mount_and_reply, with cleanup_failed_copy on error.
delete_existing_replica takes a &SourceVolumeStatus, which only fetch
can construct (private field in a child module), so the delete cannot
be called before the status RPC. Both deletes go through
delete_replica_keep_remote.

Pure refactor: call order, status codes and messages, cancellation
checks, throttling, progress reports and cleanup are unchanged.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-27 20:12:55 +08:00
hsdfatandGitHub 02353444ac fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts (#11483)
* fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts

Port of the Go VolumeCopy hardening in #11238 and #11252.

- Pick the destination disk before deleting the existing replica, counting
  the slot that replica holds as free. If no disk qualifies, the healthy
  replica is kept instead of being deleted.
- Read the source's VolumeStatus before and after the copy. When both
  succeed and the counts did not change, the mounted replica's file and
  deleted counts must match; on mismatch it is unmounted and its files
  removed. A failed "before" read skips the check; a failed "after" read
  fails the copy.

* fix(volume-rust): let a departing caller cancel VolumeCopy's post-copy status read

Go reads the source's status after the copy with stream.Context(), so the
call ends when the caller leaves. The Rust call had no such link: a
source that stalled there held the copied, unmounted files after the
caller was gone. Race it against the response channel, like the other
blocking steps, so the usual error cleanup removes the partial copy.
2026-09-27 19:40:07 +08:00
68944e83a3 volume: typed tier errors so a missing remote object answers NotFound (#11484)
remote_storage/s3_tier.rs returned Result<_, String> from every
transfer (upload_file, download_file, read_range[_blocking],
delete_file[_blocking]) and from the tier runtime helpers. The tier
move handlers could only wrap that in Status::internal, so a .dat whose
remote object is gone was indistinguishable from an I/O failure to
weed shell.

Add TierError { NotFound, Io, RuntimeUnavailable, Aborted }. Each
variant carries the existing message verbatim. NotFound follows the
rules remote_storage/s3.rs already uses: raw 404 status on HEAD,
NoSuchKey code on GET; a bare 404 on GET stays Io. A progress-callback
Err becomes Aborted. VolumeError gains a transparent Tier variant and
From<VolumeError> for Status maps Tier(NotFound) to NotFound; the tier
move handlers go through status_with_context, so their message text is
unchanged. Every other tier failure is still Internal.

The remote needle read path keeps io::Error::other, so its error kind
and vacuum's handling of it do not change.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:39:25 +08:00
00310f6588 volume server: run the vacuum compaction copy without the store lock (#11482)
* volume server: run the vacuum compaction copy without the store lock

VacuumVolumeCompact held the store write lock for the whole live-needle
copy, including every progress blocking_send on the 16-deep stream. On a
large volume that is minutes with every read, write and heartbeat on the
node parked behind it, long enough for the master to unregister the node.

Split compaction the way Go's CompactByIndex runs it. A short locked step
claims the volume's compacting flag, records the makeup_diff watermark
(index size and compaction revision) and opens fresh .dat/.idx handles.
The copy then replays .idx up to the watermark and copies from those
handles with the store lock released; writes that land meanwhile are
replayed by makeup_diff at commit, as before. The flag is an Arc<AtomicBool>
released when the job is dropped, so every exit path clears it.

Because the flag is now visible to other callers, the operations that would
pull the files out from under the copy refuse while it is set: unmount
(and VolumeConfigure, which unmounts and remounts), delete (checked before
the volume is removed from the map, which a refused destroy used to leave
unmounted), cleanup, and index relocation. A second compact and a commit
stay no-ops, as in Go. The pre-copy fsync is dropped: the copy reads its
own handles through the page cache and .cpd/.cpx are fsynced before commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume: keep a read-only in-memory index's size for the compaction copy

The unlocked copy replays .idx up to index_file_size(). A read-only
volume whose .sdx could not be built loads its index into memory without
a writer, so that size stayed 0: the copy came out empty and the commit
replaced the volume with it. CompactNeedleMap::load_from_idx now records
the rows it loaded, which is also what Go's IndexFileSize reports for a
read-only index.

The copy's index replay now stops reading at the recorded size instead
of walking rows appended since, which makeup_diff replays anyway.

Adds tests for compacting a read-only volume on both the sorted index
and the in-memory fallback, and for VolumeConfigure stopping when the
unmount is refused during a copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: stop a vacuum copy as soon as its client is gone

The progress callback only noticed a closed response stream when a report
was due, every 128 MiB. With the copy now running outside the store lock,
a copy nobody waits for keeps the volume marked compacting and so keeps
refusing unmount, delete and cleanup until that next report. Check the
stream on every callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:59 +08:00
5c9c424a84 volume server: stream ReadAllNeedles without holding the store lock (#11481)
* volume server: stream ReadAllNeedles without holding the store lock

read_all_needles held store.read() while Volume::read_all_needles read
every live needle of the volume into a Vec, and kept holding it through
the whole blocking_send loop. Memory grew with the volume, and a slow
client parked the scan in a send with the guard held; needle writes and
the heartbeat take store.write() on a writer-preferring lock, so the node
stopped serving until the client caught up.

Take a DatScanPlan (fresh .dat open, end bound) under a short guard and
walk it with the guard released, sending one needle at a time. Each
record is checked against the live needle map under a brief read guard,
as the scan reaches it, and only a live record is parsed, so a damaged
stale copy does not fail the stream. Records appended while a pass ran
are walked by a follow-up plan, so a needle overwritten during the scan
is streamed once, as its new copy. A vacuum commit or re-create of the
volume during the scan fails the stream, since the map's offsets no
longer describe the pinned file; the plan carries the volume instance
and compaction revision for that check.

DatScanPlan::scan_records yields records unparsed; scan keeps its
behaviour on top of it. Volume::read_all_needles has no caller left and
is removed; its tests move to the RPC.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: check ReadAllNeedles liveness only once the send can proceed

The per-record liveness check ran before blocking_send, so a scan parked
on a full channel held a record it had already judged live. An overwrite
landing during that park left the old copy in the stream, and the
continuation over appended records then streamed the new copy as well.

Reserve channel space first, then take the store read guard, check the
record against the needle map and enqueue it through the permit before
releasing the guard. The wait for space still happens without the lock;
the record is parsed before the guard is taken, and its parse error only
counts if the record turns out to be live.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:43 +08:00
5218e68554 volume server: collect heartbeats under the store read lock, off the runtime (#11480)
* volume server: collect heartbeats under the store read lock, off the runtime

Every pulse tick, options change and volume-state notification took
store.write() for a whole heartbeat pass, directly on the async heartbeat
task. The pass fstats every volume's .dat twice and hashes its report, so
on a server with many volumes it held the store exclusively for the whole
scan: reads and writes stalled, and with the writer-preferring RwLock a
pending pass parked every new reader too.

The pass only needs to mutate the store for a few rare actions: removing
expired EC volumes, deleting expired volumes past their removal delay, and
setting no-write on IO-quarantined volumes. It now runs under store.read(),
records those as (disk, volume id) actions, and applies them afterwards
under a short store.write() that is only taken when there is something to
do. Each action re-checks its target under the write lock, so a volume
written to, replaced or removed in between is left alone. Expired EC
volumes are still removed before the volume pass, as before, because the
EC shard count feeds the disk-space-low max volume count. Every pass runs
on the blocking pool via spawn_blocking.

The heartbeat message is unchanged for the same store state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* volume server: take has_no_ec_shards with the heartbeat's volume list

The heartbeat pass took has_no_ec_shards from the EC phase's read lock,
then built the volume list under a second one. An EC shard mounted in
between went out as "no EC shards" beside a volume list taken after the
mount, and the master clears a server's EC registrations on that flag.

has_no_ec_shards is now computed under the same read lock as the volume
list, with the EC phase's filter: not expired, not quarantined, at least
one shard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-27 19:38:23 +08:00
github-actions[bot] 00cffa028c docs: regenerate star history chart 2026-09-27 00:53:30 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
a261f90e18 vacuum: let the sweep release volumes that stay empty and quiet (#11477)
* vacuum: let the sweep release volumes that stay empty and quiet

Vacuuming reclaims bytes but not slots: a fully emptied volume stays
registered to its collection forever, and since growth is gated only on
slot count a store at 99% free disk can still refuse writes to other
collections (#11429). volume.deleteEmpty exists but is manual-only.

With -vacuumDeleteEmptyAfterSeconds (or master.vacuumDeleteEmptyAfterSeconds
under weed server/mini; default 0, off) the automatic sweep now deletes
replica copies that have stayed empty and quiet for that long, the same
rule volume.deleteEmpty applies on demand: remote-backed copies are
skipped, and every delete carries the volume server's onlyEmpty /
onlyGarbage guards so a copy written since the last report is refused
rather than removed. Copies that still hold data or were written
recently stay; only a volume whose every copy is deleted leaves the
sweep's work map, sparing a compaction of bytes that are all deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: harden empty-volume sweep against partial and racing deletes

Review follow-up on #11477:

- delete a volume only when every replica copy is a verifiable
  empty-and-quiet candidate; deleting the empty copy of a volume whose
  sibling holds live files would silently cut its replica count
  (greptile P1).
- drain the volume out of the writable list before deleting, the same
  drain the compact pass uses, so PickForWrite stops assigning it and
  pending writes settle (devin).
- bound the VolumeDelete RPC so one stalled server cannot hold the
  vacuum lock indefinitely (greptile P1, reusing allocateVolumeTimeout).

The vid2location panic scenario raised in review does not exist:
VolumeLocationList methods are nil-receiver safe and a missing vid just
fails enoughCopies, so a partially deleted volume skips compaction
instead of crashing the sweep.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: unregister deleted empty replicas and prune the sweep list

A successful VolumeDelete only updates the volume server; the master
still tracked the replica and kept it in the sweep's location list for
the compaction pass (coderabbit on #11477). Unregister the replica right
after its delete succeeds and drop it from the sweep copy, so a partially
deleted volume only compacts copies that still exist.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: pin deleting volumes out of the writable list across heartbeats

Review follow-up on #11477 (greptile): DrainAndRemoveFromWritable only
removed the volume once; a heartbeat landing between the drain and the
replica deletes re-evaluated writability and re-added it, so a client
write could reach a replica whose siblings were already gone and leave
the volume under-replicated when the last copy refused its onlyEmpty
delete.

MarkDeleting records the vid in deletingVolumes — checked inside
setVolumeWritable so heartbeat, capacity-recovery, and admin re-add
paths all hold it out — and UnmarkDeleting releases it once the sweep
finishes the copy pass. A partially deleted volume's surviving replicas
then return to writable through the normal heartbeat path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* vacuum: restore writability when a sweep delete survives

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:05:22 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
be29f44d87 s3: record requester identity before the authz verdict (#11479)
* s3: record requester identity before the authz verdict for audit

Identity was only stored in request context on the success branch, so
denied requests reached WriteErrorResponse without requester attribution
and audit entries had empty requester/requester_arn/requester_identity.
Authentication failures still resolve no identity, so unauthenticated
denials stay unattributed.

Fixes #11474

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: keep the resolved identity through authz denial in Auth

Review follow-up on #11479 (devin): authRequest discarded the identity
on every error, so a request that authenticated fine but failed the
action check still reached handleAuthResult with no identity and the
deny path could not audit a requester. Auth now calls
authRequestWithAuthType directly, the same entry AuthPostPolicy uses,
so the resolved identity reaches the error writer; a failed authN
still resolves no identity and stays unattributed. The regression test
now signs a denied request end to end through iam.Auth.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:03:01 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2864bc0fe8 s3: honor configured session bounds on AssumeRole and LDAP identity (#11478)
* sts: export CalculateSessionDuration

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: honor configured session bounds on AssumeRole and LDAP identity

prepareSTSCredentials hardcoded a one-hour session when the caller
omitted DurationSeconds, so sts.tokenDuration was ignored and
sts.maxSessionLength only clamped explicit requests: asking for 3600s
against a 20m ceiling was rejected while omitting the parameter was
granted a full hour (#11473). The two affected handlers now use the
same default-then-cap calculation as AssumeRoleWithWebIdentity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through role store copies

copyRoleDefinition rebuilt RoleDefinition field by field and dropped
MaxSessionDuration, so memory-backed role stores silently discarded the
per-role session bound on every write and read (devin on #11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: apply per-role MaxSessionDuration to resolved session durations

Review follow-up on #11478 (devin): the role bound only ever applied to
explicit DurationSeconds values — an omitted duration resolved to the
configured default and sailed past a shorter role max on every assume
path.

- capDurationByRole now resolves min(requested||tokenDuration, roleMax),
  so AssumeRoleWithWebIdentity and AssumeRoleWithCredentials cap
  defaults the same way they cap explicit values.
- prepareSTSCredentials caps the calculated duration at the named
  role's MaxSessionDuration, covering the AssumeRole and LDAP handlers;
  self-assumption has no role definition to consult.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iam: keep MaxSessionDuration through the cached role store

genericCopyRoleDefinition drops MaxSessionDuration the same way
copyRoleDefinition did, so the cached filer role store reads back a zero
maximum and every downstream duration cap is skipped (greptile on
#11478).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* sts: only materialize defaults that pass session duration validation

Review follow-up on #11478 (greptile): materializing an omitted
DurationSeconds into an explicit value could exceed the service's own
input bound (a configured tokenDuration above maxSessionLength) and turn
a previously working request into a validation error.

capDurationByRole now leaves nil anything the service can resolve
better itself, clamps a tightened default at maxSessionLengthSeconds,
and floors a role bound below 900s to the tightest issuable value.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-27 07:01:51 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
ab95d58b7c s3: keep dedicated object-lock actions pinned during action resolution (#11475)
* s3: keep dedicated object-lock actions pinned during action resolution

A coarse action that already names a dedicated operation (governance
bypass, retention, legal hold, bucket object-lock config) now resolves to
itself before request shape is consulted. Previously a synthetic
DELETE ?versionId authorization request re-resolved to
s3:DeleteObjectVersion, so the bypass check was satisfied by the
delete-version grant alone; with the pin it evaluates
s3:BypassGovernanceRetention as intended.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* s3: cover pinned object-lock actions against competing query params

Locks in the resolution for every dedicated action in the pin set, incl.
the retention and legal-hold shapes carrying versionId.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 21:42:34 +08:00
2f641a63d6 filer: honor is_moved only from ring member connections (#11456)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks

updateLocalEntry records the RemoteEntry stamp after an upload by writing the
event's entry back with UpdateEntry. The filer deletes every stored chunk
absent from an updated entry, so when the file was rewritten while its upload
was in flight (or the event is a replay), the stale snapshot deletes the
rewrite's chunks: the entry then points at the new fid with no needle behind
it, and the rewrite's own upload fails and is skipped as superseded.

The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's
chunk fids, evaluated by the filer under the path lock. A refused stamp means
the filer moved past this event; the superseding event follows in the log and
stamps the current entry, so the refusal is logged and skipped like a
superseded upload.

Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount,
filer.remote.sync; hold the remote (docker pause) so one upload stays in
flight, rewrite the file through the filer, unpause. Before: the entry's chunk
is 404 on every volume server. After: the stale stamp is refused, the rewrite's
chunk stays live and reads back after a vacuum.

* filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored

The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a
rewrite that touched inline content or metadata alone still compared
equal and the stale snapshot overwrote the live entry. The new clause
compares the whole stored entry against the event's entry under the
same path lock.

* filer: route conditional UpdateEntry to the entry's owner filer

Two filers locking the same path locally could still pass a stale
condition on the non-owner while the owner's entry had moved on. When a
condition or expected_extended precondition is set, forward the request
to the entry's owner the same way conditional CreateEntry does, with
is_moved bounding the hop.

* filer: compare IF_ENTRY_EQUAL against the normalized expected entry

FindEntry grows FileSize to the chunk extent, so a raw event entry with
FileSize still zero failed the condition on an unchanged file and the
stamp was skipped, letting a replay upload the object again.

* filer.remote.sync: classify refused stamps by gRPC status only

A FailedPrecondition substring in an unrelated error would have been
swallowed as a skipped stamp; status.FromError already unwraps.

* remote sync: keep the event entry intact for IF_ENTRY_EQUAL

* filer: honor is_moved only from ring member connections

is_moved is caller-controlled, so a request could set it to skip owner
routing and run a conditional check under a non-owner's lock. Verify the
marker against the peer's connection address and the lock ring members;
an unverified marker is ignored and the request routes like a fresh one.

* filer: refuse unverifiable is_moved at a non-owner, cache ring IPs

Follow-up fixes from review on the is_moved provenance check:

- checkMovedMarker replaces "ignore and re-forward" for markers that did
  not arrive on a ring member's connection. Re-forwarding a claimed hop
  could cycle while rings disagree; instead the request is refused with
  FailedPrecondition unless this filer is the key's owner, in which case
  applying locally is correct anyway.
- ringMemberIPs caches resolved member addresses per ring membership so
  hostname-advertising deployments do not pay a DNS lookup per forwarded
  request; failed lookups are not cached so a DNS blip self-heals.
- DistributedUnlock no longer dereferences the nil response of a failed
  next-hop RPC.

* filer: refuse unverifiable is_moved with PermissionDenied, not FailedPrecondition

A routing refusal is different in kind from a write-condition mismatch:
remote sync treats FailedPrecondition as a stale stamp and skips it, so
reusing that code let a routing failure pass as synced. Owner checks now
also run before the peer-IP lookup so the common accept path does no DNS.

* filer: expire resolved ring member IPs after 5 minutes

A member's hostname can re-resolve to a new IP while its ring address
stays unchanged; caching forever would reject its genuine forwards until
a membership change or restart.

* filer: deduplicate concurrent ring member DNS lookups

At cache expiry, parallel forwarded requests would each resolve every
member hostname serially; singleflight collapses them into one lookup
per ring membership.

* filer: detach the shared ring lookup from the caller's context

The singleflight winner's ctx is cancelled when its request ends; the
shared result would then be an incomplete member list and genuine
forwards denied. The lookup now runs on a detached context with its
own deadline so a canceled caller cannot poison it.

* filer: resolve ring member hostnames in parallel

The shared lookup gave every member one serial budget, so a few slow
resolutions could leave later members out of the cached list and reject
their genuine forwards. Each member now resolves concurrently under its
own detached deadline.

* filer: gather literal member IPs before spawning lookups

A ring mixing IP literals and hostnames raced: the literal appends ran
unlocked alongside the resolver goroutines' locked appends. Split into
two passes so only hostname results share the mutex.

---------

Co-authored-by: jsas <1351492+jsas@users.noreply.github.com>
2026-09-26 19:42:23 +08:00
Mohd Quamar TyagiGitHubcoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>Chris Lucoderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
80fd3635d2 volume: skip TTL last-write scan when it cannot fit its budget (#11472)
* volume: skip TTL last-write scan when it cannot fit its budget

* Update weed/storage/volume_checking.go

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-26 19:41:59 +08:00
Chris LuandGitHub 7129e1178e s3: evaluate bucket policy before ACL public-read for anonymous requests (#11471)
* s3: evaluate bucket policy before ACL public-read for anonymous requests

AuthWithPublicRead granted anonymous access on a public-read ACL before
consulting the bucket policy, so an explicit Deny (e.g. s3:ListBucket)
was skipped for anonymous callers while still enforced for authenticated
ones. Run the policy engine first: a matching Deny or Allow is honored,
otherwise fall through to the ACL grant as before.

* s3: defer object-level anonymous requests to the handler's policy recheck

Evaluating the bucket policy with a nil entry at middleware time makes
tag conditions like s3:ExistingObjectTag/<key> resolve against missing
values, so a conditional Deny could wrongly block anonymous Get/Head on
a public bucket whose handler recheck would permit it. Object requests
now take the ACL grant and let Get/HeadObjectHandler re-evaluate with
the fetched entry; only bucket-level requests (List, HeadBucket), which
have no such recheck, are decided by the middleware policy verdict.

Reading the bucket config first also refreshes the compiled policy on a
cache miss, so a remotely deleted policy cannot leave a stale verdict
in the engine for nonresident buckets.

* s3: recheck bucket policy before serving directory objects

handleDirectoryObjectRequest runs before the object handlers' policy
recheck, so directory content on a public-read bucket was served to
anonymous callers without any policy evaluation. Evaluate the policy
with the directory entry, matching the recheck the file path performs.
2026-09-26 17:59:57 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
3c1e8ca7a8 volume: never finish serving a needle whose data fails its CRC (#11467)
* storage: hold back last chunk until CRC verifies on whole-needle reads

Above PagedReadLimit the needle is streamed: headers and body go out
before the checksum is computed, so a corrupted needle was served as
200 with bad bytes and readers could not fall back to a replica. The
final chunk is now written only after the checksum verifies; on a
mismatch the response ends short of Content-Length and the client sees
a failed transfer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: verify needle checksum on streamed reads (parity)

Mirror the Go fix: carry the needle checksum in NeedleStreamInfo and
have StreamingBody accumulate the CRC and verify it before emitting the
last frame; a mismatch ends the body with an error so the client sees
the transfer fail rather than receiving corrupt bytes that look
complete.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* server: abort the transfer when a committed stream fails

A writeFn failure after any byte or WriteHeader call leaves the declared
status and Content-Length already sent; http.Error's text then joins the
body and can exactly fill the withheld tail of a corrupted needle read —
the client sees a complete 200 instead of a failed transfer to retry.

Track whether the response is committed (headers sent, or bytes buffered
for the deferred flush) and panic with http.ErrAbortHandler instead of
appending an error body; pre-commit failures keep the 500 path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* server: drop the response writer wrapper from the committed-response check

Counting buffered writes is enough: with no bytes buffered the status and
headers cannot have gone out, and the range branches commit via the
explicit WriteHeader call before writeFn runs. The extra ResponseWriter
wrapper added a new Write sink site that CodeQL flags.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:50:41 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
0f3ba98e11 volume: make volume.scrub report a live needle whose stored id is damaged (#11468)
* storage: scrub live needles' stored id against the index key

scrubVolumeData only compared the needle's stored id for tombstones, so
header damage on a live needle — where the data CRC cannot see it —
passed every scrub mode while reads of that needle kept failing or
serving the wrong key's data. Compare the id for every indexed needle.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: scrub live needles' stored id against the index key (parity)

Mirror the Go scrub fix: compare the stored needle id with the index
key for live needles too, not only for deleted ones.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* rust volume: cover damaged live needle id in scrub test

The tombstone test proved the index-key check fires for deleted entries;
add the live-needle mirror of Go's TestScrubVolumeDataChecksLiveNeedleId
so a regression in the live path is caught in Rust too.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:10:14 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
80a26020d7 util: serialize all ViperProxy access so startup cannot hit concurrent map read/write (#11470)
* util: serialize every ViperProxy method; stop promoting unlocked viper calls

ViperProxy embedded *viper.Viper, so only the five declared methods took
the mutex while every promoted call — GetStringMap in backend.LoadConfiguration
was the reported crash — touched viper's maps unsynchronized. `weed server`
starts the volume server (SetDefault writer) and the master (GetStringMap
reader) back to back, and a race build reports the pair on a plain start.

The wrapped viper is now a named field: a method must be declared here to
exist on the proxy, so unsynchronized access fails at compile time rather
than at runtime. Every promoted use in the tree (GetStringMap, GetUint32,
GetFloat64, GetDuration, IsSet, AllKeys, Set) gets a locked wrapper;
NewViperProxy replaces struct literals for local vipers. GetStringMap
deep-copies its result — viper hands back the internal subtree, so
iterating it after the lock is released would race the next writer.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: take the shared lock while LoadConfiguration merges a config file

viper.MergeInConfig rewrites the same maps the proxy serializes; without
the lock a merge can race a concurrent SetDefault or reader exactly like
the reported startup crash.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: deep-copy slice elements in the GetStringMap snapshot

A slice of maps inside the returned subtree still shared the inner maps —
copy elements recursively so nothing the caller mutates is viper's
internal state.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: add the missing AutomaticEnv wrapper used by tests

sse_reader_test reaches it through GetViper(); without the wrapper the
call no longer exists once the viper field stopped being embedded.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* util: return a fresh slice from GetStringSlice

A stored []string comes back uncast from viper — the backing array is
shared internal state like the GetStringMap subtree, so copy it while
holding the lock.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:09:29 +08:00
yi111GitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
5389f61cef volume server: do not finish a GET when the needle CRC mismatches (#11464)
* volume server: do not finish a GET when the needle CRC mismatches

A streamed full-needle read compared the CRC only after every page had been written. Once the response buffer flushed, the client already had a completed 200 and the corrupt bytes. Hold the last page until the checksum matches, and if an earlier page has already been flushed, abort the connection instead of calling http.Error.

Fixes #11459

* volume server: abort partial-content bodies on write error too

The non-Range path drops the unflushed tail and aborts on a mid-body
error; the single-range and multi-range paths still flushed it after
WriteHeader(206) was committed, delivering corrupt bytes as a complete
body.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: assert the started 200 is aborted in the write-error test

The test previously returned on any request error, so it passed without
verifying the abort. It now asserts the client got the committed 200
headers and then a failed body read. Also trims comments.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 16:00:22 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
8ad2f29e3e shell: let volume.deleteEmpty drop volumes with no live needles (#11437)
* shell: let volume.deleteEmpty drop volumes with no live needles

The candidate check only accepted a .dat at superblock size, so a volume
whose every needle was deleted still had to be vacuumed first — minutes
of compaction to rewrite bytes that were all garbage anyway. FileCount
counts every indexed entry and DeleteCount every entry made garbage by
overwrite or delete, so FileCount <= DeleteCount means nothing live
remains and the volume can be unlinked directly. The quietFor guard is
unchanged.

* volume server: add only_garbage VolumeDelete guard

VolumeDelete(only_empty) refuses every volume that ever held data, so a
volume whose needles are all deleted could only be removed after a
vacuum rewrote it. The new only_garbage flag deletes only when the byte
counters show nothing live: DeletedSize covering all of ContentSize, the
same all-garbage state vacuum measures. Byte counters are used because
the file/delete counts drift on index reload.

* rust volume: mirror only_garbage VolumeDelete guard

Same check as the Go server: a volume deletes under only_garbage when
its deleted bytes cover all content bytes. The grpc handler rejects
before the store drops the volume from its map, since destroy errors
after removal would still unmount it.

* volume delete: let either enabled check pass, keep onlyEmpty on the wire

An upgraded shell sending only_garbage to a pre-upgrade server would be
read as an unconditional delete (field ignored, only_empty false). The
request now keeps only_empty set so old servers check emptiness and
refuse, while new servers delete when either check passes.

* volume.deleteEmpty: skip remote-backed and protected read-only volumes

A remote-tiered replica shares its cloud object with the other replicas,
so keepRemoteData=false on one delete removes data they still reference.
Protected read-only volumes are quarantined or under maintenance, which
is exactly when a replica should not be dropped.

* volume delete: validate guarded copies across disks before deleting

* volume delete: hold copy locks across guarded validate-and-delete

CheckVolumeDeletable released each copy's locks before Destroy ran, so a
write landing on a later copy between the two passes refused its destroy
after earlier copies were already removed. Pin every copy's
dataFileAccessLock (and its location's volumesLock) across validation and
removal so a refused delete leaves all copies intact.

* volume delete: send deleted-volume notices after releasing locks

A blocking send on a full DeletedVolumesChan under volumesLock can stall
the heartbeat loop that drains it while it waits on the same locks.
Collect the notices under the lock span and send after release.

* pb: restore generated-file cosmetics to match the repo's protoc version

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 15:53:39 +08:00
Ilia DemianenkoandGitHub c58bd0dfd3 s3: honor assignment fsync in UploadWithRetry (#11449)
* fix: honor assignment fsync in UploadWithRetry

* Tests feedback
2026-09-26 12:01:52 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f7680cf812 volume: one long-lived tokio runtime for blocking tiered S3 reads (#11448)
* volume: one long-lived tokio runtime for blocking tiered S3 reads

block_on_tier_future, behind read_range_blocking and delete_file_blocking,
spawned an OS thread and built a fresh current-thread tokio runtime on every
call, then tore the runtime down. On an S3-tiered volume that is once per
needle read, per streamed 64 KiB chunk, per DatScanPlan record and per
destroy. The SDK client's pooled HTTPS connections are driven by tasks on the
runtime a request ran on, so each teardown dropped the pool and every call
re-dialed and re-handshook TLS. A panic inside the SDK was also flattened to
the fixed string "tier runtime thread panicked".

Now one process-wide runtime (OnceLock, multi_thread, 2 workers named
tier-io) drives all tier I/O; block_on_tier_future spawns onto it and parks
the caller on an mpsc channel for the JoinHandle result. Blocking the caller
is unavoidable (the storage layer is synchronous) and is what the old code
did through thread::spawn().join(). Handle::block_on is not used because the
wrappers are also reached from inside another runtime's worker, where it
panics with "Cannot start a runtime from within a runtime". JoinError panics
are downcast to &str/String and the payload is kept in the error.

Tests cover runtime reuse (Handle::id equal across calls, thread name
tier-io), calls from a std thread, from spawn_blocking, and directly from
current-thread and multi-thread runtime contexts, and the panic payload.
Against the old body 7 of 9 fail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: return a tier runtime build failure instead of panicking

Review follow-up. tier_runtime() expect'ed the runtime build, so an OS
refusing threads panicked inside Volume::destroy (after the volume left
the in-memory map, before its files were removed) and inside needle
reads, bypassing their error paths. Keep the runtime in a
Mutex<Option<Runtime>> behind tier_handle() -> Result<Handle, String>:
a failed build is returned to the caller through block_on_tier_future's
existing Result and is not cached, so a later call retries once the
pressure is gone. The lock is held only while building.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the shared tier I/O runtime

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:50 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
4299fdf578 volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest (#11446)
* volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest

Go's VolumeEcShardsDelete full teardown calls vs.store.UnloadEcVolume in
the blanket path (weed/server/volume_grpc_erasure_coding.go:488) and
location.UnloadEcVolume in the generation-fenced path (:511): each disk
that had the volume registered drops it, closes its shard descriptors
and gives back its ec_shards gauge before the artifacts are unlinked.

The Rust handler used Store::remove_ec_volume / DiskLocation::remove_ec_volume
instead, which only remove the map entry. Store::remove_ec_volume also
stops at the FIRST disk holding the vid, so on a split-disk volume
(shards on several disks) the blanket teardown left the sibling disks'
EcVolume registered with open fds while the unlink loop deleted their
files underneath it: the heartbeat kept advertising shards whose files
were gone, the inodes stayed pinned by the open descriptors, and the
VOLUME_GAUGE{collection,"ec_shards"} never came back down. The fenced
path leaked the gauge and the descriptors the same way on the one disk
it wiped.

Both paths now use the unload_ec_volume helpers from #11413 (every disk
for the blanket teardown, the strictly-older disk for the fenced one),
and the two Status::internal messages name the disk directory like Go's
"... on %s: %w".

Regression tests build a two-disk store with the same vid mounted on
each disk (the SplitDiskEcFixture, which gains a collection knob so the
gauge read is isolated from parallel tests mounting under "") and assert
that a blanket teardown leaves no EcVolume registered on any disk and
returns the gauge to its pre-mount value, and that a fenced teardown
decrements the gauge for the older disk's shard while preserving the
newer disk. Both fail against the previous handler.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments on the EC full-teardown unload path

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:23 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
3e9fc9e75b volume: reject invalid -minFreeSpace and duration flags at startup like Go (#11447)
* volume: reject invalid -minFreeSpace and duration flags at startup like Go

Symptom: the Rust volume server swallowed invalid values for three
startup flags and ran with a limit the operator did not ask for.
-minFreeSpace=150 became 150 bytes (Go: fatal, percent must be 0..=100),
-minFreeSpace=10GiBx and -minFreeSpace=50B fell through to the 1%
default (Go: fatal), and -inflightUploadDataTimeout=30sec, =abc or =30
silently became the 60 s default (Go: flag.Duration usage error, exit 2).
Go's byte-size units (1,024MB, 42 mib, 100Ki, 1eb) were also not
understood and mapped to 1%.

Go reference: weed/util/minfreespace.go:73-106 (MustParseMinFreeSpace /
ParseMinFreeSpace), weed/util/bytes.go:212-243 (ParseBytes and
bytesSizeTable), weed/command/volume.go:226-227 (flag.Duration, i.e.
time.ParseDuration grammar).

Fix: the two inflight timeout Cli fields are now std::time::Duration
with a clap value_parser that ports time.ParseDuration (units ns/us/µs/
μs/ms/s/m/h, fractions, bare 0, int64-nanosecond overflow), so clap
prints a usage error naming the value and exits 2 at the same layer as
Go's flag package. parse_min_free_spaces returns Result and is built
from Go-parity parse_min_free_space / parse_bytes; resolve_config_with_env
aborts with Go's "The value specified in -minFreeSpace not a valid
value" message, the fail-fast pattern it already uses for --max, --index
and --readMode. No new dependency. Negative durations and NaN percent
are rejected explicitly (Go quirks with no sane Rust representation).

Tests cover the Go fixtures from minfreespace_test.go and bytes_test.go,
the duration grammar, and flag-level rejection through Cli::try_parse_from
and resolve_config_with_env; the flag-level tests were confirmed to fail
against the previous parser.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: say that -minFreeSpace commas separate directories, never thousands

Review follow-up. parse_bytes keeps Go's ParseBytes thousands-comma
support, but the flag is split on commas into per-directory entries
before any entry is parsed (Go's MustParseMinFreeSpace does the same),
so -minFreeSpace=1,024MB is the two entries 1 and 024MB. The doc
comment and the single-entry tests advertised the comma form as if it
were reachable from the flag. Document the split and add a flag-level
test: two entries from parse_min_free_spaces, and resolve_config
aborting on the directory-count mismatch with one -dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the strict flag parsing

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 12:00:06 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
afce0a3dd3 docs: restore the HTTP REST API reference (#11454)
* docs: restore the HTTP REST API reference

The REST API documentation was lost in the README refactor, and the wiki
only covers the master server API. Add a complete reference for the three
HTTP surfaces — filer paths, master file-id/topology endpoints, and
volume-server content endpoints — generated from the actual handlers and
query parameters.

* docs: correct REST API details flagged in review

Tagging uses Seaweed- headers not query params, the filer recursive
delete option changes the DELETE default, omitted resize mode does not
mean fit, default file mode is 0660, the master redirect is 308, the
listing flag is -dirListLimit, TUS is enabled by default at /.tus, and
-port.public opens the separate read-only listener.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* docs: match the tagging delete example to canonical header names

Go canonicalizes Seaweed-k1 to Seaweed-K1 on write, and the delete list
is compared case-sensitively, so ?tagging=k1,k2 would not match.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:59:43 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>franchbChris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
83d44be0f3 volume: detect S3 not-found by typed SDK errors, not the "service error" string (#11444)
* volume: detect S3 not-found by typed SDK errors, not the "service error" string

remote_storage/s3.rs decided ObjectNotFound by matching the Display
output of an aws_sdk_s3 SdkError against "NoSuchKey" / "404" /
"NotFound". In the locked SDK (aws-smithy-runtime-api 1.11.6,
src/client/result.rs:487-497) that Display is a fixed string per
variant, "service error" for every S3 error, so ObjectNotFound was
unreachable: every missing remote object surfaced as
Other("s3 get object: service error") with the real cause discarded.

Go (weed/remote_storage/s3/s3_storage_client.go) uses typed checks:
HEAD (373-374): awserr.RequestFailure with StatusCode() == 404;
GET (436-437): awserr.Error with Code() == s3.ErrCodeNoSuchKey.

read_file now matches SdkError::ServiceError whose GetObjectError
is_no_such_key(); a bare 404 on GET stays a generic error, as in Go.
stat_file matches HeadObjectError::is_not_found() or a raw HTTP 404
status, Go's actual condition. Non-service errors fall through to
Other unchanged. Every SdkError message in s3.rs and s3_tier.rs is
formatted with DisplayErrorContext so the S3 error code and message
survive instead of "service error".

Six network-free unit tests drive the client through a canned
HttpClient (404 NoSuchKey, bare 404 on GET and HEAD, 404 with a
foreign body on HEAD, 403 AccessDenied on GET and HEAD). They need
aws-smithy-runtime-api as a dev-dependency; it is already in the lock
at a single version, so no new crates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: HEAD not-found is the raw 404 status alone, as in Go

Review follow-up. The HEAD arm also accepted the SDK's NotFound error
code on any status, so a 400 carrying <Code>NotFound</Code> became a
missing object. Go's stat looks only at RequestFailure.StatusCode() ==
404 (weed/remote_storage/s3/s3_storage_client.go:373); do the same. The
raw status still covers the body-less 404 the SDK turns into NotFound
and a 404 whose body names a foreign code. Regression test for the
non-404 NotFound body, which failed against the previous arm.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim comments on the typed S3 not-found checks

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: franchb <ru.balloo@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:59:09 +08:00
975cec9228 s3api: exclude marker part in listObjectParts pagination (#11463)
* s3api: exclude marker part in listObjectParts pagination

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: guard listObjectParts marker boundary and enhance pagination test

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* s3api: fold in review feedback from the parallel #11462 fix

Same core fix; this adds the explanatory comment, tightens the overflow
guard to math.MaxInt64, makes the fake filer sort entries like a real
listing, and adds the marker-exclusivity assertions alongside the
pagination walk.

Co-authored-by: yi111 <yi111@users.noreply.github.com>

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: yi111 <yi111@users.noreply.github.com>
2026-09-26 11:58:23 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
f31a026b2a master,filer: fix lock ring poisoning after leader change (#11453)
* cluster: never broadcast an empty lock ring

An empty member list is never a usable ring state, but a delayed
RemoveServer on a former leader can fire after the new leader already
broadcast the recovered ring. That late broadcast carries a newer
wall-clock version, so clients accept the empty ring and permanently
reject the good one.

Skip the broadcast entirely when the member list is empty, keeping the
last non-empty snapshot for reconnecting clients.

* cluster: periodically rebroadcast the lock ring

Ring updates are purely event-driven, so one lost or poisoned update is
permanent until the next membership change — with a single filer that may
never come. Re-arm a per-group timer after every broadcast so the current
leader keeps re-sending the ring; clients reject nothing newer than their
last accepted version, so a re-sent snapshot always heals a stale view.

* filer,s3api: reset the lock ring on master change

Ring versions are per-master monotonic — each master stamps wall-clock
nanoseconds — so a late high-version update accepted from a former leader
makes the new leader's snapshot look stale forever. Detect a leader
change across the reconnect gap (currentMaster is cleared between
attempts, so remember the last served master) and reset the ring to
bootstrap state so the new leader's view always applies.

* cluster: fail lock acquisition when no lock server exists

retryUntilLocked loops forever, so a filer reporting an empty lock ring
wedges every append write indefinitely. Bound only the "no lock server
found" case — ordinary contention is still waited out since the holder
releases eventually. The constructors now return nil on failure: the
filer append path and S3 object writes fail fast, while mounts degrade
to their existing lockless mode.

* cluster: reset only the ring version on master change

Ring versions are per-master monotonic, so a version gate reset is all a
leader change needs. Clearing the whole ring made every filer its own
write owner until the next update and dropped the prior-owner window for
keys the new leader remaps; the last ring now keeps routing until the
new leader's snapshot transitions off it.

* cluster: skip redundant ring installs and defer rebroadcasts

An unchanged member list now only bumps the accepted version instead of
installing a snapshot: periodic rebroadcasts no longer fire the
topology-change callback or restart the prior-owner window. And a
rebroadcast that lands inside a membership stabilization window yields
to the pending timer rather than publishing an intermediate ring.

* cluster,mount: bound lock unavailability, fail ops that cannot lock

Only 'lock already owned' contention retries without bound now; every
other failure — no lock server, or a dead ring member refusing
connections — shares the same unavailability budget, so a ring naming
departed filers can no longer hang a lock forever. Mount open-write,
create, and rename fail with EAGAIN when the required lock cannot be
acquired instead of proceeding without cross-mount serialization.

* cluster: check pending stabilization inside the broadcast critical section

rebroadcast released the mutex between the pending-timer check and
nextBroadcastUpdate, so a membership change arriving in the gap could arm
a stabilization timer while the rebroadcast emitted an intermediate ring.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: acquire path locks before mutating create/rename state

Create took the DLM lock only after the filer create, so a lock failure
returned EAGAIN with an eagerly persisted file left behind. Rename marked
source handles renamed before acquiring locks, so a failed acquisition
left them suppressing old-path flushes for a rename that never happened.
Both now take the locks first; the create's lock is released again if the
entry race loses to another creator and AcquireHandle takes over.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: keep the old-path lock when rename lock migration fails

The migration stopped the handle's lock before acquiring the replacement,
so a nil result left the handle writing with no lock at all. Acquiring the
new-path lock first means failure keeps the existing lock instead of
reporting success with serialization dropped.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: skip new-path rename lock when a handle already holds it

A target file open for write on this mount already carries a lock on
newPath; the lock manager does not grant a second lock to the same
owner, so the rename would wait on itself until the handle closed.
Also avoid locking twice when old and new paths coincide.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: hand the rename's target lock to the migrating handle

The rename holds a lock on newPath for its duration, so the response
migration's fresh acquisition waited on that same lock until the handle
released — under fhLockTable, blocking the handle's own close. Adopt the
rename's lock directly; nested move responses still acquire their own.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: move the replaced target's lock to the renamed handle

When the target path was already locked by an open handle on this
mount, the migrated source handle kept only its stale old-path lock —
the target's close would then release the last lock on the new path
while the renamed handle was still open. Adopt the replaced handle's
lock instead.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: stop the handle lock inside the fh lock on release

ReleaseHandle stopped fh.dlmLock before taking the fhLockTable slot, so
a rename migration holding that slot could still observe and adopt a
lock that was already stopping. Stopping under the fh lock makes the
transfer serialize against the release.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the replaced target's lock for the renamed handle

When the target path is already locked by an open handle on this mount,
adopting it at migration time keeps the renamed path protected after
that handle closes, without waiting on a lock this mount already holds.
If the handle was released mid-migration the claimed lock is stopped,
and a fresh acquire covers the case where it was already gone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: claim the target handle's lock before the rename runs

Skipping the new-path lock when a handle already holds it let that
handle's close release the lock mid-rename, leaving the path unguarded
until the response migrated it. Take over the lock at check time and
hold it for the rename's duration: the response adopts it for the
migrating handle, or it returns to the target handle / is released on
failure. The target handle lookup also falls back to the entry's stored
inode for a forgotten path mapping.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* mount: read handle locks only under the fh lock during rename

The loose dlmLock reads raced ReleaseHandle, which now mutates the lock
inside the handle lock; check and claim it under the same hold.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 11:57:00 +08:00
github-actions[bot] 4914c14982 docs: regenerate star history chart 2026-09-26 00:41:51 +00:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2f6c237238 filer: keep lazy remote reads from resurrecting deleted paths (#11452)
* filer: keep lazy remote reads from resurrecting deleted paths

Under a remote mount with filer.remote.sync as write-back, a path that
was deleted or renamed away could come back as a chunkless remote-only
entry: between the local delete and the daemon's remote delete, a store
miss made maybeLazyFetchFromRemote trust a bucket that was behind the
filer. The ghost then outlived the remote object -- HEAD answered 200,
GET failed, and nothing cleaned it up.

The filer now tombstones paths it deletes under a remote mount, learned
both synchronously from its own delete path and from peer metadata
events. The lazy fetch and the lazy listing skip a tombstoned path until
the path is written again, until the mount's persisted write-back sync
offset has passed the delete event (the remote delete has landed), or
until a generous TTL covers a mount without a daemon.

Fixes #11440

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: cover recursive remote deletes with an ancestor tombstone

A recursive delete now records the directory tombstone before walking
children, so a partial traversal or a store that drops the subtree
without listing it still leaves every descendant covered. Directory
tombstones also subsume older descendant entries on add, descendant
adds covered by a standing ancestor are skipped, and an existing
tombstone can be refreshed even at capacity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote tombstones to the deleted object's generation

A remote object whose own mtime postdates the local delete is a new
generation, not the one the tombstone hides, so a recreated directory
can surface remote writes made after its delete while old-generation
objects stay hidden. Lazy fetch now stats the remote object before
deciding, listings pass each child's remote mtime, and a sync offset
releases a tombstone once it reaches the delete's own timestamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: rebuild remote deletion tombstones after restart

In-memory tombstones are lost on restart while remote write-back
offsets persist, so a filer boot replays the persisted metadata log
from the oldest mount offset and folds deletes back into the tombstone
set through the same event handler. Lazy remote reads hold off while
the replay runs so a pending delete cannot resurrect in the gap.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: release remote tombstones only after their delete event lands

The write-back offset orders against event timestamps, but the synchronous
delete path recorded tombstones with the local clock before its event was
emitted — a later unrelated event could already have pushed the mount's
watermark past that guess, releasing the tombstone before the daemon
applied the delete. Tombstones recorded ahead of their event are now
marked pending and can only be lifted by the event confirming them or by
TTL; event-stamped tombstones release through the offset as before.

The remote-mtime generation bypass is dropped: remote and filer clocks
are independent, and a pending remote delete removes whatever object sits
at the path, so a "newer" remote object would only resurrect as a
phantom. Tombstoned lookups now skip the remote stat entirely.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: drop dir tombstone when recursive delete fails before listing

The ancestor tombstone is recorded before the child listing; if that
listing fails nothing was deleted, and the leftover tombstone would hide
still-existing remote children for the whole TTL. Tombstones for children
already deleted stay, since their remote deletes are still owed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: block lazy remote reads on startup tombstone rebuild

The rebuild gate is now a done-channel set synchronously before the
replay goroutine starts, so no lazy read can slip through in between.
Reads wait on it with context cancellation instead of returning an
empty miss that makes remote-only objects look deleted.

The replay start is floored at now-TTL: mounts without a recorded
write-back offset previously replayed the whole persisted history, and
events older than the TTL would only build already-expired tombstones.
The gate check now runs after the mount lookup so replaying the meta
log's own directory listings does not deadlock on the gate, and the
replay retries with backoff until it succeeds instead of failing open.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: mark restamped tombstone pending until its delete event lands

When a local delete raises an existing tombstone's timestamp, the new
value is only a local clock guess ahead of that delete's event. Leaving
the tombstone un-pending lets a write-back offset release it before the
event is actually consumed, reopening the resurrection window.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound tombstone replay to the tombstone TTL

Persisted-log replay retried forever, keeping lazy remote reads gated
indefinitely when the log cannot be read. Cap retries at the tombstone
TTL measured from replay start: past that point every tombstone would
have expired anyway, so opening the gate loses no protection.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: re-check deletion tombstone before persisting lazy fetch

A delete landing while StatFile is in flight passed the earlier
tombstone check but still persisted the fetched entry, resurrecting a
path whose remote delete is pending. Re-check right before CreateEntry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract a lazily persisted entry when a delete raced the insert

The pre-insert tombstone check still leaves a window between the check
and the store insert. Since deletes always record the tombstone before
removing the entry, a tombstone visible right after a successful insert
means the delete already ran: delete the entry back out so the
tombstoned path stays deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: note why the replay deadline can safely open the gate

Deletes made after startup are captured by the live delete and event
paths, so a stalled replay can only be missing pre-restart deletes, all
of which are past the tombstone TTL by the deadline.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract only the entry a lazy remote read materialized

Deleting by path after a raced delete could remove a legitimate rewrite
that replaced the fetched entry. Verify the stored entry still matches
the remote object (or the just-created directory shape) before deleting,
and apply the same post-insert check to lazy listing children.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: require full-entry equality before retracting a lazy entry

Remote-only matching still removed a write that had updated the fetched
entry, e.g. appended chunks. Compare the persisted entry against what
this read materialized; any change means a real update owns the path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 08:00:39 +08:00
Eliah RusinGitHubClaude Fable 5.1Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
df4995b894 volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock (#11443)
* volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock

make_volume_readonly looked the volume up, notified the master (step 1),
then took the store write lock (step 2) and marked the volume only `if
let Some(..)`. When the volume left the store during step 1 -- a master
round trip, during which an unmount or a heartbeat expiry can land --
the missing else meant the RPC reported success for a volume the server
no longer has, and step 3 told the master again that it is read-only.

Go's Store.MarkVolumeReadonly (weed/storage/store.go) returns
"volume %d not found" when findVolume comes back nil, and
makeVolumeReadonly (weed/server/volume_grpc_admin.go) returns that error
before the step-3 notification. The Rust step 2 now does the same:
find_volume_mut(vid) -> Status::not_found("volume {vid} not found"), and
the `?` skips step 3, as it already did for a set_read_only_persist
failure. The scrub caller already matches NotFound to skip such a
volume instead of failing the whole report; it now actually gets it.
volume_mark_writable already returns NotFound under its write lock.

The regression test opens the step-1 window deterministically: step 1
awaits the current_master_url read lock, so the test holds its write
guard, lets make_volume_readonly park there after its own lookup
succeeded, unmounts the volume, then releases the guard. With no master
configured the notification is a no-op, so the write lock in step 2 is
the only place left that can notice the volume is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments on the vanished-volume mark-readonly path

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:05:53 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
b750853c42 shell: refuse s3.bucket.create on an existing bucket (#11455)
* shell: refuse s3.bucket.create on an existing bucket

CreateEntry without o_excl replaces the bucket entry, dropping every
extended attribute: lifecycle configuration, owner, versioning and the
irreversible Object Lock flag. Send o_excl so a re-run fails with
'bucket already exists' instead of silently resetting the bucket.

* filer: fail exclusive creates when the lookup itself fails

CreateEntry discards FindEntry errors, so an o_excl create hitting a
transient store failure would take the insert path and upsert over the
entry it was meant to preserve. Propagate the lookup error when o_excl
is set; non-exclusive creates keep their existing semantics.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* shell: test s3.bucket.create requests an exclusive create

Exercises the command end to end through a fake filer gRPC server and
asserts the OExcl flag reaches the wire along with the already-exists
error path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* shell: synchronize captured requests and assert the exact bucket error

The fake filer records CreateEntry requests on the gRPC server goroutine,
so reads need the same mutex; the test also now checks for the exact
"bucket my-bucket already exists" message rather than any error that
mentions existence.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 22:04:54 +08:00
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
317e756b9a volume server: validate ext and collection in gRPC CopyFile/ReceiveFile (Rust) (#11451)
* volume server: validate ext and collection in gRPC CopyFile

Port the Go-side checks (checkVolumeFileExtension, checkVolumeCollection)
to the Rust volume server so a client-supplied collection or ext carrying
a separator or ".." cannot fold a path outside the volume directory.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: validate ext and collection in gRPC ReceiveFile

Same port on the write path: the file ReceiveFile creates is built from
client-supplied fields, so reject traversal there too. Reported through
the response error field, matching Go's SendAndClose.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-25 09:47:16 +08:00
5b79f51e3c filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks (#11435)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks

updateLocalEntry records the RemoteEntry stamp after an upload by writing the
event's entry back with UpdateEntry. The filer deletes every stored chunk
absent from an updated entry, so when the file was rewritten while its upload
was in flight (or the event is a replay), the stale snapshot deletes the
rewrite's chunks: the entry then points at the new fid with no needle behind
it, and the rewrite's own upload fails and is skipped as superseded.

The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's
chunk fids, evaluated by the filer under the path lock. A refused stamp means
the filer moved past this event; the superseding event follows in the log and
stamps the current entry, so the refusal is logged and skipped like a
superseded upload.

Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount,
filer.remote.sync; hold the remote (docker pause) so one upload stays in
flight, rewrite the file through the filer, unpause. Before: the entry's chunk
is 404 on every volume server. After: the stale stamp is refused, the rewrite's
chunk stays live and reads back after a vacuum.

* filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored

The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a
rewrite that touched inline content or metadata alone still compared
equal and the stale snapshot overwrote the live entry. The new clause
compares the whole stored entry against the event's entry under the
same path lock.

* filer: route conditional UpdateEntry to the entry's owner filer

Two filers locking the same path locally could still pass a stale
condition on the non-owner while the owner's entry had moved on. When a
condition or expected_extended precondition is set, forward the request
to the entry's owner the same way conditional CreateEntry does, with
is_moved bounding the hop.

* filer: compare IF_ENTRY_EQUAL against the normalized expected entry

FindEntry grows FileSize to the chunk extent, so a raw event entry with
FileSize still zero failed the condition on an unchanged file and the
stamp was skipped, letting a replay upload the object again.

* filer.remote.sync: classify refused stamps by gRPC status only

A FailedPrecondition substring in an unrelated error would have been
swallowed as a skipped stamp; status.FromError already unwraps.

* remote sync: keep the event entry intact for IF_ENTRY_EQUAL

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-25 09:02:34 +08:00
635f69a821 test(s3tables): add an OLake Iceberg catalog integration test (#11441)
OLake (github.com/datazip-inc/olake) is a CDC and ingestion engine that writes
Apache Iceberg. It covers two paths none of the existing catalog suites reach.

It is a strict Java Iceberg client. OLake does not write Iceberg from Go — its
Go process spawns a Java sidecar over gRPC and writes through the official
Apache Iceberg library, because the Go library has no equality deletes and CDC
needs them. That makes it the client class weed/s3api/iceberg's metadata
compliance backfill exists to serve: the one that fails with "Cannot parse
missing long current-snapshot-id" when spec-required keys are omitted.

And it produces equality deletes. Its upsert path commits operation=overwrite
with an equality-delete file and a delete manifest. ClickHouse, Doris, Trino,
Spark and DuckDB all only append, so nothing else in this directory exercises
a delete manifest at all.

Six subtests: the destination check reaches SUCCEEDED and actually loads the
REST catalog; discover enumerates the source; a full sync commits a snapshot;
PyIceberg reads back what the Java writer committed; an update plus a re-sync
records an overwrite carrying equality deletes with a delete manifest in the
current snapshot; and the catalog does not rewrite manifests the official Java
writer produced, which gives the manifest-repair path a negative test to go
with the ClickHouse positive one.

What it deliberately does not assert is a delete-applied read. PyIceberg
refuses to scan a table carrying equality deletes (apache/iceberg#6568) while
reading its metadata fine, and an engine that can apply them costs a
multi-gigabyte image. Recording the commit correctly is the catalog's
contract; applying deletes on read is the engine's. The README says so, says
the read half was verified by hand once with StarRocks, and warns against
later "upgrading" this to a PyIceberg rows read — which would either fail or,
if PyIceberg ever starts skipping deletes instead of raising, pass by not
looking.

The workflow job asserts the suite actually ran rather than trusting a green
exit, for the same reason: at least one top-level PASS and zero SKIP.

No product change — the destination config is the generic catalog_type=rest
with the standard OAuth2 client-credentials flow, and s3_path_style is not even
set, since OLake turns it on itself whenever s3_endpoint is non-empty.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-25 07:39:27 +08:00
Chris LuandGitHub 11791fad6a filer: resolve the collection a bucket delete drops (#11439)
* filer: resolve the collection a bucket delete drops

A bucket delete dropped the collection named after the bucket, which
assumes bucket name is collection name. With a collection rule the
write path honors, deleting the bucket either orphaned its collection
or, when a bucket was named after a shared collection, removed volumes
other buckets still write to.

Resolve the collection through the same rule chain the write path uses
and drop it only when no other bucket resolves there too. A listing
failure keeps the collection, the safe side of an unknown.

* filer: prove collection exclusivity across all paths before dropping it

The sibling-bucket scan missed every non-bucket writer: a broad rule like
'/' or '/buckets/', a rule under a surviving bucket, or a rule on an
unrelated path can route into the same collection. Check every storage
rule's prefix instead, and mirror the grouped gateway's explicit
<group>_<bucket> collection, which otherwise resolves a rule-named
collection the bucket never wrote to.

* s3: let the filer own the collection decision on bucket delete

Both entry points deleted a name-derived collection around the filer's
own resolved delete, bypassing its exclusivity check and wiping sibling
data. The filer now resolves the collection a bucket actually used,
including the grouped form.

* filer: keep a collection the default write route also uses

Rule-less writes outside buckets land in the filer's default collection,
so a bucket resolving there shares it with them.
2026-09-25 07:30:42 +08:00
Chris LuandGitHub 56d2f05ccd topology: wake the vacuum dispatcher when a worker frees quota (#11436)
* topology: wake the vacuum dispatcher when a worker frees quota

The dispatch loop slept a fixed 10s whenever every pending volume was
waiting for a per-server quota slot, so a sweep took volumes x 10s
regardless of how fast the compactions were. Workers now signal on a
buffered channel after crediting quota; the dispatcher waits on it with
the 10s sleep kept only as a timeout.

* master: add -vacuumIntervalSeconds to tune the automatic sweep interval

The 14-minute base interval was a literal inside the refresh loop while
every neighbouring vacuum knob was already a flag. Defaults to 840s,
unchanged.

* topology: keep the 14 minute floor on the vacuum interval

A zero-valued MasterOption or a negative -vacuumIntervalSeconds left the
sweep sleeping only its jitter, so treat non-positive intervals as the
previous default.
2026-09-25 02:14:59 +08:00
Chris LuandGitHub 8c1be63c92 ecbalancer: let a non-overflow parity shard leave a data-bearing rack (#11438)
The parity pass only queued shards past the per-type cap, so a single
parity shard sharing a rack with data was never a move candidate even
when an empty data-free rack existed (2+1 over 3 DCs settled 2/1/0).
Non-overflow candidates now move too, but only to a rack without data;
overflow shards keep the existing data-rack fallback.
2026-09-25 02:14:23 +08:00
bb9942c646 volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes (#11413)
* volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes

The Rust VolumeEcShardsGenerate went straight into write_ec_files: no unload
of an already-mounted EC volume and no stale-artifact sweep. Only .ec00..ecNN
on the encoding disk were truncated, so a retry could mix two encode runs. A
stale N.ec03 left on a sibling disk survived, reconcile later mounted it
against the new .ecx, and the new .vif made the encode_ts_ns identity guard
pass, so reads served old-run bytes at new-run offsets.

Mirror Go's VolumeEcShardsGenerate (#9880 / #9953): UnloadEcVolume on every
disk, then removeStaleEcArtifacts on every disk location before encoding.
remove_ec_volume_files_full_teardown already has removeStaleEcArtifacts'
semantics (.ec00..ec31, .ecx/.ecj/.ecsum[.vN] in both the data and idx dirs,
.vif only on a shard-only disk; never the source .dat/.idx), so reuse it. Add
Store::unload_ec_volume, which unlike remove_ec_volume does not stop at the
first disk and closes the descriptors so the unlink frees the inodes. The
store write lock covers only unload + sweep, not the encode.

The failure arm now also drops the generation-0 .ecsum, as Go's defer does.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: wake the heartbeat after VolumeEcShardsGenerate unloads shards

The pre-encode unload drops mounted EC shards from memory, but unlike every
other unmount path it did not wake the heartbeat, so the master kept routing
reads to shards this server no longer serves until the next pulse. Notify
once the store lock is released, and before the sweep error propagates: a
failed sweep has unloaded the shards too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: clean up encode artifacts when the .vif write fails too

Go's shouldCleanup defer covers every error before the .vif commits,
not just a failed encode. A serialize or write failure on the .vif left
the fresh .ecNN/.ecx/.ecsum behind, which the next generate would have
to rely on the new sweep to remove. Extract the cleanup and run it on
the .vif error paths as well.

* volume server: write the EC .vif atomically

Go's SaveVolumeInfo writes a temp file, syncs it, and renames it over
the target, so a failed write leaves the previous metadata intact and a
read-only .vif fails the save. The direct fs::write truncated the file
first, so a write or sync failure could leave an empty .vif even after
cleanup_encode removed the generated shards.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:10:54 +08:00
f0afcf904d volume: an EC volume needs an .ecx to mount, and a 0-byte stub never outranks a real index (#11415)
* volume: an EC volume needs a non-empty .ecx to mount

Two gaps against Go in how the Rust volume server treats the .ecx.

EcVolume::new mounted with no index at all. The per-shard
VolumeEcShardsMount path picks the disk by shard file alone, so a shard
whose .ecx was on no local directory still registered and was
advertised to the master; every VolumeEcShardRead then failed with
"ecx file not open", and add_shard's 0-byte guard was neutralised
because ecx_file_size stayed 0. Go's NewEcVolume returns an error
wrapping os.ErrNotExist. EcVolume::new now fails with NotFound, and
Store::mount_ec_shard looks up the .ecx owner across all disks first
(findEcxIdxDirForVolume) so a shard on a sibling disk of its index
still mounts instead of turning into a hard failure.

A 0-byte .ecx stub, as left by a failed EC distribute copy, counted as
a valid index. Go requires Size() > 0 wherever the file steers a
decision: HasEcxFileOnDisk, findEcxIdxDirForVolume, indexEcxOwners
(shared by reconcile and mirror), and VolumeEcShardsCopy removes a
copied 0-byte .ecx and fails the copy. Mirror each through one
is_usable_ecx_file helper. NewEcVolume itself still accepts a lone
0-byte .ecx as a legitimate empty index, but prefers a non-empty copy,
local directory first, over a stub in the other directory; the
resolution in EcVolume::new now follows the same order.

Tests that mounted EC volumes without any .ecx get a real fixture.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: mount_ec_shard tries every disk; reconcile ignores a 0-byte local .ecx

mount_ec_shard returned the first disk's error, so an unusable shard copy
(a 0-byte .ecNN left by an interrupted move) hid a good copy on the next
disk. Like Go's MountEcShards, keep scanning: NotFound means "not this
disk", any other failure is collected, and an all-disks-fail error names
every disk tried. "No .ecx on any local disk" is now told apart from
"shard not on this server".

The orphan-shard reconcile took its locally-mirrored fast path whenever a
local .ecx existed at all. A 0-byte stub there registered the shards against
an empty index while the owner index skipped that same stub. Go gates the
fast path on HasEcxFileOnDisk; do the same. ec_local_ecx_path loses its last
production caller and becomes test-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: match Go's mount error text and skip the owner stat on the owning disk

MountEcShards in Go skips the HasEcxFileOnDisk stat when the disk's own
directories already hold the .ecx, dedups a shared -dir.idx across
locations in findEcxIdxDirForVolume, and reports "load failures" with
the same wording. Also drop two issue-number references from comments.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:08:56 +08:00
c1ccbcda13 volume server: the write queue answers uploads with the needle's real ETag (#11414)
With SEAWEED_WRITE_QUEUE=1 every upload came back with ETag "00000000".
The upload handler built the needle with Needle::default(), so its
checksum was CRC(0), and handed a clone of it to the queue. The CRC was
only computed in the write path, on the worker's clone, and WriteResult
carries no checksum back, so n.etag() in the handler formatted the zero
checksum. The direct path writes through &mut n and was correct.

Compute the checksum in the handler while building the needle, the way
Go's CreateNeedleFromRequest does, over the same bytes the write path
hashes (the stored data, gzipped or not). The ETag and the has-name flag
are read before the write, so the needle is moved into the queue instead
of cloned, which also drops a full payload copy per queued upload.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 07:08:36 +08:00
94a68fa9b9 volume: walk_index_file keeps row alignment across short reads (#11412)
* volume: walk_index_file keeps row alignment across short reads

walk_index_file issued one Read::read per batch and decoded whatever came
back. Read::read may legally return a short count that is not a multiple
of the 17-byte entry size (FUSE and network filesystems, a BufReader whose
capacity is not a multiple of 17). The split entry at the end of the batch
was dropped with no carry and the next read started mid-entry, so every
later row was decoded from misaligned bytes and fed to the index as a
garbage key/offset/size. This function backs every in-memory index load.

Go's WalkIndexFile is immune because it reads through io.ReaderAt, which
returns a full buffer or an error. Fill the batch buffer until it is full
or the reader reports EOF, retrying ErrorKind::Interrupted, and only then
decode whole entries. Reads stay batched at ROWS_TO_READ entries.

EOF semantics are unchanged and match Go: on io.EOF Go decodes the whole
entries in the final buffer, ignores a trailing partial entry and returns
nil. A torn final entry is still skipped without an error here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume: trim walk_index_file comments

The batch-fill loop and the ShortReader test helper each carried a
paragraph where a sentence suffices.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-24 07:08:26 +08:00
Chris LuandGitHub b3a8701989 lance: authenticate the catalog with Bearer tokens and x-api-key (#11431)
* lance: accept OAuth2 bearer tokens for catalog auth

Lance and LanceDB clients can only send OAuth2 / Bearer / API-Key
headers on catalog calls, never SigV4, so behind an auth-enabled S3
gateway every namespace request failed with 403 Access Denied.

Mirror the Iceberg catalog's OAuth2 support: POST /oauth/token accepts
an S3 access key / secret key as client_id / client_secret, validates
them against IAM, and returns a signed JWT. The Auth middleware accepts
that token as a Bearer credential before falling through to SigV4.

Closes #11430

* lance: accept x-api-key header carrying an S3 credential

The Lance namespace spec's third auth scheme maps api_key onto the
x-api-key header. Accept "access_key:secret_key" there and validate it
against IAM, so clients that only hold static headers can authenticate
without minting a token first.

* lance: answer invalid_client with the Basic challenge

RFC 6749 5.2 requires a 401 from the token endpoint to carry
WWW-Authenticate matching the scheme the client used, so it knows how
to retry.

* lance: cap the token endpoint request body

/oauth/token is unauthenticated, so ParseForm needs the same size
bound decodeBody applies to every other catalog request.

* lance: keep query strings out of request logs

/oauth/token rejects a client_secret sent in the query, but the
logging middleware and the catch-all wrote RequestURI to the log
before that rejection ran. Log the path alone so a mis-sent secret
never reaches the log.

* lance: log the escaped path, not the decoded one

URL.Path decodes percent escapes, so a request like /%0aFORGED could
split log lines. EscapedPath keeps the encoding while still dropping
the query string.
2026-09-24 07:02:59 +08:00
b9ad62fc16 [Volume] Keep DAT and index state consistent after async batch Sync failure (#11425)
* fix 11400

* persist failed-recovery quarantine and harden rollback

- record the unavailable state in a .unavailable marker, fsync it, and
  re-arm it on load so a restart cannot serve an unverified pair
- quarantine the volume so heartbeats stop advertising it
- block MarkVolumeWritable while unavailable, rechecked under noWriteLock
- fail every request of a failed batch, not only the succeeded ones
- restore the needle map and truncate .dat on inline fsync rollback failure
- add truncateIndex for the sorted-file needle map
- mirror the fail-closed semantics in the Rust volume server

* volume: erase rolled-back mappings instead of leaving tombstones

A rolled-back batch or failed inline write used Delete() to undo a
needle that did not exist beforehand, leaving a tombstoned map entry
whose stale offset makes the next write to that needle fail reading a
header that no longer exists. Add removeMapping/restoreMapping to the
mappers so recovery erases entries that were absent before the batch
and reinstates the exact prior offset/size for ones that were,
including tombstones. The index row still goes through Delete so a
replay forgets the needle.

* volume: gate bulk readers on unavailable and fsync the marker's dir

- fsync_dir(&self.dir) synced the volume dir's parent, not the dir
  holding .unavailable; pass the marker path so the create survives
  a host crash
- export UnavailableError and check it in ReadAllNeedles,
  VolumeTailSender, VolumeIncrementalCopy, and IncrementalBackup so
  replica-sync paths cannot stream or append data from an unverified
  .dat/.idx pair; mirror on the Rust side via read_dat_slice,
  read_all_needles, dat_scan_plan, and the incremental-copy handler

* volume: drop issue references from comments near touched code

* volume: stop active scans when the volume becomes unavailable

The stream entry-point checks ran once per RPC, so a volume quarantined
by a failed recovery mid-scan kept serving data. Recheck availability
per needle/chunk on the detached read paths: tail scan and heartbeat,
read-all, incremental copy, incremental backup writes, and the Rust
StreamingBody chunk reads. Rust incremental copy also rejects a
quarantined volume before sync_to_disk touches the backend.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-24 06:57:44 +08:00
Chris LuandGitHub d848b8ed00 rust volume: regression test for makeup_diff replay across a 32 GiB offset boundary (#11410) (#11427)
* rust volume: test makeup_diff replay across a 32 GiB offset boundary

Issue #11410 corrupted a replayed write's index offset in Go's makeupDiff
by patching only four of the five offset bytes. The Rust makeup_diff
already encodes the whole offset through idx_entry_to_bytes and
Offset::from_actual_offset; this adds the mirror of
TestConcurrentWriteCrossesOffsetBoundary so a regression would fail here
the same way it does under -tags=5BytesOffset on the Go side.

Sparse-truncate the .dat to 64 GiB, compact, write, commit: the index
offset must equal the .cpd size and the needle must stay readable
through a second vacuum. Gated on the 5bytes feature since a 64 GiB
.dat exceeds the 32 GiB range of 4-byte offsets.

* rust volume: skip the offset-boundary replay test on Windows

Windows set_len allocates the full 64 GiB extension instead of a sparse
range, so the test fails with StorageFull on CI runners. Gate it to unix,
where set_len leaves the extension unallocated.
2026-09-23 18:08:20 +08:00
github-actions[bot] 196c71b613 docs: regenerate star history chart 2026-09-23 00:46:51 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
559ec33498 build(deps): bump github/codeql-action from 4.38.0 to 4.38.1 (#11423)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.0 to 4.38.1.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.38.0...v4.38.1)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:47 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d6397fdc50 build(deps): bump dorny/paths-filter from 3 to 4 (#11422)
Bumps [dorny/paths-filter](https://github.com/dorny/paths-filter) from 3 to 4.
- [Release notes](https://github.com/dorny/paths-filter/releases)
- [Changelog](https://github.com/dorny/paths-filter/blob/master/CHANGELOG.md)
- [Commits](https://github.com/dorny/paths-filter/compare/v3...v4)

---
updated-dependencies:
- dependency-name: dorny/paths-filter
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:38 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
264030c08c build(deps): bump docker/setup-qemu-action from 4.3.0 to 4.4.0 (#11421)
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 4.3.0 to 4.4.0.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](https://github.com/docker/setup-qemu-action/compare/v4.3.0...v4.4.0)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: 4.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:30 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
8099e71934 build(deps): bump golang.org/x/crypto from 0.56.0 to 0.57.0 (#11419)
Bumps [golang.org/x/crypto](https://github.com/golang/crypto) from 0.56.0 to 0.57.0.
- [Commits](https://github.com/golang/crypto/compare/v0.56.0...v0.57.0)

---
updated-dependencies:
- dependency-name: golang.org/x/crypto
  dependency-version: 0.57.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 16:14:21 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
63eb67c9da build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts from 1.50.0 to 1.51.0 (#11420)
build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts

Bumps [github.com/aws/aws-sdk-go-v2/service/sts](https://github.com/aws/aws-sdk-go-v2) from 1.50.0 to 1.51.0.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.50.0...service/s3/v1.51.0)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/service/sts
  dependency-version: 1.51.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:37 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
93efc64af6 build(deps): bump go.uber.org/atomic from 1.11.0 to 1.12.0 (#11416)
Bumps [go.uber.org/atomic](https://github.com/uber-go/atomic) from 1.11.0 to 1.12.0.
- [Release notes](https://github.com/uber-go/atomic/releases)
- [Changelog](https://github.com/uber-go/atomic/blob/master/CHANGELOG.md)
- [Commits](https://github.com/uber-go/atomic/compare/v1.11.0...v1.12.0)

---
updated-dependencies:
- dependency-name: go.uber.org/atomic
  dependency-version: 1.12.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:29 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f2b08e47e5 build(deps): bump github.com/tarantool/go-tarantool/v3 from 3.0.1 to 3.0.2 (#11417)
build(deps): bump github.com/tarantool/go-tarantool/v3

Bumps [github.com/tarantool/go-tarantool/v3](https://github.com/tarantool/go-tarantool) from 3.0.1 to 3.0.2.
- [Release notes](https://github.com/tarantool/go-tarantool/releases)
- [Changelog](https://github.com/tarantool/go-tarantool/blob/master/CHANGELOG.md)
- [Commits](https://github.com/tarantool/go-tarantool/compare/v3.0.1...v3.0.2)

---
updated-dependencies:
- dependency-name: github.com/tarantool/go-tarantool/v3
  dependency-version: 3.0.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:20 +09:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
5079d926e4 build(deps): bump cloud.google.com/go/kms from 1.33.0 to 1.34.0 (#11418)
Bumps [cloud.google.com/go/kms](https://github.com/googleapis/google-cloud-go) from 1.33.0 to 1.34.0.
- [Release notes](https://github.com/googleapis/google-cloud-go/releases)
- [Changelog](https://github.com/googleapis/google-cloud-go/blob/main/documentai/CHANGES.md)
- [Commits](https://github.com/googleapis/google-cloud-go/compare/kms/v1.33.0...dlp/v1.34.0)

---
updated-dependencies:
- dependency-name: cloud.google.com/go/kms
  dependency-version: 1.34.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-22 15:57:13 +09:00
github-actions[bot] a9e1b57dcd docs: regenerate star history chart 2026-09-22 00:46:32 +00:00
Chris LuandGitHub e2608edda4 volume: fix 5-byte index offset corruption in makeupDiff (#11411)
* volume: encode all offset bytes when makeupDiff replays a write

makeupDiff patched only bytes 8:12 of the index entry, so under the
5BytesOffset build the fifth byte kept the old offset's high bits and the
replayed needle's index pointed 32 GiB-aligned ranges away from its body.
A later vacuum then dropped the entry as unreadable. Rebuild the entry with
needle_map.ToBytes, the same encoder the tombstone branch just below uses.

* volume: test makeupDiff replay across a 32 GiB offset boundary

Sparse-file test: truncate the .dat to 64 GiB after one write, compact,
write a second needle, commit, and assert the index offset matches the
.compacted size and the needle stays readable through a second vacuum.
Only runs under -tags=5BytesOffset.
2026-09-21 00:59:35 -07:00
0f2ecb766f volume server: reject non-ASCII input instead of panicking (#11406)
* volume server: reject non-ASCII input instead of panicking

Three parsers sliced attacker-supplied strings by byte offset, so a
multi-byte character split inside itself and panicked the task:

  - parse_needle_id_cookie took the last 8 bytes as the cookie and the
    rest as the needle id. Reachable from VolumeServer.BatchDelete,
    whose file_ids come straight off the wire as protobuf strings;
    that handler already answers 400 per bad fid, so the guard turns a
    panicked RPC into the error it was already written to return.

  - TTL::read took the unit as the last byte and the count as
    everything before it, so "?ttl=5<multi-byte>" split mid-character.
    The HTTP upload path does TTL::read(..).ok() and drops an invalid
    TTL; AllocateVolume maps the Err to InvalidArgument.

Both now reject non-ASCII up front. Hex and a digits-plus-unit TTL are
ASCII by definition, so no accepted input changes -- covered by tests
alongside the rejection cases.

The six response-* header overrides were inserted with
parse().unwrap(). They come from the query string, so
"?response-cache-control=%0Aevil" decodes to a value HeaderValue
rejects and the unwrap panicked the connection task,
unauthenticated. They now skip the override, matching the if-let the
chunked-response path in the same file already uses.

ReplicaPlacement::from_string was reported as a fourth site but is not
one: reaching chars[2] requires chars[0] and chars[1] to be ASCII
digits, which forces the padded string to be three single-byte
characters, so a multi-byte character always lands on a to_digit()
None first. Kept as a regression test rather than a change.

Each fix was confirmed against the unfixed code first: the parser
tests panic with "byte index N is not a char boundary", and the
integration tests panic at handlers.rs:1413 and ttl.rs:88.

Not a vector, contrary to the report: the HTTP request line. The path
is not percent-decoded before parsing, so "%C3%A9" stays ASCII and
fails the length check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: fall back to needle MIME when response-content-type is invalid

Skipping an unparseable override left the response without any
Content-Type because the override had already bypassed the normal MIME
selection. Also correct a test comment that described a chars[2] panic
which cannot be reached.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
2026-09-20 23:35:05 -07:00
Chris Lu 6848cdf9e1 s3: close seaweedfs-quota policy-confusion gap (#11409)
* s3: count seaweedfs-quota as an operation subresource

PUT /bucket?policy&seaweedfs-quota was not rejected by
hasAmbiguousSubresource because operationSubresources omitted the
seaweedfs-quota key. The router then picks the policy route
(registered first) while the IAM action resolver may resolve the
request to s3:PutBucketQuota, letting a quota-only identity write a
bucket policy. Reject the combination before routing, matching the
fix for policy&tagging (#10987).

* s3: resolve seaweedfs-quota after other bucket subresources

The quota routes are registered last among the bucket subresource
routes, but the action resolver found seaweedfs-quota inside the
unordered bucketQueryActions map, so a request carrying it alongside
another selector could be authorized as the quota operation while the
router served the earlier-registered handler. Resolve it explicitly at
the end so the resolver agrees with the router, mirroring how
list-type is handled.

* s3: count resolver subresources in the ambiguity guard

hasAmbiguousSubresource only counted operationSubresources, so adding
a query parameter to the action resolver without updating that list
reopened the authorize-one-serve-another gap. Count bucketQueryActions
keys as operation selectors too, and add a test that walks the
registered routes and fails on any query key that is neither an
operation subresource nor a known modifier.
2026-09-20 23:34:57 -07:00
ca62d4297b volume: load the .ecj deletion journal in chunks, and repair a torn tail (#11408)
* volume: load the .ecj deletion journal in chunks, and repair a torn tail

Two independent defects in the EC deletion journal's load path.

1. The loader issued one NEEDLE_ID_SIZE-byte positional read per entry.

   That is fine for a healthy journal -- kilobytes -- and pathological for a
   large one. A `.ecj` is semantically a SET of deleted needle ids but is
   written as an append-only log that nothing dedupes, and several paths append
   a peer's ENTIRE journal onto the local one (VolumeEcShardsCopy with
   copy_ecj_file, EC index recovery, and ec_decode's deliberate cross-holder
   merge), so a volume whose shards are repeatedly balanced between two servers
   grows the file without bound.

   Observed in production: 1.51 TB and 1.30 TB on the two holders of one 10+4
   volume containing ~100 distinct ids. At that size the per-entry loop is
   ~188e9 syscalls, run synchronously while holding the deleted_needles write
   lock and before the HTTP port opens. The process sits at 100% of one core
   with a small RSS -- the set stays tiny because the ids repeat -- reading at a
   few MiB/s because 8-byte reads defeat readahead, logs nothing after "Adding
   storage location", and ignores SIGTERM. The master then unregisters every
   volume it holds and reads of them fail. 4.46 and 4.47 are both affected.

   Read in 1 MiB chunks and build into a local set, merging once at the end so
   the write lock is not held for the whole scan. Measured on a 256 MiB journal
   of 100 distinct ids: 33,554,500 syscalls -> 257, identical resulting set.

2. A torn tail silently corrupted later deletes.

   The journal handle is in append mode, so writes land at the physical end
   regardless of alignment. A trailing partial record therefore pushed every
   later append out of alignment: the loader skipped the partial bytes, but the
   next mount decoded them together with the leading bytes of the following
   entry, producing one garbage id and dropping the delete that came after the
   tear -- after acknowledging it.

   Truncate to a whole number of records at mount, before anything can append.
   The repair uses its own read+write (non-append) handle: on Windows,
   append(true) requests FILE_APPEND_DATA without FILE_WRITE_DATA (and
   .write(true) is subsumed by .append(true)), so SetEndOfFile through the
   journal handle fails with ERROR_ACCESS_DENIED.

   The same trap exists in journal_delete's recovery path, which calls set_len
   on the append handle to roll back a partial write whose sync failed. It is
   error-handled rather than fatal, so on Windows that rollback silently does
   not happen. Untouched here; worth a separate fix.

Bounding the journal's growth needs compaction, which is deliberately not in
this change: replacing the file under a store that can hold several EcVolume
instances for one volume id requires coordinating with the other holders, and
that belongs at the store layer. Sent separately.

Tests: a journal spanning several read chunks loads every entry; a trailing
partial record is ignored rather than panicking; a torn tail is truncated at
mount and a delete taken afterwards survives a remount.

* volume: roll back a failed .ecj append through a dedicated write handle

The append handle lacks FILE_WRITE_DATA on Windows, so the set_len
rollback after a failed sync silently did nothing and the journal could
drift one record past deleted_needles. Same trap as the torn-tail repair
in this file; fix it the same way. Also format the new tests.

* volume: mirror chunked .ecj load and torn-tail repair in Go

---------

Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:41 -07:00
4bb40732bb volume server: ReceiveFile loses bytes and hides fsync failures (#11407)
* volume server: ReceiveFile loses bytes and hides fsync failures

Three defects in one handler, all on the path that receives a pushed
.dat/.idx/.vif or EC shard:

  - `f.write(&content)` never compared the return to content.len().
    A short write (ENOSPC, NFS) counted only the bytes that landed,
    so every later chunk was written at a shifted offset and the RPC
    answered error: "" with a byte count that looked right. Go's
    os.File.Write loops. Now write_all.

  - `let _ = f.sync_all();` discarded EIO and answered success with
    the full byte count. Go omits the check too, but
    ReceiveFileResponse carries an `error` field and the caller
    renames the staged file into place on success -- so a silent
    fsync failure publishes a file whose data never reached the
    platter. Flush and fsync failures are now reported.

  - Both the per-chunk write and the final fsync were blocking
    std::fs calls inside the async fn, on the runtime worker that is
    also driving the stream. Switched to tokio::fs + BufWriter, the
    shape `drain_copy_stream_to_file` in this same file already uses
    and documents. The partial-file cleanup on the error path moves
    to tokio::fs::remove_file for the same reason.

The handler had no test at all, which is how the short-write bug
survived. Added a round-trip over a real connection with ragged chunk
boundaries, asserting the bytes on disk and not only the reported
count -- a dropped or reordered chunk changes the file even when
bytes_written still adds up.

That test guards the rewrite; it does not reproduce the original
faults. ENOSPC and EIO need fault injection that this suite has no
harness for, so the short-write and fsync paths are argued from the
code, not demonstrated by a failing test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: remove the staged file on every ReceiveFile error reply

Flush and fsync failures returned early and left the partial .copying or
shard file behind, as did the pre-existing write-error path. Route all
response-level errors through one cleanup block, matching Go's
close-and-remove on a failed write.

* volume server: tighten ReceiveFile comments

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:37 -07:00
8ff2e0777e volume server: HTTP DELETE on a distributed EC volume (#11405)
* volume server: HTTP DELETE on a distributed EC volume

The delete handler validated the cookie with EcVolume::read_ec_shard_needle,
which reads only locally-mounted shards and errors "ec shard N not available
locally" for any interval held by a peer. Every Err was mapped to 500 and no
.ecj tombstone was appended, so on a standard 10+4 spread over 14 servers an
HTTP delete of an EC needle could not succeed. The GET path already goes
through read_ec_shard_needle_distributed.

Route the delete's read through the same distributed reader. It does a
local-first pass in its snapshot phase, so the all-shards-local case costs
what it did before, and no store guard is held across the await (the reader
takes its own; RwLockReadGuard is !Send).

Two smaller corrections fall out of the new return type:

  - the reader reports both "needle not in the index" and "volume vanished
    between the has_ec check and the snapshot" as Ok(None), which collapses
    the old Some(Ok(None)) and None arms into one 404;
  - an io::ErrorKind::NotFound now answers 404 rather than 500, matching the
    GET path. Telling a caller to retry a delete that can never succeed was
    half the bug.

The cookie check and its ordering before the journal append are unchanged.

Not addressed here: Rust journals the tombstone locally while Go routes it to
the primary shard holder. That is a separate behaviour change and belongs in
its own PR against the same issue-10 checkbox.

The regression test mounts 13 of 14 shards, leaving out the one holding the
needle's interval. The distributed reader seeds its Reed-Solomon buffers from
locally mounted siblings, so with >= 10 survivors it reconstructs with no peer
fan-out -- which makes the bug reproducible on a single node. Against the
unfixed handler the test fails with 500 vs 202.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* volume server: fail the delete when the EC volume unmounts mid-request

find_ec_volume_mut returning None used to fall through to a 202 with no
.ecj tombstone written, reporting success for a delete that did not
happen. Answer 404 like the other volume-vanished arms so the caller can
retry after a remount.

* volume server: forward EC needle deletes to a primary-shard holder

Mirror Go's doDeleteNeedleFromAtLeastOneRemoteEcShards: the tombstone is
journaled on one holder of the needle's primary data shard via
VolumeEcBlobDelete (or the local journal when this server holds the
shard), falling back to any other shard holder when the primary has
none. Journaling only on the node that received the DELETE scattered
tombstones across whichever server took the request.

* volume server: route BatchDelete EC deletes through the same forwarding

BatchDelete had the same local-journal divergence as HTTP DELETE, plus a
gap the old code admitted in a comment: the .ecx index cannot supply the
needle's cookie, so EC deletes ran with no cookie check at all. A
distributed read now fills the needle for every EC entry — matching Go's
DeleteEcShardNeedle, which reads and compares the fid cookie even when
skip_cookie_check is set — and the tombstone forwards via
delete_ec_shard_needle_distributed. A needle deleted between read and
journal reports 304 like Go's ErrorDeleted; a vanished volume reports
500 so the filer retries.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: chrislusf <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <devin@cognition.ai>
2026-09-20 23:26:31 -07:00
ac876eef21 rust volume: build outgoing gRPC clients in one place and give every dial a connect timeout (#11354)
server/grpc_client.rs stopped at build_grpc_endpoint() -> Endpoint, so all 13
production call sites hand-wrote the same .connect() + X::with_interceptor()
+ two max_*_message_size() lines. Four of them -- VolumeCopy,
VolumeTailReceiver, VolumeEcShardsCopy and the HTTP chunk batch-delete fan-out
-- dialed with no timeout at all, so an unreachable peer whose TCP handshake
never completes (SYN dropped, blackholed route, host behind a silent firewall)
left the operation waiting on the kernel's own retry budget, minutes long.

Add GrpcDialOptions (unary / long / stream presets), connect_channel(), and
volume_server_client() / master_client() / filer_client() constructors that
attach the request-id interceptor and lift both message-size limits, then
route all 13 sites through them. build_grpc_endpoint is private again, so
connect_channel is the only way out of the module and no call site can dial
without picking up a bound. Each site's existing timeouts are preserved
exactly; the four bare dials gain a 5 s connect timeout and nothing else. No
per-request deadline was added to any streaming call: Endpoint::timeout is a
per-request bound on time-to-first-response-headers for every request the
channel carries, so a value picked for one short call would also be the header
deadline for the whole-volume transfer sharing the dial.

The new bound covers the TCP handshake only -- tonic hands connect_timeout to
HttpConnector::set_connect_timeout. A peer that completes the handshake and
then stalls in the TLS or HTTP/2 exchange is still unbounded at those four
sites, as are the RPCs themselves. That is why the three ping_* helpers keep
their outer tokio::time::timeout: replacing it with connect_timeout would have
narrowed a whole-connect bound they already had.

main.rs no longer re-declares GRPC_MAX_MESSAGE_SIZE and the three
keepalive/window constants; it imports them from grpc_client.rs so the
inbound server and the outgoing clients cannot drift apart.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 19:41:37 -07:00
github-actions[bot] ddf009ffe7 docs: regenerate star history chart 2026-09-21 00:50:25 +00:00
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Fable 5.1Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
818f3bb71b rust volume: share the I/O-error tracker between Volume and EcVolume (#11351)
* rust volume: share the I/O-error tracker between Volume and EcVolume

Volume and EcVolume each carried the same three fields - a mutex-held
last error, a consecutive count and a sticky quarantine flag - and the
same four methods over them, identical except for the path qualifier on
is_storage_io_error. The tolerance the count is compared against was a
fourth copy: heartbeat.rs held VOLUME_IO_ERROR_TOLERANCE for volumes,
ec_volume.rs held IO_ERROR_TOLERANCE for EC, and the volume test helper
open-coded the same 3, so the two paths could drift apart silently.

Go keeps this in one place already: weed/storage/io_error.go holds
IoErrorTracker, IoErrorTolerance and isStorageIoError, and Volume embeds
the tracker. Go's EcVolume has to re-implement it only because those
fields are unexported and EC lives in another package.

storage::io_error::IoErrorTracker now owns that state, with record /
state / should_quarantine / mark_quarantined / reset and the single
IO_ERROR_TOLERANCE. is_storage_io_error moves into the same file, so it
sits with the tracker that is now its only caller, the way io_error.go
is laid out. Both volume kinds embed one tracker and keep their existing
method names as delegates, so the ~16 internal call sites and the
readers in heartbeat.rs, store.rs and grpc_server.rs change only where
the two threshold comparisons become should_quarantine().

Volume::last_io_error and EcVolume::reset_io_error_state had no callers
and are gone.

Unchanged: what counts as a storage-media error - is_storage_io_error
changed file, not body, and is still the single predicate both volume
kinds share, where Go's EcVolume tests EIO directly and so misses the
Windows codes. Also unchanged: the tolerance value, the metric increment
on every counted error, and the sticky quarantine - a success clears the
count and the last error but never the flag, which only reset lifts. In
the heartbeat the state read moved inside the quarantine branch, so the
common path no longer takes the tracker's mutex or clones the last-error
string; should_quarantine's two relaxed loads run either way.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: hoist absolute_display_path into server

handlers.rs and ui.rs each held a byte-identical copy of the helper that
turns a configured -dir into an absolute path for display. The status
JSON and the status page are meant to show the same directory, so the
two copies had to be edited together to stay that way.

The helper now lives in server/mod.rs as pub(crate) and both callers use
it. No behaviour change: same body, same call sites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: keep EcVolume::reset_io_error_state

Moving both volume types onto the shared IoErrorTracker dropped
EcVolume's public reset while Volume kept its own, so the two sides of
the tracker drifted apart.

mark_quarantined is sticky: a later successful read clears the error
count through record(), but the quarantine flag only comes down through
reset(). Without the delegate an EC volume that hit sustained media
errors could not be returned to service in place once the storage was
repaired. Go exposes the same method as EcVolume.ResetIoErrorState
(weed/storage/erasure_coding/ec_volume.go:114).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* rust volume: name the shared tracker after Go's IoErrorTracker

- check_read_write_error, get_io_error_state, mark_io_quarantined,
  reset_io_error_state match weed/storage/io_error.go one to one
- io_error module is pub(crate) like the io module beside it
- restore EcVolume::reset_io_error_state so both volume kinds expose the
  same recovery surface
- trim comments that restate the code

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 16:57:23 -07:00
7643f4f541 rust: a seaweed-common crate for the address and TLS helpers both crates carry (#11358)
* rust: a seaweed-common crate for the address and TLS helpers both crates carry

seaweed-volume and seaweed-worker are separate cargo trees with separate
lockfiles and no root manifest, so anything both of them need has had to be
written twice. Two of those copies are a correctness risk rather than a typing
cost, and this crate is where they stop being copies.

address.rs is the HTTP<->gRPC port rule: `host:port` means gRPC on port+10000,
`host:port.grpcPort` names it outright. The two copies had already drifted —
the worker's bracketed IPv6 literals, the volume server's did not — so the rule
lives here once, returning a typed AddressError whose Display text is the volume
server's original wording, with join_host_port public beside it. A test asserts
two of those messages in full rather than by substring, because the wording is
the contract its callers hand to a Status or an io::Error; the other three end
in a std ParseIntError message, which is std's to reword. The enum is
#[non_exhaustive] so a future variant is not a breaking change for either
consumer. The tests are both crates' cases together, plus the IPv6,
already-bracketed and normalisation cases neither copy covered on its own.

tls.rs is install_default_crypto_provider. Both binaries link aws-lc-rs and ring
transitively, so rustls cannot auto-select and tonic's client TLS panics on
first use; each binary has to pin one and it has to be the same one, which is
exactly the kind of choice that should not exist twice. It is safe to share
because `cargo tree -i rustls` resolves a single rustls in each tree (0.23.37 in
seaweed-volume, 0.23.43 in seaweed-worker) and cargo unifies all
semver-compatible `rustls = "0.23"` requirements into one crate per binary, so
this crate writes the same process-wide static its consumer reads. rustls is
already in both graphs — directly in the volume server, through tonic's
tls-aws-lc in seaweed-worker-core — so the dependency adds no crate to either.

rust-version is 1.91.1, the lower of the two consumers' floors, so depending on
this crate cannot raise either tree's MSRV; verified with
`cargo +1.91.1 check --all-targets`. The lockfile is committed even though this
is a library: CI builds it directly, so a committed lock is what makes those
runs reproducible and their caches stable.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust: take the address and TLS helpers from seaweed-common

Both public signatures are kept, so no caller outside the two wrapper files
changes. parse_grpc_address stays `Result<String, String>` and maps the typed
error through Display; server_to_grpc_address stays `Option<String>` and drops
it with .ok(). Their doc comments and the volume server's 13 call sites are
otherwise untouched.

Three behaviours change, each in the direction of the copy that was already
right:

- The volume server now brackets IPv6 literals. `::1:19333` used to come back as
  `::1:29333`, which build_grpc_endpoint rejects with "invalid gRPC endpoint
  http://::1:19333: invalid authority" — an IPv6 master or EC peer could not be
  dialled at all. Two tests in grpc_client.rs pin it, one on the string and one
  on the endpoint the string builds.
- The volume server now emits the *parsed* gRPC port of the dotted form instead
  of the original text it had just validated, so `host:8080.018080` and
  `host:8080.+18080` come back as `host:18080` rather than as authorities the
  URI parser rejects. Same port either way; only malformed spellings change.
- The worker's dotted form now validates the HTTP port it discards.
  `server_to_grpc_address("host:abc.18080")` used to answer Some("host:18080");
  it now answers None, which is what the volume server's copy has always done.

install_default_crypto_provider becomes a re-export in both trees, so
`crate::security::tls::install_default_crypto_provider` and
`weed_lance_worker::tls::install_default_crypto_provider` still resolve. The
lance crate's `rustls = "0.23"` was its only direct use of rustls and goes away
with the body; seaweed-common states the same requirement, so neither the
resolved version nor the enabled features move in either lockfile.

The PEM test fixtures stay where they are. The two tests that use them are not
duplicates: the volume server's exercises build_grpc_endpoint, and the lance one
exists precisely because aws-lc-rs and ring are both linked in that crate's
graph. Only the literals are shared, and exporting test fixtures from a library
to dedupe two constants costs more than it saves.

A path dependency outside both trees means every build context that copies one
crate directory has to copy the other. The repo has one: the Rust source-build
stage of docker/Dockerfile.go_build, which now copies seaweed-common beside
seaweed-volume. Every workflow whose `paths:` filter keys on a crate directory
gains `seaweed-common/**` — the two Rust test workflows, rust_binaries_dev,
container_dev and performance. The tag- and dispatch-triggered ones
(rust_binaries_release, container_release_unified, container_latest) have no
`paths:` filter and need nothing.

The two Rust test workflows also run `cargo test` in seaweed-common, from their
unit-test job, because a path dependency is not a workspace member and neither
tree's own `cargo test` reaches it. Each step builds into its job's cached
target directory, and both cache keys now hash seaweed-common/Cargo.lock as well
so a change there invalidates the cache it would otherwise silently reuse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docker: keep go_build working for BRANCH revisions without seaweed-common

The rust_builder stage copies seaweed-common unconditionally now that seaweed-volume path-depends on it, but BRANCH can name any revision — including ones that predate the crate. Create the directory in the builder stage so the COPY always has a source; an empty dir beside an old seaweed-volume is harmless.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-20 13:30:58 -07:00
f1ed270942 rust volume: one S3 tier registry instead of two kept in sync by hand (#11357)
* rust volume: one S3 tier registry instead of two kept in sync by hand

`VolumeServerState.s3_tier_registry` and `global_s3_tier_registry()` held
the same S3 tier backends. `apply_storage_backends` — the only production
writer — registered every backend into both, and each half of the tiering
code then read a different one: the gRPC tier-move handlers resolved the
backend from the per-server field, while `Volume`'s remote mount and
destroy paths resolved it from the global registry, because a `Volume` has
no handle to the server state. Two registries that must agree, kept in
agreement by a duplicated `register_s3_backend` call and a comment in a
test constructor explaining the hand-sync.

Delete the field and let both tier-move handlers resolve from the global
registry, so `apply_storage_backends` registers once and no longer needs
the server state at all. Injecting a registry handle through `VolumeSpec`
instead was considered and rejected here: it would touch every `Volume`
constructor for no functional gain, and the process-wide registry is what
`Volume` already uses.

Behaviour is unchanged: the same names were registered in both registries,
so every lookup resolves exactly as before. The tier-down test now
registers its backend only in the global registry — before this change it
fails with `remote storage s3.tier_down_delete not found from supported:
[]`. The tier-up handler had no test at all, so it gets a cheap probe:
register a backend only in the global registry, ask for that destination,
and check the call gets past the lookup — the response is dropped straight
away, so the transfer sees a departed caller and never opens a connection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: await the tier-up probe terminal error instead of racing it

Dropping the response left it to chance whether the detached transfer saw the closed channel before its initial check; if it won that race it went on to attempt the multipart upload with no one waiting on the outcome. Hold the stream and read until the dead endpoint fails the upload — the terminal error proves the task ran and finished, so no background network work outlives the test.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-20 13:30:40 -07:00
7dbbdac030 rust volume: keep the EC shard-location map, its refresh time and stale mark under one lock (#11356)
The per-EcVolume shard-location cache was three fields under three locks:
an RwLock<HashMap> for the map, a Mutex<Option<Instant>> for the time it
was last refreshed, and a Mutex<bool> for the stale mark. Nothing tied
them together. merge_shard_locations published the merged map, released
the write lock, and only then stamped the refresh time; both readers
(scrub_ec_volume_distributed's snapshot and build_snapshot) took the two
guards one after the other. A reader landing between the two writes
paired a freshly merged map with the previous lookup's timestamp -- and
that pair is exactly what needs_refresh judges, so a read went back to
the master for a map that had just been refreshed. Go keeps the same
state in one struct behind one ShardLocationsLock. replace_shard_locations
documented itself as "a single observable step" while being two.

Fold the three fields into one ShardLocationCache behind a single RwLock.
merge_shard_locations upserts and stamps in one write section,
shard_locations_snapshot returns the map and its time from one read
section, and mark_shard_locations_stale / claim_shard_locations_refresh
move the mark's read-and-consume onto the cache. The three zero-caller
accessors -- set_shard_locations, replace_shard_locations,
get_shard_locations -- are deleted, and the field is now private, so the
invariant cannot be sidestepped from outside the module. The two test
seeding sites go through merge_shard_locations, which already produces
the state they were writing by hand.

Unchanged: the freshness rule. needs_refresh keeps its thresholds and
still judges the caller's snapshot -- the map that caller will actually
read from, not whatever is cached by the time the claim runs -- so only
the stale mark is read from under the new lock. The master lookup, the
completeness guard in write_back_shard_locations and the per-shard upsert
semantics are untouched.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:30:15 -07:00
Eliah RusinGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Claude Fable 5.1Chris LuDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
6d676eda67 rust volume: typed errors for store compaction so gRPC can answer NotFound (#11355)
* rust volume: typed errors for store compaction so gRPC can answer NotFound

The vacuum entry points on `Store` returned `Result<_, String>`, so the
gRPC layer had nothing to branch on and answered `Status::internal` for
every failure. A vacuum loop that races a volume being moved or deleted
saw the same code as a disk going bad, and `weed shell` could only tell
the two apart by matching on the message text.

`VolumeError` gains `VolumeNotFound(VolumeId)` — the existing `NotFound`
is needle-level and carries no payload — and `InsufficientSpace`, and
`compact_volume`, `commit_compact_volume`, `cleanup_compact_volume` and
`delete_collection` return it. `impl From<VolumeError> for tonic::Status`
in `server/mod.rs` maps not-found to `not_found`, read-only to
`failed_precondition`, insufficient space to `resource_exhausted`,
already-exists to `already_exists`, and everything else to `internal`;
the four RPCs prefix their own context with `status_with_context`, so a
message reads "commit compact volume 7: volume id 7 is not found". The
store-side "during compact" / "during commit compact" / "during cleaning
up" suffixes are gone, and the free-space message drops the volume id the
prefix already supplies.

`check_compact_volume` had no callers — `VacuumVolumeCheck` computes the
garbage level from its own `find_volume` — and is deleted. `compact_volume`
folded the size estimate into its first lookup, dropping the `unwrap()`
re-lookup that only existed to dodge a borrow.

`ascending_visit` on `CompactNeedleMap`, `RedbNeedleMap`,
`SortedFileNeedleMap` and the `NeedleMap` dispatch is now generic over the
visitor's error type, like `CompactMap::ascending_visit` already was. The
three signatures that can fail on their own bound `E: From<String>` to
carry those failures; the in-memory walk in `iter_entries` names
`Infallible`, which says in the type what its comment used to say in prose.

No Go shell command matches on the old error text: the strings exist only
in weed/storage/store_vacuum.go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* volume server: trim comments and answer the same codes from Go

- vacuum_volume_check reports VolumeError::VolumeNotFound like the other
  vacuum RPCs instead of its own "not found volume id" wording
- drop doc comments that restate what the code says
- Go volume server wraps ErrVolumeNotFound/ErrInsufficientSpace from
  store_vacuum.go so VacuumVolumeCheck/Compact/Commit/Cleanup and
  DeleteCollection answer NotFound/ResourceExhausted, matching the Rust
  volume server; volumeDeleteStatusError generalized to volumeStatusError

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: prefix operation context on vacuum errors

Lower-level errors forwarded by CompactVolume, CommitCompactVolume,
CommitCleanupVolume and DeleteCollection carry no volume id or operation
name. Wrap with %w so the status mapping still sees the sentinel chain,
matching the context the Rust server's status_with_context adds.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* volume server: map NotEmpty to FailedPrecondition, share mapper in VolumeDelete

Go's volumeStatusError maps ErrVolumeNotEmpty to FailedPrecondition; the
Rust Status conversion was missing it and volume_delete kept a hand-rolled
match. Route it through status_with_context like the vacuum handlers.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-20 13:30:02 -07:00
d002481037 rust volume: derive has_remote_file instead of mirroring it (#11353)
Volume carried `pub has_remote_file: bool` next to `pub volume_info`,
and the bool was only ever the answer to `!volume_info.files.is_empty()`:
outside the two constructors, `refresh_remote_write_mode` was the single
writer. Both fields being public made the pair a convention rather than
an invariant. Every caller that touched `volume_info.files` — load_vif
twice, the tier-up handler, the tier-down handler and its rollback — had
to remember to call `refresh_remote_write_mode` afterwards, and a caller
that forgot would leave the volume advertising a write mode its .vif
contradicts, or serving a remote .dat through a writable needle map.

The bool becomes `has_remote_file()`, computed from the list, so it
cannot drift. `volume_info` becomes private with a `volume_info()`
reader, and edits to the reference list go through
`update_remote_files(|files| ...)`, which applies the closure and then
refreshes the derived write mode and the needle map. With no caller left
outside the module, `refresh_remote_write_mode` is private.

Unchanged: the refresh logic itself, the order of operations in both
tier handlers, and the tier-down rollback semantics. The rollback still
snapshots the removed reference before the refresh runs, restores it on
failure, and re-refreshes unconditionally on the error path — the second
`update_remote_files` call runs with a no-op closure when there was
nothing to restore, exactly as the old code re-ran the refresh whether
or not it had re-inserted a reference.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:29:47 -07:00
1df8c05bc3 rust volume: document every unsafe block and stop mutating the process env in tests (#11352)
Three production `unsafe` blocks carried no `// SAFETY:` comment at all
(`libc::fallocate`, `libc::sysinfo`, `libc::statvfs`), and nothing made
that an error: `clippy::undocumented_unsafe_blocks` is a `restriction`
lint, allow-by-default, and appeared nowhere in either crate. Turn it on
in `seaweed-volume`'s `[lints.clippy]` and in the worker workspace's
`[workspace.lints.clippy]`, then document what each block relies on.

`memory_status.rs` and `disk_location.rs` get their blocks narrowed to
the `zeroed()` and the libc call, so each comment sits next to the
operation it justifies and the arithmetic is outside the block. Both
turn the success test into an early return on failure; the casts, the
multiplication order and the values returned on either path are
unchanged.

The bigger problem was in `config.rs`'s tests. `with_temp_env_var` and
`with_cleared_security_env` called `std::env::set_var`/`remove_var`,
claiming soundness because every caller holds `process_state_lock()`.
That mutex only serialises the fourteen annotated tests in this module.
The same lib test binary runs the `grpc_server.rs` tests, which bind a
`TcpListener`, dial loopback and drive a multi-thread tokio runtime, and
tonic/hyper/rustls/aws-sdk all read the environment lazily on those
threads — which is exactly the race Rust 2024 made these calls unsafe
for. `restore_env_var` had no SAFETY comment at all. `#[serial]` would
not have helped: it serialises annotated tests, which the mutex already
did.

So the config layer no longer reads the environment implicitly. An
`EnvLookup<'a> = &'a dyn Fn(&str) -> Option<OsString>` is threaded from
the public entry points down to every reader — `HOME`, `USERPROFILE`,
the twenty-four `WEED_*` keys and `SEAWEED_WRITE_QUEUE`. `parse_cli` and
`parse_security_config` keep their signatures and pass `process_env`, a
thin wrapper over `std::env::var_os`; `resolve_config` becomes
`resolve_config_with_env` (private, one caller). Tests build one with
`fake_env` instead, so no test touches the real environment and every
`unsafe` in the module is gone.

`process_state_lock()` stays, with a smaller job: `set_current_dir` is
safe but still process-global, so the tests that move the working
directory are still serialised against the ones that read it. Tests
naming an explicit config file never reach that search and no longer
take the lock.

No production behaviour changes: the same keys are read in the same
order with the same precedence, and `env_string` reproduces
`std::env::var(key).ok()` — absent and non-UTF-8 both read as unset.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 13:29:07 -07:00
Chris LuandChris Lu 44ba070d83 helm: document admin.allowInsecureBind and drop stale comment
- README and values.yaml now describe the allowInsecureBind escape
  hatch alongside the non-loopback bind guard
- remove a PR reference from the CI test comment
2026-09-20 12:28:53 -07:00
062238bb5c helm: let admin.allowInsecureBind satisfy the admin auth render guard
PR #11185 made `weed admin` refuse to bind a non-loopback address
without -adminPassword or mTLS. PR #11228 added -allowInsecureNoAuth
as an explicit opt-out for operators who restrict admin access some
other way (e.g. a NetworkPolicy plus an authenticating reverse proxy).

The chart's render-time guard added by #11236 (admin-statefulset.yaml,
seaweedfs.admin.authEnabled) predates -allowInsecureNoAuth and only
recognizes password-based auth, so there was no values.yaml path to
express that choice: the chart would fail(...) even though the binary
itself would start fine with a warning.

Add admin.allowInsecureBind (default false) to the seaweedfs.admin.authEnabled
helper's checks; when true it renders -allowInsecureNoAuth on the admin
command and satisfies the render guard alongside the existing
password-based checks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 12:28:53 -07:00
Chris LuandChris Lu b77c42ff32 test: use BaseEndpoint for the checksum test client 2026-09-20 01:15:11 -07:00
Chris LuandChris Lu 26fc90187e s3: accept x-amz-checksum-mode from the query string, case-insensitively
Presigned HeadObject/GetObject requests hoist x-amz-checksum-mode into the
signed query string, so a strict header-only check would withhold stored
checksums on presigned reads that AWS honors.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu ef463fe1af s3: read complete-request checksum values from headers or query
Presigned CompleteMultipartUpload requests hoist x-amz-checksum-type and
the full-object checksum header into the signed query string, so a
header-only lookup would skip BadDigest validation for them.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu 17ad5a1419 s3: accept FULL_OBJECT checksums without per-part checksums at complete
COMPOSITE uploads must still carry every part checksum in the complete
request, but FULL_OBJECT uploads may instead supply the whole-object
checksum in an x-amz-checksum-* request header. Compare that header
against the computed object checksum and return BadDigest on mismatch,
matching AWS.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu b03419ee92 s3: reject UploadPart checksum algorithms conflicting with the upload
An UploadPart that explicitly selects a different checksum algorithm than
the one declared at CreateMultipartUpload would store a checksum
CompleteMultipartUpload could never accept. Reject the conflict up front
with InvalidRequest, matching AWS.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu f849b7c823 s3: validate per-part checksums in CompleteMultipartUpload
Parse the Checksum* elements of each completed part and enforce what AWS
does for uploads created with x-amz-checksum-algorithm: every part must
carry a checksum in the complete request (InvalidRequest when missing,
BadDigest when it differs from the stored part checksum), and an
x-amz-checksum-type header must match the upload resolved checksum type
(BadDigest). Add the issue-11401 reproduction as a regression test.
2026-09-20 01:15:11 -07:00
Chris LuandChris Lu f15b980976 s3: UploadPart inherits the checksum algorithm of its multipart upload
AWS computes a checksum for every part of an upload created with
x-amz-checksum-algorithm, even when the part request carries no checksum
headers. Mirror that: when the part request specifies no algorithm, apply
the one stored on the upload entry so the part entry keeps a checksum
CompleteMultipartUpload can fold into the object checksum.
2026-09-20 01:15:11 -07:00
hsdfatandGitHub 110b485bae fix(volume): stop ScanVolumeFileFrom at a header it cannot advance past (#11398)
fix(volume): stop scans at a header they cannot advance past

A corrupt .dat header with a very negative size gives a record length
(NeedleHeaderSize + NeedleBodyLength) of zero or less: v3 sizes -43..-36
and v2 sizes -35..-28 give exactly zero, and smaller sizes give a
negative length. ScanVolumeFileFrom advanced by that length, so it
re-read the same header forever or stepped back into the record before
it. weed fix, weed export, weed compact, incremental weed backup and the
tail sender behind volume.move and volume.merge could hang on such a
volume, and weed compact could also finish with a .cpx that had dropped
every needle after the header.

Return an error wrapping needle.ErrorCorrupted instead. The check runs
after the visitor has seen the record, so the rebuild scanner still
stops quietly with io.EOF. Smaller negative sizes whose record length is
positive are still stepped over, preserving the salvage behavior
compaction relies on.

Mirror the guard into the Rust volume scans: DatScanPlan::scan and
read_all_needles fail on a non-positive record length, as does
scan_dat_head, so a corrupt header cannot stall a tail pass or leave the
repair scan walking stale offsets.
2026-09-19 22:26:56 -07:00
06dda12e4b fix(volume): validate sizes in ReadNeedleBlob and WriteNeedleBlob (#11399)
* fix(volume): reject negative sizes in ReadNeedleBlob and WriteNeedleBlob

A ReadNeedleBlob RPC with a size of -44 or below (-36 on v2 volumes)
panics in makeslice inside needle.ReadNeedleBlob. The volume gRPC server
has no recovery interceptor, so one request kills the process. Smaller
negative sizes return bytes that are not a record.

WriteNeedleBlob accepted a negative size whenever the blob header
carried the same value: it appended the blob to .dat and indexed the
needle with that size, which reads as deleted.

Reject size < 0 in both Volume methods. Size 0 still passes, since
delete records carry it. The Rust volume server got the same storage
guards in #11345.

* fix(volume): reject needle blobs whose length does not match their size

WriteNeedleBlob appends the blob as is. A blob that is not the length
its size implies leaves .dat off the 8-byte grid, and every later
ordinary write to the volume is indexed at a truncated offset and reads
back as EOF. A blob off by 8 bytes keeps the grid but leaves bytes that
a .dat scan reads as the next record.

The in-tree callers already send exact lengths. The one case this newly
refuses is a copy between volumes of different needle versions, and
that case already writes a broken record: a v3 record lands on a v2
volume with 8 extra bytes, and a v2 record on a v3 volume either fails
the timestamp check or lands 8 bytes short.

This is separate from the negative-size guards, whose Rust counterpart
is #11345. The Rust server does not check the length yet.

* fix(volume): guard the blob buffer allocation in needle.ReadNeedleBlob

Volume.ReadNeedleBlob rejected negative sizes, but needle.ReadNeedleBlob
still sized its buffer from the size and is called directly by vacuum and
other paths. Reject a deletion marker before make() there too, and use
size.IsDeleted() in the volume-level checks.

* fix(volume): mirror the blob length check in the rust volume server

write_needle_blob_and_index checked the size against the blob header but
appended the blob verbatim, so a blob that is not the length its size
implies still leaves .dat off the record grid. Match the Go check.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 21:28:37 -07:00
a93a1ab2eb fix(volume): return an error instead of 201 when a write lands on no volume (#11397)
* fix(volume): return an error instead of 201 when a write lands on no volume

ReplicatedWrite only writes locally when this server holds the volume.
For a volume id no server holds, the master lookup returns no locations,
so the write went nowhere and the upload still got 201 Created. The same
happened for a type=replicate write to a server without the volume, so
the primary, or the S3 chunk fan-out, counted a replica that was never
written.

A server without the volume still forwards the write to the replicas the
master lists. When there is nothing to forward to, fail with "volume N
not found on host:port". PostHandler returns that as 500, the status the
Rust volume server already returns here, and uploaders re-assign on 5xx.

Fixes #6609

* volume: reuse Store.HasVolume, drop issue ref from test comment

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 19:03:21 -07:00
cd1e738422 [Volume] Scrub local deletion tombstones during FULL scrub (#11396)
* fix 11388

* fix(volume): scrub validates local deletion tombstones

TombstoneFileSize (-1) is an .idx-only sentinel; the physical record it
points at carries a zero-sized body. Normalize deleted index sizes to 0
via onDiskSize before computing disk usage and calling ReadData, so
corrupted or truncated tombstone records are detected instead of
skipped. Offset-zero entries (remote logical deletes, no .dat record)
remain skipped, and the physical needle id is checked against the index
key. Mirror the behavior in the Rust volume server.

* fix(volume): scrub preserves physical size of deleted non-tombstone entries

Size.Raw()/raw() already encodes the index-to-disk mapping: tombstone
(-1) -> 0, other negative sizes -> their absolute value (the offset then
points at the original record, per the ReadDeleted path). Use it instead
of mapping every deleted size to 0.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-19 18:55:27 -07:00
github-actions[bot] 3dec359d6b docs: regenerate star history chart 2026-09-20 00:50:27 +00:00
Chris LuandGitHub f6a3286b32 fix(volume): derive needle body tail bound from the version layout (#11395)
* fix(volume): derive needle body tail bound from the version layout

The size guard in ReadNeedleBodyBytes computed the tail length as
checksum, plus timestamp only for Version3. Forks and future on-disk
formats whose tail carries more fields would silently under-check and
still panic in readNeedleTail on a truncated body. Derive the tail from
NeedleBodyLength minus data and padding so the bound stays exact for
every version.

Iterate IsSupportedVersion in the new tests instead of hardcoding
v1-v3 so downstream formats get covered automatically, and skip
versions the build cannot write rather than failing on them.

* test: skip needle write only on the unsupported-version error

A blanket skip would hide a real writer regression. Skip the version
subtest only when the writer reports the version is not supported in
this build (the error text differs between builds), and fail on any
other write error.
2026-09-19 12:20:01 -07:00
Chris LuandGitHub 01bb3b3053 s3api: add Snowflake s3compat API integration tests (#11394)
* s3api: add Snowflake s3compat API integration tests

Run the upstream snowflakedb/snowflake-s3compat-api-test-suite against a
local SeaweedFS server in CI. test/s3/snowflake/run.sh starts weed server
with S3 (-s3.autoCreateBucket=false so missing-bucket PUTs return
NoSuchBucket), prepares the fixtures the suite needs (versioned bucket,
deny-all-policy bucket, >1000-object prefix), clones the suite, patches
it to path-style addressing, and runs mvn -Dtest=S3CompatApiTest.

The suite also exposed that GetBucketLocation returned 404 NoSuchBucket
for a malformed bucket name; validate the name first and return
400 InvalidBucketName like AWS.

* test: harden snowflake s3compat runner per review

- Pin the upstream suite to a tested commit (SUITE_REV) instead of the
  moving default branch
- Bind the test server to loopback only
- Require the AccessDenied error code when verifying the denied bucket
- Fix README so go install runs in a subshell
- checkout with persist-credentials: false
- Make the concurrency group unique per PR, and widen path filters to
  the storage/operation/wdclient/cluster/pb packages the S3 stack uses

* test: advertise loopback ip for snowflake test server

-ip.bind 127.0.0.1 alone left the volume server advertising the host's
primary address, so chunk uploads were refused. Also set -ip 127.0.0.1
and disable the Iceberg/Lance listeners so the harness is loopback-only
and does not collide with other local services.
2026-09-19 03:29:51 -07:00
hsdfatandGitHub 5769057af3 fix(volume): return an error instead of panicking on a corrupt needle size (#11393)
ReadNeedleBodyBytes sliced the needle body with the size from the needle
header without checking it. A corrupted .dat header carrying size -1 still
gets a positive body length (16 bytes on v3), so vacuum compaction read
that body and panicked with "slice bounds out of range [:-1]".

Writers never put a negative size in a .dat header: a delete appends a
size-0 record, and TombstoneFileSize only lives in the .idx. Reject a size
that is negative or leaves no room for the checksum/timestamp tail with an
error wrapping ErrorCorrupted. ScanVolumeFileFrom already logs body read
errors and moves on, so compaction now skips the record like any other
corrupt needle.

Fixes #6763
2026-09-18 23:00:06 -07:00
github-actions[bot] 4160b92864 docs: regenerate star history chart 2026-09-19 00:45:06 +00:00
Chris LuandGitHub 37bf1cd91d volume: validate copy/tail source addresses before dialing (#11390)
* pb: stop exiting the process on malformed server addresses

ServerToGrpcAddress and GrpcAddressToServerAddress called glog.Fatalf
when hostAndPort could not parse the port, which os.Exit(255)ed the whole
process. A caller-supplied copy or tail source address reached this path
synchronously in the serving goroutine, so one anonymous VolumeCopy with
a non-numeric port terminated the volume server.

Log the parse error and return the input unchanged instead: the dial or
request that consumes the address then fails as an ordinary error.

* volume: validate copy and tail source addresses before dialing

VolumeCopy, VolumeEcShardsCopy and VolumeTailReceiver dial a
caller-supplied source address (SourceDataNode / SourceVolumeServer)
with no endpoint validation, so an anonymous caller could aim the volume
server at loopback, link-local (cloud metadata) or other unintended
destinations and read dial behavior back as a connectivity oracle.

Apply the same peer-target deny list FetchAndWriteNeedle uses for
replica targets: the source must be a bare host:port whose host is not
loopback, link-local or unspecified; cluster peers stay reachable on
private networks, and -volume.allowUntrustedRemoteEndpoints opts out.
The loopback-using copy tests set the flag to keep exercising the copy
path in process.

* rust volume: validate copy and tail source addresses before dialing

Mirror the Go guard on the Rust volume server: volume_copy,
volume_ec_shards_copy and volume_tail_receiver dial a caller-supplied
source address, so run it through validate_replica_target first (bare
host:port; no loopback, link-local or unspecified hosts; private peers
stay allowed). --volume.allowUntrustedRemoteEndpoints opts out; the test
fixture and the Rust test-cluster launcher set it so loopback sources in
tests keep working.

* volume: pin validated copy/tail source addresses at dial time

validateReplicaTarget resolves the source hostname once, but the gRPC
client resolved it again at connect, leaving a DNS-rebinding window for
hostname sources. The copy and tail source dials now run through the
same guardedDialerPolicy the remote-storage path uses, so every resolved
address is re-checked against the replica deny list (private peers
allowed) immediately before the TCP connect. guardedDialerPolicy also
moves to util.OutboundDialContext so the guarded path keeps the -ip.bind
source binding the default gRPC dialer had.

The Rust volume server mirrors this with connect_guarded, a tonic
connector that resolves, re-checks each address, and connects to the
first passing IP; handlers use it whenever the untrusted-endpoint
opt-out is off. A handler-level test now exercises the enabled
validation branches for all three source-taking RPCs.

* pb: return empty server address for malformed grpc addresses

GrpcAddressToServerAddress used to return the unparseable input on a
hostAndPort failure, so a malformed raft address (e.g. "host:abc")
flowed into admin dashboard master maps unchanged. Return an empty
string instead, skip empty conversions at the two raft-cluster merge
sites, and drop the now-stale comment about the fatal exit the earlier
commit removed.

* test: opt erasure-coding loopback clusters out of the remote endpoint guard

The erasure-coding suites drive VolumeEcShardsCopy / VolumeCopy between
volume servers bound to 127.0.0.1, which the copy/tail source guard now
rejects by default. Pass -volume.allowUntrustedRemoteEndpoints to the
test volume launches, matching what the volume_server framework
harnesses already do.

* admin: only claim fallback master leadership on an empty raft response

A nonempty RaftListClusterServers response whose entries were all
rejected left masterMap empty, so the fallback marked the reachable
current master as leader the same way a genuinely empty (non-raft)
response does. Track whether the successful response returned zero
servers and only promote the fallback master then.
2026-09-18 12:55:47 -07:00
Chris LuandGitHub a6d72bc272 s3api: delete orphaned chunks only when the entry is confirmed absent (#11389)
* s3api: test for chunks deleted under an entry the filer committed

Issue #11387: the filer can report a create failure after inserting the
entry (e.g. a parent-directory creation failing post-insert). The error
arrives in the response rather than as a transport status, so it maps
to a definitive error and putToFiler deletes the chunks of the live
entry.

* s3api: confirmCreateLanded also reports a confirmed-absent entry

The verification a failed create runs can answer both directions: the
entry matching the uploaded chunks proves the write landed, and an
authoritative not-found proves the uploaded chunks are orphaned. Return
both outcomes so the cleanup path can gate on the fact rather than the
error class. An empty upload can never prove a landing, so a zero-chunk
entry match no longer upgrades the outcome.

* s3api: delete orphaned chunks only when the entry is confirmed absent

A failed create no longer skips verification based on the error class: the filer can fail after inserting the entry (issue #11387) and a partially-applied routed transaction can leave it behind too, both surfacing as definitive errors. Every failed create now resolves the entry's fate, and the uploaded chunks are deleted only when the entry is confirmed absent; anything unverifiable keeps them for vacuum.

* s3api: confirm absence on every filer the create could have committed on

A lock-path create fails over across filers, so the entry can live on a replica the routed owner has not caught up to; one not-found does not prove absence. The confirmation now queries the owner, the prior owner, and the failover set, declaring absent only when none of them has the entry.

* s3api: bound the reconciliation lookups confirmCreateLanded runs

The lookups ran on context.Background() under the object write lock, so a connected filer that never replies could stall the write path. One timeout now covers the whole enumeration; an expired budget fails the remaining lookups as uncertain, which keeps the chunks.
2026-09-18 12:30:07 -07:00
Chris LuandGitHub c72eda50a8 s3: drop implicit reader cache budget that throttled S3 GETs (#11384)
* fix(filer): leave reader cache unbounded without an explicit budget

NewReaderCache silently installed a 256MiB ReaderCacheBudget when the
caller passed none. Only weed mount opts into a budget; every other
caller (S3 gateway, WebDAV, query engine, mq logstore) inherited the
cap. Under ~90 concurrent S3 GETs of medium objects, prefetch wants
far more than 64 chunk buffers, so reserve() serialized chunk fetches,
clients timed out and retried, and the retry re-downloaded chunks the
cancelled request had already fetched.

A nil budget now means unbounded, restoring the pre-4.47 behavior for
callers that never asked for a memory cap; reserve/complete/release
are nil-safe. The mount path is unchanged and still enforces
-readerCacheSizeMB.

Fixes #11380

* feat(s3): expose -s3.readerCacheSizeMB reader buffer budget

Operators who want the S3 gateway read path memory-bounded can now
opt in: -s3.readerCacheSizeMB on weed filer/server/mini and
-readerCacheSizeMB on standalone weed s3, matching the mount flag.
The default 0 keeps the unbounded pre-4.47 behavior; a positive value
installs a shared ReaderCacheBudget across in-flight and retained
chunk buffers for all S3 GETs.

* fix(filer): validate chunk size before consulting the reader budget

A nil budget returned early and skipped the negative chunkSize check,
letting a corrupted size reach mem.Allocate and panic. Also drop the
command-specific flag prefix from the S3 validation error since
standalone weed s3 exposes the option as -readerCacheSizeMB.

* filer: drop chunk buffers once fully consumed

ReaderCache retained every completed chunk buffer in the downloaders
map until the slot limit evicted it, so buffers lingered after all
readers finished with them.

Track attached readers on each SingleChunkCacher and remove the cacher
when the last reader consumes the buffer to its end. In-flight download
deduplication and the prefetch handoff are unchanged: a buffer always
survives until fully read, partial reads keep it available, and an
attached reader pins a consumed buffer until it detaches. Repeat reads
now go through the chunk cache where enabled, or refetch.

* filer: drop consumed buffers on last detach, rechecked under cache lock

Two review findings on the drop-on-consume change:

- Removal only fired when the detaching reader itself reached the chunk
  end. If the end-reaching reader finished first and the last remaining
  reader did a partial read or cancelled, the consumed buffer and its
  budget reservation lingered until eviction. Track a persistent
  consumed flag instead, so any end-reaching read marks the buffer and
  the last detach drops it.
- remove() checked only map identity, so a reader attaching between the
  reader count hitting zero and removal could attach to a cacher that
  was then deleted underneath it. removeConsumed() re-checks identity,
  readers == 0, and consumed under the ReaderCache lock; a raced attach
  keeps the cacher and its own detach retries the removal.
2026-09-18 01:02:32 -07:00
Chris LuandGitHub 87ee3b63a2 s3: abort completed multipart uploads metadata-only (#11385)
* s3: abort a completed upload's leftover directory metadata-only

A .uploads/<id> directory can outlive the object it completed into when
the commit's metadata-only removal failed or the gateway died in between;
the restored part entries then share chunks with the published object.
AbortMultipartUpload deleted the directory recursively, chunks and all,
so aborting such a leftover destroyed a committed object (#11382).

Run the same check s3.clean.uploads gained in #11375 before deleting:
when the object entry or a version file under <key>.versions carries the
upload id, remove .uploads/<id> metadata-only and answer the abort; when
the lookup cannot decide, refuse with InternalError rather than risk
live chunks.

* s3: apply the completed-upload check to lifecycle MPU abort

lifecycleAbortMPU ran the same destructive recursive delete on
.uploads/<id>. Reuse uploadCompleted so a leftover whose object entry or
version file carries the upload id is removed metadata-only, and an
undecidable lookup retries later instead of freeing live chunks.

* s3: serialize abort's upload-dir delete with the object's commit

The completed check alone leaves a race: abort can read completed=false,
then an in-flight completion publishes the object over the same part
chunks before the recursive delete frees them.

Run the check and delete inside the object write lock, which non-routed
completions hold for their whole finalize. With an owner, send the data
delete as an ObjectTransaction on the object's lock key — a routed
commit then either loses its upload-exists precondition after our delete
or has already stamped the object, which the transaction's
IF_EXTENDED_NOT_EQUAL condition detects and falls back to a
metadata-only remove. lifecycleAbortMPU shares removeUploadDir so both
callers get the same ordering.

* s3: check for an empty object before resolving its write owner

* s3: check completion at the abort's resolved object key

An upload record missing ExtMultipartObjectKey skipped the completed
check entirely even though the request's Key names the object.
2026-09-18 01:01:52 -07:00
Chris LuandGitHub 0ca1c19821 s3api: unify auth error handling across s3tables, iceberg and lance (#11381)
* s3api: fail closed when S3 Tables signature verification fails

* s3api: avoid nil Account dereference in S3 Tables auth log

* iceberg: return auth error instead of falling back to DefaultAllow

* lance: return auth error instead of falling back to DefaultAllow

* s3api: stop trusting client-supplied s3-account-id

The header is set by the server after successful authentication; scrub
inbound values alongside the other internal headers, and apply the same
admin guard to the header fallback branch of getAccountID that the
identity branch already has.

* test: cover table-catalog auth wrappers and principal resolution

* test: configure anonymous identity where catalog clients do not sign

* s3api: scrub s3-account-id after signature verification
2026-09-18 01:01:04 -07:00
Chris LuandGitHub f40687b34e s3: tighten STS session token handling (#11383)
* s3api: test that a session token must not reveal its credential

* sts: derive secret access key with HMAC keyed on the signing key

* s3api: stop accepting STS session tokens as bearer credentials

* security: reject STS session tokens on filer and admin gRPC auth

* test: sign s3/iam framework requests with the session credential

* s3api: exercise the real auth pipeline in the end-to-end harness
2026-09-17 23:50:16 -07:00
Chris LuandGitHub 15520f601f s3: commit multipart upload and remove .uploads atomically; purge completed uploads metadata-only (#11375)
* s3: commit versioned multipart upload in one transaction

CompleteMultipartUpload wrote the version file, flipped the .versions
pointer, then removed .uploads/<id> metadata-only as a best-effort
post-commit step. A filer error or gateway crash in that window left the
upload directory referencing the same chunks as the published object, and
the next s3.clean.uploads run purged it with data -- corrupting a
committed object.

Put the version file, remove the upload directory metadata-only (its
chunks are the object's chunks), and recompute the latest pointer in one
ObjectTransaction under the object's per-path lock on the owner filer.
The mutation order keeps every partial state safe: the chunks stay
referenced at all times, and a published object never coexists with the
upload directory the cleaner would purge.

Unused part entries are freed before the transaction, since the
metadata-only directory delete would otherwise leak their chunks.

* s3: remove upload directory inside the multipart object PUT

The same committed-object/stranded-upload window existed on the
suspended and non-versioned paths: writeMultipartObject committed the
object, then a best-effort rm dropped .uploads/<id>. Ride the
metadata-only removal on the routed PUT itself so the two land in one
transaction; the unrouted mkFile fallback keeps post-commit cleanup.

* shell: purge completed uploads metadata-only in s3.clean.uploads

A leftover .uploads/<id> can outlive a committed object when the
completion's metadata-only delete fails or the gateway dies in between;
its part entries then share chunks with the live object, and a recursive
purge frees them out from under it.

Before purging a stale upload, check whether it completed: the object
entry or any version file under <key>.versions carrying the upload id.
If so, delete with skipChunkDeletion. If the lookup fails, skip the
upload for this run rather than risk live chunks.

* s3: abort multipart completion when unused part cleanup fails

Deleting the upload directory metadata-only erases the only metadata
pointing at part entries whose deletion failed, orphaning their chunks.
Propagate the error so the completion fails while the upload directory
still exists and the request remains retriable.

* s3: require the upload directory to exist at multipart commit

A delete that does not take the object lock (abort, lifecycle,
s3.clean.uploads) can remove .uploads/<id> and its chunks between the
prepare step and the commit transaction. The commit now carries an
IF_EXISTS precondition on the upload directory so the race fails the
request with NoSuchUpload instead of publishing an object over freed
chunks.

* s3: keep the version file when the upload directory is gone

The finalize transaction has no rollback, so a failure at the
latest-pointer recompute leaves the version written and .uploads/<id>
removed. Deleting the version then destroys the only remaining record of
the upload, making a retried CompleteMultipartUpload return NoSuchUpload
while the version's chunks leak. Roll back only while the upload
directory survives; otherwise keep the version, which a retry resolves
through SeaweedFSUploadId and the version reconciler promotes.

* s3: keep manifests when a routed object write partially commits

For non-versioned and suspended completions the object PUT precedes the
upload-directory DELETE, so an error can mean the object entry exists
while the response reports failure. Freeing this attempt's manifest
chunks then destroys the committed object. Keep them when the object
entry survived, and after a failed null-marker finalize which always
follows a committed write.

* s3: skip the keep-version path on precondition failure

A rejected precondition means no mutation ran, so there is no version
file to preserve and this attempt's manifests are orphans the error
cleanup should free.

* s3: keep manifests when the object-existence check itself fails

A transient lookup error previously read as absent, letting the error
cleanup free manifest chunks a committed object still references.

* s3: keep the upload directory when post-commit part cleanup fails

Removing it metadata-only after a failed entry delete erases the only
reference to the leftover chunks. Leave the directory so the entries
keep their chunk references for s3.clean.uploads or manual recovery.

* pb: fix filer list entry counting on 32-bit

int(limit) wraps to -1 on 386 when limit is math.MaxUint32, so the
beyond-limit check discarded every streamed entry. Compare in uint64
instead; the semantics are unchanged on 64-bit platforms.

* shell: resolve trailing-slash object keys in s3.clean.uploads

Completion stores a key ending in / inside the directory it names
(<bucket>/dir/dir), but FullPath+DirAndName on the normalized key
looked one level too high. Deriving dir and name with path.Dir and
path.Base mirrors getEntryNameAndDir so the completed-upload check
finds the entry instead of purging its chunks.

* s3: heal a suspended completion hidden behind a delete marker

Removing .uploads/<id> inside the commit transaction means a failed
finalizeSuspendedNullWrite leaves nothing to retry against: the object
entry is committed but the marker still makes the key read as deleted,
and a retried CompleteMultipartUpload can only report NoSuchUpload.
When the upload directory is gone, check the regular path for an entry
carrying the upload id and re-run the marker finalize, so the retry
both succeeds and repairs the key. Only suspended buckets can hold
this state; anything newer owns the key.

* s3: report store errors when resuming a committed multipart object
2026-09-17 21:09:21 -07:00
Chris LuandGitHub bdc37a1e86 mount/shell: bucket allow-empty-folders toggle, mount keeps explicit false (#11370)
* mount: keep a deliberate bucket allow-empty-folders setting

* shell: s3.bucket.allowEmptyFolders toggles the empty folder cleaner

* shell: guard allow-empty-folders toggle with expected extended attrs

* filer: drop cached empty-folder policy on bucket entry update

* mount: guard allow-empty-folders write with expected extended attrs

* filer: skip caching a stale cleanup policy read across an update

* mount, shell: snapshot the full extended attributes for update preconditions

* filer: fail closed and invalidate on all bucket entry events for cleanup policy

* filer: key the cleanup policy generation by bucket

* filer: requeue cleanup when the bucket policy cannot be loaded

* filer: expire idle cleanup policy generations

* filer: skip requeueing cleanup after the cleaner stops

* filer: bound cleanup retries on repeated policy failures

* filer: cover cleanup requeue on repeated policy failures

* filer: keep cleanup policy generations while reads are in flight

* filer: exercise the cleanup queue lifecycle in the retry-cap test
2026-09-17 20:30:53 -07:00
Chris LuandGitHub 2d2619f0b4 ci: make telemetry deploy work on Oracle Linux 7 (#11377)
* telemetry/server: tidy module dependencies

* ci: make telemetry deploy work on Oracle Linux 7

* ci: install telemetry unit and logrotate on every deploy

* ci: abort telemetry deploy on install failures
2026-09-17 20:00:25 -07:00
Chris LuandGitHub ce1e0dc30a s3api: don't delete chunks when CreateEntry outcome is ambiguous (#11376)
* s3api: map ambiguous filer transport errors to retryable 503

Canceled, DeadlineExceeded and Unavailable can be returned after the
filer applied the write, so the outcome is ambiguous. Reporting them as
a 4xx tells the client not to retry; report ServiceUnavailable instead.

* s3api: verify entry existence before deleting orphaned chunks

A failed CreateEntry can still have landed on the filer when the error
is a transport failure, and entryCreated=false would tombstone chunks a
live entry references, leaving a dangling pointer that survives only
because reads pass readDeleted=true until vacuum reclaims the needle.

Before deleting, look the entry up: if it is stored with the same
chunks, the write succeeded; if the lookup cannot be answered, keep the
chunks for vacuum to reclaim; only a confirmed absence still cleans up.

* s3api: regression tests for ambiguous CreateEntry outcomes

Covers the three post-create-failure cases in putToFiler: the entry
landed despite the error (treat as success, keep chunks), the entry is
confirmed absent (delete orphans), and the outcome is unverifiable
(keep chunks, return error).

* volume: count reads served from deleted needles

A readDeleted read succeeding on a tombstoned needle is the signal that
metadata still points at deleted data. Count it under a
readDeletedNeedle handler label in both the Go and Rust volume servers
so the condition is visible before vacuum turns it into a 404.

* s3api: never delete chunks on an ambiguous create error

Review feedback on the first fix showed verification could still go
wrong in both directions: a stale or lagged lookup could report
not-found for a committed entry, a prefix object stores its chunks on a
directory entry, and filer-side manifestization rewrites the top-level
chunk ids the comparison relied on.

Rework the rule so the outcome classes are asymmetric:

- A transport-level error (anything filerErrorToS3Error maps to a
  retryable 503) is ambiguous and never deletes chunks; the lookup can
  only upgrade the write to success.
- Any other error is a definitive filer refusal and still cleans up.

confirmCreateLanded asks the write owner first, resolves the stored
entry through chunk manifests, requires an exact match of the uploaded
file ids, and on success runs the finalize callback the failed create
skipped (under the object write lock, with the same rmObject undo the
create path uses). Zero-chunk writes stay ambiguous since they cannot
be told apart by chunks.

* s3api: cover definitive refusals and stale entries in put tests

The confirmed-failure case now uses a definitive refusal so it still
exercises orphan cleanup, and a new case keeps chunks when the stored
entry belongs to an older object rather than this PUT.

* volume: count deleted-needle reads once per request

Streamed Go reads ran the deleted check in readNeedle and again in
readNeedleDataInto, and non-streamed Rust reads in stream_info and the
full-read fallback, double-counting one request. Count at the single
entry probe each implementation takes per GET: readNeedle in Go,
read_needle_stream_info in Rust.

* s3api: run recovered-write rollback under the object lock

Two follow-ups from review: ResolveChunkManifest returns traversed
manifest blobs in its manifestChunks output, so requiring it empty
rejected every manifestized landing; and the rmObject undo ran after
the object write lock was released, so a concurrent newer write could
be deleted between finalize failure and rollback. Compare only the
resolved data chunks and keep the undo inside the lock.

* s3api: verify, finalize and roll back recovered creates in one lock

A lookup done before the object write lock let a concurrent PUT replace
the entry between the chunk comparison and the finalize/rollback
section, so a failed afterCreate could rmObject a newer write. Run the
owner lookup, manifest resolution, chunk comparison, afterCreate and
the conditional undo inside a single withObjectWriteLock section.
2026-09-17 19:58:49 -07:00
github-actions[bot] d4e11a471d docs: regenerate star history chart 2026-09-18 01:41:46 +00:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
08d5daf0c1 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc from 1.44.0 to 1.45.0 (#11371)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:16:38 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
8d34433308 build(deps): bump go.opentelemetry.io/otel/exporters/zipkin from 1.36.0 to 1.45.0 (#11373)
build(deps): bump go.opentelemetry.io/otel/exporters/zipkin

Bumps [go.opentelemetry.io/otel/exporters/zipkin](https://github.com/open-telemetry/opentelemetry-go) from 1.36.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.36.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/zipkin
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-17 15:10:51 -07:00
Chris LuandGitHub 4fd67001d9 security: require go 1.26.6 and bump vulnerable deps (#11374)
* security: require go 1.26.6 and bump vulnerable deps

A dependency scan of the 4.47 release flagged the bundled toolchain
and modules:

- github.com/golang/go < 1.26.6 (CVE-2026-39821, CVE-2026-56853,
  CVE-2026-56859, CVE-2026-56862, CVE-2026-56864, CVE-2026-56865,
  CVE-2026-33818, CVE-2026-46600): raise the go directive to 1.26.6
  so every built artifact requires the fixed toolchain.
- google.golang.org/grpc (CVE-2026-84445, CVE-2026-84304): move to
  the fixed dev pseudo-version; released tags through v1.85.0-dev
  remain in the affected range.
- github.com/pelletier/go-toml/v2 <= v2.4.2 (unbounded parser
  recursion): v2.4.3.
- alpine libcrypto3/libssl3 < 3.5.8-r0 (CVE-2026-75803,
  CVE-2026-63073, CVE-2026-63075, CVE-2026-63076, CVE-2026-63072,
  CVE-2026-54874, CVE-2026-18798, CVE-2026-14456, CVE-2026-14457):
  the release images already apk-upgrade the final stage; extend the
  same to the telemetry and admin-integration images.

Same bumps applied to the test/kafka, test/sftp,
kafka-client-loadtest, and telemetry/server modules.

* telemetry: send integration test report above the 10 GiB floor

The collect endpoint keeps reports only when TotalDiskBytes >=
proto.MinDiskBytes, but the integration test still sent 1 GiB, so the
server counted the report and skipped storing it. No cluster_id series
was ever created and /metrics lacked seaweedfs_telemetry_volume_servers.
Send just above the floor (via proto.MinDiskBytes so it cannot silently
drift again) so the expected per-cluster metrics are exported.
2026-09-17 15:09:55 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c6a3280595 build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace from 1.44.0 to 1.45.0 (#11372)
build(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace

Bumps [go.opentelemetry.io/otel/exporters/otlp/otlptrace](https://github.com/open-telemetry/opentelemetry-go) from 1.44.0 to 1.45.0.
- [Release notes](https://github.com/open-telemetry/opentelemetry-go/releases)
- [Changelog](https://github.com/open-telemetry/opentelemetry-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

---
updated-dependencies:
- dependency-name: go.opentelemetry.io/otel/exporters/otlp/otlptrace
  dependency-version: 1.45.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-17 14:36:15 -07:00
799c495226 rust volume: one positional read helper; never seek a dup'd handle on Windows (#11342)
* rust volume: one positional read helper; never seek a dup'd handle on Windows

Positional read-exact was hand-rolled four times: the complete
cross-platform version in needle_map/sorted_file.rs, a Windows-only half
in volume.rs whose unix half was inlined as a
cfg(unix)/cfg(windows)/compile_error! triple at three call sites, a
byte-identical Windows-only copy in ec_volume.rs, and read_full_at in
ec_bitrot.rs. Three more sites -- EcVolumeShard::read_at,
EcLocalShard::read_at and ec_encoder::read_at_most -- hand-rolled the
short-read-permitted variant with a cfg(not(unix)) arm that
try_clone()s the handle and seeks it.

That last arm is wrong. A duplicated descriptor shares one kernel file
offset with the original, so seek-then-read is two syscalls against
state another thread can move in between: a concurrent reader or an
append repositions the offset and the read returns bytes from somewhere
else entirely. EcLocalShard::read_at documents that it must never seek,
one line above the seek. Windows seek_read carries its own offset in a
single call, so that window does not exist.

All seven now go through storage::io::{read_exact_at, read_at}, whose
module doc records why duplicating a handle is not a way to get a
private file position -- opening the file again is, as
Volume::dat_scan_plan already does. read_at_most keeps its own
fill-until-EOF loop; only the per-iteration positional read changes.

Behaviour on unix is unchanged: every unix arm was already
FileExt::read_exact_at or FileExt::read_at. The one exception is
ec_bitrot::verify_shard_blocks, which now retries on EINTR (std's
read_exact_at does; the loop it replaces did not) and, on unix, reports
the standard "failed to fill whole buffer" text instead of "short read
on shard block". The Windows arm still says "unexpected EOF in
seek_read"; both carry ErrorKind::UnexpectedEof, as before.

NeedleStreamSource::read_exact_at and Volume::read_exact_at_backend keep
their signatures; only their bodies shrink.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: retry Interrupted in Windows read_exact_at

Unix std's FileExt::read_exact_at ignores ErrorKind::Interrupted and
retries, but the Windows seek_read loop propagated it, so the shared
exact-read contract differed by platform. seek_read can surface
ERROR_OPERATION_ABORTED, which std maps to Interrupted.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-17 11:47:38 -07:00
Chris LuandGitHub 4ec564469a s3tables: hide inaccessible catalog resources (#11365)
* s3tables: hide inaccessible table buckets

* s3tables: hide inaccessible namespaces

* s3tables: hide inaccessible tables

* s3tables: hide inaccessible resources in rename and namespace delete

RenameTable/RenameView denied on the source now report the same
not-found as a missing source, and the destination name conflict is
checked only after destination authorization so a denied caller cannot
distinguish an existing destination namespace or name from a missing
one. DeleteNamespace denials use the same formatted message as a
missing namespace.
2026-09-17 11:43:52 -07:00
Chris LuandGitHub 66f1754896 s3: enforce dedicated Object Lock actions (#11362)
s3: enforce dedicated object lock actions
2026-09-16 20:34:06 -07:00
Chris LuandGitHub 994e1f7d64 admin: replace Font Awesome with MIT-licensed icons (#11364)
admin: replace Font Awesome with MIT icons
2026-09-16 20:29:14 -07:00
github-actions[bot] 74eeac6b66 docs: regenerate star history chart 2026-09-17 00:46:37 +00:00
0eb638f503 fix(ec): BatchDelete cookie fail-closed via locate_data geometry (#11348)
* fix(ec): BatchDelete cookie fail-closed via locate_data geometry

* fix(ec): honor skip_cookie_check, require full cookie header

* fix(ec): retry short cookie header reads, still fail closed on EOF

* chore(ec): trim cookie validation comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-16 16:25:03 -07:00
David ChristopherandGitHub 1a285c1334 filer: join shutdown paths before closing metadata store (#11363)
Serve can return when its listener closes while HTTP requests are still draining. The main path could then close the metadata store before those requests finish.

Make signal, context, and Serve-exit paths join one shutdown sequence. Drain gRPC and HTTP concurrently with 15-second default limits, then close the store. Test both completion orders.
2026-09-16 16:20:39 -07:00
Eliah RusinandGitHub caf3d157e6 fix(ec): encode drops tombstoned needles, last-wins replay (#11347)
* fix(ec): encode drops tombstoned needles, last-wins replay

* fix(ec): drop zero-offset rows in encode, match readNeedleMap
2026-09-16 15:40:22 -07:00
Chris LuandGitHub 3ebc05930d s3: separate Object Lock configuration permission (#11361)
* s3: separate object lock configuration permission

* test: synchronize manifest cancellation setup
2026-09-16 15:27:09 -07:00
Chris LuandGitHub 0c7beec697 server: add filer-specific disableHttp flag (#11360) 2026-09-16 14:18:57 -07:00
David ChristopherandGitHub a859f0a019 filer: preserve accepted metadata log records on shutdown (#11359)
fix: flush metadata log before closing filer store

Serialize sealed-batch handoffs with shutdown, reject late appends, and wait for log-buffer workers before closing the filer metadata store.

Cover queued writes, interval and explicit flushes, late-write rejection, and pending persistence with shutdown tests.
2026-09-16 12:23:59 -07:00
Eliah RusinandGitHub def25ca84d fix(ec): validate ShardId at gRPC boundary, reject >=32 (#11346) 2026-09-16 08:33:22 -07:00
Eliah RusinandGitHub 701e397337 fix(volume): reject negative Size, recover poisoned store lock (#11345) 2026-09-16 08:12:11 -07:00
Eliah RusinandGitHub 4fc9ada2ec ci: run seaweed-volume unit tests on Windows (#11349) 2026-09-16 08:06:21 -07:00
a73ba3adbb rust volume: parse vid/fid paths once; the proxy redirect drops the extension like Go (#11341)
handlers.rs split needle URLs in three places and the three disagreed.
Go does it once, in parseURLPath (weed/server/common.go:218-249), and
dispatches on the slash count: /vid/fid/filename takes the extension off
the filename and leaves the fid whole, /vid/fid takes it off the fid, and
the comma form splits the last segment on its last comma and dot.

Two of the Rust copies got that wrong:

- extract_file_id returned the path unchanged when it found no comma, so
  a JWT fid claim, which Go compares against vid + "," + fid for every URL
  form (volume_server_handlers.go:361-364), could never match a slash-form
  request. With a JWT key configured, every read, write or delete of
  /3/01637037d6 was a 401.
- build_proxy_request_info's slash branch had no extension handling, so a
  redirect for /3/01637037d6.jpg sent the client to /3,01637037d6.jpg.
  Go's proxyReqToTargetServer formats "%s/%s,%s" from the already-stripped
  fid (volume_server_handlers_read.go:128-137) and so emits
  /3,01637037d6. The peer still serves either form, since the comma form
  strips the extension again, so this one is parity rather than breakage.

Replace all three with one parse_needle_path returning vid, fid, ext and
filename borrowed from the path. The fid keeps its _delta suffix, as in
Go: parse_needle_id_cookie applies it and the JWT check strips it. The
leading slash stays optional, so chunk manifest fids still parse.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 01:13:06 -07:00
71f8128d75 shell: fs.verify -pruneEntries deletes entries whose needles are lost (#11338)
* shell: fs.verify -pruneEntries deletes entries whose needles are lost

* shell: harden fs.verify -pruneEntries guards; VolumeNeedleStatus returns NotFound for absent needles

* shell: resolve chunk manifests in fs.verify metadata path; require confirmed deletion before counting prunes

* shell: anchor fs.verify legacy missing-needle error matching

* shell: keep fs.verify metadata scan alive on manifest resolution failures

* shell: classify EC missing needles and keep manifest failures unverified

VolumeNeedleStatus now canonicalizes erasure_coding.NotFoundError to
codes.NotFound, so absent needles in EC volumes reach the prune path
through the same stable contract as regular volumes. The client-side
isNeedleMissingError keeps recognizing the legacy wrapped EC shape
("locate in local ec volume: ... needle not found") for mixed-version
clusters.

A chunk manifest that fails to resolve is now an entry-level
verification failure even when the raw top-level chunks are healthy:
the file is not fully readable without the manifest. Raw chunks are
still verified on a resolution failure so a missing top-level manifest
needle is classified and can be pruned. The per-entry logic is
extracted into resolveAndVerify for testability.

* shell: trim fs.verify prune comments

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 20:42:18 -07:00
github-actions[bot] 01545fc4ff docs: regenerate star history chart 2026-09-16 00:50:47 +00:00
1f037e48f9 s3: a list marker that sorts before the prefix excludes nothing (#11322)
* s3: a list marker that sorts before the prefix excludes nothing

ListObjects `marker` and ListObjectsV2 `start-after` are a plain key
cutoff: list the keys that sort after it. A marker that sorts before the
prefix and is not under it therefore excludes no key carrying the prefix,
and the listing must equal the one with no marker at all.

normalizePrefixMarker treated every marker that does not start with the
prefix as "something wrong" and the listing came back empty. Clients send
this shape routinely: docker/distribution's S3 storage driver walks
prefix "<root>/<path>/" with start-after "<root>" (its rootdirectory), so
on SeaweedFS a registry walk saw an empty bucket. zot read that as "no
repositories": /v2/_catalog was empty, GC/scrub/retention never saw a
repo, and on restart its storage parse deleted every repository's
metadata as "no longer in storage".

listFilerEntries now lists as if no marker were given when the marker
sorts before the prefix; the response still echoes the marker the client
sent. A marker that sorts after the prefix's subtree is left alone: it may
legitimately sit inside a partial-name prefix's match set, which
normalizePrefixMarker already handles, and otherwise correctly lists
nothing.

Reproduce on 4.44 and 4.47:

  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot/zot/"  # all keys
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=zot"       # KeyCount 0
  curl -s "$S/zot?list-type=2&prefix=zot/zot/&start-after=a"         # KeyCount 0

* s3: keep the prefix's own key excluded by a marker that names it

Fold the before-prefix marker rule into normalizePrefixMarker, which now also
derives prefixEndsOnDelimiter from the effective marker instead of each cursor
rebuilding the expression.

A marker equal to the prefix is no longer trimmed to a subtree cutoff:
start-after "a/b/" with prefix "a/b/" excludes only the "a/b/" key, so the walk
starts inside that directory and its children still list.

Adds a listing-level test that walks the whole path for both start-after shapes
a registry sends, and covers the new normalization cases.

* s3: leading slashes do not hide a marker that names the prefix

* s3: echo the V1 marker the client sent, not the walk's cutoff

* s3: filter only the walk's cutoff from the V1 page, not the echoed marker

* s3: skip the key an exclusive marker names as it streams

---------

Co-authored-by: Zuse <be9c90a8-c104-4be2-b7a4-9f92eb833ac8@forge.local>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 16:50:58 -07:00
563c729e70 rust volume: stream the tail scan and release the store lock (#11275)
* rust volume: add a .dat scan plan that runs without the store lock

DatScanPlan captures a fresh .dat handle, the version, the start offset
and an end bound while the caller holds a store guard, then visits one
record at a time with positional reads that never touch the Volume, the
way Go's ScanVolumeFileFrom feeds a scanner.

The handle pins the inode the offset was resolved against: a vacuum
commit renames .cpd over .dat and destroy unlinks it, and neither
rewrites the pinned bytes. The end bound is read while no writer can
hold store.write(), so the scan never meets a partial append. It is a
fresh open, not try_clone, because on Windows read_exact_at uses
seek_read, which moves a cursor a clone shares with the writer.

A header whose size is negative, or does not fit before the end bound,
ends the pass before the body length is computed or anything is
allocated. In today's scan a negative size reaches needle_body_length
and either overflows the buffer size or walks the scan from a wrong
offset. A size near i32::MAX overflows padding_length's i32 arithmetic,
which panics in debug builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: stream the tail scan with the store lock released

volume_tail_sender read every needle from the start offset to EOF into a
Vec while holding store.read(). volume.merge tails from zero, so that was
the whole volume in memory. And because needle writes and the heartbeat
take store.write() on a lock that prefers writers, the whole node stopped
serving until the scan finished: the failure #11235 fixed for EC scrub.

Each pass now runs on a blocking thread. Under one store guard it
resolves the start offset and captures a DatScanPlan, then drops the
guard and sends each needle as it is read, as Go's
VolumeFileScanner4Tailing does. This replaces the one-guard-across-
search-and-scan rule from the previous commit with a stronger invariant:
the offset, the handle and the end bound come from the same guard, and
the handle pins the inode, so a vacuum commit mid-scan cannot point the
offset into the compacted file.

A scan error now ends the stream with Status::internal instead of a
clean EOF, as Go's `streamFollow: %w` does. Once needles stream, a clean
EOF after a partial pass would let volume.move treat a truncated tail as
complete. A panic in the pass is reported the same way. A receiver that
hangs up is also noticed between skipped needles, not only on a send.

Unchanged: the append_at_ns filter, the header on every 2MB chunk, the
caught-up heartbeat without a scan, and the draining countdown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: fail the tail pass on a short read below the snapshot end

DatScanPlan::scan treated an UnexpectedEof on the header or body read as
the end of the data and returned Ok. Every byte below the captured end
existed when the plan was taken, so a short read there can only mean the
inode was truncated under the plan: an unmount followed by a VolumeCopy
of the same volume id reopens .dat with truncate(true). The pass then
reported Scanned, the next pass found the volume gone, and the stream
ended cleanly after a prefix of the planned records, which volume.move
would take as a complete tail.

Both short-read arms now fail the scan with an I/O error that names the
offset and the snapshot end, so tail_pass reports Status::internal as it
does for every other read failure. The break arms were carried over from
scan_raw_needles_from, where the whole scan ran under the store guard and
nothing could truncate the file.

Found by the Devin and Greptile reviews on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: sum the needle padding in i64 so a corrupt size cannot overflow

padding_length added the header, checksum and timestamp widths to the
needle size in i32. A size read from a corrupt header can sit near
i32::MAX, and that sum then overflows: a panic with overflow checks, a
wrapped padding without. DatScanPlan::scan bounds the size against the
bytes left before computing the body length, but that only keeps such a
size out of the arithmetic while under 2 GiB of the file remains, so on a
large volume the scan could still reach the overflow and, in release,
size a buffer from garbage.

Sum in i64 in both version branches. The result is at most
NEEDLE_PADDING_SIZE, so it still fits Size. The scan comment no longer
claims the bound check prevents the overflow.

Found by the CodeRabbit review on #11275.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* rust volume: propagate dat scan parse failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:40:06 -07:00
Chris LuandGitHub 55367afded Update README with performance details of 'weed' binary
Clarified the performance characteristics of the 'weed' binary, emphasizing O(1) complexity for read and write operations.
2026-09-15 14:24:27 -07:00
a9ecfeef45 helm: roll master pods when master config changes (#11331)
* helm: roll master pods when master config changes

The master loads master.toml once at startup (startAdminScripts reads
master.maintenance.scripts and sleep_minutes via viper with no config
watching), and the master ConfigMap is mounted with subPath, which
kubelet never refreshes in a running pod. So a change to
.Values.master.config today updates the ConfigMap but running masters
keep executing the old configuration until something else restarts
them.

Add a checksum/config annotation on the master pod template, following
the existing checksum/s3config pattern on the filer and s3 pods, so a
master config change triggers a rolling restart of the masters.

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Guard against duplicate keys

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Add checksum to deployment as well

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Always ensure the annotation is set

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Update comments

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* Soften stance

Signed-off-by: Evans Mungai <mbuevans@gmail.com>

* helm: merge pod annotations before checksums

---------

Signed-off-by: Evans Mungai <mbuevans@gmail.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-15 14:11:58 -07:00
Chris LuandGitHub beaf96a51d s3: cover object lock retention on version deletes (#11335)
* s3: cover WORM guarded version deletes

* s3: trim version delete comments
2026-09-15 13:31:32 -07:00
Chris LuandGitHub 93d4a6aefd s3: drain request body before error response (#11334)
* s3: drain request body before error response

* s3: keep oversized request bodies drainable
2026-09-15 13:09:07 -07:00
166af06a2b rust: cargo fmt both crates, with a commented-out fmt --check CI step (#11329)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-worker: run cargo fmt

Layout only; no token in the workspace changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: run cargo fmt

Layout only; no token in the crate changes. Every earlier Rust PR here
formatted only the blocks it touched so as not to drown its diff in
this one, and this commit is that debt paid in a single place. rustfmt
needed two passes to settle one block in handlers.rs; the committed
form is the fixed point, so `cargo fmt --check` is clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out cargo fmt --check step to both Rust workflows

Same shape as the commented clippy step from #11312: the check is
written out so that making formatting a gate is a one-line uncomment,
and whether to do that stays a maintainer call.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-15 09:29:22 -07:00
517f60e875 rust-volume: fold the 8–15-argument functions into parameter structs (#11328)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 09:04:16 -07:00
Eliah RusinandGitHub 49a680dd64 rust: tonic 0.14 / prost 0.14, drop the result_large_err allow (#11327) 2026-09-15 04:40:46 -07:00
87332eb60b Cloud/remote storage & tiering: configurable multipart upload/download concurrency (#11319)
* pb: add multipart concurrency fields to RemoteConf and tier move requests

RemoteConf gains upload_concurrency/download_concurrency (0 = client
default); VolumeTierMoveDatToRemote/FromRemote requests gain a
concurrency field (0 = backend default).

* remote storage: honor RemoteConf upload/download concurrency in s3 and azure clients

s3 client: ReadFile passes conf download_concurrency to the downloader,
WriteFile uses upload_concurrency for the uploader; previously
hard-coded 1 upload / 5 download parts. 0 keeps defaults. Same for
azure client.

* storage: plumb concurrency through backend interface and tier upload/download

BackendStorage.CopyFile/DownloadFile take a concurrency hint (<=0 =
backend configured default); s3 backend reads
upload_concurrency/download_concurrency from scaffold config with
parseConcurrency fallback, rclone updated to the new signature. Tier
move gRPC handlers forward the request concurrency to the backend.

* shell: -upload_concurrency/-download_concurrency for remote.configure, -concurrent for volume.tier

remote.configure exposes upload/download concurrency persisted into
RemoteConf; volume.tier move/evict commands forward -concurrent to the
tier move requests. Documented in master-cloud.toml scaffold.

* test: cover concurrency propagation in remote tier integration test

* remote.configure: merge existing config on partial update

Load the stored RemoteConf before saving so a partial update (e.g. only
-upload_concurrency) preserves credentials, endpoints, and type instead
of replacing them with new-config defaults. Only treat a confirmed
ErrNotFound as a new configuration; propagate all other load errors so a
transient filer failure does not overwrite stored settings.

On a type transition, reset backend-specific fields to the destination
type's new-config defaults rather than inheriting the old backend's
empty values. Bound configured concurrency to a sane maximum.

* remote storage: honor configured download concurrency in S3 and Azure

ReadFileWithConcurrency now resolves a zero request override against the
client's configured download_concurrency (new downloadConcurrency()
helpers), so the remote-mount/cache read path honors
RemoteConf.DownloadConcurrency instead of the hard-coded default.

Azure also clamps the resolved value to math.MaxUint16 regardless of
whether the fallback was used, preventing uint16 wraparound when a
configured value exceeds 65535.

* shell: rename -concurrent to -concurrency and validate tier transfer bounds

Rename the -concurrent flag to -concurrency across volume.tier.upload,
volume.tier.download, and volume.tier.compact to match the proto field and
RemoteConf field names. Add validateTierConcurrency to reject values that
would wrap int32 or exceed a 1024 cap before constructing the request.

* server: clamp tier move concurrency in gRPC handlers

Add clampTierConcurrency to both VolumeTierMoveDatToRemote and
VolumeTierMoveDatFromRemote handlers so a direct gRPC caller cannot spawn
an unbounded number of network workers.

* trim verbose comments added with concurrency feature

Remove redundant doc comments on the backend interface, rclone backend,
s3_backend parseConcurrency, and test helpers that restated the obvious.

* remote.configure: apply type defaults before re-parse so explicit flags win

applyTypeDefaults ran after the second flag parse, overwriting explicit
destination flags (e.g. -s3.region=eu-west-1) with new-config defaults.
Move the type-transition default reset before the re-parse so user-supplied
flags override the destination defaults.

* remote.configure: only treat explicit -type as a type transition

The first parse defaults -type to s3, so a concurrency-only update on an
existing non-S3 config captured requestedType=s3 and wrongly triggered a
type transition, resetting the stored backend to S3. Use fs.Visit to
detect whether -type was explicitly supplied; an omitted -type keeps the
stored backend.

---------

Co-authored-by: Jack Meredith <9480542+jackusm@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 22:09:08 -07:00
github-actions[bot] e735c12869 docs: regenerate star history chart 2026-09-15 00:48:17 +00:00
e4ca0d09e7 s3: preserve versions for POST policy uploads (#11316)
* s3: preserve versions for POST policy uploads

Route POST policy uploads through the existing version-aware write helpers
and validate promoted Object Lock headers before writing.

Return the generated version ID when versioning is enabled, return
x-amz-version-id: null when versioning is suspended, and omit the header
when versioning has never been enabled.

* s3: reuse versioning helpers in POST policy handler

Route the POST policy handler through the existing getVersioningState
and isObjectLockEnabled helpers instead of open-coding the object-lock
forces-versioning-enabled rule, matching the PUT path.

Drop the x-amz-version-id: null response header for suspended
versioning; the PUT handler omits it and the S3 PutObject sample
response for suspended buckets does not include it. Trim the moved
fileSize comment.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:36:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
01433e801d build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0 (#11306)
* build(deps): bump github.com/redis/go-redis/v9 from 9.21.0 to 9.22.0

Bumps [github.com/redis/go-redis/v9](https://github.com/redis/go-redis) from 9.21.0 to 9.22.0.
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/redis/go-redis/compare/v9.21.0...v9.22.0)

---
updated-dependencies:
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.22.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* test(redis_conf): track go-redis 9.22.0 default read timeout of 5s

go-redis 9.22.0 raised the default ReadTimeout from 3s to 5s (part of the
cross-SDK configuration alignment). Update TestUnsetKeepsGoRedisDefaults to
expect the new default so the bump in #11306 stops failing CI.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:16:35 -07:00
1f61097d4d helm: grant List to the generated read-only S3 identity (#11318)
* helm: grant List to the generated read-only S3 identity

The chart's anvReadOnly identity only carried the Read action, so its
credentials could GetObject and HeadObject but every ListObjects request
was denied: List is a separate action and the identity check is an exact
match. Add List so the read-only credentials can list buckets and objects.
Writes stay denied. Update the README example to match.

Bump the chart to 4.47.1. The chart label is part of the s3 and
all-in-one pod templates, so the upgrade rolls the gateways and they
reload the identity config, which is only read at startup.

Fixes #11317

* helm: roll standalone S3 and all-in-one on s3 config changes

Mirror the filer checksum/s3config pod annotation in the standalone S3
and all-in-one deployments so a changed generated S3 secret triggers a
rollout during a normal helm upgrade without relying on a chart version
bump.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 16:11:22 -07:00
Chris LuandGitHub 0e82b4e351 s3: populate Initiated timestamp in ListMultipartUploads (#11313)
* s3: populate Initiated timestamp in ListMultipartUploads

ListMultipartUploads returned each upload with only Key and UploadId,
omitting the Initiated timestamp. Clients such as GeeseFS rely on this
field to expire stale uploads and crash on its absence. Set Initiated
from the upload directory entry creation time so repeated listings
preserve the original initiation time.

* test/s3: verify Initiated timestamp in ListMultipartUploads

Add an integration test that initiates a multipart upload, lists it,
and asserts the Initiated field is populated and preserved across
repeated listings rather than reflecting the listing time.
2026-09-14 16:03:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
0bd048b76f build(deps): bump google.golang.org/api from 0.296.0 to 0.297.0 (#11307)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.296.0 to 0.297.0.
- [Release notes](https://github.com/googleapis/google-api-go-client/releases)
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.296.0...v0.297.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.297.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 15:09:14 -07:00
Chris LuandGitHub c997e54096 admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback (#11314)
* admin: extract isFlagExplicitlySet helper from applyViperFallback

No behavior change; the inline flag-visit check becomes a reusable
helper so the upcoming bind-address default can share it.

* admin: default to 0.0.0.0 for authenticated HTTP, keep worker gRPC on loopback

PR #11185 made the admin HTTP server default to 127.0.0.1 to stop
exposing the unauthenticated admin API on the network by accident.
That also locked out operators who already authenticated with
-adminPassword: their UI became unreachable from the network after
upgrade unless they added -ip=0.0.0.0 (see #11303).

An authenticated deployment is safe to expose, so auto-upgrade the -ip
default to 0.0.0.0 when -adminPassword or [https.admin] mTLS is
configured. The loopback default stays for the unauthenticated case, so
the unauthenticated API is never exposed on the network. An explicit
-ip is always honored.

The worker gRPC control plane has no password auth (only mTLS), so it
must not follow the HTTP upgrade. Give it a separate bind address that
stays on loopback unless -ip is explicit, so adminPassword no longer
re-exposes the unauthenticated worker stream.

* admin: hint loopback-only bind in startup banner

When the admin server binds to loopback (the default for the
unauthenticated case), print a one-line hint that it is not reachable
from other hosts and how to expose it. This helps operators who, after
the #11185 loopback default, can no longer reach the UI from another
machine quickly see the cause and the fix without reading the docs.

* admin: keep worker gRPC on loopback, decouple from https.admin mTLS

The worker gRPC auto-upgrade to 0.0.0.0 was gated on hasMTLS, which
reads the https.admin (HTTP) mTLS config. The worker gRPC mTLS comes
from grpc.admin + grpc.ca, a separate config, so:

- https.admin mTLS without grpc.admin mTLS widened the worker gRPC to
  0.0.0.0 unauthenticated (re-exposing the control plane), and
- grpc.admin mTLS without https.admin mTLS left the worker gRPC on
  loopback, blocking authenticated remote workers.

Drop the worker gRPC auto-upgrade entirely. The worker gRPC keeps the
raw -ip value (loopback by default), matching the pre-existing
behavior; an operator who wants remote workers sets -ip explicitly.
Only the HTTP admin listener auto-upgrades to 0.0.0.0 when
authenticated.

Addresses review feedback on #11314 from Devin and Greptile.
2026-09-14 14:04:35 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2aa6af033d build(deps): bump github.com/go-sql-driver/mysql from 1.10.0 to 1.10.1 (#11308)
Bumps [github.com/go-sql-driver/mysql](https://github.com/go-sql-driver/mysql) from 1.10.0 to 1.10.1.
- [Release notes](https://github.com/go-sql-driver/mysql/releases)
- [Changelog](https://github.com/go-sql-driver/mysql/blob/master/CHANGELOG.md)
- [Commits](https://github.com/go-sql-driver/mysql/compare/v1.10.0...v1.10.1)

---
updated-dependencies:
- dependency-name: github.com/go-sql-driver/mysql
  dependency-version: 1.10.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 14:04:10 -07:00
Chris LuandGitHub 02749c1192 s3api: configurable trusted-proxy allowlist for aws:SourceIp (#11302) (#11315)
* s3api: add TrustedProxies allowlist helper for aws:SourceIp extraction

Introduces a policy_engine.TrustedProxies type that parses a
comma-separated list of bare IPs and CIDRs (mirroring Guard.UpdateWhiteList)
and extracts the client IP for aws:SourceIp condition evaluation.

When the direct TCP peer is in the allowlist, X-Forwarded-For is walked
right-to-left skipping trusted hops (then X-Real-Ip); otherwise the direct
peer address is returned. This is the building block for restoring
configurable forwarded-header trust removed in b88156f (#11231), as
proposed in #11302.

* s3api: honor trusted-proxy allowlist in bucket/IAM policy engine

Make ExtractConditionValuesFromRequest a method on *PolicyEngine so it
can use the engine TrustedProxies when resolving aws:SourceIp. With no
allowlist configured the behavior is unchanged from b88156f: the direct
TCP peer is used and forwarded headers are ignored. When an allowlist is
configured via SetTrustedProxies, requests from a trusted peer honor
X-Forwarded-For (right-to-left) then X-Real-Ip.

Update the two call sites (auth_credentials.go, s3api_bucket_policy_engine.go)
and the engine tests to the method form, and add a regression test for the
trusted-proxy path.

* s3api: honor trusted-proxy allowlist in IAM role/session policies

Make extractRequestContext and extractSourceIP methods on
*S3IAMIntegration so they can use the integration TrustedProxies when
resolving aws:SourceIp. With no allowlist configured the behavior is
unchanged from b88156f: the direct TCP peer is used and forwarded
headers are ignored. When an allowlist is configured via
SetTrustedProxies, requests from a trusted peer honor X-Forwarded-For
(right-to-left) then X-Real-Ip.

Update the call site in isActionExplicitlyDeniedByIAM to type-assert
the integration and use the method, and add a regression test for the
trusted-proxy path.

* s3api: load [s3.trusted_proxies] from security.toml and wire to engines

Read s3.trusted_proxies.white_list (comma-separated IPs/CIDRs) from
security.toml and propagate the allowlist to the bucket policy engine,
the IAM policy engine (persisted across rebuilds via
IdentityAccessManagement.SetTrustedProxies), and the IAM integration.
Reloaded on SIGHUP alongside the JWT signing keys. Document the new
section in the scaffold security.toml.

Closes #11302.

* s3api: harden TrustedProxies parsing and X-Forwarded-For traversal

Canonicalize bare IP entries (via net.ParseIP + String) so non-canonical
IPv6 allowlist entries such as 2001:0db8::1 match peers rendered as
2001:db8::1, and log+skip unparseable bare entries instead of storing
them inertly.

When walking X-Forwarded-For right-to-left, stop at the first malformed
(non-empty, unparseable) entry instead of skipping it, and only fall
back to the leftmost valid IP when the chain was well-formed. This
prevents a malformed hop from masking a forged IP to its left.

Addresses review feedback on #11315.

* s3api: make TrustedProxies reload race-free via atomic.Pointer

Store the trusted-proxy allowlist behind sync/atomic.Pointer in
PolicyEngine and S3IAMIntegration so SIGHUP reloads (which swap the
allowlist) cannot race with concurrent request handlers reading it.
This mirrors the existing Guard guardState pattern. The
IdentityAccessManagement copy is already protected by iam.m.

Addresses review feedback on #11315.
2026-09-14 13:54:26 -07:00
ac03d3fd78 shell: warn when fs.mergeVolumes source holds only orphan needles (#11310)
* shell: warn when fs.mergeVolumes source holds only orphan needles

fs.mergeVolumes traverses filer entries, so a source volume whose
needles are all orphans — filer entries lost to a crashed write or a
wiped filer store — produces only the plan header and exits 0: no move,
no skip, no error. Operators read that as a successful merge while the
real cleanup (volume.fsck) never runs, and dat>idx volumes keep coming
back read-only after restarts.

Count the source-volume needles seen during traversal and, when a plan
source was never seen but its index still reports needles, print a
warning pointing at volume.fsck. Dry-run warns too.

* shell: make needle counting concurrency-safe and count manifest sub-chunks

TraverseBfs runs its callbacks from five workers, so the plain
needlesSeen map raced between source-heavy merges (fatal concurrent
map writes). All increments now funnel through a mutex-guarded
recordSeen closure.

Manifest sub-chunks that live on planned source volumes are now
recorded too — rewriteManifestChunk visits them (including dry-run
and capacity-skipped ones) but previously never marked their source,
which produced false 'orphan needles' warnings for sources whose
chunks were all reached through manifests.

* shell: extract sourceNeedleCounter so the concurrency test covers the production path

The orphan-warning recording was a closure local to Do, so
TestWarnUnreferencedSources_ConcurrentRecording could only exercise a
test-local copy of it — a regression in the production mutex would pass
the test. Lift the map and mutex into a sourceNeedleCounter type with
record/count methods and use it from Do and the test, so the -race test
now drives the actual recording path. Trim the verbose comments added
with the warning while here.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:48 -07:00
adaf3534fa rust: clippy-clean both crates and adopt the std APIs the 1.91 MSRV allows (#11312)
* rust: apply clippy --fix to both crates

The mechanical part of a clippy sweep: `cargo clippy --all-targets --fix`
on seaweed-volume and the seaweed-worker workspace, hand-reviewed. Both
manifests declare their MSRV (1.91.1 and 1.94.1), so every suggestion
clippy applied is within it: the collapsible_if sites become let chains
(1.88, edition 2024), `% n == 0` becomes is_multiple_of (1.87),
chunks_exact with a constant becomes as_chunks (1.88), repeat().take()
becomes repeat_n (1.82), and io::Error::new(Other, ..) becomes
io::Error::other (1.74). The rest is redundant clones, borrows, casts,
closures and field names.

Nothing here changes behaviour. The three let_and_return sites in
needle_map.rs and store_ec.rs deserve a note: the `let result = ..;
result` shape was a deliberate edition-2021 workaround to drop a redb
guard before the table it borrows. Edition 2024 drops tail-expression
temporaries before locals, which is why clippy now flags it, and the
two comments that described the workaround say so instead.

Manual edits on top of the tool output: the blocks clippy rewrote are
re-indented the way rustfmt lays them out (only those blocks — the
crate is not rustfmt-clean and a whole-crate fmt would bury this diff),
the blank lines let_and_return left behind are removed, and the CRC
legacy_value test compares against a literal worked out from the
original shift formula rather than restating rotate_right.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: clear the clippy warnings --fix cannot apply, and say why the rest stay

Hand fixes for the lints clippy only reports. Behaviour is unchanged
throughout; each rewrite is the one clippy names.

- needless_range_loop (7): index loops over shard vectors become
  iterator loops. Where the old code indexed `v[..n]` the new loop
  iterates `v[..n]` so an undersized vector still panics the same way.
- field_reassign_with_default (6): struct literals with `..Default`.
- redundant_pattern_matching (3): `if let Err(_) = guard.check()` becomes
  `.is_err()`, which also releases the read guard at the end of the
  condition instead of at the end of the block.
- manual_strip (2), manual_checked_ops, format_in_format_args,
  redundant_locals, wrong_self_convention (to_vif takes self by value,
  so it is into_vif; CompactEntry is Copy, so to_needle_value takes self).
- type_complexity (2): `OrphanShardLoad` and `RawNeedleEntry` name two
  tuples that were spelled out inline.
- new_without_default: CompactNeedleMap gets a Default that calls new().
- suspicious_open_options: a test helper spells out `.truncate(false)`,
  which is what `.create(true).write(true)` already did.

What stays, and the attribute that says so:

- too_many_arguments (10): `#[expect]` on each function. Folding 8–15
  parameters into a struct is a design change, not a lint fix.
- await_holding_lock / readonly_write_lock: one test holds the store
  write guard across a sleep on purpose, as a barrier that parks the
  copy task at the mount block. `#[expect(.., reason = ..)]` records it.
- module_inception: needle/needle.rs mirrors the Go package layout.

Two lints become crate-wide policy in `[lints.clippy]`, with the reason
next to each: result_large_err, because every RPC path returns
tonic::Status (176 bytes) and boxing it would change every handler
signature; and needless_update, because `..Default::default()` on a
protobuf message literal is what lets a proto gain a field without
touching every constructor (all 11 sites are pb messages). The worker
workspace gets the same table and its members opt in with
`lints.workspace = true`; its generated plugin.rs also allows
large_enum_variant on prost's oneof enums.

Both crates are now clean under `cargo clippy --all-targets -- -D warnings`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: use the std APIs the 1.91 MSRV already pays for

The crate declares rust-version 1.91.1, so a few things the code still
worked around are plain std now. All of them come from the 1.85–1.91
release notes; nothing here needs a newer toolchain than the manifest
already requires.

- std::sync::LazyLock (1.80) replaces the lazy_static! block in
  metrics.rs, and the lazy_static dependency goes. Every use site reads
  the same through Deref, so no caller changes.
- Duration::from_mins / from_hours (1.91) replace `from_secs(v * 60)`
  and `from_secs(v * 3600)` in the option parser and the shard-location
  refresh TTLs. One difference for the parser: an absurd count that
  overflows u64 seconds now panics in release builds too, where the
  multiplication used to wrap.
- Result::flatten (1.89) replaces `.and_then(|r| r)` on the replication
  join handle.
- OsStr::display (1.87) replaces `to_string_lossy()` where the name was
  only being formatted; the output is byte-identical.
- `#[allow]` becomes `#[expect]` (1.81) on the suppressions that are
  meant to be permanent, so a suppression that stops being needed
  becomes a warning rather than lingering. Doing that found four that
  already had: dead_code on ChunkManifest, base_name and last_io_error,
  and too_many_arguments on read_from_data_shards, which is down to
  seven parameters. Those attributes are deleted. The three allows that
  depend on cfg (a unix-only mutation, a linux-only field set, a
  profiling-only parameter) stay as allow, because expect would be
  unfulfilled on the other platforms.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out clippy step to both Rust workflows

Both crates are warning-free under `cargo clippy --all-targets
-D warnings` now. Whether that becomes a gate is a policy call, so the
step is present but commented out; uncommenting it is the whole change.
The comment points at the `[lints.clippy]` table where crate-wide
exceptions are recorded, so the gate does not become a reason to
sprinkle allows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust volume: guard parse_duration against overflow panics

Duration::from_mins/from_hours panic when the count overflows u64
seconds. Use checked_mul so an oversized CLI value falls back to the
parser default instead of crashing volume startup.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-14 11:29:29 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
49f20489e4 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.20.1 to 1.20.4 (#11305)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.20.1 to 1.20.4.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.20.1...service/mq/v1.20.4)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.20.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 10:13:14 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bdec508da9 build(deps): bump golang.org/x/image from 0.45.0 to 0.46.0 (#11304)
Bumps [golang.org/x/image](https://github.com/golang/image) from 0.45.0 to 0.46.0.
- [Commits](https://github.com/golang/image/compare/v0.45.0...v0.46.0)

---
updated-dependencies:
- dependency-name: golang.org/x/image
  dependency-version: 0.46.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:45:16 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
fd33c07843 build(deps): bump github/codeql-action from 4.37.9 to 4.38.0 (#11311)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.9 to 4.38.0.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.9...v4.38.0)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-14 09:41:36 -07:00
Chris LuandGitHub cf38c01978 admin: bind worker gRPC listener to -ip instead of wildcard (#11300)
* admin: bind worker gRPC listener to -ip instead of wildcard

The worker/plugin gRPC control plane called net.Listen("tcp", ":port")
directly, so it wildcard-bound every interface and ignored the -ip setting.
A cluster bound to loopback still exposed the unauthenticated
WorkerService/PluginControlService streams on 0.0.0.0. Bind through
util.JoinHostPort(bindIp, port) so the listener honors -ip like the
master, filer, and volume gRPC listeners.

* admin: warn when worker gRPC is exposed off loopback without mTLS

The worker gRPC stream has no password auth, so grpc.admin mTLS is the
only effective control once the listener leaves loopback. An operator who
sets -adminPassword and binds -ip=0.0.0.0 authenticates the HTTP API but
still exposes the unauthenticated worker control plane. Log a startup
warning naming the port and the mTLS knobs so the exposure is not silent.

* admin: address review on worker gRPC bind fix

- mini: reserve the admin gRPC port with util.JoinHostPort so an IPv6
  bindIp (e.g. ::1) does not form an invalid unbracketed address and
  lose the reservation.
- worker gRPC: track whether grpc.admin mTLS credentials actually loaded
  rather than only whether they were configured, and gate the
  non-loopback exposure warning on that. A cert/key that fails to load
  now still warns instead of silently suppressing.
2026-09-13 21:48:14 -07:00
Chris LuandGitHub f4bad510c9 test/ec: pin rack in seedAndSpread volume.grow to stop silent no-ops (#11299)
seedAndSpread() calls `volume.grow -dataNode X` without pinning the rack.
The master's grow picks the rack by weighted-random when -rack is unset,
and only one of the three racks holds the requested data node, so an
unpinned grow lands on the wrong rack two times out of three. The
VolumeGrow RPC swallows the "No matching data node" failure for
non-cache collections, so those grows count as success without creating
a volume. The per-server cap (maxGrowsPerServer=4) is then exhausted by
silent no-ops before the volumes ever spread, and seedAndSpread times
out with "volumes never spread across >=2 disks on all 3 nodes".

Pin -dataCenter dc1 and -rack rack{i} alongside -dataNode so every grow
reaches the target node. This removes the timing-sensitive assumption
that made TestECVacuumDuplicateShardClaimAcrossDisks flaky.
2026-09-13 21:05:55 -07:00
Chris LuandGitHub 15d9f6c6fe rust-volume: fix Windows build of find_needle_from_ecx (#11298)
* rust-volume: fix Windows build of find_needle_from_ecx

The .ecx binary-search fallback path used on non-Unix targets
(Seek + Read, both &mut self receivers) requires the ecx_file
binding to be mutable. On Unix the read_exact_at path takes &self,
so the mut would be unused there — gate that warning with
#[cfg_attr(unix, allow(unused_mut))].

Without this the build-rust-volume-windows CI job fails with
E0596 at ec_volume.rs:1033, breaking the weed-volume_windows_amd64
release asset.

* rust-volume: use positional seek_read for .ecx lookups on Windows

The previous fix (making ecx_file mut) compiled but left the Windows
fallback using Seek + Read on the shared .ecx file cursor. Concurrent
find_needle_from_ecx calls could interleave seek/read and read the
wrong index entry, corrupting the binary search (raised by Devin and
Greptile review on the PR).

Switch the Windows path to std::os::windows::fs::FileExt::seek_read,
which is positional (offset passed via OVERLAPPED, cursor untouched)
and takes &self — so the binding no longer needs mut, and concurrent
callers on the cached handle can't interfere. Mirrors the existing
read_exact_at helper in storage::volume.

Add a compile_error fallback for non-unix/non-windows targets to
match the convention in storage::volume.
2026-09-13 20:08:01 -07:00
ea179963c0 filer: clean up manifest resolve error propagation and add webdav tes… (#11297)
filer: clean up manifest resolve error propagation and add webdav test (#78)

Drop GitHub issue references from comments and trim verbose comments.
Replace the viewFromChunksOrErr helper with the existing
NonOverlappingVisibleIntervals + ViewFromVisibleIntervals at the stream
call sites, and add a WebDavFile.Read regression test for the manifest
resolution failure path.

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-13 18:34:09 -07:00
github-actions[bot] c507336000 4.47 2026-09-14 01:31:55 +00:00
github-actions[bot] 3c492b5ab1 docs: regenerate star history chart 2026-09-14 00:52:08 +00:00
Chris LuandGitHub 38c14d3c13 filer: apply SSRF guard to the lazy-remote fetch/list/delete paths (#11294)
* filer: add guarded remote-storage client builder hook for lazy fetch

The lazy-remote fetch path (maybeLazyFetchFromRemote) resolved its
remote-storage client through the unguarded shared cache, bypassing the
SSRF chokepoint (BuildGuardedRemoteStorageClient) that the CVE-2026-73080
remediation wired into the volume, filer stream and s3 stream dial paths.

Add a RemoteStorageClientBuilder hook on Filer plus conf-only lookups on
FilerRemoteStorage, and route the lazy fetch through the builder when set
(endpoint deny-list + DNS-rebinding-safe dialer), falling back to the
shared cache otherwise. The filer server wires the builder in a follow-up.

* filer: route lazy directory listing through the guarded remote client

maybeLazyListFromRemote shared the unguarded client resolution of the
fetch path, so a caller-supplied remote endpoint was dialed without the
SSRF deny-list or rebinding-safe dialer. Resolve the conf and build the
client through buildRemoteStorageClient so the same guard covers listing.

* filer: route lazy remote delete through the guarded remote client

maybeDeleteFromRemote issued outbound DELETE/RemoveDirectory requests
through the unguarded client, giving a write-side SSRF to a caller-chosen
endpoint. Resolve the conf and build the client through
buildRemoteStorageClient so the endpoint deny-list and rebinding-safe
dialer apply to the delete path as well.

* filer server: wire the guarded remote client builder into the filer

Set Filer.BuildGuardedRemoteClient to BuildGuardedRemoteStorageClient and
forward AllowUntrustedRemoteEndpoints so the lazy-remote fetch, list and
delete paths apply the same SSRF endpoint checks as the volume and
streaming read paths.

* filer: test lazy fetch honors the guarded remote client builder

Add a regression test that sets BuildGuardedRemoteClient to a rejecting
builder and asserts maybeLazyFetchFromRemote returns no entry without
reaching the remote, covering the SSRF guard wired in the prior commits.

* filer: skip remote client for local-only lazy deletes

maybeDeleteFromRemote resolved and validated the mount's remote client
before checking entry.Remote, so a local-only file (no Remote entry) under
a mount whose endpoint the guard rejects failed to delete: the guard
error aborted the metadata deletion, leaving a file that needs no remote
operation undeletable. Move the local-only check ahead of client
construction so only remote-backed files and directories pay the guard.

* filer: build the guarded remote client inside the lazy singleflight

The lazy fetch and list paths built the guarded client before their
singleflight blocks, so concurrent requests for the same key each
allocated a fresh SDK client and HTTP transport even though only one
remote operation ran. Move client construction inside the singleflight
so the deduplicated operation builds it once, matching the per-request
guard semantics of the sibling streaming paths without the duplicate
transport churn.

* filer: test guarded rejection for the lazy list and delete paths

Add regression tests that set BuildGuardedRemoteClient to a rejecting
builder and assert the lazy list does not reach the remote, a
remote-backed file delete is blocked, and a local-only file under a
rejected mount still deletes (covering the local-only fix).

* filer: decouple lazy guarded-client build from the first caller's context

Building the guarded client inside the singleflight made concurrent
fetches share the first caller's context. If that caller canceled while
endpoint DNS validation was running, the builder returned an error and
published a not-found result to other callers whose contexts were still
valid. Build with context.WithoutCancel so the guard's DNS validation
is not tied to any single caller's cancellation, matching the list
path's existing decoupling for the remote operation itself.

* filer: reject remote-storage confs that dial blocked endpoints at load

The filer's lazy-fetch / lazy-list / remote-delete paths resolve remote
storage clients by name from FilerRemoteStorage.storageNameToConf and
dial them via remote_storage.GetRemoteStorage, which bypasses the SSRF
deny-list the volume server (BuildGuardedRemoteStorageClient) and the
filer's own direct-read path apply. A RemoteConf planted under
/etc/remote with a loopback / private / IMDS S3 endpoint is reloaded into
storageNameToConf on the next metadata-change event and then dialed on
the next cache miss — server-side request forgery from the filer.

Apply the volume server's SSRF deny-list at conf load time, the single
chokepoint that populates storageNameToConf:

- Add RemoteStorageConfValidator, injected into FilerRemoteStorage by
  the filer server (the filer package cannot import the server package).
  A conf that fails validation is dropped from storageNameToConf, so the
  name-based client resolution on the lazy paths returns "not found"
  instead of dialing the blocked endpoint.
- Add ValidateRemoteConfForLoad in weed_server, which mirrors
  BuildGuardedRemoteStorageClient's gcs credential + endpoint checks
  (validateRemoteEndpoint via guardedRemoteClient) without building a
  client. allowUntrusted skips the check, mirroring the volume server
  opt-out (-filer.allowUntrustedRemoteEndpoints).
- The filer server injects the validator at construction.

A conf whose type dials a fixed provider host (no caller-supplied
endpoint) passes; only caller-influenced endpoints are denied.

* filer: skip DNS resolution in the load-time SSRF validator

ValidateRemoteConfForLoad resolved hostnames during /etc/remote reload,
so a transient DNS failure (2s timeout) dropped the conf from the fresh
map that replaces the live map, disabling a working mount until the next
metadata event. The build-time guard (BuildGuardedRemoteStorageClient)
already re-resolves and re-validates the endpoint at dial time with the
rebinding-safe dialer, so DNS at load is redundant for security.

Split the static checks (scheme, IMDS hostnames, IP-literal blocked
addresses, gcs credentials) into validateRemoteEndpointForLoad, which
does no DNS. Hostname endpoints pass at load and are caught at dial if
they resolve to a blocked address. This preserves fail-fast for
statically-blocked confs (loopback IPs, IMDS hostnames) without letting
transient DNS failures disable mounts.

* filer: accept empty S3 endpoints in the guarded remote client builder

guardedRemoteClient returned ok=true with an empty endpoint for a
standard AWS S3 config (no custom S3Endpoint), so
BuildGuardedRemoteStorageClient and ValidateRemoteConfForLoad rejected
it with "remote endpoint is empty" — breaking standard AWS S3 mounts on
the lazy paths and the sibling streaming read paths that already use the
guarded builder.

An empty endpoint is not caller-supplied: the AWS SDK derives the
regional endpoint from the region, so there is nothing for the SSRF
guard to validate. Return ok=false for empty S3-compatible endpoints so
the builder falls through to the shared unguarded cache, matching the
historical behavior for standard AWS S3.
2026-09-13 14:43:55 -07:00
Chris LuandGitHub 92c379e5b4 filer: accept gcs credentials file paths in the guarded remote client builder (#11296)
* filer: accept gcs credentials file paths in the guarded remote client builder

checkGcsCredentials rejected all filesystem paths, so a gcs mount
configured with remote.configure -gcs.appCredentialsFile (which stores
a path in GcsGoogleApplicationCredentials) was rejected by
BuildGuardedRemoteStorageClient with "gcs credentials must be inline
JSON". This broke existing gcs mounts on the volume, filer, and s3
remote-mount read paths that use the guarded builder.

Read and validate the file content instead of rejecting the path,
mirroring what the gcs client itself does in MakeWithHTTPClient. A path
that does not exist or does not contain valid gcs credentials is still
rejected before any client is built. guardedRemoteClient now reads the
file to extract the token exchange URL for the SSRF deny-list, so the
rebinding-safe dialer still guards the token endpoint.

* filer: resolve gcs credential paths and avoid leaking file existence

loadGcsCredentialsContent passed the raw credentials string to os.ReadFile,
so a documented ~/path (as written by remote.configure
-gcs.appCredentialsFile=~/...) was rejected because os.ReadFile does not
expand ~. It also wrapped the os.ReadFile error, which includes the
file path, exposing file existence to a caller who planted a conf with
an arbitrary path.

Resolve the path with util.ResolvePath, matching the gcs client's own
behavior in MakeWithHTTPClient. Return a generic sentinel error on read
failure so the path is not reflected in the error message. The credential
type validation still runs on the file content, so a path that does not
contain valid gcs credentials is rejected before any client is built.
2026-09-13 14:43:45 -07:00
Chris LuandGitHub 5d8a463b3e test/ec: fix EC interruption matrix slot exhaustion (#11295)
* test/ec: fix EC interruption matrix slot exhaustion

The EC integration test cluster (test/erasure_coding/chaos_lifecycle_test.go)
configured each disk with -max 4 and the seedAndSpread spread loop fired
volume.grow -count 4 every 2 s with no per-server cap. Because the master
topology lags the volume.grow writes, the loop re-fired before the prior
grow was visible, over-filling disks to capacity. A full disk leaves zero
free EC shard slots (failing the cluster-wide capacity check with "no
free ec shard slots") and drops the source disk below the encode's
FreeVolumeCount >= 2 health check (failing with "no healthy replicas"),
which aborted ec.encode before any phase marker printed and made every
encode scenario in TestECInterruptionMatrix fail.

Three changes to the test cluster:

1. Raise -max from 4 to 8 per disk so the source disk always retains
   FreeVolumeCount >= 2 for ec.encode's 14-shard generation (2 volume-slot
   equivalents) even after the spread loop and multiple encodes.

2. Switch the spread loop from -count 4 to -count 1 so each grow lands
   exactly one volume on the volume server's least-loaded disk, giving
   deterministic cross-disk spreading instead of relying on a single
   multi-volume grow to fan out.

3. Cap grows per server at 4 so heartbeat lag cannot run away and
   over-fill disks before the master registers the prior grow.

4. Pass -minFreeSpace 0 so the test is not falsely gated by the physical
   disk's free-space percentage on the host running CI (the EC shard slot
   calculation separately enforces a 90 % disk-usage cap via
   balancer.DiskTooFullAfter, which already guards against an over-set
   maxVolumeCount on a physically full disk).

Verified locally by running TestECInterruptionMatrix twice (all encode,
decode, and balance scenarios pass, including the previously failing
encode@Deletingoriginalvolumes).

* test/ec: only count successful grows toward the spread cap

A failed volume.grow (e.g. a transient collectTopologyInfo or VolumeGrow
RPC error) would otherwise consume one of the four permitted attempts
without creating any volume, exhausting the retry budget and leaving the
loop to only poll until the Eventually timeout. Increment the per-server
counter only when commandGrow.Do returns nil.
2026-09-13 14:41:45 -07:00
Chris LuandGitHub bea10e269f iceberg/s3tables: confine stored metadataLocation to the authorized table bucket (#11292)
* iceberg: confine commit/transaction/view-update write paths to authorized bucket

The create, register, and createView handlers already confine the client-
supplied metadata location to the caller table bucket and reject ".."
segments. The commit, create-on-commit, transaction, and view-update paths
read the stored metadataLocation back from the catalog and skipped the same
guard, so a location poisoned via the raw S3Tables UpdateTable API (which
persists metadataLocation verbatim) could escape the caller bucket through
a ".." segment that path.Join collapses in saveMetadataBlob.

Add confineMetadataLocation and apply it after parseS3Location on every
commit/update/transaction/view write path, mirroring the create/register/
createView check. Reject with 400 so a poisoned stored location fails the
commit instead of writing into another tenant bucket tree.

* s3tables: validate metadataLocation at the store layer

The raw S3Tables API (CreateTable, RegisterTable, UpdateTable, CreateView,
UpdateView) persisted the client-supplied metadataLocation verbatim with no
bucket-confinement or traversal check, so a caller could store a location
pointing outside its own bucket. The Iceberg REST gateway commit paths then
read that stored value back and wrote through it.

Add ValidateMetadataLocation and call it in every s3tables store handler
that accepts a metadataLocation, rejecting locations whose bucket differs
from the caller table bucket or whose path contains traversal segments. This
prevents a poisoned location from ever being persisted, complementing the
per-write-path guard added to the Iceberg commit handlers.

* iceberg/s3tables: validate location before repair and after idempotency check

Address review feedback:
- Move the commit-path confinement check ahead of repairManifests so a
  poisoned stored location cannot reach manifest repair I/O before the
  commit is rejected.
- Move ValidateMetadataLocation in CreateTable/CreateView to after the
  existing-resource check so idempotent retries that do not consume the
  requested location are not rejected for an unused bad location.
- Assert HTTP 400 in the cross-tenant reproduction tests so an unrelated
  failure cannot satisfy them.

* iceberg: confine staged metadata location before load in create-on-commit

The create-on-commit path parsed the staged metadata location from the
stage-create marker and called loadMetadataFile before validating that the
staged bucket/path stay within the authorized bucket. Add the same
confineMetadataLocation guard before the read so a tampered marker cannot
direct a cross-tenant metadata read.

* iceberg/s3tables: reject bucket-only metadata locations

ValidateMetadataLocation and confineMetadataLocation accepted s3://bucket
with an empty table path. metadataDirPath then maps every such table to the
shared <TablesPath>/<bucket>/metadata directory, so tables could overwrite
or read each other's metadata files. Require a non-empty table path in both
validators; the empty-location case (where the catalog derives one) is
unaffected.

* iceberg/s3tables: reject slash-only table paths in location validation

s3://bkt/// parses to tablePath="/" which passed the empty-string check
but path.Join cleans it away, mapping to the bucket-level metadata
directory shared across tables. Update isValidTablePath to require at
least one non-empty segment and mirror the same check in
ValidateMetadataLocation, closing the gap in all callers.
2026-09-13 13:48:13 -07:00
Chris LuandGitHub 10c0857476 s3: gate internal LifecycleDelete gRPC behind admin Bearer auth (#11291)
* s3/lifecycle: attach admin Bearer token on internal LifecycleDelete clients

Export credential.WithS3InternalAdminAuth (renamed from withIamCacheAdminAuth)
and use it in the worker and shell lifecycle RPC adapters so lifecycle calls
carry the same admin token the IAM-cache propagation already attaches. No-op
when jwt.filer_signing.key is unset, matching the server-side checkAdminAuth.

Prepares the internal clients for the server-side auth gate that follows.

* s3/lifecycle: gate LifecycleDelete behind admin Bearer auth

Add checkAdminAuth to LifecycleDelete, matching the SeaweedS3IamCache handlers on the same internal gRPC listener (PR #11190). No-op when jwt.filer_signing.key is unset; rejects unauthenticated callers when it is. The internal worker/shell clients already attach the token in the previous commit.
2026-09-13 13:07:10 -07:00
c462fffce6 master: name the unlabeled disk layout plainly in assign errors (#11290)
* master: name the unlabeled disk layout plainly in assign errors

When no volume server serves the layout an assign targets, the error
named the empty disk type as "hdd" (HardDriveType is the empty string),
sending operators looking for servers labeled hdd when the actual
mismatch is labeled (e.g. -disk=ssd) servers versus unlabeled clients.

- describe the layout as "default (unlabeled)" when the disk type is
  empty, keep %q naming for labeled types
- log the unserved-layout condition once per option instead of letting
  every failing write repeat an unactionable line

Observed in production: volume servers started with -disk=ssd while CSI
mounts assign with the unlabeled layout; the per-write error stream
pointed at a nonexistent hdd fleet.

* master: bound and expire the unserved-layout warning dedupe

The dedupe map retained every distinct option key permanently. Option
keys embed request-derived fields (collection, disk type), so repeated
assignments with distinct options would grow master memory without
bound, and a retained key suppressed the warning if the same option
went unserved again after the topology recovered.

Remember last-warned timestamps instead, expiring after an hour, with a
hard cap that resets the set when a client-driven key flood fills it.

* master: silence per-retry unserved-layout log and name explicit hdd

Addresses Devin Review comments on #11290.

- The unserved-layout branch already rate-limits its warning via
  assignUnservedLayoutWarning.Do, but the common epilogue still logged
  lastErr at V(0) on every retry, so the flood the dedup was meant to
  stop continued. Skip the epilogue log when the unserved-layout branch
  owns the logging; the error is still returned to the client.
- describeDiskLayout took the canonicalized option.DiskType, but
  ToDiskType folds both "" and "hdd" into HardDriveType, so an explicit
  disk=hdd request was mislabeled "default (unlabeled)". Pass the
  original request disk type instead: only an empty request is the
  unlabeled default; an explicit hdd is named "hdd".

Adds TestAssignFailsFastNamesExplicitHdd covering the explicit-hdd
wording.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-13 11:56:32 -07:00
Chris LuandGitHub 99d2479528 fix(vacuum): batch fsync in makeupDiff to prevent test timeout (#11289)
makeupDiff called dstDatBackend.Sync() (fsync) per needle in the loop
over incrementedHasUpdatedIndexEntry. With 20000 entries in
TestLDBIndexCompaction this resulted in up to 20000 fsync calls, which
on slow CI disks exceeded the 10-minute test timeout.

Batch the sync: write all needles/tombstones first, then fsync the dat
file once in the defer alongside the existing idx fsync. The durability
guarantee is unchanged — both files are still synced before CommitCompact
writes the .cpc commit marker and swaps the files.
2026-09-13 00:05:30 -07:00
8db41d0217 [Mount] Cache Chunk Manifest Resolution for Repeated File Opens (#11266)
* cache resolved chunk manifests for Mount

* Address PR review: per-mount cache, singleflight, reuse ResolveOneChunkManifest

- Own the manifest cache per WFS mount instead of a process-global
  variable, so manifests from one filer backend are never served to
  another (Devin/CodeRabbit major bug).
- Coalesce concurrent cold misses via singleflight so only one fetch
  runs during a cold burst (Greptile P2).
- Copy cached data after releasing the mutex so a large copy does not
  block concurrent hits, inserts, and evictions (CodeRabbit nitpick).
- Reuse the existing ResolveOneChunkManifest function name instead of
  introducing a new resolveOneChunkManifest wrapper.
- Validate (unmarshal) manifest bytes before caching so malformed
  manifests do not poison the cache.
- Add TestChunkGroupManifestResolutionCoalescesColdMisses covering
  the singleflight cold-miss path.

* Address round 2 review: coalesced-miss cancellation, test overlap

- Use singleflight.DoChan in fetchOrLoad and select on ctx.Done() so a
  caller whose context is canceled while waiting for an in-flight fetch
  returns ctx.Err() promptly instead of blocking for the leader's
  result (Devin BUG).
- Add TestResolveOneChunkManifestCanceledWaiterReturnsDuringCoalescedMiss
  covering the canceled-waiter path.
- Delay the cold-miss fixture response so the leader's fetch is still
  in flight when concurrent opens join the singleflight, making the
  one-fetch assertions reliable (CodeRabbit Minor).

* Address review: keep ResolveOneChunkManifest four-argument

Restore the exported ResolveOneChunkManifest to its original
four-argument signature so external callers keep compiling. Move the
cache-aware resolution into an unexported resolveOneChunkManifest
helper that accepts the per-mount ChunkManifestCache. The exported
function delegates to the helper with a nil cache, preserving the
historical uncached behavior for every non-Mount caller. The Mount
path (ChunkGroup.SetChunks) now calls the unexported helper with the
mount-owned cache. Tests and benchmarks that exercise the cache path
call the unexported helper directly.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-12 20:06:25 -07:00
github-actions[bot] bd6bcd47e3 docs: regenerate star history chart 2026-09-13 00:54:31 +00:00
Bruce ZouandGitHub eb6a7e93ca Fix mount eio on manifest resolve failure (#11287)
* mount: fail reads with error when chunk manifest resolution fails

When SetChunks fails to resolve a chunk manifest (e.g. the volume is on a
remote tier with reads disabled), the sections map stays empty and
readDataAtSequential/readDataAtParallel zero-fill every missing section as
if it were a sparse hole. Reads then return all-zero data with no error,
so a plain cp of a large manifest-based file silently produces a
completely zero-filled file.

Remember the resolve error in ChunkGroup (guarded by sectionsLock) and
return it from ReadDataAt. A later successful SetChunks clears it.

Fixes the mount path of #11286.

* filer: propagate manifest resolve errors in streaming read paths

ViewFromChunks discards the chunk manifest resolve error returned by
NonOverlappingVisibleIntervals. On failure the chunk views come back
empty, and the streaming paths zero-fill the entire requested range,
serving HTTP 200 / WebDAV 200 responses whose body is all zeros.

Propagate the error in PrepareStreamContentWithThrottler,
PrepareStreamContentWithPrefetch and the WebDAV read path so these
requests fail with 500 instead.

Fixes the filer HTTP and WebDAV paths of #11286.

* mount: fail lseek with EIO when chunk manifest resolution fails

SearchChunks still consulted the stale section map after SetChunks
recorded a manifest resolution failure, so SEEK_DATA/SEEK_HOLE would
describe the unresolved regions as sparse holes or return ENXIO.
Return the recorded error from SearchChunks and map it to EIO in
Lseek.

Also add regression tests for the stream preparation error paths.

Addresses review feedback on #11287.
2026-09-12 14:37:36 -07:00
5b2fe374fc [Volume] Scrub every disk's EC shards for a volume id, not just the first (#11258)
* storage: add Store::find_all_ec_volumes for split-disk EC lookups

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: add merge_ec_runtimes to resolve a vid's per-disk shard set

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: replace dead slots.get(14) assertion with a width-14 pin

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: build the checksum scrub plan from every per-disk runtime

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: build the local scrub plan from every per-disk runtime

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: prove the local scrub plan reaches every runtime's slots

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: make the scrub plan tests falsifiable

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: report unverifiable protection when the sidecar predates the scrubbed encode

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: commit sidecar provenance with the sidecar it describes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: scrub every disk's EC shards for CHECKSUM and LOCAL

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: run the FULL/READS parity check across split-disk shards

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: report fenced-out runtimes in FULL/READS scrubs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: tighten verify_ec_shards ordering and missing-shard coverage

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: visit each EC volume id once in node-wide scrubs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: cover split-disk scrub aggregation end to end

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* volume server: pin fenced-out disks and sibling-disk shards in EC scrubs

Three scrub behaviors shipped without a test at the RPC seam. Task 8
showed the seam exists, so close them here.

FULL/READS (mode 2|5) now marks a volume broken when the identity fence
excludes a runtime, where it previously reported clean. Pinned against a
control fixture whose two disks AGREE and scrub clean, so the test fails
on the clean->broken transition, not only on the message text. That needs
a structurally valid, tombstone-only .ecx (so the needle walk finds
nothing to complain about) and a seeded shard-location cache (so the
absent master does not short-circuit the scrub with an error of its own).

LOCAL (mode 3) and CHECKSUM (mode 4) now build their plans from every
per-disk runtime. Made observable by moving shard 0 -- the shard the
volume's single needle spans and the one the checksum sidecar is checked
against -- to the SIBLING disk, leaving shard 5 on the disk the singular
find_ec_volume lookup returns. Built from that disk alone, neither scrub
ever looks at shard 0.

The split-disk fixture grows a config struct rather than more positional
arguments; its defaults reproduce the existing layout byte for byte, so
the node-wide dedupe test is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: report fenced-out disks on a malformed sidecar too

`errors.extend(self.skipped)` sat below the whole status match, so only
`(Some(p), On)` ever reached it. The Invalid arm already returns a
non-empty error vector of its own, so the Go-parity contract that
silences the Off arm (`case BitrotOff: return 0, nil, nil`) does not
reach it -- appending the fence lines there costs nothing that contract
protects. A volume with BOTH a malformed sidecar and a disk the identity
fence excluded reported only the sidecar, hiding the unscanned disk
behind an unrelated integrity error.

Off stays byte-identical, and so does the `(None, On)` arm that is
documented as treating a missing payload defensively as protection off.
Off is now the ONLY status that drops the report, and the comment at the
On-path copy says so: that is the one place the parity constraint costs
us coverage.

Also corrects a false claim in the FULL/READS test's doc comment. It
said a fenced-out disk "is a disk this scrub did NOT read", which is true
only of the merge-driven parity half. The per-needle walk still resolves
`store.find_ec_volume` (store_ec.rs:281) and binds
`expected_encode_ts_ns` to that runtime (:311) -- position 0, the
EXCLUDED one on that fixture -- so `read_local_intervals`' generation
filter (:1204) makes it read the excluded disk and treat the anchor's
shards as non-local, the inverse of what `skipped` reports. The fixture's
tombstone-only .ecx walks nothing, so the test cannot tell the two apart;
the comment now says that rather than implying coverage it does not have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: take CHECKSUM's bitrot protection from the disk that has the sidecar

`EcChecksumScrubPlan::for_volumes` read `(prot, status)` off the ANCHOR.
The anchor is the first shard-bearing runtime at the maximum `encode_ts_ns`,
chosen with no regard for which disk holds the `.ecsum`.

That sidecar is deliberately NOT mirrored across disks -- `ec_metadata_dirs()`
exists so one authoritative copy stays reachable rather than being duplicated
-- and at mount `EcVolume::new` resolves it via `load_active_bitrot_sidecar(&[])`
with no sibling directories at all; only the `VolumeEcShardsMount` RPC ever
passes `ec_metadata_dirs()`. So after EVERY volume-server restart, the
split-disk runtime that does not physically hold the sidecar mounts
`BitrotStatus::Off`. When the one copy lives on disk 1 and the anchor is disk 0,
`run()` hit `case BitrotOff` and returned `(0, [], [])`: the whole volume
scrubbed clean, silently. That is the steady state for roughly half of all
mirrored split-disk layouts, and it is the exact failure this branch exists to
remove.

Source protection from the first MERGED runtime that has any -- `On` if one
does, else `Invalid`, else the anchor's `Off`. Two facts make that safe, and
both are load-bearing:

  - Every runtime that mounted `On` already passed the `geometry_matches` gate
    in `load_bitrot_for_generation`, so its manifest agrees with the volume's
    layout. A sidecar that contradicted it would have failed the mount.
  - All merged runtimes share the same `encode_ts_ns` by construction of the
    identity fence, so a sidecar from any of them describes the same encode run.

The `unverifiable_sidecar` provenance rule four lines down read
`anchor.bitrot_source_dir`; it now reads the SAME runtime `prot` came from.
Otherwise the two would describe different sidecars and the rule would vouch
for a manifest nobody is scanning against. One consequence worth naming: that
source dir is now non-empty by construction (a runtime with protection found a
file), where the anchor's was often "" and short-circuited the rule -- so on a
fenced volume whose anchor had no sidecar, an unverifiable-protection note now
surfaces where previously nothing was reported at all.

`run()` is untouched, and the `BitrotStatus::Off` arm still returns
`(0, [], [])` exactly, for Go parity with `case BitrotOff: return 0, nil, nil`.
`parity_shards` still comes from the anchor while `prot` may come from a
sibling; the geometry gate above makes them agree, and slot-width agreement is
handled separately.

The test drives mode 4 through the real RPC against a split-disk volume whose
sidecar exists only on dir1, and asserts up front that the anchor mounted `Off`
and the sibling `On` -- otherwise it would prove nothing. Reverting this commit's
one-line source change makes it report `[]` instead of `[0, 5]`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: pin the slot width, contain the shard-size fallback, and cover multi-disk FULL

Five findings from the whole-branch review, none of which changes what a
healthy volume reports.

Slot width was undefined and the two consumers disagreed (ec_volume.rs).
`merge_ec_runtimes` sizes `slots` to the WIDEST merged runtime, but the identity
fence keys on `encode_ts_ns` alone and never on geometry -- so two
same-generation runtimes whose `.vif`s disagree do merge. The mode 2|5 arm
truncates to the anchor's `data+parity` and silently drops the surplus slots,
while `EcChecksumScrubPlan::for_volumes` iterated the full width and emitted
"present but missing from sidecar manifest" for exactly those ids. Nothing in
the volume describes them -- the sidecar manifest and the Reed-Solomon matrix
are both the anchor's -- so that message was the width disagreement talking, not
a finding. The `slots` field doc now states the contract (the range is the
anchor's geometry; every consumer truncates to it) and CHECKSUM truncates.

The LOCAL `shard_size` fallback had grown a node-wide blast radius
(ec_volume.rs). `anchor.shard_file_size()` returns the anchor's FIRST held
shard, not a maximum. Before aggregation the plan read only that runtime's own
shards, so a truncated shard was contained to its disk; now that one value sizes
every merged sibling's shards, mis-offsetting `locate_data` and manufacturing
needle corruption across the node. Take the max over the merged slots, which is
how `verify_ec_shards` already answers the same question
(`if size > shard_size { shard_size = size }`). Only on the legacy
`dat_file_size == 0` path.

Multi-disk `all_local` had no end-to-end test (grpc_server.rs). The parity check
is gated on every shard being present, and the one all-local fixture keeps them
in a single directory, so every entry of `dirs` is the same string and a
permutation or off-by-one in the `slots` -> `dirs` mapping is invisible;
`test_verify_ec_shards_reads_shards_from_multiple_dirs` builds its `dirs` by
hand and never goes through `merge_ec_runtimes`. The new fixture is a real 10+4
encode split 0..=6 / 7..=13 across two store locations (the `.dat`/`.idx` stay
outside both, so `prune_incomplete_ec_with_sibling_dat` has nothing to act on),
driven through the real RPC: clean first, then a corrupted PARITY shard on the
SECOND disk -- which only the parity half can see, and only through a correct
mapping. Shifting that mapping by one, or computing `all_local` from the anchor
alone, both make it report `[]` instead of `[13]`.

Deleted `test_ec_volume_enumeration_is_deduped` (store_ec_reconcile.rs). It
built `raw` from `store.locations` and then applied its OWN inline
`filter(|v| seen.insert(*v))`, asserting on that -- a property of
`HashSet::insert`, never reaching the production dedupe. That path is covered by
`test_scrub_ec_volume_node_wide_dedupes_a_split_disk_volume`, which does fail
(2 != 1) when the dedupe is removed.

Corrected `test_verify_ec_shards_treats_a_none_dir_as_missing`'s docstring
(ec_encoder.rs). It claimed the unmounted shard "must not drag the shards that
ARE mounted down with it", but `dirs[5] = None` puts shard 5 in `broken_shards`
before the block loop, so every iteration takes the `read_failed` arm and the
parity comparison never runs: corrupting a mounted shard in that fixture changes
nothing about the result. The assertions are unchanged; the docstring now states
what they actually establish.

Also refreshed two comments that cited `shard_file_size() - 1` as the reason
`merge_ec_runtimes` prefers a shard-bearing anchor -- true before this commit,
stale after it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: correct the Fix 1 rationale and truncate the shard-size scan

The safety argument attached to `EcChecksumScrubPlan::for_volumes`'s protection
selection was false as written, and it is the argument a reviewer reads first.
`geometry_matches` compares a sidecar against the MOUNTING runtime's own
data/parity/block size, not the anchor's, and returns true vacuously when
`ec_shard_config` is `None` -- so it establishes agreement only when all merged
runtimes share one geometry, which an `encode_ts_ns`-only fence does not
guarantee and which `test_checksum_scrub_truncates_slots_to_the_anchors_geometry`
constructs a counterexample to. The second clause was weaker than stated too: a
`.ecsum` records no encode identity at all, so merged runtimes agreeing on
`encode_ts_ns` does not transfer to the sidecar.

Replace it with the property that is true, checkable from the selection itself,
and stronger for what actually matters. `anchor` is an element of `merged`, so
the `.unwrap_or(anchor)` fallback is reached only when no merged runtime is `On`
and none is `Invalid` -- in which case the anchor is necessarily `Off`. The
status can therefore only move `Off -> On`, `Off -> Invalid` or
`Invalid -> On`; never `On -> Off`, never `Invalid -> Off`. This selection
cannot stop a volume that was being scanned from being scanned, and cannot turn
a reported integrity error into silence: every change it makes is toward more
verification. The comment now also states what it does NOT establish -- geometry
agreement is not guaranteed -- and names geometry fencing as the follow-up that
would close it.

Second, `EcLocalScrubPlan::for_volumes`'s `shard_size` max scanned the FULL slot
width, violating the `slots` contract documented in the same commit that
introduced the max: the volume's shard-id range is the anchor's geometry and
every consumer must truncate to it. Pre-fix that input could not exist, because
`anchor.shard_file_size()` read only the anchor's own anchor-sized vector -- so
the max opened a new, narrow path to the same node-wide mis-sizing it exists to
close (same-generation runtimes with disagreeing `.vif`s, the wider one holding
an out-of-geometry shard larger than the in-geometry ones, `dat_file_size == 0`).
`.take(anchor.data_shards + anchor.parity_shards)` mirrors the truncation
already applied to the CHECKSUM shard scan.

The sibling `shards:` vector is left untruncated on purpose: every access in
`EcLocalScrubPlan::run` is `shards.get(sid)` with `sid < data_shards`, so the
surplus entries are inert.

No behavior change for any healthy volume, and no test added -- the suite is
unchanged at 575 passing, 0 failing, 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VUf2cmVKHNhAZPTNv39rDE

* ec: aggregate split-disk runtimes in Go scrubs, mirroring Rust

Go volume scrubs previously used FindEcVolume (first runtime only), so a
volume whose EC shards are split across multiple disks was scrubbed against
just one disk's shards and the others were silently skipped. Node-wide
ScrubEcVolume also appended each disk's EcVolumeIds without deduplication,
scrubbing a split-disk volume once per disk.

Add MergedEcRuntimes/MergeEcRuntimes (Go counterpart to Rust's
merge_ec_runtimes): select the maximum EncodeTsNs as the anchor generation,
fence out runtimes whose encode generation or geometry (DataShards,
ParityShards, BlockSize) disagrees with the anchor, merge shard handles by
shard ID, and report excluded runtimes rather than dropping them. Wire it
into every scrub mode:

- INDEX: scrub the anchor's index, report skipped runtimes.
- LOCAL: aggregate local shards across all merged runtimes via a synthetic
  EcVolume built from the merged shard slots.
- FULL/READS: resolve the runtime matching the anchor's encode generation
  (not the first match) so the needle walk and parity phase inspect one
  encode run; report skipped runtimes.
- CHECKSUM: take bitrot protection from the first merged runtime that has a
  valid sidecar (On, else Invalid, else anchor's Off), preserve invalid
  sidecar errors from every other merged runtime, and report skipped
  runtimes.

Deduplicate EC volume IDs in node-wide ScrubEcVolume so each volume is
scrubbed exactly once.

Refactor ScrubEcVolume to share the per-needle walk via scrubEcVolumeWalk,
called by both the legacy first-runtime path and the new merged path.

Add Go regression tests covering split-disk deduplication, encode-generation
fencing, geometry fencing, sibling-disk LOCAL reach, and merge anchor
selection.

Rust: keep the previously-landed merge/fence/checksum changes intact; revert
incidental cargo-fmt drift from unrelated files so the diff stays focused.

* ec: fence merged CHECKSUM on sidecar encode generation and fix legacy shard size

Address two review findings on the Go merged-runtime scrub:

1. Sidecar provenance: a merged runtime can load a bitrot sidecar from a
   sibling metadata directory (ReloadBitrotSidecar), and the merge fence may
   then exclude the runtime owning that directory. Generation-0 sidecars do
   not identify the encode run, so geometry validation alone cannot prove the
   borrowed manifest describes the anchor shards. If the sidecar records a
   non-zero EncodeTsNs that disagrees with the anchor, refuse the scan
   instead of applying stale checksums to current shards and reporting false
   corruption.

2. Legacy shard size: for volumes without datFileSize in .vif,
   LocateEcShardNeedleInterval derives the shard size from Shards[0].ecdFileSize.
   The merged shard set is compacted in shard-ID order, so a truncated
   lowest-ID shard would shrink every interval and misread intact sibling
   shards. Synthesize a datFileSize from the maximum mounted shard size when
   the anchor lacks one, so the datFileSize>0 path uses the largest shard
   size across all merged runtimes.

* ec: fix copylocks, legacy shard boundary, and encode-aware Rust lookups

Address review findings from CodeRabbit and Devin:

Go (ec_volume_merge.go):
- Remove bitrotLock copy from the synthetic EcVolume: copying a sync.RWMutex
  is a go vet copylocks error. The synthetic volume uses its own zero-value
  mutex; bitrot/bitrotStatus are set directly before ChecksumScrub reads them
  via BitrotProtection(), so no concurrent access occurs.
- Fix legacy shard-size boundary: synthesize datFileSize from
  (maxShardSize - 1) * DataShards, not maxShardSize * DataShards, to match
  the legacy fallback in LocateEcShardNeedleInterval (ecdFileSize - 1). An
  exact large-block boundary is ambiguous; the unadjusted size would select
  an extra large row and misread intact sibling shards.

Rust (store_ec.rs):
- Add find_ec_volume_for_scrub helper that resolves by encode generation
  (not first-match find_ec_volume) and use it in scrub_snapshot_under_lock,
  write_back_shard_locations, and the post-refresh shard-location read.
  Previously the encode-aware lookup was only used for the initial runtime
  selection; the cache write-back and per-needle snapshot still used
  first-match, so a split-disk volume whose first runtime was from an older
  encode run would write to and read from the wrong runtime's shard-location
  cache and falsely abort with 'remounted as a different encode run'.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-12 14:19:18 -07:00
Chris LuandGitHub 3b4a681e53 test(fuse_failover): dump chunk list and hex on append corruption (#11285)
* test(fuse_failover): dump chunk list and hex on append corruption

The failover append test (TestAppendWhileVolumeServerRestarts) failed
in CI with an 8-byte NUL region at offset 632 that appeared in both
the writer mount and the filer own view, but the failure message
only showed a quoted-string window around the divergence. That is
not enough to tell which chunk covered the zeroed bytes or which
volume server held it, so the next recurrence would be just as
unattributable.

Add a FileChunkList helper that reads the filer resolved chunk
list, and on failure dump:
  - every chunk fid, offset, size, volume id, and current master
    holders, flagging the chunk that covers the first divergence;
  - a hex+ASCII dump of the writer mount around the divergence so
    the exact zero-filled region is visible byte-for-byte.

No production code is touched; this only makes the test fail louder.

* test(fuse_failover): preserve diagnostic collection errors

Address review feedback from CodeRabbit and Greptile on PR #11285:

- FileChunkList now returns the wrapped ParseUint error when
  fid.volume_id is zero and the file_id prefix is invalid, matching
  FileVolumeIds instead of silently keeping vid=0 (which would
  query /dir/lookup?volumeId=0 and report the wrong holders).
- dumpChunkList captures the VolumeHolders error and renders it as
  'lookup failed: ...' so a failed master request is distinguishable
  from a successful lookup with no holders (both previously showed
  'holders=[unknown]').
- runChaosAppend captures the writer-mount read error and includes
  it in the failure message so an unavailable writer view is not
  mistaken for corrupted content.
2026-09-11 23:08:39 -07:00
Chris LuandGitHub c46f82d29a fix(master): stop goraft server on shutdown and bump raft to v1.2.1 (#11284)
MasterServer.Shutdown only stopped the Hashicorp raft implementation;
when using the default goraft backend, the raft event-loop goroutine
(leaderLoop/followerLoop) kept running after the master shut down. In
the in-process test harness this leaked goroutines across sequential
test runs, and a stale event occasionally reached a leader at term 0
and tripped the goraft "leader.elected.at.same.term" assertion,
crashing the whole test binary (CI run 34670959967, PR 11279).

Stop the goraft server in Shutdown() so its goroutines exit cleanly,
and bump seaweedfs/raft to v1.2.1 which replaces that assertion with a
graceful step-down to Follower instead of a panic.
2026-09-11 22:25:05 -07:00
Chris LuandGitHub 5a0e017457 s3: reject virtual-host bucket retargeting via X-Forwarded-Host (#11281)
* s3: reject virtual-host bucket retargeting via X-Forwarded-Host

SigV4 verification tries the client-supplied X-Forwarded-Host as a
signed host candidate, while routing and IAM select the bucket from
the actual Host header.  A presigned URL for one virtual-host bucket
could therefore be retargeted to another bucket accessible to the same
signing identity by changing Host and adding X-Forwarded-Host.

After the signature matches a host candidate, extract the bucket that
the candidate implies (via the configured virtual-host domains) and
compare it with the bucket the router selected.  Reject when they
differ, before returning success.

* test(s3api): cover virtual-host presigned URL retargeting

Add unit tests for bucketFromVirtualHost and end-to-end tests that
reproduce the X-Forwarded-Host retargeting attack for both presigned
and signed requests, plus a negative test confirming the legitimate
same-bucket case still verifies.

* s3: harden bucketFromVirtualHost for case and overlapping domains

Compare host and domain suffixes case-insensitively so a mixed-case
X-Forwarded-Host cannot bypass the consistency check.  Only treat the
exact path-style domain as non-virtual-host; subdomains of a path-style
domain still match the virtual-host router pattern and must be checked.
2026-09-11 22:23:26 -07:00
Chris LuandGitHub 210afacd12 s3: close list-type / ownership-controls routing mismatch (#11280)
* s3: reject list-type paired with another operation subresource

?list-type=2&ownershipControls= routes to ListObjectsV2 (the list-type
route is registered first) while the IAM action resolver resolves the
ownershipControls selector to s3:GetBucketOwnershipControls. A principal
denied s3:ListBucket but allowed s3:GetBucketOwnershipControls would
therefore list the bucket. list-type selects an operation just like the
other keys in operationSubresources, so add it there and reject the
combination before routing, matching the fix for policy&tagging (#10987).

* s3: resolve list-type to s3:ListBucket ahead of bucket subresources

The router registers the ListObjectsV2 route ahead of the bucket
subresource routes, so the action resolver should resolve list-type the
same way. Without this, a request carrying list-type and another operation
selector resolves to the subresource action (e.g. s3:GetBucketOwnershipControls)
while being served by ListObjectsV2. The ambiguity guard rejects such
combinations before routing, but resolving list-type to s3:ListBucket keeps
the resolver aligned with the router, mirroring how versions is handled.

* s3: match list-type=2 exactly in action resolver

The router selects ListObjectsV2 only for list-type=2; other values fall
through to the subresource routes. Resolve the same way so the action
matches the handler for every list-type value, not just 2.
2026-09-11 22:21:58 -07:00
Chris LuandGitHub 9f6feef299 feat(s3api): add bucket quota S3 extension via ?seaweedfs-quota (#11279)
* feat(s3api): add bucket quota S3 extension via ?seaweedfs-quota

Add a SeaweedFS-specific S3 subresource for bucket quota management:

  PUT /{bucket}?seaweedfs-quota   — set bucket quota (s3:PutBucketQuota)
  GET /{bucket}?seaweedfs-quota   — get bucket quota (s3:GetBucketQuota)

The request/response body is JSON:
  {"quota_size": 100, "quota_unit": "GB", "quota_enabled": true}

Quota is stored on the bucket's filer entry (positive = enabled,
negative = disabled but retained, zero = no quota), matching the
existing admin REST API behavior. When quota is cleared, the bucket's
read-only flag is also lifted.

Authentication uses the existing S3 SigV4 flow — no new global secret
is needed. Authorization uses two new dedicated IAM permissions:
  s3:PutBucketQuota
  s3:GetBucketQuota

This allows integrations like Apache CloudStack to manage per-bucket
quotas through the S3 endpoint with a scoped credential, without
exposing the broad admin REST API or requiring a separate admin token.
The credential can be limited to s3:PutBucketQuota/s3:GetBucketQuota
only, preventing bucket deletion, user management, or cluster topology
changes.

The coarse-grained ACTION_PUT_BUCKET_QUOTA/ACTION_GET_BUCKET_QUOTA
constants are added to s3_constants, and the action resolver maps the
seaweedfs-quota query parameter to the fine-grained s3: actions for
policy evaluation.

* docs: update design for S3 ?seaweedfs-quota extension approach

Replace the broad admin REST API + bearer-token design with the narrow,
scoped S3 ?seaweedfs-quota extension. Update quota, usage reporting, and
SeaweedFS-side changes sections to reflect PR #11279.

* fix(s3api): address review comments on quota handler

Fix four issues identified by Devin, Greptile, and CodeRabbit reviews:

1. Integer overflow in convertQuotaToBytes: large quota_size values
   (e.g. 8388608 TB) could overflow int64, wrapping to negative and
   being silently treated as zero quota. Now returns an error when
   size * multiplier would exceed math.MaxInt64.

2. Disabled quotas returned negative sizes in GET: the GET handler
   returned entry.Quota directly, which is negative for disabled-but-
   retained quotas. Now returns the absolute magnitude as quota_size
   and derives quota_enabled from the sign, making the response
   round-trippable.

3. Missing buckets returned 500 instead of NoSuchBucket: the PUT
   handler treated all lookup failures as internal errors. Now
   distinguishes filer_pb.ErrNotFound and returns ErrNoSuchBucket.

4. Trailing JSON was silently accepted: the decoder read only the
   first JSON object without checking for trailing data. Now
   requires EOF after the object, rejecting malformed payloads.

Also add tests for overflow detection and trailing data rejection.

* fix(s3api): cast math.MaxInt64 to int64 for 32-bit vet

On 32-bit platforms, math.MaxInt64 is an untyped int constant that
overflows int (32-bit) when used directly in fmt.Errorf with %d.
Cast to int64 explicitly to fix Go Vet 32-bit.

* docs: reconcile design doc with implementation and add AWS tools note

- Resolve open question about IAM endpoint path: driver accepts optional
  iamUrl and defaults to <s3Url>/iam
- Add note explaining ?seaweedfs-quota is not callable by standard AWS tools
  (aws s3api, s3cmd, rclone), and how this compares to MinIO and Ceph quota
  APIs which also live outside the standard S3 API

* docs: fix IAM endpoint default — SeaweedFS IAM is at POST / on S3 endpoint

SeaweedFS registers its embedded IAM API at POST / on the same S3
endpoint (UnifiedPostHandler), not under /iam. The design doc
previously said the driver defaults iamUrl to <s3Url>/iam, which would
send IAM operations to an unregistered path. Correct the default to
s3Url.

Found by Greptile review on PR #11279.

* docs: fix credential model, signer, and GET response shape in design doc

Three issues found by CodeRabbit review on PR #11279:

1. Credential-scope contradiction: the doc claimed the service credential
   is scoped to only s3:PutBucketQuota/s3:GetBucketQuota, but the
   implementation uses it as the admin credential for all operations
   (bucket CRUD, IAM user provisioning, quota). Document the actual
   model.

2. S3Signer -> AWSS3V4Signer: the doc said 'S3Signer for SigV4 signing'
   but S3Signer is legacy SigV2. Correct to AWSS3V4Signer.

3. GET response shape: the doc showed a single JSON example with 'GB'
   for both PUT and GET, but GET always returns quota_unit 'B' and the
   absolute byte count. Document PUT input and GET response separately.
2026-09-11 22:17:30 -07:00
Chris LuandGitHub 79994b69af s3: fail closed on unsupported bucket-policy condition operators (#11283)
* s3: support StringEqualsIgnoreCase and related condition operators

The S3 bucket-policy condition engine rejected StringEqualsIgnoreCase
(and StringNotEqualsIgnoreCase, StringLikeIgnoreCase,
StringNotLikeIgnoreCase), which AWS and the IAM policy engine both
accept. Add evaluators and register them in GetConditionEvaluator so
valid policies using these operators evaluate correctly instead of
being skipped.

* s3: reject bucket policies with unsupported condition operators

validateStatement did not check Condition operators, so a policy with
an unknown operator (e.g. a typo or unsupported key) was accepted at
upload time and only surfaced at evaluation, where it was silently
skipped. Reuse GetConditionEvaluator to reject unknown operators when
a policy is parsed or stored, failing closed at the entry point
instead of relying on evaluation-time handling.

* s3: fail closed on unsupported condition operators at evaluation

EvaluateConditions skipped statements whose condition operator was
unsupported, logging a warning and continuing. With no remaining
conditions to fail, the function returned true, so an Allow statement
conditioned on an unrecognized operator became unconditional and
granted access to private objects. Return false instead so an
unrecognized operator fails the condition block and the statement does
not match, matching the fail-closed behavior of the IAM policy engine.

* s3: validate condition operators at upload time only, not load time

Validating condition operators in validateStatement rejected the whole
policy document from ParsePolicy, which SetBucketPolicy uses when loading
stored bucket policies. A legacy policy saved before this change could
contain an unsupported operator, and rejecting it at load time dropped
the entire policy - including unrelated explicit Deny statements - so
the bucket lost its protections. Move the operator check into
ValidateBucketPolicy, which only the PutBucketPolicy handler and admin
UI run at upload time, so legacy policies still load and EvaluateConditions
fails the unsupported statement closed instead.

* s3: drop non-AWS StringLikeIgnoreCase and StringNotLikeIgnoreCase operators

AWS defines StringEqualsIgnoreCase and StringNotEqualsIgnoreCase but
not StringLikeIgnoreCase or StringNotLikeIgnoreCase (StringLike and
StringNotLike are case-sensitive only). Registering the wildcard
IgnoreCase variants made the engine accept operators AWS rejects. Keep
only the two AWS-defined IgnoreCase operators and add a test asserting
the wildcard IgnoreCase names are unsupported.
2026-09-11 22:17:11 -07:00
42b0ca7850 s3 sink: report the source read error the SDK hides (#11277)
* s3 sink: report the source read error the SDK hides

filer.backup stops for good on an event whose chunks are gone from the
volume servers: the uploader reads the body, the read fails with the
volume's 404, and the AWS SDK returns "ContentLength=N with Body length 0"
without the cause. isIgnorable404 would skip such an event, but it never
sees the 404, so the event is retried forever and the checkpoint never
advances.

ChunkStreamReader keeps its first source failure and the s3 sink returns
it when the upload fails.

* s3 sink: trim verbose comments on source error propagation

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-11 19:44:30 -07:00
github-actions[bot] 9b902a7662 docs: regenerate star history chart 2026-09-12 00:46:11 +00:00
d8aa7ecf04 fix(vacuum): stop comparing compact size against the live needle map (#11263)
* fix(vacuum): stop comparing compact size against the live needle map

CompactByIndex's post-copy integrity check compared bytes written to
the .cpd against v.nm.ContentSize()-DeletedSize(), the live map that
keeps mutating for as long as the volume stays writable during the
copy. Any write landing after the point-in-time index snapshot was
loaded made the live map's tally exceed what got copied, aborting
compaction with "unexpected new data size" — even though
CommitCompact's makeupDiff exists specifically to reconcile writes
that land mid-copy. On a busy volume this can fail every vacuum cycle.

Tally the expected live size from oldNm, the same frozen snapshot the
copy loop reads from, instead of the live map. This keeps the check's
original protection (destination smaller than what should have been
copied signals real data loss) while removing the false positive from
ordinary concurrent traffic.

* fix(vacuum): stop double-subtracting skipped bytes from the size check

Unreadable needles return before reaching the expectedLiveBytes tally,
so it already excludes them. Subtracting skippedDataBytes again on top
loosened the integrity check's margin by that same amount, letting a
.cpd short of the true expected size slip past undetected — the exact
failure mode the check exists to catch. Flagged independently by three
automated PR reviewers (Devin, Greptile, CodeRabbit).

Extract the comparison into exceedsExpectedCompactedSize and drop the
subtraction entirely; add TestExceedsExpectedCompactedSize to pin the
threshold to expectedLiveBytes alone.

* fix(vacuum): trim verbose integrity-check comment

Reduce the 8-line block comment to a concise 3-line rationale. No
behavior change.

* fix(vacuum): mirror compact integrity check in Rust volume server

Mirror the Go fix in the Rust volume server's do_compact_by_index:
tally expected_live_bytes from the frozen index snapshot (not the live
needle map) and compare the compacted .dat against it after the copy.
Unreadable needles already return before the tally, so no skipped-byte
adjustment is needed. Adds exceeds_expected_compacted_size and two
regression tests.

* fix(vacuum): exercise makeup_diff in Rust concurrent-write test

Address CodeRabbit review: write a needle after compaction (before
commit), then call commit_compact() and assert the late write survives
via makeup_diff. This actually exercises the concurrent-write path
rather than just confirming the integrity check passes.

---------

Co-authored-by: chrislusf <chris.lu@gmail.com>
2026-09-11 17:08:37 -07:00
Chris LuandGitHub 2ebfeabfce mount: rebuild expired directory cache on entry lookup (#11268)
* test: reproduce expired directory cache degrading lookup to N RPCs

After cacheMetaTtlSec elapses the kernel can still serve a directory
listing from its page cache, so ReadDir never runs and EnsureVisited is
not called. Metadata lookups then fall through to one LookupEntry RPC
per entry instead of rebuilding the directory cache once.

Issue #11262

* mount: add expired-directory rebuild predicate with cooldown to InodeToPath

ShouldRebuildExpiredDir distinguishes a TTL-expired cached directory from
a never-cached, invalidated, evicted, or read-through one (those clear
isChildrenCached, while a plain TTL expiry keeps it set). It also gates
retries on a cooldown since the last failed rebuild attempt, recorded by
MarkRebuildAttempt, so a transient listing failure does not trigger a
full rebuild on every later lookup.

Issue #11262

* mount: rebuild expired directory cache on entry lookup

When the kernel still serves a directory listing from its page cache past
cacheMetaTtlSec, ReadDir never runs and EnsureVisited is not called, so
lookupEntry issues one LookupEntry RPC per entry. Rebuild the expired
directory once via ensureDirectoryVisited before the cache-hit check so
later lookups are served locally. The EnsureVisited singleflight
deduplicates concurrent rebuilds.

On a non-oversized rebuild failure, record the attempt so the cooldown
suppresses repeated rebuilds while the listing keeps failing; once it
elapses a later lookup retries, recovering without waiting for ReadDir.
Oversized dirs are already marked read-through by ensureDirectoryVisited.

Issue #11262

* test: cover concurrent rebuild dedup and rebuild-cooldown fallback

Add a test that runs concurrent lookups into the same expired directory
behind a gated listing, asserting they share one rebuild via the
EnsureVisited singleflight. Add a test that a failed rebuild records the
attempt so an immediate retry is suppressed (per-entry RPC fallback), and
that once the cooldown elapses and the filer recovers a later lookup
rebuilds the cache.

Issue #11262

* mount: wait for pending async flush before rebuilding parent cache

The rebuild lists the parent directory from the filer, so a pending
async flush of the target entry must land first; otherwise the rebuilt
cache captures pre-flush metadata and the cache-hit path returns it
without the wait that guards the filer-fallback path. waitForPendingAsync
Flush is a no-op when no flush is pending, so the common case is unaffected.

Issue #11262
2026-09-11 11:54:10 -07:00
Chris LuandGitHub a3638e479e fix(s3api/audit): surface OIDC identity claim in audit log for STS sessions (#11269)
* Add ResolveIdentityClaim helper for OIDC audit identity

ComputeParentUser derives a stable per-identity hash from (sub, iss) for
internal keying, but it is opaque and not human-readable. Audit logs for
STS-assumed OIDC sessions currently surface that opaque value (or the
random session id) as the requester, leaving no authoritative trace of the
federated user.

Add ResolveIdentityClaim next to ComputeParentUser to recover a
human-readable, server-asserted identity attribute from the STS request
context populated at federation time. It walks a priority list
(preferred_username, email, name, sub) so a federated session always
audits against a stable OIDC claim rather than a client-supplied role
session name.

For #11264

* Surface authoritative OIDC identity claim in S3 audit log

For STS-assumed sessions minted from an OIDC web identity, the audit log
requester field is the opaque session subject, which cannot be traced back
to the federated user who performed the operation. The OIDC identity claims
(preferred_username, email, sub) are already carried in the session request
context and reach the auth layer as identity.Claims, but they were never
surfaced to the audit log.

Add a requester_identity field to the S3 access audit log, populated from
the authoritative OIDC identity claim resolved via ResolveIdentityClaim.
The claim is propagated through the shared identity holder (the same
mechanism the requester name and principal ARN already use) so it survives
the request-context copy that hides auth-set values from the outer audit
middleware.

The existing requester field is left unchanged for backward compatibility;
requester_identity is empty for non-federated sessions, where requester
already carries the real username.

For #11264

* Gate OIDC audit identity on federation marker and harden resolver

Address review feedback (Devin Review, Greptile) on the initial
implementation:

- Non-federated STS sessions no longer gain a false requester_identity.
  ValidateJWTWithClaims merges the JWT registered sub claim (the opaque
  session id) into RequestContext for sessions without an explicit request
  context, so the previous ResolveIdentityClaim fallback to sub surfaced
  that session id as an authoritative identity. Resolution is now gated on
  SessionInfo.ParentUser, which is set only for OIDC-federated sessions in
  AssumeRoleWithWebIdentity. The claim is resolved from the original
  sessionInfo.RequestContext (not the local claims map, whose sub the bearer
  path overwrites with the session subject) so SigV4 and bearer sessions
  surface the same identity.

- ResolveIdentityClaim now trims whitespace and treats whitespace-only
  claims as absent, so a blank preferred_username no longer masks a usable
  email or sub.

The resolved claim is carried on Identity.IdentityClaim (and IAMIdentity for
the bearer path) rather than re-derived in recordIdentityInContext, making
the federation gate explicit at the auth boundary.

For #11264

* Resolve OIDC identity claim for external bearer tokens

The external OIDC bearer path (a raw OIDC JWT presented directly, not via
STS) populates Claims with preferred_username/email/name/sub from the
validated token but did not set IdentityClaim, so requester_identity stayed
blank for that authentication path. Resolve the claim there too — sub is the
real OIDC subject on this path (not an STS session id), so no federation
gate is needed.

Also drop an ineffectual ctx assignment flagged by ineffassign in the audit
test.

For #11264
2026-09-11 11:29:46 -07:00
Chris LuandGitHub 80dae68dbf fix: write the new key when a remote-synced file is renamed (#11270)
* refactor: extract update event handling into processUpdateEvent

Pull the OldEntry/NewEntry update branch of the remote sync event
processor into its own function so the rename skip logic can be
exercised by tests with stub clients. No behavior change.

* test: reproduce remote sync rename dropping the new key

A rename under a remote mount arrives as an update whose NewEntry
inherits the source RemoteEntry. shouldSendToRemote returns false
for it, so processUpdateEvent skipped the event without writing the
new key, while the filer had already deleted the old object. The test
runs such an event through processUpdateEvent and expects both a
delete of the old key and a write of the new one. Fails before the
fix. See #11261.

* fix: write the new key when a remote-synced file is renamed

A rename under a remote mount arrives as an update whose NewEntry
inherits the source RemoteEntry, so shouldSendToRemote returns false
(RemoteMtime >= Mtime) and processUpdateEvent skipped the event.
That skip is only valid when the destination key is unchanged; a path
change always needs a write, and the delete-old/write-new handling
below the early return is exactly what a rename needs. Guard the skip
with proto.Equal(oldDest, dest) so a rename falls through to it.

Fixes #11261.

* fix: skip empty upload when renaming a remote-only entry

A remote-only entry (no local chunks or content, data lives only on
the remote object) carries a positive RemoteSize but nothing for
NewFileReader to read. After the previous commit lets a rename fall
through to the delete-old/write-new path, such a rename would upload
EOF and create a zero-byte object at the new key, then stamp it as
synced. Guard the write so a path change on a remote-only entry skips
the upload instead of replacing the file with zero bytes. The filer
has already deleted the old object, so the data is gone regardless;
this avoids leaving a misleading empty object behind.

* fix: propagate old-key delete errors except already-deleted

When deleting the old key on a rename fails for a non-multipart entry,
the error was swallowed and the write proceeded, which could leave both
remote keys. Return the error so MetadataProcessor retries the event.

The filer deletes the source remote object synchronously during the
rename, so the sync delete is redundant and the object may already be
gone. GCS reports that as ErrRemoteObjectNotFound (unlike S3/Azure,
whose deletes are idempotent), so treat it as a successful deletion and
continue to retriedWriteFile rather than pinning the sync offset.
2026-09-11 10:47:37 -07:00
Chris LuandGitHub 5ff49909a0 fix(s3api/iam): avoid transient AccessDenied from full reloads on single IAM file changes (#11271)
* fix(s3api/iam): fail config snapshot on empty or malformed IAM files

A full IAM reload reads every identity/policy/service-account/group file
from the filer. When an external secrets tool rewrites a file, a reload
that reads it mid-rewrite sees empty or partially-written content. The
identity, policy and service-account loaders silently skipped such files
(``continue``), so the snapshot was missing entries that still existed
on disk. The atomic swap then installed an incomplete identity set while
``isAuthEnabled`` stayed on, denying unrelated clients mid-reload
(#11259).

The group loader and the read-error paths already fail the snapshot in
this situation (a skipped entry reads as deleted). Apply the same
behavior to empty content and unmarshal failures across the identity,
policy, service-account and group loaders, so a transient mid-rewrite
fails the reload (preserving the last known-good state) instead of
silently dropping entries.

* fix(s3api/iam): coalesce burst IAM config reloads through the reload queue

onIamConfigChange did a full synchronous reload for every identity/policy
file change event. When several independently-refreshing credentials
rewrite their files within the same second, that produced a burst of
dozens of back-to-back full reloads, each reading the whole store and
widening the window where a mid-rewrite file is observed (#11259).

Route every IAM config change through the existing coalescing reload
queue (scheduleReload/reloadRetryLoop) instead. A burst of N events now
collapses into a single reload (plus one tail reload for events that
arrived while one was in flight). scheduleReload gains a reason argument
for the existing log line; the reloadRetryLoop already retries failed
reloads, so the per-event failure handoff is no longer needed.

Tests that asserted on the synchronous reload now wire up the queue
(centralized in newTestS3ApiServerWithMemoryIAM) and poll via
waitForIdentity/waitForIdentityGone. Adds TestOnIamConfigChangeCoalescesBurstReloads
showing 50 events coalesce into <=3 reloads.

* fix(s3api/iam): skip non-JSON auxiliary files before failing IAM snapshot

Per review: the multi-file loaders unmarshal every entry in an IAM
directory, so a non-JSON auxiliary file (README, .DS_Store, a migration
backup such as identity.json.old) would hit the new empty/malformed
errors and reject the whole snapshot, blocking all later IAM reloads.

Only *.json files are IAM objects (SeaweedFS writes identities,
policies, service accounts and groups as <name>.json, and other call
sites already gate on the .json suffix). Skip non-.json entries at the
top of each loader loop, before reading content, so auxiliary files are
ignored while empty/malformed .json files still fail the snapshot.

Adds TestLoadConfigurationIgnoresNonJsonAuxiliaryFiles.

* fix(s3api/iam): reject IAM files with empty identifiers and skip aux in listing

Per review:

- ListPolicyNames listed every regular entry in the policies directory as a
  policy name, including non-JSON auxiliary files, but GetPolicy cannot
  retrieve them. Apply the same .json suffix filter used by the loader so
  the list only exposes retrievable policies.

- json.Unmarshal accepts `{}` and unknown fields. The identity and group
  loaders merge by the decoded Name (not the file name), so a `{}` file
  could install an empty-key record and displace a real one; the
  service-account loader accepted an empty Id. Validate Identity.Name,
  Group.Name and ServiceAccount.Id (via validateServiceAccountId) after
  unmarshal and fail the snapshot on empty identifiers.

Adds TestFilerEtcStoreListPolicyNamesSkipsNonJsonAuxiliary and
empty-identifier regression tests for identity, group and service-account
files.
2026-09-11 10:42:19 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
bc0efa4d10 build(deps): bump github.com/rclone/rclone from 1.75.0 to 1.75.1 (#11274)
Bumps [github.com/rclone/rclone](https://github.com/rclone/rclone) from 1.75.0 to 1.75.1.
- [Release notes](https://github.com/rclone/rclone/releases)
- [Changelog](https://github.com/rclone/rclone/blob/master/RELEASE.md)
- [Commits](https://github.com/rclone/rclone/compare/v1.75.0...v1.75.1)

---
updated-dependencies:
- dependency-name: github.com/rclone/rclone
  dependency-version: 1.75.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-11 09:42:47 -07:00
3ae9e332ec rust volume: honour is_last in the tail sender instead of rescanning the whole volume (#11273)
* rust volume: honour is_last in the tail sender instead of rescanning

volume_tail_sender discarded the is_last flag from
binary_search_by_append_at_ns:

    Ok((offset, _is_last)) => {
        if offset.is_zero() { Ok(sb_size) } ...

is_last means the caller is already caught up. Go answers that with a
heartbeat and does not scan at all (volume_grpc_tail.go, `if isLastOne`).
Dropping it is expensive rather than untidy, because the branches interact:
when the search reports caught-up it returns Offset::default(), which is
zero, so the start offset falls back to sb_size -- the beginning of the
data -- and scan_raw_needles_from materialises every needle from there to
EOF into a Vec. The timestamp filter discards all of it, the loop sleeps
2s, and it happens again.

A volume being moved is marked read-only before the copy, so it is ALWAYS
caught up during the tail phase. Measured on one volume.move of a 2.15 GB
volume, sampling the source's cgroup anon every 2s against the move's own
phase output:

  copying   16 -> 37 MB          CopyFile streams correctly, stays bounded
  tailing   904 -> 2166 -> 629 -> 2166 -> 342 -> 2173 -> 2179 MB
  deleting  46 MB

Six full-volume allocate/free cycles in 35s, peak 2179 MB against a volume
of 2147 MiB. The destination never exceeded 35 MB, so this is entirely
source-side. Under a per-process memory cap it OOM-kills the source
whenever the volume exceeds the cap.

The ordering here is the whole fix and is easy to get wrong: resolve the
start offset and is_last under a brief lock, return the heartbeat
immediately when caught up, and only then reach the scan. An earlier cut
set the flag correctly but placed the early return after the block that
performs the scan -- the heartbeat fired and the destination received
nothing, yet every iteration still read the whole volume and discarded it.
Production showed no improvement (1770 MB across five cycles), which is
what caught it. The binary search is over the .idx and costs nothing; the
scan is the expensive part and must not run speculatively.

Three tests, and the last two matter as much as the first: a fix that
always reported "caught up" would make tailing silently lose needles, a
worse bug than the one being fixed. One asserts is_last for a caller at or
beyond the newest append_at_ns; one asserts NOT is_last for a caller that
is behind, so real tail data is still scanned and shipped; one asserts NOT
is_last when the only newer record is a delete, and that scanning from the
returned offset ships exactly that tombstone.

Left deliberately unfixed, and worth separate changes: the scan still
collects into a Vec rather than streaming through a visitor as Go's
ScanVolumeFileFrom does, and it runs while holding store.read(), the same
lock-across-a-large-read shape as #11235. Both are latent once the rescan
is gone, since remaining scans are bounded by genuinely new data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MFr2v4BUqrXdgj4LEUAwVF
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: resolve and scan the tail under one store guard

The tail sender took store.read() once for the binary search and again
for the scan. A vacuum commit takes the store write lock and swaps
.dat/.idx, so it could land between the two: the offset resolved against
the old files would then be applied to the new ones and start the scan
inside an unrelated record. The code before the is_last fix held a
single guard for both. Restore that, and scan only when the caller is not
caught up, so the caught-up heartbeat still skips the scan and is sent
outside the lock.

Also pin the compacted-volume boundary raised in review. Compaction
writes .idx in needle-id order in both Go and Rust, so the search can
report caught-up while an earlier row is newer; such a caller's since_ns
is the last row's timestamp, so those rows were in the files it copied.
A write made afterwards is appended as the final row, which the search
cannot step past. The new test asserts it still reaches the scan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018VF7E9SHPihG1jC1grU9H3

* rust volume: make the compaction tail test a genuine overwrite

The compaction regression test's second id=1 write reused the first
write's data, so write_needle's dedup short-circuit (is_file_unchanged)
returned without appending or updating append_at_ns. Compaction then
kept key 1's original (older) timestamp, so the test passed without
exercising the overwrite it describes -- key 2 was the final row only
because key 1 was never actually newer.

Give the overwrite distinct data so it appends a new record, and assert
key1_ns > key2_ns up front so a future dedup regression fails the test
instead of silently hollowing it out. Trim the verbose comments on the
tail sender and the binary-search tests to their essentials.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-11 09:42:20 -07:00
Chris LuandGitHub 0de9c1f231 fix(s3api/sts): respect MaxSessionLength config in DurationSeconds validation (#11267)
* Refactor parseDurationSeconds into a STSHandlers method

Convert the parseDurationSeconds wrapper from a package-level function
into a method on STSHandlers so it can reach the configured STS service.
No behavior change; the three AssumeRole* handlers now invoke it via
their receiver.

* Respect MaxSessionLength config in STS DurationSeconds validation

parseDurationSeconds validated DurationSeconds against a hardcoded
43200s (12h) ceiling, so raising maxSessionLength in iam.json above
12h had no effect on AssumeRole, AssumeRoleWithWebIdentity, or
AssumeRoleWithLDAPIdentity — requests were rejected at the handler
before reaching the service layer.

Derive the upper bound from the configured STS MaxSessionLength,
falling back to maxDurationSeconds (43200s) when unset. The service
layer (calculateSessionDuration) already caps the issued duration
at MaxSessionLength, so this only relaxes the input-validation gate.

* Add tests for STS DurationSeconds MaxSessionLength bound

Cover the configured MaxSessionLength upper bound, rejection above
it, fallback to the 43200s default when STS config is unset, the
900s minimum, and the empty-parameter nil path.

* Refactor validateSessionDurationSeconds into a STSService method

Convert validateSessionDurationSeconds from a package-level function
into a method on STSService so it can reach the configured STS config.
No behavior change; the three assume-role entry points in the service
(AssumeRoleForPrincipal, validateAssumeRoleWithWebIdentityRequest,
validateAssumeRoleWithCredentialsRequest) now invoke it via their
receiver.

* Respect MaxSessionLength config in STS service DurationSeconds validation

The STS service validateSessionDurationSeconds rejected DurationSeconds
above a hardcoded 43200s (12h) ceiling, so even after the handler
accepted a longer duration it was rejected again in the service layer
for AssumeRoleForPrincipal, AssumeRoleWithWebIdentity, and
AssumeRoleWithCredentials.

Derive the upper bound from the configured MaxSessionLength, falling
back to DefaultMaxSessionLength (43200s) when unset. The issued
duration is still capped at MaxSessionLength by calculateSessionDuration.

* Add tests for STS service DurationSeconds MaxSessionLength bound

Cover the configured MaxSessionLength upper bound, rejection above
it, fallback to the 43200s default when STS config is unset, the
900s minimum, and the nil DurationSeconds path.

* Preserve capping when MaxSessionLength is below the API minimum

Deriving the DurationSeconds upper bound directly from MaxSessionLength
created an empty valid range when MaxSessionLength is configured below
the 900s API minimum, rejecting every explicit DurationSeconds that the
old code silently capped via calculateSessionDuration.

Only apply the configured MaxSessionLength as the upper bound when it is
at least minDurationSeconds; otherwise keep the default bound and let
calculateSessionDuration enforce the shorter configured limit.

* Add tests for sub-minimum MaxSessionLength capping behavior

Verify that a MaxSessionLength below the 900s API minimum keeps the
default upper bound so explicit DurationSeconds within the default
range are still accepted (and later capped by calculateSessionDuration).
2026-09-10 23:03:32 -07:00
github-actions[bot] b3aace2a08 docs: regenerate star history chart 2026-09-11 00:47:23 +00:00
4f9bbd51cb rust volume: stop glibc retaining freed EC buffers as unreturnable heap (#11255)
* rust volume: stop glibc retaining freed EC buffers as unreturnable heap

A Rust volume server doing EC work accumulates hundreds of MB of resident
anonymous memory that it never gives back, and under a hard cgroup
MemoryMax that ends in an OOM kill while most of the resident set is
free-but-unreturned.

It is not a leak. glibc serves allocations >= M_MMAP_THRESHOLD with mmap
and munmaps them on free, but the threshold is ADAPTIVE: freeing an
mmap'd block raises it toward that block's size, up to 32 MiB. EC
reconstruction and needle reassembly allocate large short-lived buffers,
so the first few train the threshold upward and every later buffer is
carved from the heap instead. Heap pages only return to the OS from the
top of the arena, so they stay resident for the life of the process --
reusable, but anonymous, and anonymous pages cannot be reclaimed under
pressure the way page cache can. The retained footprint is exactly the
headroom a burst of maintenance work needs.

Measured on a 17-node cluster (EC 10+4, --index=redb), one node, two
identical `ec.scrub -mode full` rounds over 10912 EC files each, same
unit restarted with and without a pinned threshold:

                       baseline  round 1  round 2  60s idle
  default (adaptive)      10 MB    84 MB    88 MB     88 MB
  pinned threshold        10 MB    13 MB    14 MB     14 MB

78 MB retained versus 4 MB for identical work. On heavier mixed scrub
workloads the same effect reached ~600 MB per volume server against a
3 GiB cap, and restarting the process was the only way to release it.

Calling mallopt(M_MMAP_THRESHOLD, ...) sets the threshold and disables
the dynamic adjustment. Pin it to glibc's own default rather than
inventing a value: the goal is to stop the adaptation, not to second-guess
the default. MALLOC_MMAP_THRESHOLD_ still wins if an operator sets it,
glibc-only, and a failed mallopt is logged rather than fatal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MFr2v4BUqrXdgj4LEUAwVF

* Address PR review: validate env overrides, honour GLIBC_TUNABLES, fix non-glibc test compile

Three review-bot findings on seaweed-volume/src/malloc_tuning.rs:

1. (CodeRabbit) The test used cfg!(...), which keeps both branches in
   compilation. On non-glibc targets DEFAULT_MMAP_THRESHOLD is undefined,
   so the test failed to compile. Split into #[cfg]-gated tests so each
   branch only references items defined for that target.

2. (Greptile) MALLOC_MMAP_THRESHOLD_ was checked by presence only. An
   empty or non-numeric value makes glibc ignore the override while we still
   skipped mallopt, leaving the adaptive threshold enabled -- exactly the
   behaviour this module exists to prevent. Now we defer only when the value
   is non-empty and parses as an integer; otherwise we fall through to
   pinning.

3. (Codex) The modern GLIBC_TUNABLES=glibc.malloc.mmap_threshold=... tunable
   was missed, so mallopt could overwrite an operator's explicit tunable. Now
   we detect that tunable (with the same validation) and defer to it.

The override check moved into the glibc-gated inner function, so off glibc
pin_mmap_threshold() always reports NotApplicable regardless of any
allocator env vars that happen to be set. The startup log for DeferredToEnv
is reworded to cover both override sources. Added tests for the override
parsers and the off-glibc no-op.

* Address round-2 review: match glibc's actual override parsing

Three follow-up review-bot findings after the first round of fixes, all
rooted in our validation not matching how glibc actually parses the
overrides:

1. (Greptile, P1) parse::<i64>() accepted negative values like "-1" and
   returned DeferredToEnv, but glibc's threshold is unsigned and rejects
   negatives — so we skipped mallopt while glibc also ignored the override,
   leaving the adaptive threshold enabled. Now we reject negatives and
   zero.

2. (Devin, BUG) glibc parses thresholds as unsigned (strtoul for tunables,
   atoi for the legacy var). Values above i64::MAX are valid for glibc but
   were rejected by parse::<i64>(), so we pinned 128 KiB over the operator's
   explicit setting. Now we parse as u64, accepting the full unsigned range.

3. (CodeRabbit, Major) Two issues in usable_glibc_tunable_threshold:
   a. A malformed sibling entry (e.g. glibc.malloc.check=2=2:...) makes
      glibc reject the entire GLIBC_TUNABLES string, but our per-entry scan
      still returned true for the valid-looking mmap_threshold entry. Now
      we validate every entry (exactly one '=') before accepting any.
   b. Hex values (0x20000) are accepted by glibc's strtoul but were rejected
      by parse::<i64>(). Now parse_strtoul_threshold handles 0x-prefixed hex.
      MALLOC_MMAP_THRESHOLD_ stays decimal-only (atoi), matching glibc.

Added regression tests for negatives, zero, >i64::MAX, hex tunables, and
malformed mixed GLIBC_TUNABLES entries. Verified: clippy clean and tests
pass on macOS (non-glibc); glibc-gated code type-checks for
x86_64-unknown-linux-gnu.

* Address round-3 review: match glibc's actual override parsing

Three follow-up review-bot findings (Greptile P1, Devin BUG, CodeRabbit
Major) all on the same issue: the round-2 fix rejected negative and zero
override values, but glibc actually accepts them.

Verified against the glibc source (malloc/malloc.c, malloc/arena.c,
elf/dl-tunables.c, elf/dl-misc.c):

- do_set_mmap_threshold(size_t value) does NO clamping — it just sets
  mp_.mmap_threshold = value and mp_.no_dyn_threshold = 1.
- MALLOC_MMAP_THRESHOLD_: glibc calls atoi(value) then mallopt, which
  always sets the threshold and disables dynamic adjustment — even for
  empty, negative, or non-numeric values (atoi returns 0). So ANY
  presence of the variable means the operator's override is in effect.
  Reverted to presence-only check for the legacy variable. The round-1
  Greptile comment claiming glibc "cannot apply the override" for
  empty/malformed values was incorrect.
- GLIBC_TUNABLES: glibc parses values with _dl_strtoul (elf/dl-misc.c),
  which accepts decimal, 0x hex, 0 octal, an optional sign (negatives
  wrap to unsigned long), and requires the entire value consumed
  (tunable_parse_num checks endptr == strval + len). Replaced
  parse_strtoul_threshold with dl_strtoul_consumes_all that replicates
  _dl_strtoul's parsing and checks full consumption. Now accepts -1
  (wraps to SIZE_MAX), 0, 0x20000, 010 (octal), and values above
  i64::MAX.

The duplicate-= validation for GLIBC_TUNABLES (from round 1) is kept —
glibc's parse_tunables_string returns -1 if any entry's value contains
a duplicate =, rejecting the entire string.

Added dl_strtoul_consumes_all tests covering decimal, hex, octal,
negative, zero, empty, whitespace, trailing garbage, and sign-only
inputs. Updated usable_glibc_tunable_threshold tests to accept
negative, zero, and empty values. Verified: clippy clean and tests
pass on macOS (non-glibc); glibc-gated code type-checks and clippy
clean for x86_64-unknown-linux-gnu.

* Address round-4 review: add overflow detection, fix sign-only test assertions

Two Greptile P1 findings:

1. Overflowing tunables bypass threshold pinning: dl_strtoul_consumes_all
   consumed every digit and returned true for values like
   18446744073709551616 (u64::MAX + 1), but glibc's _dl_strtoul stops at
   the overflowing digit (sets endptr there, returns UINT64_MAX), so
   tunable_parse_num rejects the value (endptr != strval + len). Added
   overflow detection matching glibc's cutoff/cutlim logic — on overflow,
   the parser stops and returns false.

2. Sign-only parser assertions fail: the test asserted
   !dl_strtoul_consumes_all("-") and !dl_strtoul_consumes_all("+"), but
   _dl_strtoul skips the sign, finds no digit, sets endptr to the position
   after the sign (== end of string), and returns 0. tunable_parse_num
   sees endptr == strval + len → true. So glibc accepts sign-only strings
   as value 0. Fixed the test assertions to expect true.

Also fixed "0x" with no hex digits: _dl_strtoul parses "0" as octal, then
stops at "x" (not an octal digit), so endptr != end of string → rejected.
The base-detection now requires a hex digit after "0x" before switching
to hex; otherwise "0" is parsed as octal and "x" stops the parser.

Added overflow regression tests: 18446744073709551616 (u64::MAX + 1),
99999999999999999999 (20 nines), 0x10000000000000000 (2^64). Verified:
clippy clean and tests pass on macOS (non-glibc); glibc-gated code
type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Address round-5 review: accept bare 0x prefix, remove unused helper

Two review-bot findings (Devin BUG + CodeRabbit Major) on the same issue:
the round-4 fix required a hex digit after "0x" before switching to hex
base, but glibc's _dl_strtoul unconditionally advances past "0x"/"0X"
when the first char is '0' and the next is 'x'/'X' — even if no hex digit
follows. In that case the digit loop breaks immediately, endptr reaches
the end, and the value is 0. tunable_parse_num accepts it.

Removed the is_digit_in_base lookahead from the base-detection condition
and the now-unused is_digit_in_base helper. Updated the test assertions
for "0x" and "0X" to expect true (accepted as value 0).

The Greptile P1 overflow comment is invalid: glibc's _dl_strtoul rejects
18446744073709551616 (u64::MAX + 1) — on overflow it sets endptr to the
overflowing digit (not end of string) and returns UINT64_MAX, so
tunable_parse_num sees endptr != strval + len and rejects. My
implementation correctly returns false for this value, matching glibc.

Verified: clippy clean and tests pass on macOS (non-glibc); glibc-gated
code type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Address round-6 review: rewrite tunable parser to match glibc exactly

Two Greptile P1 comments (3975151906, 3975151911) both invalid, but
investigation revealed a real bug in the split(':')-based parser:

Bug: usable_glibc_tunable_threshold used split(':') which loses the
distinction between an entry terminated by ':' (glibc skips it) and one
terminated by '\0' with no '=' (glibc rejects the entire string). Examples:
  - "glibc.malloc.mmap_threshold=262144:glibc.cpu.x" (no '=' at end):
    glibc rejects entire string, old code accepted it.
  - "glibc.malloc.mmap_threshold=262144:" (trailing ':'):
    glibc rejects entire string, old code accepted it.

Fix: replaced split(':') with a character-by-character parser matching
glibc's parse_tunables_string exactly. The parser tracks position in the
original string and correctly handles all three terminators ('=', ':', '\0')
for both name and value scanning.

Comment 3975151906 (near-maximum values): Invalid. Verified against
_dl_strtoul: for 18446744073709551615 (u64::MAX), cutoff = u64::MAX/10,
cutlim = u64::MAX%10 = 5. After 19 digits result == cutoff. 20th digit 5:
overflow check (digval > cutlim) is 5 > 5 = false → no overflow. glibc
accepts u64::MAX. Added regression test asserting it's accepted.

Comment 3975151911 (later malformed entry): Invalid. Verified against
parse_tunables (elf/dl-tunables.c): when parse_tunables_string returns -1,
parse_tunables prints a warning and returns immediately without applying
ANY tunable — including ones already parsed into the array. Added
regression test for "threshold=262144:check=2=2" (threshold before
malformed sibling) asserting it's rejected.

Added regression tests: u64::MAX accepted, threshold-before-malformed
rejected, no-'=' at end rejected, trailing ':' rejected, leading ':'
accepted. Verified: clippy clean and tests pass on macOS; glibc-gated
code type-checks and clippy clean for x86_64-unknown-linux-gnu.

* Fix CI: correct hex trailing-garbage test assertion

The test asserted !dl_strtoul_consumes_all("0x20000abc"), but in hex
mode a-f are valid digits — "0x20000abc" is a valid hex number
(0x20000abc = 536874044), not trailing garbage. _dl_strtoul consumes
the entire string and tunable_parse_num accepts it. The assertion
failed on Linux CI where the glibc-gated test actually runs.

Replaced with "0x20000g" — 'g' is not a hex digit, so _dl_strtoul
stops at 'g' and tunable_parse_num rejects the value.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-10 09:29:31 -07:00
e919bec9d1 fix(volume): Harden Volume Copy Validation and Failure Handling (#11252)
* fix(volume): harden volume copy validation

* fix(volume): use stream context for ReadVolumeFileStatus in VolumeCopy

ReadVolumeFileStatus ran on context.Background() while the adjacent
VolumeStatus call used stream.Context(), an inconsistency left over
from the context revert in #11252. Use stream.Context() consistently
so the source status check is cancelled with the VolumeCopy stream.

* fix(volume): reserve destination before deleting existing replica

FindFreeLocation now runs before DeleteVolume so a full target fails
without destroying the existing replica. Previously, when the initial
VolumeStatus check failed (advisory) but ReadVolumeFileStatus
succeeded, the existing replica was deleted before a destination was
reserved, risking data loss if no location had enough free space.

Add a regression test verifying the existing replica survives when
the destination is full and the initial status check fails.

* fix(volume): count replaced replica slot in FindFreeLocation

FindFreeLocation now accepts the volume being replaced so its slot is
treated as available. Without this, a location at its MaxVolumeCount
limit could not replace its sole replica even though deleting it would
free the slot. VolumeCopy passes the volume ID so destination selection
succeeds before the existing replica is deleted.

Add TestVolumeCopyReplacesReplicaAtSlotLimit covering a single-slot
location that must replace its only replica.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 23:33:05 -07:00
2cd6c36c54 filer: end local-only metadata subscriptions when remote peers appear (#11251)
* filer: end local-only metadata subscriptions when remote peers appear

SubscribeMetadata delegates to SubscribeLocalMetadata whenever the
MetaAggregator knows no remote peers at stream setup. Peer discovery is
asynchronous with the gRPC server accepting streams: the master announces
filers after Filer.Init, via ListExistingPeerUpdates and OnPeerUpdate.
A subscriber that connects inside that window is pinned to a filer-local
stream for its whole life, silently missing every other filer's writes.
For filer.remote.sync in a multi-filer cluster this means the remote tier
permanently stops receiving writes served by other filers (#11247).

End the delegated local stream when the first remote peer appears, so the
client reconnects into the aggregated stream. The end surfaces as an
error, not a clean EOF: RetryUntil-driven followers (mount meta cache,
s3api IAM) treat a clean end as following finished and stop
reconnecting. The arrival channel is armed under the same lock as the
peer check in RemotePeerArrivedChan, so a peer learned in between sends
the stream straight to the aggregated path instead of parking on a
channel that would never fire.

A standalone filer is unaffected: no peer ever appears, the channel
never fires, and the local stream serves indefinitely.

Fixes #11247

* filer: interrupt disk replay on peer arrival, trim comments

Check upgradeOnRemotePeer inside eachLogEntryFn and chunkDiskPass so a
peer arriving during a backlog replay stops the stream before the
cursor advances past older remote events. Wrap errAggregationUpgrade
with StopReadingError so LoopProcessLogData does not log it. Remove
issue references from comments and trim verbose commentary.

* filer: check upgrade signal between ref batches

Pass upgradeOnRemotePeer to sendRefsBatched so a peer arriving while
refs are shipped to a slow client is detected between batches, not
only after the full batch completes.

* filer: interrupt gap park on peer arrival

Pass upgradeOnRemotePeer through gapPass to parkOnGap so a peer
arriving during a gap park ends the stream immediately instead of
waiting for the retry timer (up to one minute).

---------

Co-authored-by: Tyagiquamar <Tyagiquamar@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 20:16:50 -07:00
Chris LuandGitHub 7fa2f75f30 s3: bucket-policy Allow must not override an identity explicit Deny (#11256)
* s3: add isActionExplicitlyDeniedByApplicablePolicies helper

Add a helper that reports whether any applicable identity-side policy
(attached IAM policies, enabled-group policies, or the IAM-integration
session policy) explicitly denies an action. It reuses the existing
evaluateAttachedIAMPolicies, resolveS3AuthTarget, buildPrincipalARN, and
isActionExplicitlyDeniedByIAM helpers, and fails closed on evaluation
errors. A nil identity has no identity-side policy plane, so the helper
returns false to keep the bucket policy authoritative for anonymous
access. No behavior change yet; the next commits apply it to the two
bucket-policy Allow short-circuits.

* s3: enforce identity explicit Deny before bucket-policy Allow

authRequestWithAuthType short-circuits on a matching bucket-policy Allow
and skips VerifyActionPermission, so an explicit Deny in an authenticated
identity attached, group, or session policy is bypassed. A non-admin
principal with s3:PutBucketPolicy can install a bucket-policy Allow for
itself and read an object its identity policy explicitly denies.

Before honoring a bucket-policy Allow, check the applicable identity-side
policies for a matching explicit Deny via the new
isActionExplicitlyDeniedByApplicablePolicies helper, and fail closed.
The cross-account behavior is preserved: a bucket Allow still supplies the
Allow an identity policy omits (implicit denial), and a nil identity keeps
the bucket policy authoritative for anonymous access.

Regression tests cover the explicit-Deny override, the implicit-deny
Allow preservation, and the unmatched-key fall-through control.

* s3: enforce identity explicit Deny in secondary object-key auth

authorizeObjectKeyAction authorizes keys the request URL does not name
(CopySource, DeleteObjects body keys, POST Object form keys) and shares
the same bucket-policy Allow short-circuit as the primary path, so an
explicit Deny in the identity, group, or session policy is bypassed the
same way when a bucket policy allows the secondary key.

Apply isActionExplicitlyDeniedByApplicablePolicies before accepting the
bucket-policy Allow, mirroring the primary path. A regression test covers
AuthorizeCopySource for both the explicit-Deny override and the
implicit-deny Allow preservation.
2026-09-09 20:15:43 -07:00
github-actions[bot] 5061a16b12 docs: regenerate star history chart 2026-09-10 02:39:04 +00:00
Chris LuandGitHub 3c9a4bbdda rust: prevent phantom volumes + validate collection hint in mount_volume_by_id (#11254)
* rust: prevent phantom volumes + validate collection hint in mount_volume_by_id

The collection-hint path (and the find_volume_file_base fallback) called
create_volume on any matching .vif/.idx sidecar. create_volume ->
Volume::new -> load(create_dat_if_missing=true) writes an empty .dat and
registers a phantom normal volume, which can shadow a real EC volume whose
.ecx lives on a sibling disk. This reintroduces the phantom-volume bug the
codebase explicitly guards against in load_existing_volumes.

Apply the same guard load_existing_volumes uses to both paths: only mount
when a real .dat is present or the .vif references a remote-tiered file;
otherwise skip the candidate (no phantom). Also reject path-bearing
collection hints ('/', '\\', '..') so the shortcut cannot route .dat
creation outside the storage directory, falling back to the safe scan.

Adds 3 regression tests; all 336 storage:: tests pass.

Addresses Devin + Greptile review comments on PR #11249.

* rust: address review — .note guard, multi-candidate scan, foo..bar hint

Address the four review comments on #11254:

1. Greptile (P1): contains("..") rejected valid collections like "foo..bar".
   Replaced with collection != ".." — volume_file_name joins with "_" so a
   ".." inside a name is part of the filename, not a parent reference. Only
   the exact ".." name is rejected. Added a test that "foo..bar" mounts.

2. Devin #0001 (bug): mount_volume_by_id did not check the .note marker, so
   an interrupted VolumeCopy could mount as a live (truncated) volume. Added
   a .note check before create_volume in both the collection-hint path and
   the fallback — a candidate with .note is skipped (matches
   load_existing_volumes). Added a test covering both paths.

3. Devin #0002 + CodeRabbit (major): find_volume_file_base returned only the
   first matching candidate, so a lone sidecar on disk 0 hid a real .dat on
   disk 1 (the split-disk EC layout the phantom guard protects against).
   Added find_volume_file_bases (plural) that collects all candidates; the
   fallback now iterates every candidate and mounts the first with a real
   .dat or remote .vif. find_volume_file_base delegates to it for
   configure_volume. Added a two-disk test: sidecar on disk 0, real .dat on
   disk 1 — mount succeeds from disk 1.

All 339 storage:: tests pass (6 mount_volume_by_id tests).

* rust: continue past create_volume failure in mount_volume_by_id

Address Devin review comment on #11254: when create_volume fails on an
earlier candidate (e.g. an unreadable .dat), mount_volume_by_id returned
the error immediately instead of trying later candidates. A valid volume
on another disk remained unmounted.

Both the collection-hint loop and the find_volume_file_bases fallback now
remember the last error and continue scanning. A successful mount returns
immediately; if no candidate succeeds, the last error (or NotFound) is
returned. Matches DiskLocation::open_volumes and Go Store.mountVolume.

Added test_mount_volume_by_id_continues_past_open_failure (chmod 000 .dat
on disk 0, real volume on disk 1, mounts from disk 1).

All 340 storage:: tests pass.
2026-09-09 14:54:09 -07:00
Feng ShaoandGitHub 13bf056a15 Mount req with collection (#11249)
* volume mount req support specify collection

* rust mirror change
2026-09-09 12:55:47 -07:00
Nguyễn Đăng Minh LựcandGitHub c968084b34 iceberg: fix OAuth token expiry handling (401 + token-exchange + configurable TTL) (#11242)
* iceberg: return 401 for invalid or expired Bearer tokens

BUG-0001: when the OAuth JWT expired, Server.Auth fell through to the S3
SigV4 authenticator, which rejects the "Authorization: Bearer" scheme
with NotImplemented — a 501. Iceberg clients (Java OAuth2Manager,
pyiceberg) only refresh tokens on 401, so they retried the dead token
forever: RisingWave sinks stalled and Doris catalog queries failed every
token TTL (1h) until the client process was restarted.

A request carrying a Bearer header is an Iceberg REST client: answer 401
(+ WWW-Authenticate: Bearer, RFC 6750) when the token fails, and only
fall through to the S3 authenticator when no Bearer header is present.

* iceberg: make OAuth token TTL configurable via ICEBERG_OAUTH_TOKEN_EXPIRY

BUG-0001 follow-up: production evidence shows Iceberg Java 1.10.x
clients (RisingWave connector node, Doris FE) never re-fetch tokens on
401 — the sink stalled again on token expiry even with the 501→401 fix,
and no POST /v1/oauth/tokens appeared in server logs across dozens of
retries. 401 is necessary but not sufficient for these clients.

The TTL was hardcoded to 3600 with no knob. Read the expiry (seconds)
from ICEBERG_OAUTH_TOKEN_EXPIRY, defaulting to 3600, so deployments can
issue longer-lived tokens (e.g. 86400) to survive client restart cycles.

* iceberg: support OAuth token exchange (RFC 8693) for client refresh

Decompiling the Iceberg Java 1.10.1 client bundled with Doris FE showed
the missing half of BUG-0001: OAuth2Manager refreshes via token-exchange
(AuthConfig.exchangeEnabled defaults to true — the client_credentials
re-fetch branch only runs with exchange disabled), so a server that only
accepts client_credentials leaves Iceberg clients unable to ever refresh
their token, regardless of 401 correctness.

Accept grant_type=urn:ietf:params:oauth:grant-type:token-exchange on
POST /v1/oauth/tokens: verify the subject_token signature against the
issuing credential, allow exchange within a recovery grace window
(max(2*TTL, 1h), capped 24h) so clients holding tokens that expired
while the grant was unsupported recover without a restart, and mint a
fresh access token with the configured TTL.

* iceberg: harden OAuth token exchange and Bearer matching per review

- match the Bearer scheme case-insensitively (RFC 7235), like
  authenticateBearer already does
- accept optional client authentication on the token-exchange grant
  (Basic or form credentials, bound to the subject token's client);
  expired subject tokens now require it. Iceberg Java's proactive
  refresh sends Bearer-only headers, so the grant cannot require it
- reject subject tokens without an exp claim, and re-check the issuer
  on the verified claims
- unauthenticated exchange cannot extend the lifetime past the
  subject token's own expiry (no chain-refresh from a leaked token)
- return 400 invalid_grant per RFC 6749 §5.2 (was 401)
- include issued_token_type on exchange responses (RFC 8693)
- clamp ICEBERG_OAUTH_TOKEN_EXPIRY to 365d so Duration math cannot
  overflow into already-expired tokens

* iceberg: give authenticated token exchanges a fresh full TTL

The remaining-lifetime cap only guards unauthenticated (Bearer-only)
exchanges; an authenticated client renewing a live token must get the
full configured TTL, matching client_credentials.

* iceberg: reject token exchange when no lifetime remains

A Bearer-only exchange with under a second of subject lifetime would
mint a token with expires_in: 0. Reject with invalid_grant instead.

* iceberg: pin near-expiry test token to the next second boundary

jwt/v5 serializes exp at one-second precision, so a 300 ms offset can
round into the current second and route the test through the expired
branch instead of the ttlSeconds<=0 guard. Mint the subject with the
next whole-second expiry: live at exchange time, deterministically
under a second of remaining lifetime.

* iceberg: drop internal ticket reference from comments

* iceberg: clamp oversized OAuth TTLs on 32-bit platforms

strconv.Atoi on an int-sized value fails with ErrRange on 386, so an
oversized ICEBERG_OAUTH_TOKEN_EXPIRY silently fell back to the default
instead of clamping. Parse in 64-bit space and clamp, then narrow.

* iceberg: make OAuth TTL narrowing explicit

* iceberg: disable legacy OAuth in PyIceberg integration tests
2026-09-09 10:54:39 -07:00
516e251f9e rust volume: move the crate to edition 2024 (#11244)
* rust volume: move the crate to edition 2024

Edition 2024 turns three things in this crate into hard errors, and changes
drop order in a further 34 places without changing compilation. The compiler
errors are fixed here; the silent changes were audited against
`RUSTFLAGS='-W rust-2024-compatibility' cargo check --all-targets` output
captured before the flip, since edition 2024 stops reporting them.

`std::env::set_var`/`remove_var` are unsafe as of 2024 because they race with
concurrent readers. All six call sites are safe by construction rather than by
assertion, and the SAFETY comments say why: the build script runs
single-threaded before anything else in the process, and every test reaching
the `config.rs` helpers holds `process_state_lock()` for the duration.

The two `ref` bindings in handlers.rs sit in patterns that already borrow
implicitly, so removing the modifier leaves both bindings at `&String`.

On the 34 drop-order sites: no lock guard's scope is extended anywhere, and
`volume.rs` has none. Most are moved-from `Option`/`Result` husks — `if let
Some(v) = map.remove(&k)`, `while let Some(m) = stream.next().await` — where
the value is moved into the binding and the temporary has nothing left to drop;
where closing order actually matters these paths already call `v.close()`,
`ec_vol.destroy()` or `drop(writer)` explicitly. Two sites get strictly better
ordering: the metrics read guard in `run_metrics_push_loop` shrinks to the end
of its initializer block (it never crossed an `.await` either way), and an EC
test now closes the volume's descriptors before the `TempDir` removes the
directory.

No `rust-version` is declared. Edition 2024 needs rustc 1.85, but that is not
the binding constraint — the dependency tree already requires 1.91.1 through
the `aws-sdk-s3`/`aws-smithy-*` family, so `cargo +1.85 check` fails on the
deps regardless. CI builds on `dtolnay/rust-toolchain@stable`.

`vendor/reed-solomon-erasure` is a separate package and keeps edition 2021.
Cargo.lock is unchanged despite edition 2024 implying resolver 3.

Verified: `cargo test` 551 passed / 0 failed, `cargo test
--no-default-features` 550 passed / 0 failed (the two feature sets produce an
identical migration site list), `cargo build --release` clean. No automated
test covers shutdown ordering, so the channel and runtime sites in `main.rs`,
`write_queue.rs` and `grpc_server.rs` were read individually.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust volume: address edition-2024 review feedback

Three fixes from review of the edition bump.

Serialize the two environment-reading tests. The SAFETY comments on the
`env::set_var`/`remove_var` helpers claim every test touching the environment
holds `process_state_lock()`, but `test_resolve_config_defaults_dir_to_platform_temp_dir`
and `test_resolve_config_index_accepts_redb_and_leveldb_aliases` called
`resolve_config` — which reads HOME/USERPROFILE, SEAWEED_WRITE_QUEUE and the
WEED_* set — without taking it. `set_var` is unsafe precisely because a
concurrent *reader* is UB, not only a concurrent writer, so the comment was
overclaiming. An audit of the module found exactly these two; every other
environment-touching test already held the lock. The race predates edition
2024, which only made the requirement explicit.

Declare `rust-version = "1.91.1"`. The edition needs 1.85, but that was never
the binding constraint: `cargo +1.90 check --all-targets` fails on the
`aws-sdk-s3`/`aws-smithy-*` family, and 1.91.1 checks clean. Declaring the
verified floor turns a wall of per-dependency errors into one clear message.
Cargo.lock is unchanged despite this making the resolver MSRV-aware.

Update the README, which advertised "Rust 1.75+ (2021 edition)". 1.75 was
already stale before this branch — the tree has needed 1.91 for a while.

Verified: `cargo test` 551 passed / 0 failed, `cargo test --no-default-features`
550 passed / 0 failed, `cargo build --release` clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust volume: state the exact MSRV patch release in the README

The README said "Rust 1.91+", which reads as 1.91.0 and is wrong by one patch
release: `cargo +1.91.0 check --all-targets` fails on the aws-sdk-s3 family,
`cargo +1.91.1` passes. Say 1.91.1+, matching `rust-version` in Cargo.toml, and
call out that the patch component is load-bearing so nobody installs 1.91.0 and
hits the same wall.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nty5Rj7ssMQdFxHHjZgDC

* rust worker: move the workspace to edition 2024

Moves the seaweed-worker workspace (core, lance, sort) from edition 2021 to
2024, the same migration seaweed-volume just got in this branch.

Edition 2024 turns exactly one thing in this workspace into a hard error. The
baseline came from RUSTFLAGS='-W rust-2024-compatibility' cargo check
--all-targets, run before the flip; unlike seaweed-volume's 34 silent +
8 hard sites, the worker reports only the one hard site and no
tail_expr_drop_order or if_let_rescope sites at all. The worker is a much
smaller crate and none of its expressions hold a guard or temporary whose drop
order the edition changes, so there is nothing to audit on the silent side.

Fixed (1 site):

std::env::set_var is unsafe as of 2024 because it races with concurrent
readers. The single call is in crates/core/build.rs, which sets PROTOC from
protoc_bin_vendored the way seaweed-volume's build script does. A build
script's main runs single-threaded before anything else in the process, so
no other thread can be reading the environment concurrently; the SAFETY comment
says so. There are no config.rs-style test helpers here -- the worker's tests
do not mutate the environment -- so unlike the volume crate there are no
process_state_lock() callers to audit.

No redundant ref bindings to clean up: a grep for ref across the three
crates finds none.

MSRV:

rust-version = "1.94.1", verified rather than inferred. Edition 2024 only
needs 1.85, but the dependency tree needs more: lance's aws feature pulls in
a newer cut of the same aws-sdk-*/aws-smithy-* family that sets
seaweed-volume's 1.91.1 floor, and that newer cut requires 1.94.1.
cargo +1.94.0 check --all-targets fails on that family; cargo +1.94.1
check --all-targets is clean. The worker's floor is therefore higher than
the volume's, and moves with lance and the AWS SDK rather than with the
edition. CI builds on dtolnay/rust-toolchain@stable, so nothing changes
there.

The edition is set once in [workspace.package] and inherited by each member
via edition.workspace = true; rust-version is added the same way. The
workspace keeps its explicit resolver = "2" -- edition 2024 would default to
resolver 3, but the pin is deliberate and Cargo.lock is unchanged by this
commit either way.

The README gains a "Requires Rust 1.94.1+ (2024 edition)" line in its Building
section, matching the one seaweed-volume's README now carries, and calling out
that the patch release is load-bearing (1.94.0 does not build) so nobody
installs 1.94.0 and hits the same wall.

Verification:

* cargo check --all-targets -- clean, zero warnings (default toolchain 1.97)
* cargo +1.94.1 check --all-targets -- clean
* cargo +1.94.0 check --all-targets -- fails on the AWS SDK, as claimed
* cargo test --all-targets -- 40 passed, 0 failed
  (core 13, sort 11, lance lib 3, lance bin 2, compaction 6, lifecycle 1,
  sort integration 4)
* Cargo.lock unchanged

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-09 10:54:06 -07:00
Chris LuandGitHub 01fc31cb71 fix(filer): use path.Split instead of filepath.Split for filer paths (#11246)
FullPath.DirAndName() and FullPath.Name() used filepath.Split, which is
OS-dependent: on Windows it treats backslash as a path separator,
corrupting filer paths that contain literal backslashes. Filer paths
always use "/" as the separator, so switch to path.Split and path.Join
which only split on "/" regardless of the host OS.

This fixes the backslash case from #11243 where a file saved as
/test/special\reverseslash4.jpg was stored with a corrupted path on
Windows filer builds. The #, ?, and % cases from the same issue are
client-side URL-encoding problems (the server never receives the raw
characters), but once the client properly percent-encodes them the
server now handles the decoded path correctly on all platforms.
2026-09-09 10:43:16 -07:00
ssshr-66andGitHub 966692fa23 [Volume] Validate record counts after volume copy (#11238)
* validate Volume Copy record counts

* Delete s3api_object_versioning_bench_test.go

* reply ai comments
2026-09-09 02:18:18 -07:00
github-actions[bot] 168b9c39f8 docs: regenerate star history chart 2026-09-09 05:10:20 +00:00
Chris LuandGitHub f1f6886d0e fix: rebase on latest master before pushing star history chart (#11240)
* fix: rebase on latest master before pushing star history chart

The daily star history workflow git push was rejected with a
non-fast-forward error because new commits landed on master between
the checkout and the push. Fetch full history (fetch-depth: 0) and
rebase the generated commit on top of the latest remote branch before
pushing so the workflow no longer fails when master has moved.

* fix: retry rebase-and-push to handle concurrent master updates

Address review feedback: a one-shot rebase still races if master advances
between the rebase and the push. Match the bounded retry loop used by
java_release.yml — push first, and on rejection rebase and retry up to
five times before failing.

* fix: serialize runs and ensure every rebase is followed by a push

Address review feedback:
- Devin (line 54-55): the old loop rebased after the 5th failed push but
  never pushed the rebased commit. Restructure so each rebase (attempts
  2-5) is followed by a push attempt, with a clear 5-attempt cap.
- Greptile: overlapping runs could conflict on the SVG during rebase.
  Add a concurrency group (cancel-in-progress: true, matching the repo
  convention) so only one chart regeneration runs at a time.

* fix: scope concurrency by ref and guard rebase against transient failures

Address review feedback:
- Devin (line 16): the global star-history concurrency group let a manual
  run on another branch cancel an in-flight daily master update. Scope
  the group by github.ref so only same-branch runs cancel each other.
- Greptile (line 58): git pull --rebase runs under the fail-fast shell,
  so a transient fetch error or conflict aborted the whole step before
  remaining attempts ran. Guard the rebase so a failure aborts the
  in-progress rebase and continues to the next attempt instead.
2026-09-08 21:43:24 -07:00
Chris LuandGitHub 2ffa696809 fix(volume): handle faulty storage media (Go + Rust) (#11233)
* fix(volume): track EC shard read errors and unmount on faulty media

Extract the volume EIO tracker into a reusable IoErrorTracker and add the
same tracking to EcVolume. Sustained EIO on .ecx lookups or .ecd shard
reads now unmounts the EC volume in the heartbeat (without deleting
files) so the master re-replicates from healthy peers, mirroring the
existing volume replica quarantine.

Closes #11227 (EC shard unmount).

* rust(volume): mirror EC shard read error tracking and unmount

Add EIO tracking to the Rust EcVolume mirroring Go: a streak counter
with IO_ERROR_TOLERANCE, a sticky quarantine flag, and unmount (not
file deletion) in the heartbeat so the master re-replicates from
healthy peers.

* feat(metrics): expose storage IO error counter and quarantine gauge

Add a storage_io_error_total counter incremented on every EIO recorded
by the volume or EC shard tracker, and an io_quarantine gauge labelled
by kind (volume/ec_shard) reflecting the count of replicas suppressed
in the heartbeat. Mirrored in Go and Rust.

* feat(healthz): report 503 when local replicas are IO-quarantined

Add Store.HasIoQuarantine (Go) / Store::has_io_quarantine (Rust) and
have /healthz return 503 when any local volume or EC shard is
quarantined due to sustained storage-media EIO, so a load balancer
can drain a server whose underlying media is faulty. Mirrored in Go
and Rust.

* fix(volume): keep quarantined EC volumes in memory and reset EIO on success

Address review feedback: instead of unloading quarantined EC volumes
(which discards the quarantine state healthz needs), keep them in
memory and just skip them from heartbeat reporting, mirroring the
regular volume quarantine. Also clear the EIO streak on successful
.ecx reads in Rust so a transient error does not accumulate, and add
an ec_shard label to the io_quarantine gauge in both Go and Rust.

* fix(volume): exclude quarantined EC shards from heartbeat and add Rust volume tolerance

Address review feedback:
- Filter quarantined EC volumes from CollectErasureCodingHeartbeat
  (Go) and collect_ec_shard_delta_messages / collect_live_ec_shards
  (Rust) so the master stops advertising faulty shards and
  re-replicates from healthy peers.
- Add consecutive EIO count and sticky quarantine to the Rust
  regular Volume, mirroring Go IoErrorTracker: a single EIO no
  longer deletes the replica; the heartbeat quarantines after the
  tolerance threshold and keeps the volume in memory.
- Use the quarantine flag (not last_io_error) in has_io_quarantine
  so /healthz reflects sustained, not transient, failures.

* fix(volume): make Rust quarantined volumes read-only and wire recovery

Address Devin review:
- Set no_write_or_delete on Rust volumes when quarantined in the
  heartbeat, so cached or direct clients cannot mutate a faulty
  replica after the master removes it (mirrors Go).
- Wire reset_io_error_state into Volume::set_writable so an operator
  making a volume writable again clears the sticky quarantine and
  the volume re-enters heartbeat rotation.

* fix(volume): clear EC quarantine on shard re-mount for operator recovery

Address Greptile review: re-mounting EC shards (Go loadEcShardWithIdxDir
/ Rust mount_ec_shards_with_idx_dir) now calls ResetIoErrorState on the
existing EcVolume, giving operators a documented recovery path that
clears the sticky quarantine and returns the EC volume to heartbeat
rotation. Mirrored in Go and Rust.

* fix(volume): do not clear EC quarantine on routine shard mounts

Address review feedback: clearing the EC IO quarantine on every mount
(including duplicate, retry, sibling-shard, and reconciliation mounts)
is too aggressive and can re-advertise known-bad shards before the
storage media has been validated. Remove the automatic reset from the
mount path; quarantine clears naturally on restart or full unmount
when a fresh EcVolume is created with clean state.

* test(volume): update Rust IO error test for quarantine semantics

The heartbeat now quarantines a volume with sustained EIO (keeps it
mounted, makes it read-only, omits it from heartbeat) instead of
deleting it. Update test_collect_heartbeat_deletes_io_error_volume to
assert the volume stays in the store with no_write_or_delete set, and
update set_last_io_error_for_test to set the consecutive error count
at the tolerance threshold so the test reflects a sustained error.

* fix(volume): reset EIO streak after full write and match Windows media errors

Move the success-side EIO reset from append_needle (after write_all only)
to the end of do_write_request, after flush_dat/flush_idx complete, so a
successful write_all followed by a failed fsync no longer resets the
counter before the EIO is recorded. Repeated fsync EIOs now accumulate
toward the quarantine threshold as intended.

Recognize Windows storage-media failure codes ERROR_CRC (23) and
ERROR_IO_DEVICE (1117) in addition to Unix EIO (errno 5), so quarantined
heartbeat behavior is preserved on Windows. Mirrors the change in both
Go and Rust volume servers.

* fix(volume): preserve checkpoint EIO and clear streak on successful delete

maybe_checkpoint_index now returns whether the checkpoint succeeded;
the success-side EIO reset in do_write_request and do_delete_request
only fires when it did, so a checkpoint media failure is no longer
erased by the unconditional reset that followed it. do_delete_request
also gains the success reset that was lost when append_needle stopped
clearing the streak, so a successful delete still clears an earlier
failure streak.

is_storage_io_error now uses libc::EIO on Unix instead of a hard-coded
5, and the ECX binary-search read path gains a Windows fallback
(seek + read_exact) so the buffer is no longer zeroed on non-Unix
targets.
2026-09-08 21:42:56 -07:00
Chris LuandGitHub 0ce5ca42ea helm: supply admin auth in CI renders that enable admin (#11239)
PR #11236 added a render-time guard that fails the chart when admin.ip
is non-loopback (default 0.0.0.0) and admin auth is not configured,
since weed admin 4.46 refuses to bind a non-loopback address without
authentication. Several pre-existing helm_ci.yml test cases enable
admin.enabled=true as part of "everything on" renders without a
password, so helm template now exits non-zero and the Verify template
rendering step fails.

Add admin.secret.adminPassword to the four render calls that turn on
admin without auth (IAM gRPC opt-in, NetworkPolicy EVERYTHING, egress
without kubeApiServer.cidrs, and the license ALL_ON dict), using the
same key ci/admin-values.yaml already uses.
2026-09-08 20:51:53 -07:00
9b12d13934 volume server: release the store lock before scrubbing EC volumes (#11235)
* volume server: release the store lock before scrubbing EC volumes

`ec.scrub` makes a Rust volume server stop serving for the duration of the
scrub, and then kills its own gRPC connection:

    error: rpc error: code = Unavailable desc = keepalive ping failed to
    receive ACK within timeout

Measured on a 4.46 cluster (17 Rust volume servers on one host, ~520 volumes
and 53 EC volumes, --index=redb, EC 10+4). It reproduces against a SINGLE
node in 30-70s, in checksum, index and local modes, at -maxParallelization 1.

## Cause

The CHECKSUM arm of scrub_ec_volume reads every byte of every local shard
while holding the caller's store.read() guard:

    let store = self.state.store.read().unwrap();
    let ecv = store.find_ec_volume(vid)...?;
    let (blocks, broken, errs) = ecv.checksum_scrub();   // GBs of I/O, lock held

VolumeServerState::store is a std::sync::RwLock, which is write-preferring.
The periodic heartbeat's collect_heartbeat_with_snapshot takes store.write()
and blocks; once that writer is pending, every later store.read() queues
behind it. Every HTTP handler takes store.read(), so the node serves nothing,
stops heart-beating, and cannot answer the scrub RPC's own keepalive - the
scrub kills the connection it is running on.

The INDEX and LOCAL arms have the same shape, and the node-wide scrub_volume
loop is worse: it held ONE guard across every volume on the node.

## Evidence

offcputime, off-CPU stacks >1s in a 30s window during a scrub:

    futex_wait
      seaweed_volume::server::heartbeat::collect_heartbeat_with_snapshot
      - tokio-rt-worker
        27967020        <- 27.97s blocked, of a 30s window

A single HTTP /status request issued 12s into a scrub, with 180s of patience,
was accepted and queued for 120 seconds, then served once the scrub released.
Thread states throughout: 1 D + 48 S. One thread working, 48 idle - not
executor starvation and no thread pileup, which is what a single lock holder
looks like.

Memory was tested and ruled out as the cause: the same scrub was run at
MemoryMax 3G, 8G and unlimited. With no limit there is no reclaim at all,
page cache grows freely to 22 GB, and the node still goes unresponsive at
t+30s. anon stays flat at 48-86 MB in every run.

## Fix

checksum_scrub, scrub_index and scrub_local gain plan types -
EcChecksumScrubPlan, EcIndexScrubPlan and EcLocalScrubPlan - snapshotted from
the volume under a brief guard. The handler builds a plan, drops the guard,
and runs the scan in spawn_blocking, off the async workers, since it is
synchronous CPU + file I/O either way.

A plan captures DESCRIPTORS, not paths. Resolving a path again after the
guard is dropped would let a writer that legitimately unlinks the files - the
heartbeat's delete_expired_ec_volumes, which reaches EcVolume::destroy(), or
volume_ec_shards_delete - surface an intentional removal as "scrub read
error: No such file or directory" and put the volume in broken_volume_ids. A
descriptor outlives the name.

For the shards it duplicates the handle the mounted EcVolumeShard already
holds (try_clone_file), which is what Go does: ChecksumScrub reads through
shard.ReadAt (weed/storage/erasure_coding/ec_volume_scrub.go:71), never
through a path. That also inherits open_volume_file's O_NOATIME and drops a
dead branch - the old code built {base}.ec{id}.v{gen} for a non-zero
generation, a name nothing in this tree writes. dup shares the kernel offset,
so shard reads stay positional; the .ecx gets a fresh open instead, since
check_index_file seeks.

FULL/READS is unchanged here: it already released the guard across the index
walk, and still re-takes it per needle in store_ec::scrub_snapshot_under_lock
for that needle's local shard intervals - short holds, many of them.

scrub_volume now takes the read guard PER VOLUME instead of across the whole
loop, so the heartbeat can land between volumes. Its per-volume work still
runs under the guard; Volume needs an equivalent plan to fix that properly,
left as a follow-up and noted in the code.

## A failed scrub task must not take the whole RPC down

Moving the scans into spawn_blocking changed where a panic lands. It no
longer unwinds inside the handler's own future; it comes back as a JoinError
at the .await, and all four join points sat behind a `?`. So one bad volume
out of six hundred returned Err from the entire handler: the
broken_volume_ids, broken_shard_infos and details already gathered for the
other 599 were dropped, and emit_scrub_metrics - the only writer of
SCRUB_LAST_TIME_SECONDS, SCRUB_VOLUME_FAILURES and SCRUB_SHARD_FAILURES - was
never reached, so the staleness alert kept firing while real corruption went
unreported.

And there is a reachable panic behind it. EcLocalScrubPlan::run() sized its
reassembly buffer with

    Vec::with_capacity(get_actual_size(size, version) as usize)

which for any negative size that is not the -1 tombstone skipped above is a
capacity-overflow abort. Mode 3 (LOCAL) is the default of `weed shell
ec.scrub`, and a scrub is what you point at an index you already suspect, so
an arbitrary i32 in a .ecx size field is in-scope input. The buffer is
Rust-only - Go appends to a nil slice and has no capacity hint here. Guard on
`want <= 0` and fall through with an empty buffer: locate_data returns no
intervals for a non-positive size, read stays 0, and the existing
`read != want` error reports the row exactly as Go does.

Each join point now records the failure against its own volume and continues.
A panic is evidence about the volume and counts as broken; a non-panic
JoinError is not - spawn_blocking only reports one when the runtime is going
down, the volume was never scanned, and counting it would put a false
corruption into SCRUB_VOLUME_FAILURES. total_volumes moves before the join in
modes 1, 3 and 4 (2|5 already counted there) so a failed join cannot silently
shrink it. Mode 2|5's verify_ec_shards join is the one that must not
`continue`: the needle walk above has already produced findings for that
volume.

The tombstone guard stays is_tombstone() on purpose. ScrubLocal in
ec_volume_scrub.go:228 skips only IsTombstone(), while the distributed walk
in store_ec.go:516 skips all IsDeleted() - the asymmetry is Go's, and both
Rust walks mirror their own counterpart.

## Both servers: a node-wide scrub skips a volume that vanished mid-run

Releasing the lock makes the volume set legitimately mutable during a scrub,
so a node-wide run can reach a volume that has since been unmounted. That is
not a scrub failure. A node-wide run now logs and skips it; an explicitly
requested volume id still returns NotFound. The Go server is changed the same
way, so both implementations answer the same shell command identically.
mark_broken_volumes_readonly tolerates the same teardown one step later,
instead of throwing away the whole scrub report.

## Test

test_scrub_plans_are_self_contained_and_match_direct_call drops the EcVolume
and runs both plans on another thread, asserting the results match the direct
calls. A plan that borrowed from EcVolume could do neither, so the test stops
compiling if the snapshot regresses to a borrow.

test_scrub_plans_survive_files_removed_after_snapshot unlinks every shard and
the .ecx after the plans are built, then asserts the results still equal the
direct call. Against a path-resolving version it fails with all 14 shards
reported as "No such file or directory".

test_local_scrub_plan_reports_negative_size_ecx_row rewrites a .ecx row's
size to -1000 and runs the local plan on another thread, so the join is the
assertion - that thread is the spawn_blocking whose panic used to fail the
RPC. Without the capacity guard it fails with "capacity overflow"; with it,
the row is reported.

The Go tests cover both halves of the vanished-volume rule for volumes and EC
volumes.

517 lib tests pass, plus 34 across the other targets (`cargo test`).
`go test ./weed/server -run Scrub` passes.

## Known remaining, not fixed here

`ec.scrub -volumeId=N` is still fanned out to every node, and a node that
holds no shard of N returns NotFound, so the shell command errors even when
the nodes that do hold shards scrub cleanly. That is a shell-side fan-out
question rather than a volume-server one, and both servers keep the existing
behaviour for an explicitly requested id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DvHoW85w6SNKNBPvrqLMmK

* scrub: discard checksum block count from total_files; capture .ecx fd for FULL walk

Two review fixes:

1. CHECKSUM arm: plan.run() returns blocks scanned, not a file count.
   Go discards it (_, shardInfos, serrs = v.ChecksumScrub()) so TotalFiles
   stays a needle/file count. The Rust arm was adding it to total_files,
   inflating the count. Discard it to match Go.

2. FULL/READS (scrub_ec_volume_distributed): the needle walk reopened the
   .ecx by PATH after the store guard was released, so a concurrent teardown
   that unlinks or replaces the .ecx (heartbeat delete_expired_ec_volumes,
   volume_ec_shards_delete) could surface an intentional removal as a scrub
   error or mix index generations within one scrub. Capture a second .ecx
   descriptor under the guard (the index plan handle is consumed by its own
   structural walk, and both seek) and read through it instead -- the same
   descriptor-outlives-name invariant the checksum plan shard handles use.

* scrub: bind FULL/READS walk to one encode generation

Address Devin review: after capturing the .ecx descriptor under the guard,
scrub_snapshot_under_lock still re-resolves the volume by id per needle, so
a teardown-and-remount of the same vid between two rows would apply the
captured .ecx offsets to a replacement volume's shards -- falsely reporting
corruption.

Capture the volume's encode_ts_ns (encode-run identity) in Phase A and pass
it to scrub_snapshot_under_lock. If the mounted volume's encode_ts_ns no
longer matches, abort the walk like a mid-scan unmount instead of mixing
generations within one scrub.

* scrub: run FULL/READS index scan in the blocking pool

Address CodeRabbit review (5147767192): index_plan.run() reads the whole
.ecx synchronously, so running it on the async executor worker could block
unrelated RPC work handled on the same executor. Move it into spawn_blocking,
matching the treatment the CHECKSUM/LOCAL arms already give their plans. A
join failure (panic/cancellation) is reported as a seed error so the
per-volume findings below are not silently dropped.

* scrub: move ecx walk to blocking pool, classify join errors, guard encode_ts_ns==0

Three CodeRabbit review fixes (5148034447):

1. Move the FULL/READS needle walk (walk_index_file over the captured ecx
   descriptor) into spawn_blocking. It reads the full .ecx synchronously and
   was still running on the async executor worker, the same blocker the
   index_plan.run() fix in the previous commit addressed.

2. Preserve JoinError classification in both spawn_blocking join points in
   scrub_ec_volume_distributed. A panic is evidence about the volume and
   counts as broken; a cancellation only happens at runtime shutdown, the
   volume was never scanned, and returning it as an error would put a false
   corruption into broken_volume_ids (the FULL/READS arm marks the volume
   broken on any non-empty errs). Panics return an error; cancellations
   return clean.

3. Do not treat encode_ts_ns == 0 as a verified generation match. The .vif
   assigns 0 when it carries no encode-run identity (legacy/pre-feature
   volumes), so 0 == 0 would accept a teardown-and-remount and apply the old
   .ecx offsets to the replacement volume's shards. Only enforce the
   generation check when the captured identity is non-zero; when it is zero,
   fall back to the pre-check behavior (no generation binding) rather than
   aborting a scrub that was already running without the guard.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-08 19:12:30 -07:00
Chris LuandGitHub 723f473f02 filer: widen metadata-subscription readahead buffers (#11237)
The metadata-subscription readahead channels were sized for a
low-throughput era and now bottleneck replay catch-up:

- ReadPersistedLogBuffer's readaheadSize was 1024 entries: the
  background visitor fills the channel, then blocks on the consumer's
  gRPC Send, so volume-server I/O for the next log file never overlaps
  with delivery of the current one. Each disk pass takes longer, and
  the subscribe loop re-lists log files (ListDirectoryEntries on the
  filer store) more often to drain the same backlog. Raised to 8192
  so the reader stays ahead of the consumer through a full log file's
  worth of entries.

- readFilersMerged's logEntryChannelSize was 512 entries per filer
  stream: the same serialization on the client side, where weed mount
  (chunk mode) reads persisted log chunks directly from volume
  servers. A small channel means the producer stalls on the merge
  consumer's processEventFn, and the next log file's chunks are never
  fetched ahead. Raised to 4096 so volume I/O overlaps with event
  delivery.

The wider buffers keep the producer goroutines reading through a full
log file while the consumer is still processing the previous one,
turning serial read→process→read into pipelined read∥process. This
cuts the per-pass wall time that drives filer store listings and
volume-server round-trips, reducing filer workload under backlog
catch-up (e.g. CSI deployments where ~200 mounts reconnect on filer
restart).
2026-09-08 16:58:09 -07:00
5f77a0b67e admin: allow insecurely binding to any IP if -allowInsecureNoAuth is set (#11228)
* admin: allow insecurely binding to any IP if -allowInsecureNoAuth is set

* admin: rename -allowInsecureNoAuth to -allowInsecureBind

The new flag name is shorter and clearer: it describes what is being
allowed (an insecure bind to a non-loopback address) without the
redundant "NoAuth" suffix.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-09-08 16:57:44 -07:00
Chris LuandGitHub fd4fa72289 helm: pass -ip to admin so StatefulSet becomes Ready on 4.46 (#11236)
* helm: pass -ip to admin so StatefulSet becomes Ready on 4.46

Since weed admin 4.46 changed its default listen address from all
interfaces to loopback (127.0.0.1), the admin StatefulSet template never
passed -ip, so the admin server bound to loopback only. The chart's httpGet
readiness/liveness probes dial the pod IP, not loopback, so the probes
never succeeded and the admin StatefulSet stayed 0/1 forever — breaking
upgrades with helm --wait or GitOps controllers.

Add an admin.ip value (default "0.0.0.0", restoring the pre-4.46
behaviour) and render it as -ip. A non-loopback bind requires
authentication, so fail at render time when admin.ip is non-loopback and
neither admin.secret.adminPassword nor admin.secret.existingSecret is
set, instead of letting the pod crash-loop. Document the value and add a
chart-testing CI values file.

Bumps chart to 4.46.1.

Fixes #11234

* helm: address review feedback on admin bind validation

Align the chart's loopback classification with weed admin's isLoopbackIp
(net.ParseIP + IsLoopback): the whole 127.0.0.0/8 range and ::1 are
loopback; localhost and wildcard addresses are non-loopback, matching
the binary. Previously the exact-string check rejected valid loopback
addresses like 127.0.0.2 while permitting localhost (which the binary
treats as non-loopback).

Recognize WEED_ADMIN_PASSWORD supplied via admin.extraEnvironmentVars /
admin.secretExtraEnvironmentVars as authentication, since weed admin
picks it up through viper's AutomaticEnv. Previously such deployments
were wrongly rejected at render time.

Remove [https.admin] mTLS from the validation message and docs: the
chart only generates [grpc.admin] (gRPC mTLS), not [https.admin] (HTTP
mTLS), so mentioning it as an alternative was misleading.

Document that the -ip flag requires SeaweedFS 4.46 or newer, so pinning
admin.imageOverride to an older image is not supported with this chart.

Extracted the loopback and auth checks into reusable helpers
(seaweedfs.admin.isLoopbackIp, seaweedfs.admin.authEnabled) following
the existing seaweedfs.filer.mysqlEnabled pattern.

* helm: tighten loopback classification to reject malformed 127.x addresses

Use regexMatch instead of hasPrefix for the IPv4 loopback check so
malformed values like "127.not-an-ip" are not accepted as loopback
(net.ParseIP returns nil for them, so weed admin treats them as
non-loopback). Also recognize the expanded IPv6 loopback form
"0:0:0:0:0:0:0:1" in addition to "::1", matching net.ParseIP behavior
for the two common representations.
2026-09-08 16:18:27 -07:00
Chris LuandGitHub cb9fcd39d2 filer/postgres: create filemeta table on startup via createTable config (#11229)
* filer/postgres: create default filemeta table on startup

The postgres filer store hardcoded CreateTableSqlTemplate to empty and
never created the filemeta table, unlike postgres2/mysql2/sqlite which
all create it during Initialize. Users had to create the table manually
or the filer would crash loop with "relation filemeta does not exist".

Read the createTable config option (same as postgres2), default to
DefaultCreateTableQuery when unset, and execute CREATE TABLE IF NOT EXISTS
on the default table after the connection pool is established.

SupportBucketTable stays false so per-bucket table creation remains a
no-op; only the shared filemeta table is created, via a direct ExecContext
since AbstractSqlStore.CreateTable short-circuits without bucket support.

* filer/postgres: accept boolean createTable = true/false

viper reads a TOML boolean as the string "true"/"false" via GetString, so
createTable = true was being used as a SQL template and failed. Add
ResolveCreateTableQuery to normalize the value: true and empty select the
default template, false disables table creation, anything else is a custom
template. Both postgres and postgres2 now use it, and both skip the CREATE
TABLE call when the resolved template is empty.

* scaffold: document createTable option for postgres filer store

Replace the commented-out CREATE TABLE SQL in the [postgres] scaffold with
a createTable config hint, matching the [postgres2] section. Users no longer
need to manually create the filemeta table before starting the filer.

* filer/postgres: make createTable opt-in for postgres, keep postgres2 default

The previous commit defaulted postgres to create the filemeta table even
when createTable was unset, which could break existing deployments whose
DB user lacks CREATE TABLE privileges. ResolveCreateTableQuery now returns
empty for an unset value so postgres only creates the table when
createTable is explicitly true or a custom template — preserving the
prior no-DDL behaviour for existing configurations.

postgres2 keeps its existing always-create default: it defaults an empty
resolved value to DefaultCreateTableQuery, and only skips when
createTable is explicitly false.

* filer/postgres2: simplify createTable handling, document all modes

Drop the false opt-out from postgres2 — it only skipped the default table
while per-bucket CreateTable still ran, leaving restricted DB roles broken
on bucket access. postgres2 now accepts true the same way (defaulting to
DefaultCreateTableQuery) and keeps its existing always-create behaviour
for every other value, matching the original semantics.

The scaffold comment now documents true/false/custom for the postgres
section so users know false (or unset) is the backward-compatible default.

* filer/postgres2: normalize false via ResolveCreateTableQuery

postgres2 only handled "" and "true", leaving createTable = false as the
literal string "false" which CreateTable then executed as invalid SQL.
Route it through ResolveCreateTableQuery (which maps false to empty) and
default the empty result to DefaultCreateTableQuery, so false is treated
the same as unset for the bucket-aware store.

* filer/postgres2: honor createTable = false for default table

postgres2 treated false the same as unset and always created the default
filemeta table, failing startup for restricted DB roles that explicitly
opted out. Track the original false value before ResolveCreateTableQuery
collapses it to empty, and skip the default CreateTable call when set.

Per-bucket table creation is unaffected — it is a runtime requirement of
the bucket-aware store. Users who need to suppress all DDL should use the
postgres (non-bucket) store with createTable unset.

* filer/postgres2: disable bucket tables when createTable = false

Setting SupportBucketTable = false when createTable is explicitly false
makes AbstractSqlStore.CreateTable a no-op (it already returns nil when
SupportBucketTable is false), so neither the default filemeta table nor
per-bucket tables are created. The template stays empty and no DDL runs,
honouring the opt-out for restricted DB roles. All data routes to the
pre-provisioned filemeta table, matching the postgres (non-bucket) store.

* filer: suppress DDL without disabling bucket routing

Setting SupportBucketTable = false when createTable = false also disabled
per-bucket routing, hiding objects in pre-provisioned per-bucket tables.
Keep SupportBucketTable true and instead skip the CREATE TABLE execution
when the resolved template is empty.

GetSqlCreateTable now returns empty for both postgres and mysql SQL
generators when CreateTableSqlTemplate is empty, and
AbstractSqlStore.CreateTable skips the ExecContext call when the SQL is
empty. This preserves bucket routing while suppressing all DDL for users
who explicitly set createTable = false and pre-provision their tables.

* filer: add SkipDDL to suppress CREATE and DROP without disabling routing

createTable = false with SupportBucketTable = true preserved bucket
routing but deleteTable still executed DROP TABLE on bucket deletion,
dropping externally managed tables. CanDropWholeBucket also returned
true, so the S3 layer tried whole-table drops instead of row-by-row
deletes.

Add a SkipDDL flag to AbstractSqlStore, independent of SupportBucketTable.
CreateTable and deleteTable both skip when SkipDDL is set, and
CanDropWholeBucket returns false so bucket deletion falls back to
row-by-row metadata deletes. postgres2 sets SkipDDL when createTable is
explicitly false — bucket routing is preserved, no DDL runs.

* filer: fall back to row-by-row delete when CanDropWholeBucket is false

DeleteFolderChildren took the whole-table drop path whenever the path was
a bucket root, even when SkipDDL made deleteTable a no-op. The no-op
returned nil, the caller returned early, and rows inserted after the
recursive enumeration survived the bucket deletion.

Gate the whole-table drop on CanDropWholeBucket so the row-by-row
DeleteFolderChildren SQL runs when SkipDDL is set, removing all metadata
without issuing DROP TABLE.
2026-09-08 16:17:02 -07:00
Chris LuandGitHub 8782749f26 admin: reject IAM policy deletion while still attached to a user/group (#11230)
* admin: add IsPolicyAttached helper to detect user/group attachments

Introduces AdminServer.IsPolicyAttached, which lists the users and groups
that still have a managed policy attached, reusing the existing credential
manager ListUsers / ListAttachedUserPolicies / ListGroups / GetGroup
methods. This is the building block for rejecting policy deletion while a
policy is still referenced, so deleted policy names stop lingering in a
user attached policy names list (issue #11225).

* admin: reject IAM policy deletion while still attached

Guards AdminServer.DeletePolicy with the new IsPolicyAttached check and
returns the typed ErrPolicyStillAttached error when the policy is still
referenced by a user or group. This matches AWS IAM and the existing IAM
API handler behavior, fixing the stale reference where a deleted policy
name kept showing up in a user attached policy names list (issue #11225).

* admin: return 409 Conflict when deleting an attached IAM policy

The admin UI DeletePolicy handler now maps ErrPolicyStillAttached to HTTP
409 Conflict instead of 500, so the dashboard can surface the attachment
conflict to the user rather than reporting a generic server error.

* admin: skip vanished groups when checking policy attachments

IsPolicyAttached now treats a group that disappears between ListGroups
and GetGroup (credential.ErrGroupNotFound) as no longer attached instead
of failing the whole deletion with HTTP 500, matching the IAM API handler
which skips vanished groups.

* test: assert policy state after deletion paths

Strengthen GetPolicy assertions in the policy deletion tests to check the
returned policy is non-nil after a rejected deletion and nil after a
successful one, not just that no lookup error occurred (GetPolicy returns
nil, nil when a policy is absent).
2026-09-08 16:16:01 -07:00
Chris LuandGitHub b88156fe6b fix(s3api): evaluate aws:SourceIp from the direct TCP peer, not forwarded headers (#11231)
* fix(s3api): use direct peer IP for aws:SourceIp in bucket policy engine

extractSourceIP in the bucket-policy engine trusted X-Forwarded-For and
X-Real-Ip whenever the TCP peer looked private (loopback/RFC1918/link-local),
with no configurable trusted-proxy allowlist. In containerized deployments
the gateway peer is almost always private, so any caller reaching it directly
or from a co-located workload could spoof aws:SourceIp and bypass
IpAddress/NotIpAddress bucket-policy restrictions.

Always return the direct peer address (r.RemoteAddr), matching AWS S3
semantics. Remove the now-unused isPrivateIP helper and header-trust branch.

Update TestExtractConditionValuesFromRequestSourceIPPrecedence to assert the
peer IP is used regardless of forwarding headers, and add regression tests
TestExtractSourceIP_IgnoresForwardedHeaders and
TestExtractSourceIP_EnforcesIPRestrictionPolicy.

* fix(s3api): use direct peer IP for aws:SourceIp in IAM role/session policies

The IAM middleware's extractSourceIP trusted X-Forwarded-For and X-Real-IP
whenever the TCP peer looked private (loopback/RFC1918/link-local), with no
configurable trusted-proxy allowlist. In containerized deployments the gateway
peer is almost always private, so any caller reaching it directly or from a
co-located workload could spoof aws:SourceIp and bypass IpAddress/NotIpAddress
conditions on role and session policies (IsPrincipalActionExplicitlyDenied).

Always return the direct peer address (r.RemoteAddr), matching AWS S3
semantics. Remove the now-unused isPrivateIP helper, privateNetworks table,
and its init().

Update TestRequestContextExtraction and TestIPBasedPolicyEnforcement to assert
the peer IP is enforced regardless of forwarding headers, and add regression
test TestUserInlinePolicySourceIpCondition_IgnoresForwardedHeaders.
2026-09-08 15:09:31 -07:00
Chris LuandGitHub 557fffa350 iam: preserve native Admin when IAM policies are attached (#11226) (#11232)
* iam: expose tri-state result from attached policy evaluation

evaluateIAMPolicies returned a bool that collapsed explicit Deny and
no-match into a single false, so the authorization path could not tell
"policies forbid this" from "policies say nothing". Introduce
evaluateAttachedIAMPolicies returning Allow/Deny/NoMatch and keep
evaluateIAMPolicies as a bool projection for existing callers. This is
preparation for unioning native permissions with attached policies while
preserving deny-always-wins.

* iam: preserve native Admin when IAM policies are attached

Attaching an IAM policy routed authorization exclusively to the attached
policies, dropping the identity native permissions. A user with native
Admin lost all access after attaching a non-granting policy, and stayed
locked out if that policy was deleted without being detached first
(#11226).

Treat a native bare Admin grant as a permission floor that survives
attached policies: when the attached policies do not explicitly allow,
fall back to isAdmin() on the attached-policy path, and on the IAM
integration path allow unless an attached policy explicitly denies.
Explicit Deny still wins on both paths.

Only bare Admin is consulted because inline policies flatten lossily into
Actions (dropping conditions), so scoped actions are not unambiguously
native and must keep flowing through the policy engine.

* iam: regression tests for native Admin surviving attached policies

Reproduces issue #11226:

- TestNativeAdminSurvivesAttachedPolicy: a user with native Admin keeps
  Write access after attaching a policy that does not grant it.
- TestNativeAdminSurvivesDeletedPolicy: the same user keeps Write access
  after the attached policy is deleted without being detached.
- TestAttachedPolicyExplicitDenyOverridesNativeAdmin: an explicit Deny in
  an attached policy still constrains a native admin (deny-always-wins).

* iam: apply native Admin floor before IAM principal validation

The native Admin floor in authorizeWithIAM ran after the auth-path
switch, which denies when no session principal or PrincipalArn is
present. An Admin identity without a PrincipalArn (no session token)
was therefore denied before the floor executed. Move the floor ahead of
the switch and derive the principal for its explicit-deny check with
buildPrincipalARN, which already handles identities without a
PrincipalArn. Adds a regression case for an Admin identity with an
empty PrincipalArn.

Addresses CodeRabbit review feedback on PR #11232.
2026-09-08 14:39:43 -07:00
Chris LuandGitHub 4a1d65939f fix(mount): bound reader cache memory across open files (#11220)
* fix(filer): bound retained reader cache buffers by bytes

* test(filer): keep in-flight downloads during cache trimming

* feat(mem): expose pooled allocation capacity for byte reservations

* fix(mount): share a configurable reader buffer budget across files

* fix(filer): release failed prefetch slots and memory reservations

* feat(mount): expose a soft Go runtime memory limit

* docs(filer): restore shared-download rationale in startCaching

The one-line comment replacing the original context.Background() explanation was too thin for readChunkAt to cross-reference shared resource semantics. Restore a concise note on why request cancellation must not abort a download shared by concurrent readers.

* test(filer): loosen reader cache test deadlines to 5s

Three tests used 1-second deadlines that can flake on CI under load:
TestReaderCacheBudgetInFlight, TestReaderCacheEvictionDoesNotHoldCacheLock,
and TestReaderCacheFailedPrefetchReleasesBudget. Increase to 5 seconds.

* test(filer): cover re-read after reader cache eviction

Add TestReaderCacheReReadAfterEviction: reads chunk 'a', reads chunk 'b'
(evicting 'a' via budget pressure), then re-reads 'a' and asserts a
fresh download returns correct data. Verifies the core correctness
property that eviction never exposes missing or stale data to readers.
2026-09-08 10:51:28 -07:00
ssshr-66andGitHub c6b330be2b [Mount] Add ChunkGroup seeking tests and fix boundary handling (#11223)
* fix issue 11221

* reply ai comments
2026-09-08 10:50:41 -07:00
213f4c5d5c release: wait for Go proxy propagation (#11219)
Allow normal post-tag proxy propagation before dispatching downstream releases, while preserving the check that prevents them from pinning the previous commit.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-07 21:38:49 -07:00
690 changed files with 59692 additions and 13055 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.9
uses: github/codeql-action/init@v4.38.2
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.9
uses: github/codeql-action/autobuild@v4.38.2
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.9
uses: github/codeql-action/analyze@v4.38.2
+2 -1
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'weed/**'
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'seaweed-worker/**'
- 'docker/**'
- 'go.mod'
@@ -152,7 +153,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
run: |
@@ -129,7 +129,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
+10 -6
View File
@@ -156,7 +156,7 @@ jobs:
runs-on: ubuntu-latest
strategy:
matrix:
platform: [amd64, arm64, arm, 386]
platform: [amd64, arm64, arm, 386, ppc64le, s390x]
variant: ${{ fromJSON(needs.setup.outputs.variants) }}
steps:
@@ -236,7 +236,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
- name: Set up QEMU
if: matrix.platform != 'amd64'
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
run: |
cat > /tmp/buildkitd.toml <<EOF
@@ -405,7 +405,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.2
if: always()
with:
sarif_file: trivy-results.sarif
@@ -505,7 +505,9 @@ jobs:
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-amd64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-ppc64le \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-s390x
# The copy and the signature below use the digest this run pushed, not whatever the tag points at by then.
DIGEST=$(jq -er '."containerimage.descriptor".digest' /tmp/manifest.json)
echo "digest=${DIGEST}" >> "$GITHUB_OUTPUT"
@@ -549,13 +551,15 @@ jobs:
echo "Using skopeo to copy..."
retry_with_backoff skopeo copy --all docker://ghcr.io/chrislusf/seaweedfs@${DIGEST} docker://chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}
else
echo "Using docker buildx imagetools (pulling 4 images from Docker Hub)..."
echo "Using docker buildx imagetools (pulling 6 images from Docker Hub)..."
# Fallback: create manifest directly on Docker Hub (pulls from Docker Hub - rate limited)
retry_with_backoff docker buildx imagetools create -t chrislusf/seaweedfs:${BASE_TAG}${SUFFIX} \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-amd64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm64 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-arm \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-386 \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-ppc64le \
ghcr.io/chrislusf/seaweedfs:${BASE_TAG}${SUFFIX}-s390x
fi
- name: Sign
@@ -46,7 +46,7 @@ jobs:
org.opencontainers.image.vendor=Chris Lu
-
name: Set up QEMU
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
-
name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
@@ -149,12 +149,16 @@ jobs:
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/arm64, arch: arm64, runner: ubuntu-24.04-arm, qemu: false }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/arm/v7, arch: armv7, runner: ubuntu-latest, qemu: true }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/386, arch: i386, runner: ubuntu-latest, qemu: false }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/ppc64le, arch: ppc64le, runner: ubuntu-latest, qemu: true }
- { variant: normal, tag_suffix: "", dockerfile: ./docker/Dockerfile.go_build, build_args: "", rust_variant: normal, platform: linux/s390x, arch: s390x, runner: ubuntu-latest, qemu: true }
# Large disk - multi-arch
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/amd64, arch: amd64, runner: ubuntu-latest, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/arm64, arch: arm64, runner: ubuntu-24.04-arm, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/arm/v7, arch: armv7, runner: ubuntu-latest, qemu: true }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/386, arch: i386, runner: ubuntu-latest, qemu: false }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/ppc64le, arch: ppc64le, runner: ubuntu-latest, qemu: true }
- { variant: large_disk, tag_suffix: _large_disk, dockerfile: ./docker/Dockerfile.go_build, build_args: TAGS=5BytesOffset, rust_variant: large-disk, platform: linux/s390x, arch: s390x, runner: ubuntu-latest, qemu: true }
# Full tags - multi-arch
- { variant: full, tag_suffix: _full, dockerfile: ./docker/Dockerfile.go_build, build_args: "TAGS=elastic,gocdk,rclone,sqlite,tarantool,tikv,ydb", rust_variant: normal, platform: linux/amd64, arch: amd64, runner: ubuntu-latest, qemu: false }
@@ -231,7 +235,7 @@ jobs:
- name: Set up QEMU
if: (github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant) && matrix.qemu
uses: docker/setup-qemu-action@v4.3.0
uses: docker/setup-qemu-action@v4.4.0
- name: Create BuildKit config
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -456,7 +460,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.38.2
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
@@ -85,7 +85,7 @@ jobs:
echo "seaweedfs_ref=$seaweed" >> "$GITHUB_OUTPUT"
- name: Set up QEMU
uses: docker/setup-qemu-action@1f40c72289eff860ee54a304f1438e3cff362e0a # v1
uses: docker/setup-qemu-action@99012661954931238ded8c8b007157a8430204e1 # v1
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@4d04d5d9486b7bd6fa91e7baf45bbb4f8b9deedd # v1
+45 -35
View File
@@ -35,11 +35,48 @@ jobs:
cd telemetry/server
go mod tidy
echo "Building telemetry server..."
GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
CGO_ENABLED=0 GOOS=linux GOARCH=amd64 go build -o ../../telemetry-server .
cd ../..
ls -la telemetry-server
echo "Build completed successfully"
- name: Generate Service Configuration
if: github.event_name == 'workflow_dispatch' && (inputs.setup || inputs.deploy)
env:
REMOTE_USER: ${{ secrets.TELEMETRY_USER }}
run: |
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/bin/sh -c 'exec /home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353 >>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log 2>>/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log'
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
- name: First-time Server Setup
if: github.event_name == 'workflow_dispatch' && inputs.setup
env:
@@ -61,40 +98,6 @@ jobs:
touch ~/seaweedfs-telemetry/logs/telemetry.log ~/seaweedfs-telemetry/logs/telemetry.error.log && \
chmod 644 ~/seaweedfs-telemetry/logs/*.log"
# Create systemd service file
echo "
[Unit]
Description=SeaweedFS Telemetry Server
After=network.target
[Service]
Type=simple
User=$REMOTE_USER
WorkingDirectory=/home/$REMOTE_USER/seaweedfs-telemetry
ExecStart=/home/$REMOTE_USER/seaweedfs-telemetry/bin/telemetry-server -port=8353
Restart=always
RestartSec=5
StandardOutput=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.log
StandardError=append:/home/$REMOTE_USER/seaweedfs-telemetry/logs/telemetry.error.log
[Install]
WantedBy=multi-user.target" > telemetry.service
# Setup logrotate configuration
echo "# SeaweedFS Telemetry service log rotation
/home/$REMOTE_USER/seaweedfs-telemetry/logs/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
create 644 $REMOTE_USER $REMOTE_USER
postrotate
systemctl restart telemetry.service
endscript
}" > telemetry_logrotate
# Copy configuration files
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
@@ -137,11 +140,18 @@ jobs:
scp -i ~/.ssh/deploy_key telemetry/grafana-dashboard.json $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
scp -i ~/.ssh/deploy_key telemetry/prometheus.yml $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Copy updated service and logrotate files
scp -i ~/.ssh/deploy_key telemetry.service telemetry_logrotate $REMOTE_USER@$REMOTE_HOST:~/seaweedfs-telemetry/
# Check if service exists and deploy accordingly
ssh -i ~/.ssh/deploy_key $REMOTE_USER@$REMOTE_HOST "
if systemctl list-unit-files telemetry.service >/dev/null 2>&1; then
echo 'Service exists, performing update...'
set -e
sudo systemctl stop telemetry.service
sudo mv ~/seaweedfs-telemetry/telemetry.service /etc/systemd/system/
sudo mv ~/seaweedfs-telemetry/telemetry_logrotate /etc/logrotate.d/seaweedfs-telemetry
sudo systemctl daemon-reload
mkdir -p ~/seaweedfs-telemetry/bin
mv ~/seaweedfs-telemetry/tmp/telemetry-server ~/seaweedfs-telemetry/bin/
chmod +x ~/seaweedfs-telemetry/bin/telemetry-server
+24 -1
View File
@@ -116,6 +116,20 @@ jobs:
grep -q "security-config" /tmp/security.yaml
echo "Security configuration renders correctly"
echo ""
echo "=== Testing admin.allowInsecureBind satisfies the admin auth render guard ==="
helm template test $CHART_DIR --set admin.enabled=true --set admin.allowInsecureBind=true \
> /tmp/admin-allow-insecure-bind.yaml
grep -q -- "-allowInsecureBind" /tmp/admin-allow-insecure-bind.yaml
echo "admin.allowInsecureBind renders -allowInsecureBind and passes the render guard"
if helm template test $CHART_DIR --set admin.enabled=true > /tmp/admin-no-auth.yaml 2>/tmp/admin-no-auth.err; then
echo "FAIL: admin.enabled=true with no auth configured should fail to render"
exit 1
fi
grep -q "admin.allowInsecureBind" /tmp/admin-no-auth.err
echo "admin with no auth configured still fails the render guard, and the guard mentions admin.allowInsecureBind"
echo ""
echo "=== Testing JWT expiration overrides ==="
helm template test $CHART_DIR \
@@ -227,6 +241,9 @@ jobs:
out = render({
"global.seaweedfs.securityConfig.jwtSigning.filerWrite": "true",
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
})
cm = configmap(out, "test-seaweedfs-security-config")
if cm is None:
@@ -1141,6 +1158,9 @@ jobs:
"s3.enabled": "true",
"sftp.enabled": "true",
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
"worker.enabled": "true",
"cosi.enabled": "true",
"s3.createBuckets[0].name": "b",
@@ -1337,7 +1357,7 @@ jobs:
# Which means egress on its own must render for a release that runs
# neither COSI nor a resize: no component of it reaches the API server,
# so nothing may demand a CIDR for one.
for label, values in {"defaults": {}, "admin": {"admin.enabled": "true"}}.items():
for label, values in {"defaults": {}, "admin": {"admin.enabled": "true", "admin.secret.adminPassword": "ci-admin-password"}}.items():
try:
render(dict(values, **{"networkPolicy.enabled": "true",
"networkPolicy.egress.enabled": "true"}))
@@ -1398,6 +1418,9 @@ jobs:
ALL_ON = {
"admin.enabled": "true",
# admin.ip defaults to 0.0.0.0 (non-loopback), which weed admin 4.46
# refuses to bind without authentication.
"admin.secret.adminPassword": "ci-admin-password",
"s3.enabled": "true",
"sftp.enabled": "true",
"worker.enabled": "true",
+1
View File
@@ -8,6 +8,7 @@ on:
- 'go.mod'
- 'go.sum'
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/perf/**'
- '.github/workflows/performance.yml'
workflow_dispatch:
+3 -2
View File
@@ -206,8 +206,9 @@ jobs:
# The dispatched workflow pins seaweedfs with `go get -u ...@latest`, so
# wait until the proxy serves the release commit as the tip. Asking for
# the commit by name is what makes the proxy fetch it.
for _ in $(seq 30); do
# the commit by name is what makes the proxy fetch it. The proxy can
# take longer than five minutes to refresh @latest after a new tag.
for _ in $(seq 120); do
curl -sf "https://proxy.golang.org/${MODULE}/@v/${SHA}.info" >/dev/null || true
TIP=$(curl -sf "https://proxy.golang.org/${MODULE}/@latest" | jq -r '.Origin.Hash // ""' || true)
[ "$TIP" = "$SHA" ] && break
+88 -3
View File
@@ -5,6 +5,7 @@ on:
branches: [ master ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/volume_server/**'
- 'weed/pb/volume_server.proto'
- 'weed/pb/volume_server_pb/**'
@@ -13,6 +14,7 @@ on:
branches: [ master, main ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- 'test/volume_server/**'
- 'weed/pb/volume_server.proto'
- 'weed/pb/volume_server_pb/**'
@@ -27,6 +29,29 @@ permissions:
jobs:
changes:
name: Detect changed paths
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
outputs:
rust: ${{ steps.filter.outputs.rust }}
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
fetch-depth: 0
- name: Filter changed paths
id: filter
uses: dorny/paths-filter@v4
with:
filters: |
rust:
- 'seaweed-volume/**'
- '.github/workflows/rust-volume-server-tests.yml'
rust-unit-tests:
name: Rust Unit Tests
runs-on: ubuntu-22.04
@@ -52,13 +77,73 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
- name: Build Rust volume server
run: cd seaweed-volume && cargo build --release
# The crate is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[lints.clippy]` in
# seaweed-volume/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-volume && cargo clippy --all-targets -- -D warnings
# The crate is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-volume && cargo fmt --check
# seaweed-common is a path dependency of this crate, not a member of its
# workspace, so the run below does not reach its own tests. It builds into
# this job's cached target directory, and the cache key above covers the
# shared crate's lock, so the aws-lc-sys that rustls pulls in is restored
# with the cache instead of compiled from scratch on every run.
- name: Run shared-crate unit tests
env:
CARGO_TARGET_DIR: ${{ github.workspace }}/seaweed-volume/target
run: cd seaweed-common && cargo test
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
- name: Run Rust unit tests (redb experimental cursor)
run: cd seaweed-volume && cargo test --features redb-experimental-cursor --lib storage::needle_map
rust-unit-tests-windows:
name: Rust Unit Tests (Windows)
runs-on: windows-latest
timeout-minutes: 30
needs: [changes]
if: needs.changes.outputs.rust == 'true'
defaults:
run:
shell: bash
steps:
- name: Checkout code
uses: actions/checkout@v7
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# No glibc on Windows: key the cache on the toolchain and OS only.
- name: Fingerprint build toolchain
id: toolchain
run: echo "fingerprint=windows-rustc-$(rustc -V | awk '{print $2}')" >> "$GITHUB_OUTPUT"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-windows-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-windows-${{ steps.toolchain.outputs.fingerprint }}-
- name: Run Rust unit tests
run: cd seaweed-volume && cargo test
@@ -95,7 +180,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
@@ -177,7 +262,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
key: rust-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-${{ steps.toolchain.outputs.fingerprint }}-
+25 -1
View File
@@ -5,12 +5,14 @@ on:
branches: [ master ]
paths:
- 'seaweed-worker/**'
- 'seaweed-common/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
push:
branches: [ master, main ]
paths:
- 'seaweed-worker/**'
- 'seaweed-common/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
@@ -49,7 +51,7 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/release
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock', 'seaweed-common/Cargo.lock') }}
restore-keys: |
rust-worker-${{ steps.toolchain.outputs.fingerprint }}-
@@ -73,6 +75,28 @@ jobs:
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The workspace is warning-free under clippy as of the sweep that added
# this step. Uncomment to make that a gate; `[workspace.lints.clippy]`
# in seaweed-worker/Cargo.toml is where crate-wide exceptions live.
# - name: Clippy
# run: cd seaweed-worker && cargo clippy --workspace --all-targets -- -D warnings
# The workspace is rustfmt-clean as of the PR that added this step.
# Uncomment to keep it that way.
# - name: Check formatting
# run: cd seaweed-worker && cargo fmt --all --check
# seaweed-common is a path dependency of core and lance, not a member of
# this workspace, so `--workspace` below does not reach its own tests.
# Release and this job's cached target directory, and the cache key above
# covers the shared crate's lock. That lock pins the same rustls and
# aws-lc-sys this workspace resolves, so the release build above has
# already paid for them.
- name: Run shared-crate unit tests
env:
CARGO_TARGET_DIR: ${{ github.workspace }}/seaweed-worker/target
run: cd seaweed-common && cargo test --release
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
+1
View File
@@ -5,6 +5,7 @@ on:
branches: [ master ]
paths:
- 'seaweed-volume/**'
- 'seaweed-common/**'
- '.github/workflows/rust_binaries_dev.yml'
permissions:
+106
View File
@@ -0,0 +1,106 @@
name: "Snowflake S3Compat API tests"
on:
push:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/iam/**'
- 'weed/command/**'
- 'weed/storage/**'
- 'weed/operation/**'
- 'weed/wdclient/**'
- 'weed/cluster/**'
- 'weed/pb/**'
- 'test/s3/snowflake/**'
- 'go.mod'
- 'go.sum'
- '.github/workflows/s3-snowflake-tests.yml'
pull_request:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/iam/**'
- 'weed/command/**'
- 'weed/storage/**'
- 'weed/operation/**'
- 'weed/wdclient/**'
- 'weed/cluster/**'
- 'weed/pb/**'
- 'test/s3/snowflake/**'
- 'go.mod'
- 'go.sum'
- '.github/workflows/s3-snowflake-tests.yml'
concurrency:
group: ${{ github.event.pull_request.number || github.ref }}/s3-snowflake-tests
cancel-in-progress: true
permissions:
contents: read
jobs:
snowflake-s3compat-tests:
name: Snowflake S3Compat API tests
runs-on: ubuntu-22.04
timeout-minutes: 30
env:
WORK_DIR: /tmp/seaweedfs-snowflake-tests
steps:
- name: Check out code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Set up Java
uses: actions/setup-java@v6
with:
java-version: '17'
distribution: 'temurin'
cache: 'maven'
- name: Install SeaweedFS
run: |
cd weed
go install -buildvcs=false
weed version
- name: Run Snowflake S3Compat API tests
timeout-minutes: 20
run: |
# Starts weed server, creates the buckets/objects the suite needs,
# clones the upstream suite, and runs mvn -Dtest=S3CompatApiTest.
bash test/s3/snowflake/run.sh
- name: Show logs on failure
if: failure()
run: |
echo "=== SeaweedFS Server Log ==="
tail -200 "$WORK_DIR/weed.log" || echo "No server log"
echo ""
echo "=== Surefire results ==="
cat "$WORK_DIR"/snowflake-s3compat-api-test-suite/s3compatapi/target/surefire-reports/*.txt 2>/dev/null || echo "No surefire reports"
- name: Upload test results
if: always()
uses: actions/upload-artifact@v7
with:
name: snowflake-s3compat-surefire-reports
path: /tmp/seaweedfs-snowflake-tests/snowflake-s3compat-api-test-suite/s3compatapi/target/surefire-reports/
retention-days: 14
- name: Cleanup
if: always()
run: |
pkill -9 -f "weed server" || true
rm -rf "$WORK_DIR" || true
+111
View File
@@ -439,6 +439,117 @@ jobs:
path: test/s3tables/catalog_clickhouse/test-output.log
retention-days: 3
olake-iceberg-catalog-tests:
name: OLake Iceberg Catalog Integration Tests (${{ matrix.tag }})
runs-on: ubuntu-22.04
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
include:
# Pinned baseline, and latest so new OLake releases are exercised
# without a code change. OLake's Iceberg writer is a Java sidecar
# whose Iceberg version moves independently of the Go release, so
# the latest leg is the one that catches library drift.
- olake-image: olakego/source-postgres:v0.10.1
tag: "v0.10.1"
- olake-image: olakego/source-postgres:latest
tag: latest
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull ${{ matrix.olake-image }}
pull postgres:16
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Install SeaweedFS
run: |
go install -buildvcs=false ./weed
- name: Run OLake Iceberg Catalog Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_olake
env:
OLAKE_IMAGE: ${{ matrix.olake-image }}
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting OLake Iceberg Catalog Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "OLake Iceberg catalog integration tests failed"
exit 1
}
# The suite skips itself when Docker is unavailable, so a green job is not
# by itself evidence that anything ran. Assert execution explicitly.
- name: Assert the suite actually ran
working-directory: test/s3tables/catalog_olake
run: |
log=test-output.log
if [ ! -f "$log" ]; then
echo "::error::no test-output.log; the suite did not run"
exit 1
fi
passes=$(grep -c '^--- PASS' "$log" || true)
skips=$(grep -c '^--- SKIP' "$log" || true)
echo "top-level PASS=$passes SKIP=$skips"
if [ "$skips" -gt 0 ]; then
echo "::error::the OLake suite skipped $skips top-level test(s); the environment it needs was not provisioned, so this job proves nothing"
grep '^--- SKIP' "$log" | head -20
exit 1
fi
if [ "$passes" -lt 1 ]; then
echo "::error::the OLake suite recorded no passing top-level test"
exit 1
fi
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_olake
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|olake|postgres)" || true
echo "=== Containers ==="
docker ps -a | head -30 || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: olake-iceberg-catalog-test-logs-${{ matrix.tag }}
path: test/s3tables/catalog_olake/test-output.log
retention-days: 3
polaris-integration-tests:
name: Polaris Integration Tests
runs-on: ubuntu-22.04
+4
View File
@@ -289,6 +289,7 @@ jobs:
s3tests/functional/test_s3.py::test_object_write_check_etag \
s3tests/functional/test_s3.py::test_object_write_cache_control \
s3tests/functional/test_s3.py::test_object_write_expires \
s3tests/functional/test_s3.py::test_object_content_encoding_aws_chunked \
s3tests/functional/test_s3.py::test_object_write_read_update_read_delete \
s3tests/functional/test_s3.py::test_object_metadata_replaced_on_put \
s3tests/functional/test_s3.py::test_object_write_file \
@@ -311,6 +312,7 @@ jobs:
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_good \
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_object_ifunmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_checksum_object_attributes \
s3tests/functional/test_s3.py::test_bucket_head \
s3tests/functional/test_s3.py::test_bucket_head_notexist \
s3tests/functional/test_s3.py::test_object_raw_authenticated \
@@ -1149,6 +1151,7 @@ jobs:
s3tests/functional/test_s3.py::test_object_write_check_etag \
s3tests/functional/test_s3.py::test_object_write_cache_control \
s3tests/functional/test_s3.py::test_object_write_expires \
s3tests/functional/test_s3.py::test_object_content_encoding_aws_chunked \
s3tests/functional/test_s3.py::test_object_write_read_update_read_delete \
s3tests/functional/test_s3.py::test_object_metadata_replaced_on_put \
s3tests/functional/test_s3.py::test_object_write_file \
@@ -1171,6 +1174,7 @@ jobs:
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_good \
s3tests/functional/test_s3.py::test_get_object_ifmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_object_ifunmodifiedsince_failed \
s3tests/functional/test_s3.py::test_get_checksum_object_attributes \
s3tests/functional/test_s3.py::test_bucket_head \
s3tests/functional/test_s3.py::test_bucket_head_notexist \
s3tests/functional/test_s3.py::test_object_raw_authenticated \
+29 -1
View File
@@ -9,6 +9,14 @@ on:
permissions:
contents: write
concurrency:
# Only one chart regeneration per branch at a time; a newer run on the same
# branch cancels an in-flight one so overlapping runs never conflict on
# note/star_history.svg during rebase. Scoped by ref so a manual run on
# another branch can't cancel the daily master update.
group: star-history-${{ github.ref }}
cancel-in-progress: true
jobs:
render:
name: Regenerate star history chart
@@ -17,6 +25,9 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history so the chart commit can rebase onto a moved master.
fetch-depth: 0
- name: Set up Python
uses: actions/setup-python@v7
@@ -43,4 +54,21 @@ jobs:
fi
git add note/star_history.svg
git commit -m "docs: regenerate star history chart"
git push
# Rebase and retry so a concurrent push to master doesn't lose the chart.
for attempt in 1 2 3 4 5; do
if [ "$attempt" -gt 1 ]; then
# Guard the rebase: a transient fetch error or conflict must not
# abort the fail-fast shell before the remaining attempts run.
if ! git pull --rebase origin "$GITHUB_REF_NAME"; then
echo "rebase failed (attempt ${attempt}); aborting and retrying"
git rebase --abort || true
continue
fi
fi
if git push origin HEAD:"$GITHUB_REF_NAME"; then
exit 0
fi
echo "push rejected (attempt ${attempt}); will rebase and retry"
done
echo "::error::could not push star history chart after retries"
exit 1
+2 -3
View File
@@ -17,10 +17,11 @@ SeaweedFS is a simple and highly scalable distributed file system. There are two
1. to store billions of files!
2. to serve the files fast!
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently.
One `weed` binary serves an S3 object store, a POSIX file system, and a lakehouse with S3 Tables, all over the same data. Each blob is one disk read away, capacity grows by starting another volume server, and cloud storage can be cached or tiered transparently. Both read and write operations have O(1) complexity and can run at the full speed supported by the underlying hardware.
- [Download Binaries for different platforms](https://github.com/seaweedfs/seaweedfs/releases/latest)
- [Wiki Documentation](https://github.com/seaweedfs/seaweedfs/wiki)
- [HTTP REST API](REST_API.md) for the filer, master, and volume servers
- Community: [Slack](https://join.slack.com/t/seaweedfs/shared_invite/enQtMzI4MTMwMjU2MzA3LTEyYzZmZWYzOGQ3MDJlZWMzYmI0OTE4OTJiZjJjODBmMzUxNmYwODg0YjY3MTNlMjBmZDQ1NzQ5NDJhZWI2ZmY), [Twitter](https://twitter.com/SeaweedFS), [Telegram](https://t.me/Seaweedfs), [Reddit](https://www.reddit.com/r/SeaweedFS/), [Mailing List](https://groups.google.com/d/forum/seaweedfs)
- [SeaweedFS White Paper](https://github.com/seaweedfs/seaweedfs/wiki/SeaweedFS_Architecture.pdf) and introduction slides: [2025.5](https://docs.google.com/presentation/d/1tdkp45J01oRV68dIm4yoTXKJDof-EhainlA0LMXexQE/edit?usp=sharing), [2021.5](https://docs.google.com/presentation/d/1DcxKWlINc-HNCjhYeERkpGXXm6nTCES8mi2W5G0Z4Ts/edit?usp=sharing), [2019.3](https://www.slideshare.net/chrislusf/seaweedfs-introduction)
@@ -400,8 +401,6 @@ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
The text of this page is available for modification and reuse under the terms of the Creative Commons Attribution-Sharealike 3.0 Unported License and the GNU Free Documentation License (unversioned, with no invariant sections, front-cover texts, or back-cover texts).
[Back to TOC](#table-of-contents)
# Sponsors #
+344
View File
@@ -0,0 +1,344 @@
# SeaweedFS HTTP REST API
SeaweedFS exposes three HTTP surfaces:
| Service | Default port | Addressing |
|---------|--------------|------------|
| Filer | 8888 | File system paths (`/dir/name`) |
| Master | 9333 | File id assignment and cluster topology |
| Volume server | 8080 | File content by file id (`vid,fid`) |
Most clients only need the filer API (paths) or the S3 API. The master and
volume APIs are the lower-level blob store interface.
Conventions applying to all three:
- Responses are JSON unless noted otherwise. Append `&pretty=y` to pretty-print.
- A file id (`fid`) has the form `volumeId,fileKeyCookie`, e.g. `3,01637037d6`.
An optional suffix selects a reserved id from a `count` assignment
(`3,01637037d6_1`, `_2`, ...), and an optional extension
(`3,01637037d6.jpg`) sets the content type on reads.
- `replication` is a 3-digit replica placement `xyz`: `x` copies in other
data centers, `y` on other racks in the same data center, `z` on other
volume servers on the same rack. `000` = no replication, `001` = one copy
on the same rack, `010` = one copy on a different rack, `100` = one copy in
another data center, `200` = two copies in two other data centers, `110` =
one copy in another data center plus one on another rack.
- `ttl` units: `m` minute, `h` hour, `d` day, `w` week, `M` month, `y` year.
## Filer API (port 8888)
The filer presents a POSIX-like namespace over the volume servers.
### Upload a file
```bash
# PUT the raw body to the target path
curl -T /home/chris/myphoto.jpg "http://localhost:8888/dir/myphoto.jpg"
# or POST as multipart form (the part filename becomes the entry name)
curl -F file=@/home/chris/myphoto.jpg "http://localhost:8888/dir/"
```
Response `201 Created`:
```json
{"name":"myphoto.jpg","size":43234,"eTag":"0x6c656...","mtime":"...","chunks":[...]}
```
Query parameters:
| Parameter | Description | Default |
|-----------|-------------|---------|
| `collection` | collection name | empty |
| `replication` | replica placement code | filer default |
| `ttl` | file expiration, e.g. `3d` | never |
| `disk` | disk type to store on | filer default |
| `fsync` | `true` fsyncs on the volume server | false |
| `dataCenter` | preferred data center | empty |
| `rack` | preferred rack | empty |
| `dataNode` | preferred volume server | empty |
| `saveInside` | store small content inside the metadata instead of a volume | false |
| `maxMB` | split the upload into chunks of this many MB | filer `-maxMB` |
| `mode` | unix permission bits, e.g. `0644` | `0660` |
| `op` | `append` appends to an existing file | overwrite |
| `skipCheckParentDir` | `true` skips the parent-directory existence check | false |
### Create a directory
```bash
curl -X POST "http://localhost:8888/dir/newdir/"
```
A POST to a path ending in `/` with no content creates the directory,
including missing parents.
### Read a file
```bash
curl "http://localhost:8888/dir/myphoto.jpg"
```
Supports `Range` requests (`Accept-Ranges: bytes`), `ETag`, and the
`If-None-Match` / `If-Modified-Since` conditional headers. `HEAD` returns
headers only. Entry headers stored as extended attributes are echoed back,
minus internal `Seaweed-` and `xattr-` keys.
Entry metadata instead of content:
```bash
curl "http://localhost:8888/dir/myphoto.jpg?metadata=true"
```
`metadata=true&resolveManifest=true` additionally resolves chunked-manifest
entries into their real chunk list.
### List a directory
```bash
curl -H "Accept: application/json" "http://localhost:8888/dir/?limit=10&lastFileName=a.jpg"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `limit` | max entries per page | filer `-dirListLimit` |
| `lastFileName` | resume listing after this entry name | empty |
| `namePattern` | include only names matching the wildcard | empty |
| `namePatternExclude` | exclude names matching the wildcard | empty |
The JSON response carries `Path`, `Entries`, `Limit`, `LastFileName`,
`ShouldDisplayLoadMore`, and `EmptyFolder`. Without the `Accept` header the
filer renders its HTML browser.
### Move and copy
```bash
curl -X POST "http://localhost:8888/dir/newname.jpg?mv.from=/dir/myphoto.jpg"
curl -X POST "http://localhost:8888/dir/copy.jpg?cp.from=/dir/myphoto.jpg"
```
`mv.from` renames or moves the source to the request path (`204 No Content`).
`cp.from` copies it.
### Append
```bash
curl -T chunk2.bin "http://localhost:8888/dir/file.bin?op=append"
```
### Delete
```bash
curl -X DELETE "http://localhost:8888/dir/myphoto.jpg"
curl -X DELETE "http://localhost:8888/dir/?recursive=true"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `recursive` | delete a non-empty directory tree | false; when the filer runs with `filer.options.recursive_delete=true`, deletes are recursive unless `recursive=false` |
| `ignoreRecursiveError` | keep deleting remaining entries after an error | false |
| `skipChunkDeletion` | remove only the metadata, keep volume data | false |
### Tagging
Tags are carried as `Seaweed-`-prefixed request headers, not query
parameters; `?tagging` selects the tagging handler and `?tagging=K1,K2`
lists the keys to remove. Header names are canonicalized on write
(`Seaweed-k1` is stored as `Seaweed-K1`), and the delete list is matched
case-sensitively against the stored names.
```bash
curl -X PUT -H "Seaweed-k1: v1" -H "Seaweed-k2: v2" "http://localhost:8888/dir/file.jpg?tagging"
curl -X DELETE "http://localhost:8888/dir/file.jpg?tagging=K1,K2"
```
### Read by file id
```bash
curl "http://localhost:8888/?proxyChunkId=3,01637037d6"
```
The filer proxies the chunk read to the right volume server, so only the
filer port needs to be exposed.
### Resumable uploads
The filer serves the [TUS protocol](https://tus.io/) for resumable uploads
(`POST`, `PATCH`, `HEAD` on upload URLs). It is enabled by default at
`/.tus`; `-tusBasePath` changes the endpoint base path.
### Health
`GET /healthz` and `GET /readyz` return `200 OK`.
## Master API (port 9333)
Write-affecting endpoints are automatically proxied to the current leader, so
any master in the quorum can serve them.
### Assign a file id
```bash
curl "http://localhost:9333/dir/assign?count=1&replication=001&collection=turbo&dataCenter=dc1&ttl=3d&disk=ssd"
{"count":1,"fid":"3,01637037d6","url":"127.0.0.1:8080","publicUrl":"localhost:8080"}
```
Upload the file content to `http://<url>/<fid>` afterwards. With `count>1`,
use `<fid>_1`, `<fid>_2`, ... for the additional ids.
| Parameter | Description | Default |
|-----------|-------------|---------|
| `count` | file ids to reserve | 1 |
| `collection` | collection name | empty |
| `dataCenter` | preferred data center | empty |
| `rack` | preferred rack | empty |
| `dataNode` | preferred volume server | empty |
| `replication` | replica placement | master `-defaultReplication` |
| `ttl` | file expiration, e.g. `3d` | never |
| `disk` | disk type | empty |
| `dataSize` | expected file size in bytes | 0 |
| `preallocate` | bytes to preallocate for new volumes | master `-volumePreallocate` |
| `writableVolumeCount` | grow this many volumes when none are writable | master default |
| `memoryMapMaxSizeMb` | memory-mapped file size (Windows) | 0 |
### Look up a volume or file id
```bash
curl "http://localhost:9333/dir/lookup?volumeId=3"
{"locations":[{"url":"localhost:8080","publicUrl":"localhost:8080"}]}
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `volumeId` | volume id; a full `vid,fid` is accepted too | required |
| `fileId` | like `volumeId`, but also returns a write JWT when security is on | empty |
| `collection` | speeds up the lookup | empty |
| `read` | `yes` generates a read JWT instead of a write JWT | empty |
### Store a file in one call
```bash
curl -F file=@/home/chris/report.pdf "http://localhost:9333/submit?collection=turbo&replication=001"
{"fileName":"report.pdf","fid":"3,01637037d6","fileUrl":"localhost:8080/3,01637037d6","size":43234,"eTag":"0x6c656..."}
```
`POST /submit` accepts multipart file data plus the `dir/assign` placement
parameters (`count`, `collection`, `dataCenter`, `rack`, `replication`,
`ttl`, `disk`), assigns a file id, uploads to the volume server, and returns
the result.
### Redirect to a file
```bash
curl -v "http://localhost:9333/3,01637037d6"
```
`GET /{fileId}` answers `308 Permanent Redirect` to a volume server holding
the file, preserving the query string (e.g. image-resize parameters).
### Cluster status
```bash
curl "http://localhost:9333/dir/status?pretty=y" # full topology tree
curl "http://localhost:9333/vol/status?pretty=y" # every volume on every node
curl "http://localhost:9333/collection/info?collection=turbo"
curl "http://localhost:9333/collection/info?collection=turbo&detail=true"
```
`collection/info` returns aggregated `TotalSize`, `FileCount`, `UsedSize`,
`VolumeCount`; `detail=true` splits them per volume layout.
### Grow volumes
```bash
curl "http://localhost:9333/vol/grow?count=4&replication=001&collection=turbo&ttl=5d&disk=ssd&dataCenter=dc1&rack=rack1"
{"count":4}
```
`count` is required; the placement parameters match `dir/assign`. One volume
serves one write at a time, so pre-allocated volumes raise write concurrency.
### Vacuum deleted space
```bash
curl "http://localhost:9333/vol/vacuum?garbageThreshold=0.4"
```
| Parameter | Description | Default |
|-----------|-------------|---------|
| `garbageThreshold` | minimum deleted-bytes ratio before a volume is compacted | master `-garbageThreshold` (0.3) |
Vacuuming makes a volume read-only, copies live needles to a new volume, and
swaps it in.
### Delete a collection
```bash
curl "http://localhost:9333/col/delete?collection=benchmark"
```
Deletes all volumes of the collection, including erasure-coded shards.
`204 No Content` on success.
### Health
```bash
curl -I "http://localhost:9333/healthz" # liveness
curl -I "http://localhost:9333/readyz" # readiness
curl "http://localhost:9333/" # web UI
```
## Volume server API (port 8080)
The volume server stores file content by file id. Clients normally get the
volume URL from `dir/assign` or `dir/lookup`.
### Upload
```bash
curl -F file=@/home/chris/myphoto.jpg "http://127.0.0.1:8080/3,01637037d6"
{"name":"myphoto.jpg","size":43234,"eTag":"0x6c656...","mime":"image/jpeg","contentMd5":"..."}
```
PUT or POST the body (or a multipart `file` part) to `/{vid},{fid}`.
`204 No Content` is returned when the content is unchanged. `?ts=<unix>`
sets the stored modification time.
### Read
```bash
curl "http://127.0.0.1:8080/3,01637037d6"
curl "http://127.0.0.1:8080/3,01637037d6.jpg" # sets Content-Type from the extension
```
Supports `Range` and `HEAD`. Image files can be resized server-side:
| Parameter | Description |
|-----------|-------------|
| `width`, `height` | resize bounds in pixels |
| `mode` | `fit` (contain) or `fill` (cover); omitted resizes to `width`/`height` |
| `crop_x1`, `crop_y1`, `crop_x2`, `crop_y2` | explicit crop rectangle |
| `cm` | `false` returns the chunk-manifest blob instead of resolving it |
| `readDeleted` | `true` reads soft-deleted needles |
| `collection` | passed through redirects for the right volume |
### Delete
```bash
curl -X DELETE "http://127.0.0.1:8080/3,01637037d6"
{"size":43234}
```
`?ts=<unix>` sets the deletion timestamp. Replicated volumes propagate the
delete to every replica.
### Status
```bash
curl "http://localhost:8080/status?pretty=y" # disk and volume inventory
curl -I "http://localhost:8080/healthz" # liveness/readiness
```
`OPTIONS` preflights answer CORS headers. When `-port.public` differs from
`-port`, the volume server opens a separate read-only public listener on
that port; `-publicUrl` sets the address it advertises to clients.
+418
View File
@@ -0,0 +1,418 @@
# SeaweedFS as an Apache CloudStack Object Storage Provider
A CloudStack ObjectStore plugin that makes SeaweedFS a first-class object storage
backend inside Apache CloudStack, alongside the existing MinIO and Ceph RGW
providers. This is a collaboration with proIO (Swen), who builds private clouds on
CloudStack and wants SeaweedFS as a storage option.
## The request
> We can only add MinIO and Ceph as object storage [in CloudStack] today. I want
> to get SeaweedFS into this project... What we need is to build a provider which
> does the communication between Cloudstack and SeaweedFS.
This is **not** a SeaweedFS-side feature. The work lives in the Apache CloudStack
repo (Java): a new plugin under `plugins/storage/object/seaweedfs/` that implements
CloudStack's ObjectStore plugin framework and talks to SeaweedFS over its S3 and
IAM APIs. SeaweedFS itself needs no changes for the core to work — its S3 API
already covers every bucket operation CloudStack requires, and its IAM API covers
user/credential management.
## How the CloudStack ObjectStore framework works
CloudStack 4.18+ introduced an Object Storage framework. An admin registers an
object storage pool via `addObjectStoragePool` (URL + provider + credentials);
tenants then create and manage buckets on it through CloudStack APIs. CloudStack
manages pool and bucket lifecycle; the underlying provider handles the actual
object protocol.
A provider is a plugin module implementing three interfaces:
### 1. `ObjectStoreProvider` — registration
`MinIOObjectStoreProviderImpl` is the reference. It is a Spring `@Component` that:
- Returns a provider name (`"MinIO"`)
- Returns `DataStoreProviderType.OBJECT`
- In `configure()`, injects the lifecycle and driver implementations and calls
`storeMgr.registerDriver(name, driver)`
### 2. `ObjectStoreLifeCycle` — pool add/remove
`MinIOObjectStoreLifeCycleImpl.initialize()` reads the URL, name, and
`accesskey`/`secretkey` details from the `addObjectStoragePool` call, tests the
connection by listing buckets, and persists an `ObjectStoreVO` via
`ObjectStoreHelper`. The other methods (attachCluster/Host/Zone, maintain,
deleteDataStore) are no-ops for object storage.
### 3. `ObjectStoreDriver` — bucket + user operations
`ObjectStoreDriver` (in `engine/storage/.../object/ObjectStoreDriver.java`) extends
`DataStoreDriver` and defines the bucket/user contract. Every provider must
implement:
| Method | Purpose |
| --- | --- |
| `createBucket(Bucket, boolean objectLock)` | Create a bucket |
| `listBuckets(long storeId)` | List all buckets |
| `deleteBucket(BucketTO, long storeId)` | Delete a bucket |
| `createUser(long accountId, long storeId)` | Provision a user + credentials for a CloudStack account |
| `setBucketPolicy` / `getBucketPolicy` / `deleteBucketPolicy` | Bucket policy CRUD |
| `setBucketEncryption` / `deleteBucketEncryption` | SSE config |
| `setBucketVersioning` / `deleteBucketVersioning` | Versioning enable/suspend |
| `setBucketQuota(BucketTO, long storeId, long size)` | Per-bucket quota |
| `getAllBucketsUsage(long storeId)` | Usage map for billing/accounting |
| `getBucketAcl` / `setBucketAcl` | ACLs (MinIO/Ceph return null / no-op) |
`BaseObjectStoreDriverImpl` provides no-op defaults for the `DataStoreDriver`
methods (`createAsync`, `deleteAsync`, `copyAsync`, `canCopy`, `resize`,
`getTO`, `getStoreTO`), so object-store providers only implement the bucket/user
methods above.
## How the four existing providers differ (and where SeaweedFS lands)
CloudStack ships four object-store providers. Three are relevant; the simulator
is a test stub.
| Concern | MinIO | Ceph RGW | Cloudian HyperStore | SeaweedFS |
| --- | --- | --- | --- | --- |
| Bucket CRUD | `MinioClient` (S3) | `AmazonS3` (AWS SDK v1) | `AmazonS3` (AWS SDK v1) | `AmazonS3` (AWS SDK v1) |
| Bucket policy | `MinioClient` | `AmazonS3` | `AmazonS3` | `AmazonS3` |
| Versioning | `MinioClient` | `AmazonS3` | `AmazonS3` | `AmazonS3` |
| Encryption | `MinioClient` | not implemented | `AmazonS3` | `AmazonS3` |
| **User creation** | `MinioAdminClient` | `RgwAdmin` | **`AmazonIdentityManagement`** | **`AmazonIdentityManagement`** |
| **Per-bucket quota** | `MinioAdminClient` | `RgwAdmin` | **not supported** (throws) | **S3 extension** (`PUT /{bucket}?seaweedfs-quota`, SigV4, `s3:PutBucketQuota`) |
| **Usage reporting** | `MinioAdminClient` | `RgwAdmin` | Cloudian admin API | S3 `ListObjectsV2` (MVP); Prometheus / SOSAPI `capacity.xml` (recommended) |
**Cloudian HyperStore is the direct precedent.** It is an S3-compatible store
that, like SeaweedFS, manages users via the **standard AWS IAM API** using the
AWS IAM Java SDK (`com.amazonaws.services.identitymanagement`). Its driver
(`CloudianHyperStoreObjectStoreDriverImpl`) and util
(`CloudianHyperStoreUtil`) are the template this design follows almost line for
line. Cloudian even validates the quota limitation the same way this design
proposes for the MVP: `setBucketQuota` throws for any non-zero size and only
accepts `0` (no quota).
The SeaweedFS plugin is therefore a **simpler Cloudian** — same AWS S3 + IAM SDK
clients, same store-details keys (`s3Url`, `iamUrl`, `accesskey`, `secretkey`),
same IAM-user-with-restricted-policy pattern, but with no proprietary admin
client at all (Cloudian has its own `CloudianClient` for its admin API; SeaweedFS
needs only S3 + IAM). For quota, the plugin uses a narrow SeaweedFS S3 extension
(see below); for usage reporting, it falls back to S3 `ListObjectsV2` in the MVP
and recommends Prometheus or SOSAPI `capacity.xml` for production scale.
### Quota via the S3 `?seaweedfs-quota` extension
SeaweedFS supports bucket quota natively (server-side enforcement via a
read-only flag when usage exceeds the limit). Rather than exposing the broad
admin REST API (which would require a global bearer token and grant cluster-wide
admin access), the integration uses a **narrow, scoped S3 subresource**:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
**PUT request body** (JSON):
```json
{"quota_size": 100, "quota_unit": "GB", "quota_enabled": true}
```
**GET response body** (JSON):
```json
{"quota_size": 107374182400, "quota_unit": "B", "quota_enabled": true}
```
Note: GET always returns `quota_unit: "B"` and the absolute byte count, not
the original unit. A disabled-but-retained quota returns a positive
`quota_size` with `quota_enabled: false`.
Quota is stored on the bucket's filer entry (positive = enabled, negative =
disabled but retained, zero = no quota), matching the existing admin REST API
behavior. When quota is cleared, the bucket's read-only flag is also lifted.
**Authentication** uses the existing S3 SigV4 flow — no new global secret is
needed. The CloudStack service credential (the `accesskey`/`secretkey` on the
object store) is the admin credential used for all driver operations: bucket
CRUD, IAM user provisioning, and quota management. It must have broad S3 and
IAM permissions. The per-account IAM users created by `createUser` are the
ones with restricted permissions (full S3 access except bucket
creation/deletion). A future hardening could split quota management onto a
separate credential scoped to only `s3:PutBucketQuota`/`s3:GetBucketQuota`,
but the MVP uses the single admin credential for simplicity, matching how
the MinIO and Ceph providers work.
The plugin's `setBucketQuota` signs and sends the `PUT /{bucket}?seaweedfs-quota`
request using the AWS SDK v1 `AWSS3V4Signer` for SigV4 signing, then sends the
signed request via `java.net.http.HttpClient` (the AWS S3 SDK doesn't natively
support custom subresources, so we sign manually and send the request
ourselves). The `seaweedfs-quota` query parameter is included in the signed
canonical query string.
### Usage reporting
`getAllBucketsUsage` must return a `Map<String, Long>` of bucket name → size.
MinIO uses `MinioAdminClient.getDataUsageInfo`; Ceph uses
`RgwAdmin.listBucketInfo`. SeaweedFS has no admin rollup endpoint, so the MVP
plugin computes it by listing buckets and summing object sizes via S3
`ListObjectsV2` — expensive for large stores.
For production scale, SeaweedFS already exposes per-bucket size in:
- **Prometheus metrics** (`bucket_size_bytes` gauge, refreshed every minute)
- **SOSAPI `capacity.xml`** (reports capacity, available space, and usage
through the S3 endpoint)
Operators should consume one of those instead of S3 list-based aggregation for
large deployments. The MVP's list-based approach is correct but slow; flag it as
a known limitation.
## SeaweedFS API surface (what the plugin relies on)
SeaweedFS exposes two relevant APIs, both AWS-compatible:
### S3 API (`weed s3`)
Full S3-compatible surface. Confirmed against the SeaweedFS S3 wiki and code:
- `CreateBucket`, `HeadBucket`, `ListBuckets`, `DeleteBucket`
- `PutBucketPolicy`, `GetBucketPolicy`, `DeleteBucketPolicy`
- `PutBucketVersioning` (Enabled / Suspended), `GetBucketVersioning`
- `PutBucketEncryption`, `GetBucketEncryption`, `DeleteBucketEncryption`
- `PutBucketAcl`, `GetBucketAcl`
- `ListObjectsV2`, `HeadObject`, `GetObject`, `PutObject`, `DeleteObject`
- Bucket quota via extended attributes / `s3.bucket.quota` (enforced server-side,
surfaced as a read-only state when exceeded — see PR #10224)
### IAM API (`weed iam` / `iamapi`)
AWS IAM-compatible REST endpoints, implemented in `weed/iamapi/`. Confirmed by
the test suite which uses the **AWS IAM SDK** (`aws-sdk-go/service/iam`) against
the same handlers CloudStack would call:
- `CreateUser`, `DeleteUser`, `ListUsers`, `GetUser`
- `CreateAccessKey`, `DeleteAccessKey`, `ListAccessKeys`
- `PutUserPolicy`, `GetUserPolicy`, `DeleteUserPolicy`
- `AttachUserPolicy`, `ListAttachedUserPolicies`
This means the CloudStack plugin can manage SeaweedFS users with the **AWS IAM
Java SDK** (`com.amazonaws.services.identitymanagement.AmazonIdentityManagement`),
exactly the way the AWS IAM Go SDK is used in SeaweedFS's own tests. No proprietary
admin client is needed. **Cloudian HyperStore already does exactly this** in the
CloudStack tree — the SeaweedFS plugin follows the same pattern.
## Design
### Module layout
New CloudStack plugin module, mirroring `plugins/storage/object/cloudian/`
(the closest precedent — same AWS S3 + IAM SDK approach):
```
plugins/storage/object/seaweedfs/
pom.xml
src/main/java/org/apache/cloudstack/storage/datastore/
driver/SeaweedFSObjectStoreDriverImpl.java
lifecycle/SeaweedFSObjectStoreLifeCycleImpl.java
provider/SeaweedFSObjectStoreProviderImpl.java
util/SeaweedFSObjectStoreUtil.java
src/test/java/org/apache/cloudstack/storage/datastore/
driver/SeaweedFSObjectStoreDriverImplTest.java
provider/SeaweedFSObjectStoreProviderImplTest.java
src/main/resources/META-INF/cloudstack/storage-object-seaweedfs/
module.properties
spring-storage-object-seaweedfs-context.xml
```
### `SeaweedFSObjectStoreProviderImpl`
Direct copy of `MinIOObjectStoreProviderImpl` with `providerName = "SeaweedFS"`,
injecting the SeaweedFS lifecycle and driver. Registers via
`storeMgr.registerDriver`.
### `SeaweedFSObjectStoreLifeCycleImpl`
Copy of `MinIOObjectStoreLifeCycleImpl`. `initialize()` reads `url`, `name`,
`accesskey`, `secretkey` from the `addObjectStoragePool` details map, tests the
connection by calling `AmazonS3.listBuckets()` against the SeaweedFS S3 endpoint,
and persists the `ObjectStoreVO`. No proprietary client needed — the AWS S3 SDK
is enough for the health check.
### `SeaweedFSObjectStoreDriverImpl`
The substantive class. Uses two AWS SDK v1 clients (same dependency Ceph already
pulls in, so no new CloudStack dependency):
- `AmazonS3` for bucket operations (path-style, endpoint-pinned, `us-east-1`
region placeholder — same as Ceph's `getS3Client`)
- `AmazonIdentityManagement` for user/credential operations, pointed at the
SeaweedFS IAM endpoint
#### Bucket operations — straightforward S3
| Interface method | Implementation |
| --- | --- |
| `createBucket` | `s3.createBucket(name)`; reject if `doesBucketExistV2`; persist access/secret key + URL on `BucketVO` (same as Ceph) |
| `listBuckets` | `s3.listBuckets()` → wrap as `BucketObject` (same as Ceph) |
| `deleteBucket` | `s3.deleteBucket(name)` (same as Ceph) |
| `setBucketPolicy` | `s3.setBucketPolicy(...)` with the same public/private JSON the MinIO/Ceph drivers build |
| `getBucketPolicy` / `deleteBucketPolicy` | `s3.getBucketPolicy` / `s3.deleteBucketPolicy` |
| `setBucketVersioning` | `s3.setBucketVersioningConfiguration(Enabled)` |
| `deleteBucketVersioning` | `s3.setBucketVersioningConfiguration(Suspended)` |
| `setBucketEncryption` | `s3.setBucketEncryptionConfiguration(SSE-S3 rule)` |
| `deleteBucketEncryption` | `s3.deleteBucketEncryptionConfiguration` |
| `getBucketAcl` / `setBucketAcl` | no-op / null (same as MinIO and Ceph) |
#### User creation — the key difference
MinIO calls `MinioAdminClient.addUser`; Ceph calls `RgwAdmin.createUser`. SeaweedFS
exposes the standard AWS IAM API, so the plugin calls:
```java
AmazonIdentityManagement iam = getIamClient(storeId);
String userName = "acs-" + account.getUuid();
// CreateUser (idempotent — check GetUser first, like Ceph does)
iam.createUser(new CreateUserRequest(userName));
// CreateAccessKey → returns the access key + secret key to persist
CreateAccessKeyResult result = iam.createAccessKey(
new CreateAccessKeyRequest().withUserName(userName));
AccessKey key = result.getAccessKey();
// Persist per-account, same pattern as Ceph's CEPH_ACCESS_KEY/CEPH_SECRET_KEY
details.put(SEAWEEDFS_ACCESS_KEY, key.getAccessKeyId());
details.put(SEAWEEDFS_SECRET_KEY, key.getSecretAccessKey());
_accountDetailsDao.persist(accountId, details);
```
This is the cleanest mapping of the three providers: no proprietary admin client,
just the AWS IAM SDK that CloudStack already has access to. The IAM endpoint URL
is provided as `iamUrl` in the store details. If `iamUrl` is omitted, the driver
defaults it to `s3Url` — SeaweedFS registers its embedded IAM API at `POST /` on
the same S3 endpoint (`UnifiedPostHandler` in `s3api_server.go`), so the IAM
endpoint is the same as the S3 endpoint unless the deployment runs a separate
`weed iam` server.
#### Bucket quota — S3 `?seaweedfs-quota` extension
This is the one genuine gap. MinIO and Ceph both have an admin API to set a
per-bucket quota that the backend enforces. SeaweedFS enforces bucket quota
server-side, but the configuration path was **not exposed over a standard S3 or
IAM API** — it was only set via the admin REST API or shell commands.
The integration adds a **narrow S3 subresource** to SeaweedFS:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
This is implemented in SeaweedFS PR #11279. It uses SigV4 authentication and
dedicated IAM permissions, so the CloudStack service credential can be scoped
to quota management only — no global admin token, no cluster-wide admin access.
The enforcement already exists (PR #10224); this PR only adds the HTTP
configuration surface.
An earlier approach (PR #11278, closed) added bearer-token auth to the broad
admin REST API. After review, that was unnecessary for this integration —
static S3 config plus standard S3 APIs plus one scoped quota mutation API is
sufficient and far safer.
> **Note on AWS tools compatibility.** `?seaweedfs-quota` is a SeaweedFS-specific
> S3 subresource, not part of the AWS S3 API. Standard AWS tools (`aws s3api`,
> `s3cmd`, `rclone`) cannot call it directly. This is the same limitation MinIO
> and Ceph have — MinIO quota lives behind a separate admin API (`mc admin
> bucket quota`), and Ceph quota lives behind the Admin Ops API
> (`radosgw-admin quota set`). Neither is callable via `aws s3api` either.
> SeaweedFS's approach is the closest to standard S3 because it uses the same
> endpoint and same SigV4 credentials, just with a custom query parameter.
> Interactive quota management remains available via `weed shell`; the S3
> extension exists for programmatic integration (CloudStack) where the
> integrator can sign SigV4 requests but cannot run shell commands.
#### Usage reporting
`getAllBucketsUsage` must return a `Map<String, Long>` of bucket name → size.
MinIO uses `MinioAdminClient.getDataUsageInfo`; Ceph uses
`RgwAdmin.listBucketInfo`. SeaweedFS has no admin rollup endpoint, so the MVP
plugin computes it by listing buckets and summing object sizes via S3
`ListObjectsV2` — expensive for large stores. Better options exist in
SeaweedFS already:
- **Prometheus metrics** (`bucket_size_bytes` gauge, refreshed every minute)
- **SOSAPI `capacity.xml`** (reports capacity, available space, and usage
through the S3 endpoint — note: the current "return zero on backend error"
behavior should be validated before using it for billing)
For the MVP, `listBuckets` + per-bucket size via the S3 API is correct but slow;
flag it as a known limitation. Operators should consume Prometheus or SOSAPI
for production-scale usage reporting.
### Spring wiring
`spring-storage-object-seaweedfs-context.xml` registers the provider bean,
identical to the MinIO one. `module.properties` sets
`name=storage-object-seaweedfs`, `parent=storage`.
### `pom.xml`
Depends on `aws-java-sdk-s3` and `aws-java-sdk-iam` — both already in the
CloudStack dependency tree (Ceph uses the S3 SDK; the IAM SDK is the standard AWS
bundle). No new third-party dependency, unlike MinIO which pulls in the MinIO
Java client.
## What changes on the SeaweedFS side
**One narrow S3 extension is required for quota management.** SeaweedFS PR #11279
adds the `?seaweedfs-quota` S3 subresource:
- `PUT /{bucket}?seaweedfs-quota` — set bucket quota (IAM permission `s3:PutBucketQuota`)
- `GET /{bucket}?seaweedfs-quota` — get bucket quota (IAM permission `s3:GetBucketQuota`)
This is authenticated via standard S3 SigV4 and authorized via dedicated IAM
permissions, so no global admin token is needed. The enforcement already exists
(PR #10224); this PR only adds the HTTP configuration surface.
One follow-up improvement on the SeaweedFS side would close the usage reporting
gap:
1. **Validate SOSAPI `capacity.xml` usage calculation** — the current "return
zero on backend error" behavior should be validated before using it for
billing. If reliable, CloudStack can consume it directly instead of
list-based aggregation.
## Open questions for proIO / Swen
1. **IAM endpoint path.** ~~Where does `weed iam` listen relative to the S3
endpoint in a typical proIO deployment?~~ **Resolved.** SeaweedFS registers
its embedded IAM API at `POST /` on the same S3 endpoint
(`UnifiedPostHandler`), so the driver defaults `iamUrl` to `s3Url`. A
separate `iamUrl` is only needed if the deployment runs a standalone
`weed iam` server on a different host/port.
2. **Quota requirements.** Do proIO's customers need server-enforced per-bucket
quotas, or is CloudStack-side accounting sufficient for the first release?
The `?seaweedfs-quota` S3 extension (PR #11279) provides server-enforced
quotas via a scoped credential; this is the recommended path.
3. **Object Lock.** `createBucket` takes an `objectLock` boolean. MinIO supports
it; Ceph ignores it. SeaweedFS has Object Lock support. Should the plugin pass
it through?
4. **Contribution model.** Does proIO want to submit the PR to
`apache/cloudstack` themselves (with SeaweedFS maintainers as reviewers), or
the reverse? Apache CloudStack requires an ICLA for non-trivial contributions.
## Files
All in the `apache/cloudstack` repo (new module):
| File | Purpose |
| --- | --- |
| `plugins/storage/object/seaweedfs/pom.xml` | Maven module |
| `.../datastore/util/SeaweedFSObjectStoreUtil.java` | S3 + IAM client builders, constants, URL validators |
| `.../datastore/provider/SeaweedFSObjectStoreProviderImpl.java` | Spring provider registration |
| `.../datastore/lifecycle/SeaweedFSObjectStoreLifeCycleImpl.java` | Pool add/health-check |
| `.../datastore/driver/SeaweedFSObjectStoreDriverImpl.java` | Bucket + user ops via S3 + IAM SDK |
| `.../resources/META-INF/cloudstack/storage-object-seaweedfs/module.properties` | Module name |
| `.../resources/META-INF/cloudstack/storage-object-seaweedfs/spring-storage-object-seaweedfs-context.xml` | Spring bean |
| `plugins/pom.xml` | Register `storage/object/seaweedfs` module |
No files in `seaweedfs/seaweedfs` for the MVP.
### SeaweedFS-side changes (PR #11279)
| File | Purpose |
| --- | --- |
| `weed/s3api/s3_constants/s3_action_strings.go` | Add `S3_ACTION_PUT_BUCKET_QUOTA` and `S3_ACTION_GET_BUCKET_QUOTA` |
| `weed/s3api/s3_constants/s3_actions.go` | Add coarse-grained `ACTION_PUT_BUCKET_QUOTA` and `ACTION_GET_BUCKET_QUOTA` |
| `weed/s3api/s3_action_resolver.go` | Map `seaweedfs-quota` query param to fine-grained s3: actions |
| `weed/s3api/s3api_bucket_quota_handlers.go` | New — `PutBucketQuotaHandler` and `GetBucketQuotaHandler` |
| `weed/s3api/s3api_bucket_quota_handlers_test.go` | New — tests for unit conversion, validation, and error paths |
| `weed/s3api/s3api_server.go` | Register the two routes in the bucket subrouter |
+6
View File
@@ -14,6 +14,9 @@ RUN cd /go/src/github.com/seaweedfs/seaweedfs && \
git checkout $BRANCH) || \
(echo "ERROR: Branch/commit $BRANCH not found in repository" && \
echo "Available branches:" && git branch -a && exit 1))
# seaweed-common only exists on revisions that have it; a BRANCH predating it
# still needs the directory so the COPY into rust_builder below never fails.
RUN mkdir -p /go/src/github.com/seaweedfs/seaweedfs/seaweed-common
ARG TARGETOS TARGETARCH TARGETVARIANT
RUN cd /go/src/github.com/seaweedfs/seaweedfs/weed \
&& export LDFLAGS="-X github.com/seaweedfs/seaweedfs/weed/util/version.COMMIT=$(git rev-parse --short HEAD)" \
@@ -31,6 +34,9 @@ ARG TAGS
COPY weed-volume-prebuilt/ /prebuilt/
COPY weed-worker-prebuilt/ /prebuilt-worker/
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-volume /build/seaweed-volume
# seaweed-common is a path dependency of seaweed-volume that lives beside it,
# so the source build below needs it in the same relative position.
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-common /build/seaweed-common
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/weed /build/weed
WORKDIR /build/seaweed-volume
RUN if [ -f "/prebuilt/weed-volume-${TARGETARCH}" ]; then \
+5 -2
View File
@@ -1,11 +1,14 @@
FROM alpine:latest
# Install required packages
RUN apk add --no-cache \
RUN apk upgrade --no-cache && \
apk add --no-cache \
ca-certificates \
fuse \
curl \
jq
jq \
libcrypto3 \
libssl3
# Copy our locally built binary
COPY weed-local /usr/bin/weed
+2
View File
@@ -30,3 +30,5 @@ sleep_minutes = 17 # sleep minutes between each script execution
bucket = "volume_bucket" # an existing bucket
endpoint = "http://server2:8333"
storage_class = "STANDARD_IA"
# upload_concurrency = 5 # concurrent multipart part uploads per volume (volume.tier.upload -concurrent overrides)
# download_concurrency = 5 # concurrent multipart part downloads per volume (volume.tier.download -concurrent overrides)
+55 -55
View File
@@ -1,6 +1,6 @@
module github.com/seaweedfs/seaweedfs
go 1.26
go 1.26.6
require (
cloud.google.com/go v0.123.0 // indirect
@@ -14,7 +14,7 @@ require (
github.com/coreos/go-semver v0.3.1 // indirect
github.com/coreos/go-systemd/v22 v22.7.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/dustin/go-humanize v1.0.1
github.com/dustin/go-humanize v1.1.0
github.com/eapache/go-resiliency v1.6.0 // indirect
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 // indirect
github.com/eapache/queue v1.1.0 // indirect
@@ -25,7 +25,7 @@ require (
github.com/facebookgo/subset v0.0.0-20200203212716-c811ad88dec4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/go-redsync/redsync/v4 v4.17.0
github.com/go-sql-driver/mysql v1.10.0
github.com/go-sql-driver/mysql v1.10.1
github.com/go-zookeeper/zk v1.0.4 // indirect
github.com/golang/protobuf v1.5.4
github.com/golang/snappy v1.0.0
@@ -65,12 +65,12 @@ require (
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/seaweedfs/goexif v2.0.0+incompatible
github.com/seaweedfs/raft v1.2.0
github.com/seaweedfs/raft v1.2.1
github.com/sirupsen/logrus v1.9.4 // indirect
github.com/spf13/afero v1.15.0 // indirect
github.com/spf13/cast v1.10.0 // indirect
github.com/spf13/viper v1.21.0
github.com/stretchr/testify v1.11.1
github.com/stretchr/testify v1.12.1
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203
github.com/syndtr/goleveldb v1.0.1-0.20190318030020-c3a204f8e965
github.com/tidwall/gjson v1.18.0
@@ -90,18 +90,18 @@ require (
gocloud.dev v0.46.0
gocloud.dev/pubsub/natspubsub v0.46.0
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.55.0
golang.org/x/crypto v0.57.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.45.0
golang.org/x/image v0.46.0
golang.org/x/net v0.58.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/oauth2 v0.37.0
golang.org/x/sys v0.48.0
golang.org/x/text v0.42.0 // indirect
golang.org/x/tools v0.49.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.296.0
google.golang.org/api v0.297.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7
google.golang.org/protobuf v1.36.12
gopkg.in/inf.v0 v0.9.1 // indirect
modernc.org/b v1.0.0 // indirect
@@ -111,7 +111,7 @@ require (
)
require (
cloud.google.com/go/kms v1.33.0
cloud.google.com/go/kms v1.35.0
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0
github.com/DATA-DOG/go-sqlmock v1.5.2
github.com/Jille/raft-grpc-transport v1.6.1
@@ -122,10 +122,10 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2 v1.47.0
github.com/aws/aws-sdk-go-v2/config v1.32.35
github.com/aws/aws-sdk-go-v2/credentials v1.20.1
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.0
github.com/aws/aws-sdk-go-v2/credentials v1.20.4
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
github.com/fluent/fluent-logger-golang v1.10.1
@@ -144,23 +144,23 @@ require (
github.com/parquet-go/parquet-go v0.32.0
github.com/pkg/sftp v1.13.11
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rclone/rclone v1.75.0
github.com/rclone/rclone v1.75.1
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/redis/go-redis/v9 v9.22.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.7
github.com/tarantool/go-option v1.1.0
github.com/tarantool/go-tarantool/v3 v3.0.1
github.com/tarantool/go-tarantool/v3 v3.0.2
github.com/testcontainers/testcontainers-go v0.44.0
github.com/tikv/client-go/v2 v2.0.7
github.com/twmb/avro v1.8.0
github.com/twmb/avro v1.9.0
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
go.uber.org/atomic v1.12.0
golang.org/x/sync v0.23.0
golang.org/x/tools/godoc v0.1.0-deprecated
google.golang.org/grpc/security/advancedtls v1.0.0
)
@@ -185,7 +185,7 @@ require (
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
@@ -262,8 +262,8 @@ require (
github.com/pquerna/otp v1.5.0 // indirect
github.com/pterm/pterm v0.12.83 // indirect
github.com/quic-go/qpack v0.6.0 // indirect
github.com/rclone/Proton-API-Bridge v1.0.4 // indirect
github.com/rclone/go-proton-api v1.0.3 // indirect
github.com/rclone/Proton-API-Bridge v1.0.5 // indirect
github.com/rclone/go-proton-api v1.0.4 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd // indirect
github.com/ryanuber/go-glob v1.0.0 // indirect
@@ -283,34 +283,34 @@ require (
github.com/xeipuuv/gojsonreference v0.0.0-20180127040603-bd5ef7bd5415 // indirect
github.com/xo/terminfo v0.0.0-20220910002029-abceb7e1c41e // indirect
github.com/zeebo/xxh3 v1.1.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/mock v0.5.2 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/mod v0.38.0 // indirect
go.yaml.in/yaml/v3 v3.0.5 // indirect
golang.org/x/mod v0.41.0 // indirect
gonum.org/v1/gonum v0.17.0 // indirect
)
require (
cel.dev/expr v0.25.2 // indirect
cel.dev/expr v0.25.3 // indirect
cloud.google.com/go/auth v0.23.2 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
cloud.google.com/go/iam v1.12.0 // indirect
cloud.google.com/go/monitoring v1.30.0 // indirect
filippo.io/edwards25519 v1.2.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.8.0
github.com/Azure/azure-sdk-for-go/sdk/storage/azfile v1.7.0 // indirect
github.com/Azure/go-ntlmssp v0.1.1 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0 // indirect
github.com/Files-com/files-sdk-go/v3 v3.3.194 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/resourcemapping v0.57.0 // indirect
github.com/IBM/go-sdk-core/v5 v5.23.1 // indirect
@@ -326,21 +326,21 @@ require (
github.com/andybalholm/cascadia v1.3.4 // indirect
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.16 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.28 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.36 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0
github.com/aws/smithy-go v1.28.1
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
@@ -363,7 +363,7 @@ require (
github.com/elastic/gosigar v0.14.3 // indirect
github.com/emersion/go-message v0.18.2 // indirect
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.37.0 // indirect
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be // indirect
github.com/envoyproxy/protoc-gen-validate v1.3.3 // indirect
github.com/fatih/color v1.18.0 // indirect
github.com/felixge/httpsnoop v1.1.0 // indirect
@@ -388,7 +388,7 @@ require (
github.com/gogo/protobuf v1.3.2 // indirect
github.com/golang-jwt/jwt/v4 v4.5.2 // indirect
github.com/google/s2a-go v0.1.9 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.20 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.21 // indirect
github.com/gorilla/schema v1.4.1 // indirect
github.com/gorilla/securecookie v1.1.2 // indirect
github.com/gorilla/sessions v1.4.0
@@ -438,7 +438,7 @@ require (
github.com/oracle/oci-go-sdk/v65 v65.121.0 // indirect
github.com/panjf2000/ants/v2 v2.12.1 // indirect
github.com/patrickmn/go-cache v2.1.0+incompatible // indirect
github.com/pelletier/go-toml/v2 v2.4.1 // indirect
github.com/pelletier/go-toml/v2 v2.4.3 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.29
@@ -489,9 +489,9 @@ require (
go.etcd.io/bbolt v1.5.0 // indirect
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 // indirect
go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.45.0 // indirect
@@ -499,9 +499,9 @@ require (
go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
golang.org/x/term v0.46.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
+108 -108
View File
@@ -6,8 +6,8 @@ atomicgo.dev/keyboard v0.2.9 h1:tOsIid3nlPLZ3lwgG8KZMp/SFmr7P0ssEN5JUsm78K8=
atomicgo.dev/keyboard v0.2.9/go.mod h1:BC4w9g00XkxH/f1HXhW2sXmJFOCWbKn9xrOunSFtExQ=
atomicgo.dev/schedule v0.1.0 h1:nTthAbhZS5YZmgYbb2+DH8uQIZcTlIrd4eYr3UQxEjs=
atomicgo.dev/schedule v0.1.0/go.mod h1:xeUa3oAkiuHYh8bKiQBRojqAMq3PXXbJujjb0hw8pEU=
cel.dev/expr v0.25.2 h1:K6j46C81hXtZQfuX60cVWQFBJahKSE2gfRbNuvr5bFs=
cel.dev/expr v0.25.2/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cel.dev/expr v0.25.3 h1:A2jO8jwOugrrovveCWfj0KEZOfqiLgAcwjpHPhzIGw0=
cel.dev/expr v0.25.3/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4=
cloud.google.com/go v0.26.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.34.0/go.mod h1:aQUYkXzVsufM+DwF1aE+0xfcU+56JwCaLick0ClmMTw=
cloud.google.com/go v0.38.0/go.mod h1:990N+gfupTy94rShfmMCWGDn0LpTmnzTp2qbd1dvSRU=
@@ -298,8 +298,8 @@ cloud.google.com/go/kms v1.4.0/go.mod h1:fajBHndQ+6ubNw6Ss2sSd+SWvjL26RNo/dr7uxs
cloud.google.com/go/kms v1.5.0/go.mod h1:QJS2YY0eJGBg3mnDfuaCyLauWwBJiHRboYxJ++1xJNg=
cloud.google.com/go/kms v1.6.0/go.mod h1:Jjy850yySiasBUDi6KFUwUv2n1+o7QZFyuUJg6OgjA0=
cloud.google.com/go/kms v1.9.0/go.mod h1:qb1tPTgfF9RQP8e1wq4cLFErVuTJv7UsSC915J8dh3w=
cloud.google.com/go/kms v1.33.0 h1:pG0X78m212b2pv9N4fdMoUO69LuZGQ9kSvn8sHBOFAo=
cloud.google.com/go/kms v1.33.0/go.mod h1:CSGvW6GnMQbY+1nOHcIzhMtHSbExXlOmCKjWtYVjcpA=
cloud.google.com/go/kms v1.35.0 h1:nJ/ktaqspx1nPM9vIcO0SHbhqCAm8nvAxL1siuVgKm0=
cloud.google.com/go/kms v1.35.0/go.mod h1:0++71pIHvJL+GmMa8K4jOWFq7gNOX3jm2PRMSJwTKJw=
cloud.google.com/go/language v1.4.0/go.mod h1:F9dRpNFQmJbkaop6g0JhSBXCNlO90e1KWx5iDdxbWic=
cloud.google.com/go/language v1.6.0/go.mod h1:6dJ8t3B+lUYfStgls25GusK04NLh3eDLQnWM3mdEbhI=
cloud.google.com/go/language v1.7.0/go.mod h1:DJ6dYN/W+SQOjF8e1hLQXMF21AkH2w9wiPzPCJa2MIE=
@@ -553,10 +553,10 @@ gioui.org v0.0.0-20210308172011-57750fc8a0a6/go.mod h1:RSH6KIUZ0p2xy5zHDxgAM4zum
git.sr.ht/~sbinet/gg v0.3.1/go.mod h1:KGYtlADtqsqANL9ueOFkWymvzUvLMQllU5Ixo+8v3pc=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6 h1:He8afgbRMd7mFxO99hRNu+6tazq8nFF9lIwo9JFroBk=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6/go.mod h1:8o94RPi1/7XTJvwPpRSzSUedZrtlirdB3r9Z20bi2f8=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0 h1:aokoqcHvaGjiM3VpjKDfMMnF/8epJ+Q1HLJ7CudztqE=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0/go.mod h1:/WYEx9pcM9Y+Dd/APJaNlSvVSvzl54rrMdZT5+Oi2LM=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0 h1:CU4+EJeJi3TKYWEcYuSdWsjzw0nVsK/H0MSQOiPcymU=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0/go.mod h1:q0+UTSRvShwUCrR/s5HtyInYphN7Wvxb7snFM3u+SLA=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1 h1:zvXfGJCWvywnCA814d8ZiVyt+fm9nnTE8xSb99zRyfo=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.23.1/go.mod h1:iptorS+VYKFL2N6PnebpS91dubG35eAOEERnT4PJbQU=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1 h1:u93s+zU2JD62im61Bm5CZIc1ZrOJaIAWEg0WOrMVkEo=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.1/go.mod h1:oXtinPO4OLj9d1DOTrqrL1oRwGhcqadvAmrl6wTeGlk=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.4.0 h1:xFaZZ+IubdftrDHnGGwZ6QvQ3KHTtWl2MCK+GMt2vxs=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.4.0/go.mod h1:mCBhUhlMjLLJKr5aqw2TNS/VqJOie8MzWq3DAMJeKso=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 h1:fhqpLE3UEXi9lPaBRpQ6XuRW0nU7hgg4zlmZZa+a9q4=
@@ -577,8 +577,8 @@ github.com/Azure/go-ntlmssp v0.1.1 h1:l+FM/EEMb0U9QZE7mKNEDw5Mu3mFiaa2GKOoTSsNDP
github.com/Azure/go-ntlmssp v0.1.1/go.mod h1:NYqdhxd/8aAct/s4qSYZEerdPuH1liG2/X9DiVTbhpk=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1 h1:WJTmL004Abzc5wDB5VtZG2PJk5ndYDgVacGqfirKxjM=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1/go.mod h1:tCcJZ0uHAmvjsVYzEFivsRTN00oz5BEsRgQHu5JZ9WE=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 h1:RHK7bS+HQMslb1sZpAokUt+zTVmue0hKSs2C791hhzU=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2/go.mod h1:HKpQxkWaGLJ+D/5H8QRpyQXA1eKjxkFlOMwck5+33Jk=
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0 h1:Nljr4q1GRA/5vCrMONS+g4u4LRHNgOXVSh3O43J2CnI=
github.com/AzureAD/microsoft-authentication-library-for-go v1.8.0/go.mod h1:Y33QHnf0FfdVewFFISOGe20mkZbxX4H839o955/PoeI=
github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU=
github.com/BurntSushi/xgb v0.0.0-20160522181843-27f122750802/go.mod h1:IVnqGOEym/WlBOVXweHU+Q+/VP0lqqI8lqeDx9IjBqo=
github.com/Codefor/geohash v0.0.0-20140723084247-1b41c28e3a9d h1:iG9B49Q218F/XxXNRM7k/vWf7MKmLIS8AcJV9cGN4nA=
@@ -593,8 +593,8 @@ github.com/FilenCloudDienste/filen-sdk-go v0.0.39 h1:tgV5jYL6dsXop9TpDTIQU6UwJjw
github.com/FilenCloudDienste/filen-sdk-go v0.0.39/go.mod h1:0cBhKXQg49XbKZZfk5TCDa3sVLP+xMxZTWL+7KY0XR0=
github.com/Files-com/files-sdk-go/v3 v3.3.194 h1:dtOFxSTWWRpkmvXa6ycNiw8dVDu1wkgzcXyVV1VafNc=
github.com/Files-com/files-sdk-go/v3 v3.3.194/go.mod h1:rl0WumSN9gSo775DgvQv+wMQ8rlb0ES/1hU5jkMtLXg=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0 h1:yzIYdwuro811Z27D3T80Wkd3rqZzb0K43nner7Eh1yE=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.34.0/go.mod h1:pJTkW8hEUIIi3Pf65lPZOnn4Y81yCllX6IWk2jNXdkM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0 h1:bN1gA3of5bXtbnLsRPrwfmbbe7A5UWFlcTHseujLnpc=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.35.0/go.mod h1:Yj5vHEz/aAepZGliRJsA6uvHAVAQyEwajq9ORCHPxzM=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0 h1:jLdiS1vO+XJFyDSWRHBx56r4s/NNtcl5J6KyCcWUX/w=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.57.0/go.mod h1:8lmpHY+1VRoteiOwyrQMDt1YGXOrFKCz+1wJW7n3ODY=
github.com/GoogleCloudPlatform/opentelemetry-operations-go/internal/cloudmock v0.57.0 h1:cSjUzZ7KU8hicTgzaSv9NmSyM9fTVK3y5lsBUl3wOis=
@@ -710,48 +710,48 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.16 h1:aiuaKlDweRC5qExJondpWjOgyzMHpofpwspGXUtwn4c=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.16/go.mod h1:nG/LOlmox9BDe9HvQnXWzgcK8uKbgBMZ/Hp5pVt/21I=
github.com/aws/aws-sdk-go-v2 v1.47.0 h1:0jsHallhJCeaU0Ko48c/3FK1ctOQ7NpzggxriJOQ8MQ=
github.com/aws/aws-sdk-go-v2 v1.47.0/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18 h1:LAfOuhAH331fmOjTQpAaOlH+Ftn7RzSDJ2VFwjdMMy4=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.18/go.mod h1:4e5xhuXHx1e4U9EthvbPP1r/DIMp5c2823OL8karzcM=
github.com/aws/aws-sdk-go-v2/config v1.32.35 h1:UEzXuET8E42lxBPijuACu/tEK7v5lFPlk0Q+GT5WD9E=
github.com/aws/aws-sdk-go-v2/config v1.32.35/go.mod h1:KaMtJpFa2JlL2BStjjHQVwQpzZEmw+ND/EgVrfFoo2g=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1 h1:Z8GRNEx0u9sDkZOq4PUnN8mjGwbUQGRzMSXpvt3d8xQ=
github.com/aws/aws-sdk-go-v2/credentials v1.20.1/go.mod h1:uBIK00kFo95dnemqfFMTWx0X8YRqsh6ecIoCjjOkZqM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1 h1:YIEBqcqRnpi4Pfv0YHImtgi6czGCwKHANC7SwmUAVD0=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.19.1/go.mod h1:imEf0oufgAo8KAkCHhrOdqGEC0YWx1PPBQH82shSxGw=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4 h1:hTvrJJseKbvw32kmiE0G+u/9ZqpqscjDrTigHIXP2qs=
github.com/aws/aws-sdk-go-v2/credentials v1.20.4/go.mod h1:gWp9O1ZBWwpcIrgV+mVHk4gZUurAEDkgypu/OXOlIaw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0 h1:AM4hHjww+PSFtt6E+UrBrPlZkWsePCLEt9AjkfQX+yM=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.0/go.mod h1:3x/yXezeQjpOvBb4jEMxrS8SXvpdvJ5abv6l5c1gWM8=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11 h1:eBXB8KZgzQ8A9QB4iJS4aw/u6+4OY3i2hQXPABeAIOg=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.3.11/go.mod h1:N9+5pG27Fy61GUL5YXVLXDTLmUudMrgwsuDbgBMNLxQ=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1 h1:pc138gM1CW+XPc60rEwUlwwuwWFQK16CI1T7v1F9Oec=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.1/go.mod h1:1+koxpPIbfBdfzP6vojm5/zTpTQ/micYwlxIiNB3TxI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1 h1:K0JsbZQj+1h208Ro1zHeA4l7bMp0NvRffHQ91q8Ol1s=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.1/go.mod h1:W3/vL6EtCIatICGy9ab29QhMuae+cOKPWcMxv02CO+Q=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1 h1:yhw5KD1phVyP9vijxOUzDfEtJx+bt+L63k+VfuiYFAA=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.1/go.mod h1:ZW2e0d7DYlRxlS9hEiMXE47gTdX5KRN4byUiNbUpG+Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3 h1:Hp/VgjP0BysR3OgLlR057Vz2LcbbVnoWeJ+3qWiS/fY=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.3/go.mod h1:nwGV5qw7F1IZPgxCvA/ph8N2TAuz+BkRG/bXn808qMA=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3 h1:MUaM4f+kj1ZIBPZfUS8cxP1GKXXZtHJjAthy93AN7SM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.3/go.mod h1:6YmVmEVRI5ZZzRjCSsb9SryKH0hAlMRdgA7kG9aDvBU=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3 h1:fuSCw4Z2qfRCztMPO3GXJNSiEp6Wee+WOLwrHHUMy9c=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.3/go.mod h1:6SxcHheD1pPR5+kWm1wGvjlL/YqUsh267sAfEmN4K7A=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.28 h1:Q1TF1J9jVD+vFo0LzNnmNdQ9EAt52TS+MQlq9Ir+Yxo=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.28/go.mod h1:4KqXXC/p1hrotmouDFbrRoWaLy962b9PMUReCG6+uWo=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1 h1:RmmWQPREQdk9U+PfqeHW3MqZaBaNK7TpV9W3RY+b+7g=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.1/go.mod h1:0A3W4F+68ZnNk5XcNL/e9HFMwnP8RlEicFfy6eOEDyw=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.36 h1:EUIwBoN+q7UmhAejxgD27APiRjh1vwCFo53gSqdT0BM=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.36/go.mod h1:6u00gmlTGR6W0b2k9NBrld7MnOEmf1Spqx0VVt6AqyE=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.0 h1:OkYV+1171za+ab9otU1tGxMXhx6uZvwVEtVddjLuYTg=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.0/go.mod h1:5FTZoQxhmLEiCAtYVk6V+t0iS/B5yGZVLZ3Wq5FDJZI=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1 h1:mdMtSVKdQ3+mzBh+l0ogrFYZVQUCg6pJZOirA2ARsYE=
github.com/aws/aws-sdk-go-v2/service/signin v1.7.1/go.mod h1:9IqUlsJDbUPcg6cgx3WEzXdjrbWzLDQrak0aaSqlTcI=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31 h1:uZOinZb+h7lZw8IYzP1z1IuEnueB76/EFkcf/fEW4Ag=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.31/go.mod h1:NRtwAM/p5VRt03TlEUs0pH3TeWamWdf4YyJpSrzPYLc=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3 h1:bON1rJf67TSTDCKg816AAIE4xSTtoo9tl0XRkO72R+I=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.3/go.mod h1:c5BBpjJcQXpfeq9iASyVKA3T6vX6B6LEXY4mL/gklDY=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39 h1:HLPAVrlLDaN2boN0xJx7MgaQDNEO3Q+c9L6kl/8m47Q=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.39/go.mod h1:Pg/dVfsNkm1hsIDK/gMvCKtmyNfNTV12mrgHqVE/6Oo=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3 h1:IKoCZqfWfZzSBi16QFQ+QcbQ3LRQ7QgB1S5tDAyPBQQ=
github.com/aws/aws-sdk-go-v2/service/s3 v1.107.3/go.mod h1:RBpRcXiM4s2pOInVs32GsBonnje+fiAj4mcrStRmlCA=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0 h1:ZD5qFpWcaOKdTuhBi431pIDkCgrMkMlMT6jlpSPoIRI=
github.com/aws/aws-sdk-go-v2/service/signin v1.10.0/go.mod h1:8Nuuf+tR346PjJ3MvZPh9pekbLiLQFWJhzMXfwy7alA=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1 h1:B6WFn91tobD6gG4724ONHaqrpKsoETGnv98LHe/yIGM=
github.com/aws/aws-sdk-go-v2/service/sso v1.35.1/go.mod h1:tWuiVBUtPBr8/rgRiYS8Uf85sHcAN+G7XS3D3CEoUh8=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1 h1:6yeYCWFvgbI2TI3K6jr9LtBNhXgJ7g4xqD+DEiaDDmM=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.40.1/go.mod h1:naFe83jSMuYkH+QjQPX8n1MLhBkeCFM5Lsnh5m5wz3c=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1 h1:Sv2xPnRHlThSUtVujYuUBPI/Il8si6UPHXL8DMiB/F0=
github.com/aws/aws-sdk-go-v2/service/sts v1.47.1/go.mod h1:mKo/CzaCz8qytGW70NG4vIIGAx1HXTlb5lHNkC5k3lk=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0 h1:JGeeBcMlhg1xtOXYpeCaTQBZObtXMPQCUqBcmr65NRA=
github.com/aws/aws-sdk-go-v2/service/sso v1.38.0/go.mod h1:XwteswG9EOMRFm73UT0t+MbTwyLxMrEXkU6e+v92Lzo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0 h1:obhahQXDEdVEv8y5bTKXR30LVaxYe1kyYM0L7l2Iq+k=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.0/go.mod h1:6twZZ/aXHNy1vXUO8koUbp++MYzMASkOgEBdkbJYmO0=
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0 h1:Zpnqa6XtrNzXZnwbdCqHOXpXhMsa01ql/pcRQ1sb4hk=
github.com/aws/aws-sdk-go-v2/service/sts v1.51.0/go.mod h1:/8JRcdTt//hG0Q4BTmGbuOplT7ABe+5rdtqUHqXvYIM=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
@@ -953,8 +953,8 @@ github.com/dropbox/dropbox-sdk-go-unofficial/v6 v6.4.0/go.mod h1:gDXhl0OElhzYoDs
github.com/dsnet/try v0.0.3 h1:ptR59SsrcFUYbT/FhAbKTV6iLkeD6O18qfIWRml2fqI=
github.com/dsnet/try v0.0.3/go.mod h1:WBM8tRpUmnXXhY1U6/S8dt6UWdHTQ7y8A5YSkRCkq40=
github.com/dustin/go-humanize v1.0.0/go.mod h1:HtrtbFcZ19U5GC7JDqmcUSB87Iq5E25KnS6fMYU6eOk=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/dustin/go-humanize v1.1.0 h1:dbKTrvD0klcbBV/h4AWJdMuZogJACoMlvWIWZ5b2xWg=
github.com/dustin/go-humanize v1.1.0/go.mod h1:hc1CvRkJMsgxqjmjMQF3QNRAZBwY8AXBAzKYoSX9sFI=
github.com/eapache/go-resiliency v1.6.0 h1:CqGDTLtpwuWKn6Nj3uNUdflaq+/kIPsg0gfNzHton30=
github.com/eapache/go-resiliency v1.6.0/go.mod h1:5yPzW0MIvSe0JDsv0v+DvcjEv2FyD6iZYSs1ZI+iQho=
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 h1:Oy0F4ALJ04o5Qqpdz8XLIpNA3WM/iSIXqxtqo7UGVws=
@@ -987,8 +987,8 @@ github.com/envoyproxy/go-control-plane v0.10.3/go.mod h1:fJJn/j26vwOu972OllsvAgJ
github.com/envoyproxy/go-control-plane v0.11.0/go.mod h1:VnHyVMpzcLvCFt9yUz1UnCwHLhwx1WguiVDV7pTG/tI=
github.com/envoyproxy/go-control-plane v0.14.0 h1:hbG2kr4RuFj222B6+7T83thSPqLjwBIfQawTkC++2HA=
github.com/envoyproxy/go-control-plane v0.14.0/go.mod h1:NcS5X47pLl/hfqxU70yPwL9ZMkUlwlKxtAohpi2wBEU=
github.com/envoyproxy/go-control-plane/envoy v1.37.0 h1:u3riX6BoYRfF4Dr7dwSOroNfdSbEPe9Yyl09/B6wBrQ=
github.com/envoyproxy/go-control-plane/envoy v1.37.0/go.mod h1:DReE9MMrmecPy+YvQOAOHNYMALuowAnbjjEMkkWOi6A=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be h1:SWe0x6yfglnxuvOiYgTTnNq7QD/yvqthOh1RBI8Bj8w=
github.com/envoyproxy/go-control-plane/envoy v1.39.1-0.20260819172001-e6e3fd93e4be/go.mod h1:PYEOlng9XcrulfyWpm49jECTPV0LT4q8cO7fLW/xwgk=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0 h1:/G9QYbddjL25KvtKTv3an9lx6VBE2cnb8wp1vEGNYGI=
github.com/envoyproxy/go-control-plane/ratelimit v0.1.0/go.mod h1:Wk+tMFAFbCXaJPzVVHnPgRKdUdwW/KdbRt94AzgRee4=
github.com/envoyproxy/protoc-gen-validate v0.1.0/go.mod h1:iSmxcyjqTsJpI2R4NaDN7+kN2VEUnK/pcBlmesArF7c=
@@ -1109,8 +1109,8 @@ github.com/go-redsync/redsync/v4 v4.17.0 h1:FFJ+uxZs44y4Sq10//IFKic9T94AYl+u3Sog
github.com/go-redsync/redsync/v4 v4.17.0/go.mod h1:CKVA6qwT07S/916i+Yd9h1/8YFQhCCpPYTQhvvYytJo=
github.com/go-resty/resty/v2 v2.17.2 h1:FQW5oHYcIlkCNrMD2lloGScxcHJ0gkjshV3qcQAyHQk=
github.com/go-resty/resty/v2 v2.17.2/go.mod h1:kCKZ3wWmwJaNc7S29BRtUhJwy7iqmn+2mLtQrOyQlVA=
github.com/go-sql-driver/mysql v1.10.0 h1:Q+1LV8DkHJvSYAdR83XzuhDaTykuDx0l6fkXxoWCWfw=
github.com/go-sql-driver/mysql v1.10.0/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/go-stack/stack v1.8.0/go.mod h1:v0f6uXyyMGvRgIKkXu+yp6POWl0qKG85gN/melR3HDY=
github.com/go-task/slim-sprig v0.0.0-20230315185526-52ccab3ef572 h1:tfuBGBXKqDEevZMzYi5KSi8KkcZtzBcTgAUUtapy0OI=
github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI=
@@ -1262,8 +1262,8 @@ github.com/googleapis/enterprise-certificate-proxy v0.1.0/go.mod h1:17drOmN3MwGY
github.com/googleapis/enterprise-certificate-proxy v0.2.0/go.mod h1:8C0jb7/mgJe/9KK8Lm7X9ctZC2t60YyIpYEI16jx0Qg=
github.com/googleapis/enterprise-certificate-proxy v0.2.1/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.2.3/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.3.20 h1:t/xL64VUoN69MuMRQuJETqYGOw4Z9mSRJK9epIEtwFk=
github.com/googleapis/enterprise-certificate-proxy v0.3.20/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/enterprise-certificate-proxy v0.3.21 h1:OFdQ3tnCX/zaQ0Cedur3D3z7kI6HiLX9g3TiAN4/DFU=
github.com/googleapis/enterprise-certificate-proxy v0.3.21/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/gax-go/v2 v2.1.0/go.mod h1:Q3nei7sK6ybPYH7twZdmQpAd1MKb7pfu6SK+H1/DsU0=
@@ -1652,8 +1652,8 @@ github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144T
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
github.com/patrickmn/go-cache v2.1.0+incompatible/go.mod h1:3Qf8kWWT7OJRJbdiICTKqZju1ZixQ/KpMGzzAfe6+WQ=
github.com/pborman/getopt v0.0.0-20170112200414-7148bc3a4c30/go.mod h1:85jBQOZwpVEaDAr341tbn15RS4fCAsIst0qp7i8ex1o=
github.com/pelletier/go-toml/v2 v2.4.1 h1:j5OMOImsH+j2k7GJ5YO+RxfWwohNiH6t5zB/+h3bagc=
github.com/pelletier/go-toml/v2 v2.4.1/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pelletier/go-toml/v2 v2.4.3 h1:GTRvJQutkOSftxIFD5xw9aepkYNuPWmVJpffdDPYVpY=
github.com/pelletier/go-toml/v2 v2.4.3/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 h1:XeOYlK9W1uCmhjJSsY78Mcuh7MVkNjTzmHx1yBzizSU=
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14/go.mod h1:jVblp62SafmidSkvWrXyxAme3gaTfEtWwRPGz5cpvHg=
github.com/peterh/liner v1.2.2 h1:aJ4AOodmL+JxOZZEL2u9iJf8omNRpqHc/EbrK+3mAXw=
@@ -1760,18 +1760,18 @@ github.com/quic-go/quic-go v0.59.0 h1:OLJkp1Mlm/aS7dpKgTc6cnpynnD2Xg7C1pwL6vy/SA
github.com/quic-go/quic-go v0.59.0/go.mod h1:upnsH4Ju1YkqpLXC305eW3yDZ4NfnNbmQRCMWS58IKU=
github.com/rabbitmq/amqp091-go v1.14.0 h1:RSaT7aOKt/OrkVUyswPDW29lnRz9psuGmfZFBmLqLek=
github.com/rabbitmq/amqp091-go v1.14.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rclone/Proton-API-Bridge v1.0.4 h1:uGQJRjQC1hVLd5kqLsXc6CWO6oqrVeLoKQYoHapEZDg=
github.com/rclone/Proton-API-Bridge v1.0.4/go.mod h1:VTPBYZotKAeDLlAzxU2O/s14NXk9FxUt9hn1jhH2iY8=
github.com/rclone/go-proton-api v1.0.3 h1:3gBTzR+j0dYiTwtj9yKIdN/aV3W2a8KIPKp0GArojyQ=
github.com/rclone/go-proton-api v1.0.3/go.mod h1:QAlkFfswzrBuxvCORWV8rZdddg52hahMN98CFWoFW1E=
github.com/rclone/rclone v1.75.0 h1:3ARHem4jXWltvl+b0PvDAG8s6J/inHd5BRzfwMRb3W8=
github.com/rclone/rclone v1.75.0/go.mod h1:PGLJUW/WSIJCysALqUcxmaCFyfMXUevf8CbuoOwsAdU=
github.com/rclone/Proton-API-Bridge v1.0.5 h1:K1++Qtk3PvgkiCCiv6Pahju1TMOzKY6VSwiwT7XLAVc=
github.com/rclone/Proton-API-Bridge v1.0.5/go.mod h1:vCeOPhlXzevN0AFojgh1zsjhetiShy/ArvJ/xkFUDWk=
github.com/rclone/go-proton-api v1.0.4 h1:AJW0e9pB4j0hVK4WqyGErFwaI+5MUQWPCtj5FYYxtPg=
github.com/rclone/go-proton-api v1.0.4/go.mod h1:QAlkFfswzrBuxvCORWV8rZdddg52hahMN98CFWoFW1E=
github.com/rclone/rclone v1.75.1 h1:kIxQcoDLj2Gke/gMSHK7OnxhX1Gu1cJBLP1kJZoaFp0=
github.com/rclone/rclone v1.75.1/go.mod h1:4zmMjGatCkSJPRZDpo+7y3xOl8S29EMUyKvZop5mHr4=
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 h1:N/ElC8H3+5XpJzTSTfLsJV/mx9Q9g7kxmchpfZyxgzM=
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475/go.mod h1:bCqnVzQkZxMG4s8nGwiZ5l3QUCyqpo9Y+/ZMZ9VjZe4=
github.com/rdleal/intervalst v1.5.0 h1:SEB9bCFz5IqD1yhfH1Wv8IBnY/JQxDplwkxHjT6hamU=
github.com/rdleal/intervalst v1.5.0/go.mod h1:xO89Z6BC+LQDH+IPQQw/OESt5UADgFD41tYMUINGpxQ=
github.com/redis/go-redis/v9 v9.21.0 h1:FPBE4hhbAke+TLmcY3WkpbDffJEomdqPn3HYiqAtL9E=
github.com/redis/go-redis/v9 v9.21.0/go.mod h1:v/M13XI1PVCDcm01VtPFOADfZtHf8YW3baQf57KlIkA=
github.com/redis/go-redis/v9 v9.22.0 h1:laDvpYXTJtZLloinw1fA5Kqd6HAEH2XKxOkG/PDq2F0=
github.com/redis/go-redis/v9 v9.22.0/go.mod h1:y2g0Wj8rQvuK0ELM+oxSudcLtC09JScs98I/X9gRWY4=
github.com/redis/rueidis v1.0.76 h1:RdDWuvlYBSp+bTrBvaXqJnNEL3VVzsnjo+0psPFgLc4=
github.com/redis/rueidis v1.0.76/go.mod h1:UsfHPSbomB6QAVMk4iiFkzRy0nh9o7scDGa+SitvBY4=
github.com/redis/rueidis/rueidiscompat v1.0.76 h1:7LikbiqCQqCsZXeZ+akgZMnjIV/J0VHih9PIX4gGZC4=
@@ -1821,8 +1821,8 @@ github.com/seaweedfs/go-fuse/v2 v2.9.4 h1:ACyloiuopdhRSjdLLeSWbsVaemMPskORaRF01T
github.com/seaweedfs/go-fuse/v2 v2.9.4/go.mod h1:zABdmWEa6A0bwaBeEOBUeUkGIZlxUhcdv+V1Dcc/U/I=
github.com/seaweedfs/goexif v2.0.0+incompatible h1:x8pckiT12QQhifwhDQpeISgDfsqmQ6VR4LFPQ64JRps=
github.com/seaweedfs/goexif v2.0.0+incompatible/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/raft v1.2.0 h1:Ez4Hw9ifBbTT7wg54DvGHBjw1vRlTb4roH0TKl0Oj9Y=
github.com/seaweedfs/raft v1.2.0/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/seaweedfs/raft v1.2.1 h1:QgFl/aaPnagpUxYB6Bx+fFss1NyetVVcJmraMmnaQ5Q=
github.com/seaweedfs/raft v1.2.1/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/secure-systems-lab/go-securesystemslib v0.11.0 h1:iuCR9kcMFD4QurdKrGvPLoKZLv9YvwPYVr0473BdtFs=
github.com/secure-systems-lab/go-securesystemslib v0.11.0/go.mod h1:+PMOTjUGwHj2vcZ+TFKlb1tXRbrdWE1LYDT5i9JC80Q=
github.com/sergi/go-diff v1.0.0/go.mod h1:0CfEIISq7TuYL3j771MWULgwwjU+GofnZX9QAmXWZgo=
@@ -1897,8 +1897,8 @@ github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o
github.com/stretchr/testify v1.8.2/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4=
github.com/stretchr/testify v1.8.3/go.mod h1:sz/lmYIOXD/1dqDmKjjqLyZ2RngseejIcXlSw2iwfAo=
github.com/stretchr/testify v1.8.4/go.mod h1:sz/lmYIOXD/1dqDmKjjqLyZ2RngseejIcXlSw2iwfAo=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/stretchr/testify v1.12.1 h1:EuwCh5fleGS7H32xRwO3wRGT7DxrDhLAT6FF8MpWDWE=
github.com/stretchr/testify v1.12.1/go.mod h1:MDEgiDPPsNp5cuIrHPPCyornHKgEVbtFUmoNlxoYthg=
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203 h1:QVqDTf3h2WHt08YuiTGPZLls0Wq99X9bWd0Q5ZSBesM=
github.com/stvp/tempredis v0.0.0-20181119212430-b82af8480203/go.mod h1:oqN97ltKNihBbwlX8dLpwxCl3+HnXKV/R0e+sRLd9C8=
github.com/subosito/gotenv v1.6.0 h1:9NlTDc1FTs4qu0DDq7AEtTPNw6SVm7uBMsUCUjABIf8=
@@ -1918,8 +1918,8 @@ github.com/tarantool/go-iproto v1.1.0 h1:HULVOIHsiehI+FnHfM7wMDntuzUddO09DKqu2Wn
github.com/tarantool/go-iproto v1.1.0/go.mod h1:LNCtdyZxojUed8SbOiYHoc3v9NvaZTB7p96hUySMlIo=
github.com/tarantool/go-option v1.1.0 h1:ShoOhNsdL41sRpm4hXCRDjV8H0WzPkd4UnKhLKbW//w=
github.com/tarantool/go-option v1.1.0/go.mod h1:hMr9z2JXOWlgdCBpCPSL2nwp8718GKYvNBJ+ZuzJbCo=
github.com/tarantool/go-tarantool/v3 v3.0.1 h1:vaUX4xmVmXh2dIJ/LqlX1MXK3iYqAqV6YiE54Wwl/qg=
github.com/tarantool/go-tarantool/v3 v3.0.1/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/tarantool/go-tarantool/v3 v3.0.2 h1:9ZtHllun80QX7KS9tqd3dXa+QAu2BfqFtX1ZWcNnaW8=
github.com/tarantool/go-tarantool/v3 v3.0.2/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/testcontainers/testcontainers-go v0.44.0 h1:/Fwh6HY1mIikhnm9e7HwoxGycx0lzRAE0f5VQpjFxzI=
github.com/testcontainers/testcontainers-go v0.44.0/go.mod h1:IcnwQrYTO86xHXu5bvMaBH7ATlbS3Qn1M1QWW3c66rE=
github.com/testcontainers/testcontainers-go/modules/compose v0.44.0 h1:8YcW51jhgpkkiRVe10Wj9TCBthJmoNpU2fK5WSf7TQ8=
@@ -1965,8 +1965,8 @@ github.com/tsuna/gohbase v0.0.0-20201125011725-348991136365/go.mod h1:zj0GJHGvyf
github.com/tv42/httpunix v0.0.0-20150427012821-b75d8614f926/go.mod h1:9ESjWnEqriFuLhtthL60Sar/7RFoluCcXsuvEwTV5KM=
github.com/twitchyliquid64/golang-asm v0.15.1 h1:SU5vSMR7hnwNxj24w34ZyCi/FmDZTkS4MhqMhdFk5YI=
github.com/twitchyliquid64/golang-asm v0.15.1/go.mod h1:a1lVb/DtPvCB8fslRZhAngC2+aY1QWCk3Cedj/Gdt08=
github.com/twmb/avro v1.8.0 h1:UMWLg+nH4P3yad5Om7yFSohYLy2RG1s7BcFFiOvmK9Q=
github.com/twmb/avro v1.8.0/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/avro v1.9.0 h1:JSiqewo3AANj7vlVCQXBsIZTA3LiiBble7Xp8T4lbtA=
github.com/twmb/avro v1.9.0/go.mod h1:X0fT1dY2xcbV4YuCE4mYro+qljHl4kUF5uA/2z1rgSk=
github.com/twmb/murmur3 v1.1.8 h1:8Yt9taO/WN3l08xErzjeschgZU2QSrwm1kclYq+0aRg=
github.com/twmb/murmur3 v1.1.8/go.mod h1:Qq/R7NUyOfr65zD+6Q5IHKsJLwP7exErjN6lyyq3OSQ=
github.com/twpayne/go-geom v1.6.1 h1:iLE+Opv0Ihm/ABIcvQFGIiFBXd76oBIar9drAwHFhR4=
@@ -2103,30 +2103,30 @@ go.opencensus.io v0.24.0 h1:y73uSU6J157QMP2kn2r30vwW1A2W2WFwSCGnAVxeaD0=
go.opencensus.io v0.24.0/go.mod h1:vNK8G9p7aAivkbmorf4v+7Hgx+Zs0yY+0fOtgBfjQKo=
go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64=
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 h1:NmLfL734pJhM0JKaYd2Y28+nY9dPRWYAAbxhRCrKXPw=
go.opentelemetry.io/contrib/detectors/gcp v1.44.0/go.mod h1:tNAsgd8avTGke1+MndXlU5Cru4PQ9Ai/cCNWQv/ZJ/s=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0 h1:9jR0ZPRok9ryaOQ2Wx8rg5F7Aon59mxrqbVI60/vlBk=
go.opentelemetry.io/contrib/detectors/gcp v1.45.0/go.mod h1:VSme3o2fvSg5bVg0dRzyHaj4Z5EVhG+g2Fde6LKzmQA=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0 h1:2yEATaop1/a1I4psnSLgWVPLWwCzkqWakgJy7xTDVy0=
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.69.0/go.mod h1:D7J12YRapIekYyPWgGPlA/23pRmpSEZC5xJC/TTLI9U=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0 h1:MCcYL7J6Vt/X0kjqbMZkekCmwsurbQRbL69vkiye2lk=
go.opentelemetry.io/contrib/instrumentation/net/http/httptrace/otelhttptrace v0.69.0/go.mod h1:3jnStNwSufK+f5ktjL4EPcwtig4rtd81NS70lqHuXl8=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0 h1:LMuyCAyfalSjDyjdC65nK6N0zoTT63+E/u95X0JovZI=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.70.0/go.mod h1:085m8qbm4hgc8rZWGDEa4vmyyo2c3nPxUslYUKUIU04=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0 h1:SUplec5dp06reu1zaXmOXdvqH398taqrDXqUl99jxSc=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetricgrpc v1.44.0/go.mod h1:ho2g4N+ane+swq5I/VBkKWnRDY4kUINH3FuqyZqX/Ug=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0 h1:RuynHbfU8JUEw7DyONgkVYg2SVtsoF28y0LGIr69jgA=
go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp v1.44.0/go.mod h1:qZF+/lBs71APw8mlnEZcqZHMzqrYrsFiJOv83lX1OGo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0 h1:lgh3PiVrRUWMLOVSkQicxzZll5NjF1r+AtsX1XRIHw0=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp v1.44.0/go.mod h1:5Cnhth3m/AgOeTgE3ex12pPmiu/gGtZit03kSzx9X7s=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0 h1:hqxVTu/GtBF+vJ8d1fzW7fRxZFvgoDjWcxwwCaFDYpU=
go.opentelemetry.io/otel/exporters/stdout/stdoutmetric v1.44.0/go.mod h1:z5fVEF4X5v0ESvlJqBrrFlBVoj5EQuefZpzsu7R+x5Q=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0 h1:s0n95ya5tOG03exJ5JySOdJFtwGo4ZQ+KeY7Zro4CLI=
go.opentelemetry.io/otel/exporters/zipkin v1.36.0/go.mod h1:m9wRxtKA2MZ1HcnNC4BKI+9aYe434qRZTCvI7QGUN7Y=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0 h1:KN3btaILMTxR4QDHVGAO87lq5ButzK7l+kIfLuxQ1oA=
go.opentelemetry.io/otel/exporters/zipkin v1.45.0/go.mod h1:yNcodmUclM4InyWoOwX/YW4Jri0Gj5FWAlM+NqCrtqY=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/metric/x v0.67.0 h1:PcicCNZFkZ4bXfSooXdo3WN7RBOVOtjVdo1wD358Uns=
@@ -2140,13 +2140,13 @@ go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbi
go.opentelemetry.io/proto/otlp v0.7.0/go.mod h1:PqfVotwruBrMGOCsRd/89rSnXhoiJIqeYNgFYFoEGnI=
go.opentelemetry.io/proto/otlp v0.15.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v0.19.0/go.mod h1:H7XAot3MsfNsj7EXtrA2q5xSNQ10UqI405h3+duxN4U=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk=
go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/atomic v1.6.0/go.mod h1:sABNBOSYdrvTF6hTgEIbc7YasKWGhgEQZyfxyTvoXHQ=
go.uber.org/atomic v1.7.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
go.uber.org/atomic v1.9.0/go.mod h1:fEN4uk6kAWBTFdckzkM89CLk9XfWZrxpCo0nPH17wJc=
go.uber.org/atomic v1.11.0 h1:ZvwS0R+56ePWxUNi+Atn9dWONBPp/AUETXlHW0DxSjE=
go.uber.org/atomic v1.11.0/go.mod h1:LUxbIzbOniOlMKjJjyPfpl4v+PKK2cNJn91OQbhoJI0=
go.uber.org/atomic v1.12.0 h1:BvcXdFKuviU4fTL/f+SxdQ5qJX/Jix8pAkgdUcb3XOE=
go.uber.org/atomic v1.12.0/go.mod h1:I6c4cg+6HCxRjfjSsYtApoFILnpc0CGUdGkXVqbYVNk=
go.uber.org/goleak v1.1.10/go.mod h1:8a7PlsEVH3e/a/GLqe5IIrQx6GzcnRmZEufDUTk4A7A=
go.uber.org/goleak v1.1.12/go.mod h1:cwTWslyiVhfpKIDGSZEM2HlOvcqm+tG4zioyIeLoqMQ=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
@@ -2163,8 +2163,8 @@ go.uber.org/zap v1.27.1 h1:08RqriUEv8+ArZRYSTXy1LeBScaMpVSTBhCeaZYfMYc=
go.uber.org/zap v1.27.1/go.mod h1:GB2qFLM7cTU87MWRP2mPIjqfIDnGu+VIO4V/SdhGo2E=
go.yaml.in/yaml/v2 v2.4.4 h1:tuyd0P+2Ont/d6e2rl3be67goVK4R6deVxCUX5vyPaQ=
go.yaml.in/yaml/v2 v2.4.4/go.mod h1:gMZqIpDtDqOfM0uNfy0SkpRhvUryYH0Z6wdMYcacYXQ=
go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
go.yaml.in/yaml/v3 v3.0.5 h1:N6y/pJk8buWs9NY5ERU2HSMfm+IuD/OtfdAnq6kESPw=
go.yaml.in/yaml/v3 v3.0.5/go.mod h1:HVTZu1O7/Vkt2N+BFy8Zza+lnLsABggaTM2ZpNIGuKg=
go.yaml.in/yaml/v4 v4.0.0-rc.6 h1:1h7H1ohdUh93/FyE4YaDa1Zh64K6VVbjF4K6WUxMtH4=
go.yaml.in/yaml/v4 v4.0.0-rc.6/go.mod h1:aZqd9kCMsGL7AuUv/m/PvWLdg5sjJsZ4oHDEnfPPfY0=
gocloud.dev v0.46.0 h1:niIuZwSjMtBx8K+ITB2s5kZullB13PGOS2ZoQPZxQ4Q=
@@ -2193,8 +2193,8 @@ golang.org/x/crypto v0.6.0/go.mod h1:OFC/31mSvZgRz0V1QTNCzfAI1aIRzbiufJtkMIlEp58
golang.org/x/crypto v0.7.0/go.mod h1:pYwdfH91IfpZVANVyUOhSIPZaFoJGxTFbZhFTx+dXZU=
golang.org/x/crypto v0.13.0/go.mod h1:y6Z2r+Rw4iayiXXAIxJIDAJ1zMW4yaTpebo8fPOliYc=
golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf4=
golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M=
golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis=
golang.org/x/crypto v0.57.0 h1:3ZVCjf8Ggz7zneR/EHRVx68Ctf+2pmIMP2UFhh9cC6M=
golang.org/x/crypto v0.57.0/go.mod h1:Fdz0i5U6CoizGwLda9DttjSk6qlZo25zYNtR+ycvuZA=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2225,8 +2225,8 @@ golang.org/x/image v0.0.0-20210607152325-775e3b0c77b9/go.mod h1:023OzeP/+EPmXeap
golang.org/x/image v0.0.0-20210628002857-a66eb6448b8d/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20211028202545-6944b10bf410/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.0.0-20220302094943-723b81ca9867/go.mod h1:023OzeP/+EPmXeapQh35lcL3II3LrY8Ic+EFFKVhULM=
golang.org/x/image v0.45.0 h1:FMb1nTbH5H9vF55SriQHgFw5GnNL9Jg6L25BwXKzhB0=
golang.org/x/image v0.45.0/go.mod h1:n62x/7RqlwXDvGsSU4u6IUTUf6KghUZ9Bt7cG/T9Fx4=
golang.org/x/image v0.46.0 h1:b1+oYj0Jbp6K5MDT4i4/eZpYlk3V8SJhhDKh6LBHAyQ=
golang.org/x/image v0.46.0/go.mod h1:3B3W05VGVQyuXucLINLjXKrqISASfi4Xj+iCVkLMwew=
golang.org/x/lint v0.0.0-20181026193005-c67002cb31c3/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
golang.org/x/lint v0.0.0-20190227174305-5b3e6a55c961/go.mod h1:wehouNa3lNwaWXcvxsM5YxQ5yQlVC4a0KAMCusXpPoU=
golang.org/x/lint v0.0.0-20190301231843-5614ed5bae6f/go.mod h1:UVdnD1Gm6xHRNCYTkRU2/jEulfH38KcIWyp/GAMgvoE=
@@ -2258,8 +2258,8 @@ golang.org/x/mod v0.8.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.9.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.12.0/go.mod h1:iBbtSCu2XBx23ZKBPSOrRkjjQPZFPuis4dIYUhu/chs=
golang.org/x/mod v0.13.0/go.mod h1:hTbmBsO62+eylJbnUtE2MGJUyE7QWk4xUqPFrRgJ+7c=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c=
golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o=
golang.org/x/net v0.0.0-20180724234803-3673e40ba225/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180826012351-8a410e7b638d/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
@@ -2354,8 +2354,8 @@ golang.org/x/oauth2 v0.0.0-20221014153046-6fdb5e3db783/go.mod h1:h4gKUeWbJ4rQPri
golang.org/x/oauth2 v0.4.0/go.mod h1:RznEsdpjGAINPTOF0UH/t+xJ75L18YO3Ho6Pyn+uRec=
golang.org/x/oauth2 v0.5.0/go.mod h1:9/XBHVqLaWO3/BRHs5jbpYCnOZVjj5V0ndyaAM7KB4I=
golang.org/x/oauth2 v0.6.0/go.mod h1:ycmewcwgD4Rpr3eZJLSB4Kyyljb3qDh40vJ8STE5HKw=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/oauth2 v0.37.0 h1:JUlcxA8oAtauLfiH8FX2/FkAWHAdi0QtGCGc+hofE98=
golang.org/x/oauth2 v0.37.0/go.mod h1:IxwZNxUULJmpBFf9K/9NTMSIfZZuvuTy1gGxhigP/58=
golang.org/x/sync v0.0.0-20180314180146-1d60e4601c6f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20181108010431-42b317875d0f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20181221193216-37e7f081c4d4/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
@@ -2374,8 +2374,8 @@ golang.org/x/sync v0.0.0-20220929204114-8fcdb60fdcc0/go.mod h1:RxMgew5VJxzue5/jJ
golang.org/x/sync v0.1.0/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.3.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.4.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.0.0-20180810173357-98c5dad5d1a0/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180830151530-49385e6e1522/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180905080454-ebe1bf3edb33/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
@@ -2477,8 +2477,8 @@ golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.8.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.12.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.13.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210220032956-6a3ed077a48d/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210615171337-6886f2dfbf5b/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
@@ -2491,8 +2491,8 @@ golang.org/x/term v0.6.0/go.mod h1:m6U89DPEgQRMq3DNkDClhWw02AUbt2daBVO4cn4Hv9U=
golang.org/x/term v0.8.0/go.mod h1:xPskH00ivmX89bAKVGSKKtLOWNx2+17Eiy94tnKShWo=
golang.org/x/term v0.12.0/go.mod h1:owVbMEjm3cBLCHdkQu9b1opXd4ETQWc3BhuQGKgXgvU=
golang.org/x/term v0.13.0/go.mod h1:LTmsnFJwVN6bCy1rVCoS+qHT1HhALEFxKncY3WNNh4U=
golang.org/x/term v0.45.0 h1:NwWyBmoJCbfTHpxrWoZ9C6/VxOf7ic219I8xZZFdrf0=
golang.org/x/term v0.45.0/go.mod h1:9aqxs0blBcrm/n0L9QW0aRVD+ktan8ssZromtqJC43w=
golang.org/x/term v0.46.0 h1:3+OXuTbaKDgwk8jTi3aSLHRlmWqHEUDUtxnbFigO4YE=
golang.org/x/term v0.46.0/go.mod h1:+K02xbkittuwc0Am4abfA3Fc+XRGXkvBXNO88NCXPoc=
golang.org/x/text v0.0.0-20170915032832-14c0d48ead0c/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.1-0.20180807135948-17ff2d5776d2/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
@@ -2511,8 +2511,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI=
golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2589,8 +2589,8 @@ golang.org/x/tools v0.6.0/go.mod h1:Xwgl3UAJ/d3gWutnCtw505GrjyAbvKui8lOU390QaIU=
golang.org/x/tools v0.7.0/go.mod h1:4pg6aUX35JBAogB10C9AtvVL+qowtN4pT3CGSQex14s=
golang.org/x/tools v0.13.0/go.mod h1:HvlwmtVNQAhOuCjW7xxvovg8wbNq7LwfXh/k7wXUl58=
golang.org/x/tools v0.14.0/go.mod h1:uYBEerGOWcJyEORxN+Ek8+TT266gXkNlHdJBwexUsBg=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
golang.org/x/tools/godoc v0.1.0-deprecated h1:o+aZ1BOj6Hsx/GBdJO/s815sqftjSnrZZwyYTHODvtk=
golang.org/x/tools/godoc v0.1.0-deprecated/go.mod h1:qM63CriJ961IHWmnWa9CjZnBndniPt4a3CK0PVB9bIg=
golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
@@ -2668,8 +2668,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.296.0 h1:Nn5EHeKdGx70MFClaV/II0gsWUm6xhEjb0xYLylVvaA=
google.golang.org/api v0.296.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/api v0.297.0 h1:WktxTsnnx0yZNnsR6j0q6hR21RnnK81FHTOPy/ux4OE=
google.golang.org/api v0.297.0/go.mod h1:S4m8x0M6OkQpkOzGk1y9JG2sm4fFQrMh6dxzjCTszhE=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2805,8 +2805,8 @@ google.golang.org/genproto v0.0.0-20230222225845-10f96fb3dbec/go.mod h1:3Dl5ZL0q
google.golang.org/genproto v0.0.0-20230306155012-7f2fa6fef1f4/go.mod h1:NWraEVixdDnqcqQ30jipen1STv2r/n24Wb7twVTGR4s=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d h1:C9v1o0/4quuhOAfmRXA2j+we0PqZIp8traLdeogF3Ms=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d/go.mod h1:Wz2wFJntZFmLGo7pLDXZ3wYk5hyc0Mb+SkHhDDXT+lU=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d h1:QwnJwPte4XXAkhPu26LTDIahnsMSUV0kK8HkxbC+Pc4=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d/go.mod h1:WRrQ7/7N19PypuT0fxLOL5Lq0waoiRri4FbtHDEKrGE=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1 h1:lrupDmKL3p5kEX1M92oan027eCKcouzjuPbH6YBK+Rs=
google.golang.org/genproto/googleapis/api v0.0.0-20260817212433-ac3dfec99bb1/go.mod h1:q/3oV3jAi5vwelxsVAprMBC8BcM2zmNe+IjRGd+9/ks=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 h1:cYNAzI2sUwhmCcoj9TxvihSrqsxt6uIkj3rDRhSDmW4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
@@ -2849,8 +2849,8 @@ google.golang.org/grpc v1.51.0/go.mod h1:wgNDFcnuBGmxLKI/qn4T+m5BtEBYXJPvibbUPsA
google.golang.org/grpc v1.52.0/go.mod h1:pu6fVzoFb+NBYNAvQL08ic+lvB2IojljRYuun5vorUY=
google.golang.org/grpc v1.53.0/go.mod h1:OnIrk0ipVdj4N5d9IUoFUx72/VlD7+jUsHwZgwSMQpw=
google.golang.org/grpc v1.55.0/go.mod h1:iYEXKGkEBhg1PjZQvoYEVPTDkHo1/bjTnfwTeGONTY8=
google.golang.org/grpc v1.85.0-dev h1:HxkDyKIIZPpFnroC56tQv5gNuKTmVvi0t7TzOf5zt7g=
google.golang.org/grpc v1.85.0-dev/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7 h1:5+EEM1fC0yjOZID0NUZVrE2+8M/+1TclNrSz/l1xMYs=
google.golang.org/grpc v1.85.0-dev.0.20260915183914-4e49413dcab7/go.mod h1:Ovl0ECo4xx5r4kn/6d4BPSNB7OIFuu6EAjOzjtVAKaM=
google.golang.org/grpc/cmd/protoc-gen-go-grpc v1.1.0/go.mod h1:6Kw0yEErY5E/yWrBtf03jp27GLLJujG4z/JK95pnjjw=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6 h1:ExN12ndbJ608cboPYflpTny6mXSzPrDLh0iTaVrRrds=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6/go.mod h1:6ytKWczdvnpnO+m+JiG9NjEDzR1FJfsnmJdG7B8QVZ8=
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.46"
appVersion: "4.48"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.46.0
version: 4.48.0
+25 -1
View File
@@ -286,7 +286,7 @@ metadata:
app.kubernetes.io/component: s3
stringData:
# this key must be an inline json config file
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read","List"]}]}'
```
#### Source S3 credentials from an existing Secret
@@ -363,6 +363,30 @@ If `adminPassword` is empty or not set, the admin interface runs without authent
As an alternative, a kubernetes Secret can be used (`admin.secret.existingSecret`).
### Admin listen address
Since SeaweedFS 4.46, `weed admin` defaults to listening on loopback (`127.0.0.1`).
The chart's httpGet readiness/liveness probes dial the pod IP, so the admin
server must bind a non-loopback address for the probes to succeed. The chart
therefore passes `-ip={{ .Values.admin.ip }}`, defaulting `admin.ip` to `0.0.0.0`
(the pre-4.46 behaviour of listening on all interfaces).
Binding a non-loopback address requires authentication: `weed admin` refuses to
start on a non-loopback address without `-adminPassword`, so the chart fails at
render time if `admin.ip` is non-loopback and authentication is not configured via
`admin.secret.adminPassword`, `admin.secret.existingSecret`, or
`WEED_ADMIN_PASSWORD` supplied through `admin.extraEnvironmentVars` /
`admin.secretExtraEnvironmentVars`. Setting `admin.allowInsecureBind` renders
`-allowInsecureBind` and bypasses this guard; it leaves the admin API
unauthenticated on the network, so use it only when access is otherwise
restricted (e.g. network policies). The whole `127.0.0.0/8` range and `::1` are
treated as loopback (matching `weed admin`); `localhost` is treated as
non-loopback. Set `admin.ip` to a loopback address only if you also replace the
httpGet probes (e.g. with an `exec` probe that checks `127.0.0.1`).
The `-ip` flag requires SeaweedFS 4.46 or newer; pinning `admin.imageOverride`
to an older image is not supported with this chart version.
### Admin Data Persistence
The admin component can store configuration and maintenance data. You can configure storage in several ways:
@@ -0,0 +1,7 @@
# Admin install: exercises the admin StatefulSet, which passes -ip (default
# 0.0.0.0) and therefore requires authentication to bind a non-loopback address.
admin:
enabled: true
secret:
adminUser: "admin"
adminPassword: "ci-admin-password"
@@ -6,6 +6,11 @@
{{- if and (not .Values.admin.masters) (not .Values.global.seaweedfs.masterServer) (not .Values.master.enabled) }}
{{- fail "admin.masters or global.seaweedfs.masterServer must be set if master.enabled is false" -}}
{{- end }}
{{- $adminAuthEnabled := include "seaweedfs.admin.authEnabled" . }}
{{- $adminIp := .Values.admin.ip | default "0.0.0.0" }}
{{- if and (not (include "seaweedfs.admin.isLoopbackIp" $adminIp)) (ne $adminAuthEnabled "true") }}
{{- fail (printf "admin.ip is set to %q (non-loopback) but admin authentication is not configured. Since `weed admin` 4.46 refuses to bind a non-loopback address without authentication, the admin container would exit on startup. Set admin.secret.adminPassword or admin.secret.existingSecret, or supply WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars / admin.secretExtraEnvironmentVars, or set admin.ip to a loopback address such as 127.0.0.1 (note: a loopback bind makes the chart's httpGet readiness/liveness probes fail), or set admin.allowInsecureBind to true to opt out via -allowInsecureBind (INSECURE: exposes the admin API unauthenticated on the network)." $adminIp) -}}
{{- end }}
apiVersion: apps/v1
kind: StatefulSet
metadata:
@@ -162,6 +167,7 @@ spec:
-v={{ .Values.global.seaweedfs.loggingLevel }} \
{{- end }}
admin \
-ip={{ .Values.admin.ip | default "0.0.0.0" }} \
-port={{ .Values.admin.port }} \
-port.grpc={{ .Values.admin.grpcPort }} \
{{- if or (eq .Values.admin.data.type "hostPath") (eq .Values.admin.data.type "persistentVolumeClaim") (eq .Values.admin.data.type "emptyDir") (eq .Values.admin.data.type "existingClaim") }}
@@ -170,14 +176,17 @@ spec:
-dataDir={{ .Values.admin.dataDir }} \
{{- end }}
{{- if .Values.admin.masters }}
-masters={{ .Values.admin.masters }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ .Values.admin.masters }} \
{{- else if .Values.global.seaweedfs.masterServer }}
-masters={{ .Values.global.seaweedfs.masterServer }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ .Values.global.seaweedfs.masterServer }} \
{{- else }}
-masters={{ range $index := until (.Values.master.replicas | int) }}${SEAWEEDFS_FULLNAME}-master-{{ $index }}.${SEAWEEDFS_FULLNAME}-master.{{ $.Release.Namespace }}:{{ $.Values.master.port }}{{ if lt $index (sub ($.Values.master.replicas | int) 1) }},{{ end }}{{ end }}{{- if or $urlPrefix .Values.admin.extraArgs }} \{{ end }}
-masters={{ range $index := until (.Values.master.replicas | int) }}${SEAWEEDFS_FULLNAME}-master-{{ $index }}.${SEAWEEDFS_FULLNAME}-master.{{ $.Release.Namespace }}:{{ $.Values.master.port }}{{ if lt $index (sub ($.Values.master.replicas | int) 1) }},{{ end }}{{ end }} \
{{- end }}
{{- if $urlPrefix }}
-urlPrefix={{ $urlPrefix }}{{- if .Values.admin.extraArgs }} \{{ end }}
-urlPrefix={{ $urlPrefix }} \
{{- end }}
{{- if .Values.admin.allowInsecureBind }}
-allowInsecureBind \
{{- end }}
{{- range $index, $arg := .Values.admin.extraArgs }}
{{ $arg }}{{- if lt $index (sub (len $.Values.admin.extraArgs) 1) }} \{{ end }}
@@ -37,13 +37,17 @@ spec:
{{- with .Values.allInOne.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.allInOne.podAnnotations | default dict) }}
{{- $existingS3ConfigSecret := or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret }}
{{- if $existingS3ConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace $existingS3ConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.allInOne.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.allInOne.restartPolicy }}
{{- if .Values.allInOne.affinity }}
@@ -25,6 +25,10 @@ spec:
organizations:
- "SeaweedFS CA"
dnsNames:
- '{{ include "seaweedfs.fullname" . }}-admin'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc'
- '{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc.cluster.local'
- '*.{{ include "seaweedfs.fullname" . }}-admin'
- '*.{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}'
- '*.{{ include "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}.svc'
@@ -43,19 +43,15 @@ spec:
{{- with .Values.filer.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
annotations:
{{- with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.filer.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.filer.podAnnotations | default dict) }}
{{- if .Values.filer.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
checksum/s3config: {{ $configSecret | toYaml | sha256sum }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.filer.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
checksum/s3config: {{ include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.filer.restartPolicy }}
{{- if .Values.filer.affinity }}
@@ -43,13 +43,10 @@ spec:
{{- with .Values.master.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.master.podAnnotations | default dict) }}
{{- $_ := set $podAnnotations "checksum/master-config" (include (print .Template.BasePath "/master/master-configmap.yaml") . | sha256sum) }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.master.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.master.restartPolicy }}
{{- if .Values.master.affinity }}
@@ -35,13 +35,15 @@ spec:
{{- with .Values.s3.podLabels }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- $podAnnotations := mergeOverwrite (deepCopy (.Values.podAnnotations | default dict)) (.Values.s3.podAnnotations | default dict) }}
{{- if .Values.s3.existingConfigSecret }}
{{- $configSecret := (lookup "v1" "Secret" .Release.Namespace .Values.s3.existingConfigSecret) | default dict }}
{{- $_ := set $podAnnotations "checksum/s3config" ($configSecret | toYaml | sha256sum) }}
{{- else }}
{{- $_ := set $podAnnotations "checksum/s3config" (include (print .Template.BasePath "/s3/s3-secret.yaml") . | sha256sum) }}
{{- end }}
annotations:
{{ with .Values.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- with .Values.s3.podAnnotations }}
{{- toYaml . | nindent 8 }}
{{- end }}
{{- toYaml $podAnnotations | nindent 8 }}
spec:
restartPolicy: {{ default .Values.global.seaweedfs.restartPolicy .Values.s3.restartPolicy }}
{{- if .Values.s3.affinity }}
@@ -60,7 +60,7 @@ stringData:
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read","List"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
{{ toJson .Values.filer.s3.auditLogConfig | nindent 4 }}
@@ -88,6 +88,43 @@ true
{{- end -}}
{{- end -}}
{{/* Classify an admin bind address as loopback, mirroring weed admin's
isLoopbackIp (net.ParseIP + IsLoopback). Helm templates cannot call
net.ParseIP, so we approximate: valid IPv4 addresses in 127.0.0.0/8
(validated via regex to reject malformed values like "127.not-an-ip")
and the IPv6 loopback "::1" / its expanded form "0:0:0:0:0:0:0:1" are
loopback. Hostnames (e.g. "localhost") and wildcard addresses
("0.0.0.0", "::") are non-loopback, matching the binary, which
treats unparseable hostnames as non-loopback to be safe. Other IPv6
loopback representations are not matched; the binary's own runtime
validation is the authoritative guard. */}}
{{- define "seaweedfs.admin.isLoopbackIp" -}}
{{- $ip := toString . -}}
{{- if or (regexMatch "^127\\.[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}$" $ip) (eq $ip "::1") (eq $ip "0:0:0:0:0:0:0:1") -}}
true
{{- end -}}
{{- end -}}
{{/* Whether the admin non-loopback bind guard is satisfied: admin.secret
(adminPassword or existingSecret), WEED_ADMIN_PASSWORD via
extraEnvironmentVars / secretExtraEnvironmentVars, or
admin.allowInsecureBind. A secret-backed entry counts as enabled even
though the chart cannot read its value. */}}
{{- define "seaweedfs.admin.authEnabled" -}}
{{- if or .Values.admin.secret.existingSecret .Values.admin.secret.adminPassword .Values.admin.allowInsecureBind -}}
true
{{- else -}}
{{- $merged := dict -}}
{{- $_ := include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.admin "target" $merged) -}}
{{- $envPassword := index $merged "WEED_ADMIN_PASSWORD" -}}
{{- if or (kindIs "map" $envPassword) (hasKey (.Values.admin.secretExtraEnvironmentVars | default dict) "WEED_ADMIN_PASSWORD") -}}
true
{{- else if and $envPassword (ne (toString $envPassword) "") -}}
true
{{- end -}}
{{- end -}}
{{- end -}}
{{/* Return the proper filer image */}}
{{- define "seaweedfs.filer.image" -}}
{{- if .Values.filer.imageOverride -}}
+29
View File
@@ -190,6 +190,8 @@ master:
podLabels: {}
# Annotations to be added to the master pods
# The chart sets checksum/master-config on master pods; other checksum/* keys
# can be used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the master resources
@@ -773,6 +775,8 @@ filer:
podLabels: {}
# Annotations to be added to the filer pods
# The chart sets checksum/s3config on filer pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the filer resource
@@ -1078,6 +1082,8 @@ s3:
podLabels: {}
# Annotations to be added to the s3 pods
# The chart sets checksum/s3config on s3 pods; other checksum/* keys can be
# used for custom rollouts.
podAnnotations: {}
# Annotations to be added to the s3 resources
@@ -1328,8 +1334,26 @@ admin:
replicas: 1
port: 23646 # Default admin port
grpcPort: 33646 # Default gRPC port for worker connections
# IP address the admin server listens on. Since `weed admin` 4.46 defaults to
# loopback (127.0.0.1), the chart must bind a non-loopback address for the
# kubelet's httpGet readiness/liveness probes (which dial the pod IP) to ever
# succeed. "0.0.0.0" restores the pre-4.46 behaviour of listening on all
# interfaces. A non-loopback address requires authentication: set
# admin.secret.adminPassword or admin.secret.existingSecret, supply
# WEED_ADMIN_PASSWORD via admin.extraEnvironmentVars /
# admin.secretExtraEnvironmentVars, or opt out with admin.allowInsecureBind;
# otherwise the admin container will exit with a clear error rather than
# silently staying unready. The whole
# 127.0.0.0/8 range and ::1 are treated as loopback (matching weed admin).
# Set to a loopback address only if you also replace the httpGet probes.
# Note: the -ip flag requires SeaweedFS 4.46 or newer; pinning
# admin.imageOverride to an older image is not supported with this chart.
ip: "0.0.0.0"
loggingOverrideLevel: null
# INSECURE: allow binding a non-loopback ip without authentication.
allowInsecureBind: false
# Admin authentication
secret:
# Name of an existing secret containing admin credentials. If set, adminUser and adminPassword below are ignored.
@@ -1774,6 +1798,8 @@ allInOne:
initContainers: "" # Init containers
sidecars: "" # Sidecar containers
annotations: {} # Annotations for the deployment
# The chart sets checksum/master-config and checksum/s3config on all-in-one
# pods; other checksum/* keys can be used for custom rollouts.
podAnnotations: {} # Annotations for the pods
podLabels: {} # Labels for the pods
@@ -1908,6 +1934,9 @@ certificates:
# Labels to be added to all the created pods
podLabels: {}
# Annotations to be added to all the created pods
# The chart sets checksum/master-config and checksum/s3config on pods whose
# rendered ConfigMaps or Secrets should trigger rollouts. Other checksum/* keys
# can be used for custom rollout annotations.
podAnnotations: {}
networkPolicy:
+1006 -1059
View File
File diff suppressed because it is too large Load Diff

Before

Width:  |  Height:  |  Size: 54 KiB

After

Width:  |  Height:  |  Size: 53 KiB

+293
View File
@@ -0,0 +1,293 @@
# This file is automatically @generated by Cargo.
# It is not intended for manual editing.
version = 4
[[package]]
name = "aws-lc-rs"
version = "1.18.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ce2b2dcc879c3bae0d371e77c99f2238400ef24ec001394befa67b6e543add9e"
dependencies = [
"aws-lc-sys",
"zeroize",
]
[[package]]
name = "aws-lc-sys"
version = "0.44.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f09fae7be8bb3174e05c6afdb34199e6dc0c7c04ba9fa237b1967adfbde27483"
dependencies = [
"cc",
"cmake",
"dunce",
"fs_extra",
"pkg-config",
]
[[package]]
name = "cc"
version = "1.4.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "509591b7bcd67f4ef775afad7662703b4935daaa6ec0e5605cfb1090b32a2b6d"
dependencies = [
"find-msvc-tools",
"jobserver",
"libc",
"shlex",
]
[[package]]
name = "cfg-if"
version = "1.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
[[package]]
name = "cmake"
version = "0.1.58"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c0f78a02292a74a88ac736019ab962ece0bc380e3f977bf72e376c5d78ff0678"
dependencies = [
"cc",
]
[[package]]
name = "dunce"
version = "1.0.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "92773504d58c093f6de2459af4af33faa518c13451eb8f2b5698ed3d36e7c813"
[[package]]
name = "find-msvc-tools"
version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d45db016d36b838f563236e9193d0ee6ce38f3f68b6c94e914b4929c96bbb890"
[[package]]
name = "fs_extra"
version = "1.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "42703706b716c37f96a77aea830392ad231f44c9e9a67872fa5548707e11b11c"
[[package]]
name = "getrandom"
version = "0.2.17"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0"
dependencies = [
"cfg-if",
"libc",
"wasi",
]
[[package]]
name = "getrandom"
version = "0.4.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "300e883d756b2e4ec94e02791f39b04b522276138852cfc41d9fb7e904106099"
dependencies = [
"cfg-if",
"libc",
"r-efi",
]
[[package]]
name = "jobserver"
version = "0.1.35"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1c00acbd29eabad4a2392fa0e921c874934dbbf4194312ad20f04a0ed67a3cb3"
dependencies = [
"getrandom 0.4.3",
"libc",
]
[[package]]
name = "libc"
version = "0.2.189"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3eaf3ede3fee6db1a4c2ee091bf8a8b4dccdc6d17f656fb07896ee72867612f2"
[[package]]
name = "log"
version = "0.4.33"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad"
[[package]]
name = "once_cell"
version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "pkg-config"
version = "0.3.34"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f6b464fbc74e149a392436b17d523f769e057cb6877f6a5c4618bc6f11800548"
[[package]]
name = "r-efi"
version = "6.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f8dcc9c7d52a811697d2151c701e0d08956f92b0e24136cf4cf27b57a6a0d9bf"
[[package]]
name = "ring"
version = "0.17.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7"
dependencies = [
"cc",
"cfg-if",
"getrandom 0.2.17",
"libc",
"untrusted",
"windows-sys",
]
[[package]]
name = "rustls"
version = "0.23.43"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
dependencies = [
"aws-lc-rs",
"log",
"once_cell",
"rustls-pki-types",
"rustls-webpki",
"subtle",
"zeroize",
]
[[package]]
name = "rustls-pki-types"
version = "1.15.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96"
dependencies = [
"zeroize",
]
[[package]]
name = "rustls-webpki"
version = "0.103.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0527518605e68109d875e248ea259b6758801cf165e4b2c2733ae3b51f12535a"
dependencies = [
"aws-lc-rs",
"ring",
"rustls-pki-types",
"untrusted",
]
[[package]]
name = "seaweed-common"
version = "0.1.0"
dependencies = [
"rustls",
]
[[package]]
name = "shlex"
version = "2.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f8fadd59c855ef2080decdef8ff161eb6661b86933c9d82e5ba29dc602a55aba"
[[package]]
name = "subtle"
version = "2.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292"
[[package]]
name = "untrusted"
version = "0.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "wasi"
version = "0.11.1+wasi-snapshot-preview1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b"
[[package]]
name = "windows-sys"
version = "0.52.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "282be5f36a8ce781fad8c8ae18fa3f9beff57ec1b52cb3de0789201425d9a33d"
dependencies = [
"windows-targets",
]
[[package]]
name = "windows-targets"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9b724f72796e036ab90c1021d4780d4d3d648aca59e491e6b98e725b84e99973"
dependencies = [
"windows_aarch64_gnullvm",
"windows_aarch64_msvc",
"windows_i686_gnu",
"windows_i686_gnullvm",
"windows_i686_msvc",
"windows_x86_64_gnu",
"windows_x86_64_gnullvm",
"windows_x86_64_msvc",
]
[[package]]
name = "windows_aarch64_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32a4622180e7a0ec044bb555404c800bc9fd9ec262ec147edd5989ccd0c02cd3"
[[package]]
name = "windows_aarch64_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09ec2a7bb152e2252b53fa7803150007879548bc709c039df7627cabbd05d469"
[[package]]
name = "windows_i686_gnu"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8e9b5ad5ab802e97eb8e295ac6720e509ee4c243f69d781394014ebfe8bbfa0b"
[[package]]
name = "windows_i686_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0eee52d38c090b3caa76c563b86c3a4bd71ef1a819287c19d586d7334ae8ed66"
[[package]]
name = "windows_i686_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "240948bc05c5e7c6dabba28bf89d89ffce3e303022809e73deaefe4f6ec56c66"
[[package]]
name = "windows_x86_64_gnu"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "147a5c80aabfbf0c7d901cb5895d1de30ef2907eb21fbbab29ca94c5b08b1a78"
[[package]]
name = "windows_x86_64_gnullvm"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "24d5b23dc417412679681396f2b49f3de8c1473deb516bd34410872eff51ed0d"
[[package]]
name = "windows_x86_64_msvc"
version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec"
[[package]]
name = "zeroize"
version = "1.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e"
+27
View File
@@ -0,0 +1,27 @@
[package]
name = "seaweed-common"
version = "0.1.0"
edition = "2024"
# The lower of the two consumers' floors (seaweed-volume 1.91.1,
# seaweed-worker 1.94.1), so depending on this crate cannot raise either
# tree's MSRV. Verified with `cargo +1.91.1 check --all-targets`.
rust-version = "1.91.1"
description = "Helpers shared by the SeaweedFS Rust volume server and the Rust plugin workers"
# There is no root manifest: seaweed-volume and seaweed-worker are separate
# cargo trees with their own lockfiles, and this crate is a path dependency of
# both rather than a member of either. Keeping the lint policy identical in all
# three manifests is what stops them drifting.
[lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
[dependencies]
# The same requirement both consumers already write. Cargo unifies all
# semver-compatible `rustls = "0.23"` requirements into one crate per binary,
# which is what makes `install_default_crypto_provider` write the same
# process-wide static the consuming crate reads. rustls is already in both
# trees (the volume server directly, seaweed-worker-core through tonic's
# `tls-aws-lc`), so this adds no crate to either graph.
rustls = "0.23"
+307
View File
@@ -0,0 +1,307 @@
//! SeaweedFS server addresses, the way the Go tree does them.
//!
//! An operator gives a SeaweedFS process an HTTP address and the gRPC port is
//! derived from it rather than asked for separately: `host:port` means gRPC on
//! `port + 10000`, and the explicit `host:port.grpcPort` form names it outright.
//! Dialling the HTTP port by mistake fails as "frame with invalid size", which
//! reads like a protocol bug rather than a wrong port, so the rule is worth its
//! own module. Mirrors `pb.ServerToGrpcAddress` in
//! `weed/pb/grpc_client_server.go`.
//!
//! The volume server and the workers each had their own copy of this and the
//! copies had drifted: the worker's bracketed IPv6 literals and the volume
//! server's did not, so `::1:19333` produced `::1:29333`, which the HTTP
//! authority parser rejects. One implementation, two thin wrappers.
use std::fmt;
use std::num::ParseIntError;
/// SeaweedFS's HTTP↔gRPC port-offset convention.
pub const GRPC_PORT_OFFSET: u16 = 10000;
/// Why an address could not be turned into a gRPC address.
///
/// The `Display` text is the volume server's original wording, because its
/// `parse_grpc_address` wrapper hands it straight to callers that put it in a
/// `Status` or an `io::Error`.
#[derive(Debug, Clone, PartialEq, Eq)]
#[non_exhaustive]
pub enum AddressError {
/// No `:` at all, so there is no port to translate.
MissingPort(String),
/// The HTTP port of the `host:port.grpcPort` form is not a `u16`. It is
/// validated even though it is then discarded, so that a malformed address
/// is rejected here instead of failing later as an opaque connect error.
InvalidHttpPort { port: String, source: ParseIntError },
/// The gRPC port of the `host:port.grpcPort` form is not a `u16`.
InvalidGrpcPort { port: String, source: ParseIntError },
/// The port of the `host:port` form is not a `u16`.
InvalidPort { port: String, source: ParseIntError },
/// `port + GRPC_PORT_OFFSET` leaves the TCP port range, e.g. `host:60000`.
/// Without the check the cast would wrap silently.
ImplicitGrpcPortOutOfRange(u16),
}
impl fmt::Display for AddressError {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
match self {
Self::MissingPort(address) => write!(f, "cannot parse address: {address}"),
Self::InvalidHttpPort { port, source } => {
write!(f, "invalid http port {port:?}: {source}")
}
Self::InvalidGrpcPort { port, source } => {
write!(f, "invalid grpc port {port:?}: {source}")
}
Self::InvalidPort { port, source } => write!(f, "invalid port {port:?}: {source}"),
Self::ImplicitGrpcPortOutOfRange(port) => write!(
f,
"implicit grpc port out of range: {port} + {GRPC_PORT_OFFSET} = {}",
u32::from(*port) + u32::from(GRPC_PORT_OFFSET)
),
}
}
}
impl std::error::Error for AddressError {
fn source(&self) -> Option<&(dyn std::error::Error + 'static)> {
match self {
Self::InvalidHttpPort { source, .. }
| Self::InvalidGrpcPort { source, .. }
| Self::InvalidPort { source, .. } => Some(source),
Self::MissingPort(_) | Self::ImplicitGrpcPortOutOfRange(_) => None,
}
}
}
/// Turn a SeaweedFS server address (`"host:port.grpcPort"` or `"host:port"`)
/// into the `host:grpcPort` form the endpoint builders expect.
///
/// With the trailing `.grpcPort` segment that segment *is* the gRPC port;
/// without it the gRPC port is `port + GRPC_PORT_OFFSET`. An unbracketed IPv6
/// literal comes back bracketed, because otherwise the port reads as part of
/// the address.
pub fn to_grpc_address(server: &str) -> Result<String, AddressError> {
// rfind, not find: an IPv6 literal is full of colons and the port is after
// the last one.
let colon_idx = server
.rfind(':')
.ok_or_else(|| AddressError::MissingPort(server.to_string()))?;
let host = &server[..colon_idx];
let port_part = &server[colon_idx + 1..];
// rfind again rather than split_once: the host may be an IPv4 address, and
// only the part after the last colon is being split here anyway.
if let Some(dot_idx) = port_part.rfind('.') {
let http_port = &port_part[..dot_idx];
let grpc_port = &port_part[dot_idx + 1..];
http_port
.parse::<u16>()
.map_err(|source| AddressError::InvalidHttpPort {
port: http_port.to_string(),
source,
})?;
let grpc_port =
grpc_port
.parse::<u16>()
.map_err(|source| AddressError::InvalidGrpcPort {
port: grpc_port.to_string(),
source,
})?;
return Ok(join_host_port(host, grpc_port));
}
let port: u16 = port_part
.parse()
.map_err(|source| AddressError::InvalidPort {
port: port_part.to_string(),
source,
})?;
let grpc_port = port
.checked_add(GRPC_PORT_OFFSET)
.ok_or(AddressError::ImplicitGrpcPortOutOfRange(port))?;
Ok(join_host_port(host, grpc_port))
}
/// Join a host and a port, bracketing an IPv6 literal that is not bracketed
/// already. Public because the address rule is not the only place that has to
/// put a host and a port back together.
pub fn join_host_port(host: &str, port: u16) -> String {
// An IPv6 literal has to keep its brackets or the port reads as part of it.
if host.contains(':') && !host.starts_with('[') {
format!("[{host}]:{port}")
} else {
format!("{host}:{port}")
}
}
#[cfg(test)]
mod tests {
use super::{AddressError, GRPC_PORT_OFFSET, join_host_port, to_grpc_address};
// ---- the volume server's cases -------------------------------------
#[test]
fn dotted_form_states_the_grpc_port() {
assert_eq!(
to_grpc_address("127.0.0.1:8080.18080").unwrap(),
"127.0.0.1:18080"
);
assert_eq!(
to_grpc_address("192.168.1.66:8080.18080").unwrap(),
"192.168.1.66:18080"
);
}
#[test]
fn implicit_form_adds_the_offset() {
assert_eq!(
to_grpc_address("127.0.0.1:8080").unwrap(),
"127.0.0.1:18080"
);
assert_eq!(
to_grpc_address("192.168.1.66:8080").unwrap(),
"192.168.1.66:18080"
);
assert_eq!(
to_grpc_address("localhost:9333").unwrap(),
"localhost:19333"
);
}
#[test]
fn the_dotted_grpc_port_comes_back_normalised() {
// The volume server's copy validated this segment as a u16 and then
// emitted the original text, so a padded or signed port produced an
// authority the URI parser rejects. The parsed value is emitted now.
assert_eq!(to_grpc_address("host:8080.018080").unwrap(), "host:18080");
assert_eq!(to_grpc_address("host:8080.+18080").unwrap(), "host:18080");
}
#[test]
fn an_ipv4_host_is_not_confused_with_the_dotted_port() {
// Regression: a naive split on '.' breaks on IP addresses.
assert_eq!(
to_grpc_address("10.0.0.1:8080.18080").unwrap(),
"10.0.0.1:18080"
);
assert_eq!(to_grpc_address("10.0.0.1:8080").unwrap(), "10.0.0.1:18080");
}
#[test]
fn rejects_a_non_numeric_http_port_in_the_dotted_form() {
let err = to_grpc_address("host:abc.18080").unwrap_err();
assert!(
matches!(err, AddressError::InvalidHttpPort { .. }),
"{err:?}"
);
assert!(err.to_string().contains("invalid http port"), "{err}");
}
#[test]
fn rejects_a_non_numeric_grpc_port_in_the_dotted_form() {
let err = to_grpc_address("host:8080.xyz").unwrap_err();
assert!(
matches!(err, AddressError::InvalidGrpcPort { .. }),
"{err:?}"
);
assert!(err.to_string().contains("invalid grpc port"), "{err}");
}
#[test]
fn rejects_an_implicit_port_that_leaves_the_tcp_range() {
let err = to_grpc_address("127.0.0.1:60000").unwrap_err();
assert!(
matches!(err, AddressError::ImplicitGrpcPortOutOfRange(60000)),
"{err:?}"
);
assert!(err.to_string().contains("out of range"), "{err}");
}
#[test]
fn the_messages_are_the_volume_servers_wording_verbatim() {
// parse_grpc_address hands these straight to callers that put them in a
// Status or an io::Error, so the whole string is the contract, not just
// the substring the older tests match on. Only the two variants whose
// text is entirely ours are pinned exactly; the other three end in a
// std ParseIntError message, which is std's to reword.
assert_eq!(
to_grpc_address("127.0.0.1:60000").unwrap_err().to_string(),
"implicit grpc port out of range: 60000 + 10000 = 70000"
);
assert_eq!(
to_grpc_address("hostname").unwrap_err().to_string(),
"cannot parse address: hostname"
);
}
#[test]
fn rejects_an_address_without_a_port() {
for source in ["hostname", "no-colon", "localhost"] {
let err = to_grpc_address(source).unwrap_err();
assert!(matches!(err, AddressError::MissingPort(_)), "{err:?}");
assert!(err.to_string().contains("cannot parse"), "{err}");
}
}
// ---- the worker's cases --------------------------------------------
#[test]
fn derives_the_grpc_port() {
assert_eq!(
to_grpc_address("localhost:23646").unwrap(),
"localhost:33646"
);
assert_eq!(
to_grpc_address("127.0.0.1:9333").unwrap(),
"127.0.0.1:19333"
);
}
#[test]
fn honours_an_explicit_grpc_port() {
assert_eq!(
to_grpc_address("localhost:23646.33999").unwrap(),
"localhost:33999"
);
}
#[test]
fn rejects_what_it_cannot_parse() {
let err = to_grpc_address("localhost:notaport").unwrap_err();
assert!(matches!(err, AddressError::InvalidPort { .. }), "{err:?}");
assert!(err.to_string().contains("invalid port"), "{err}");
}
// ---- IPv6, which only the worker's copy handled --------------------
#[test]
fn brackets_ipv6_literals() {
assert_eq!(to_grpc_address("::1:23646").unwrap(), "[::1]:33646");
assert_eq!(to_grpc_address("::1:9333").unwrap(), "[::1]:19333");
assert_eq!(
to_grpc_address("fe80::1:9333.19333").unwrap(),
"[fe80::1]:19333"
);
}
#[test]
fn leaves_an_already_bracketed_literal_alone() {
assert_eq!(to_grpc_address("[::1]:9333").unwrap(), "[::1]:19333");
assert_eq!(to_grpc_address("[::1]:9333.19333").unwrap(), "[::1]:19333");
}
#[test]
fn join_host_port_brackets_only_unbracketed_literals() {
assert_eq!(join_host_port("127.0.0.1", 19333), "127.0.0.1:19333");
assert_eq!(join_host_port("localhost", 19333), "localhost:19333");
assert_eq!(join_host_port("::1", 19333), "[::1]:19333");
assert_eq!(join_host_port("[::1]", 19333), "[::1]:19333");
}
// ---- the offset itself ---------------------------------------------
#[test]
fn the_offset_is_the_seaweedfs_convention() {
assert_eq!(GRPC_PORT_OFFSET, 10000);
}
}
+11
View File
@@ -0,0 +1,11 @@
//! Helpers the SeaweedFS Rust volume server and the Rust plugin workers both need.
//!
//! `seaweed-volume` and `seaweed-worker` are separate cargo trees with separate
//! lockfiles and no root manifest, so anything both of them need was, until this
//! crate existed, written twice. The two things in here are the ones where a
//! second copy is a correctness risk rather than a typing cost: the HTTP↔gRPC
//! address rule, which two copies had already drifted on, and the process-wide
//! rustls provider, which only works if every binary installs the same one.
pub mod address;
pub mod tls;
+28
View File
@@ -0,0 +1,28 @@
//! The process-wide rustls crypto provider.
//!
//! Both binaries link aws-lc-rs and ring transitively — in the volume server
//! through the AWS SDK and reqwest, in the lance worker through lance's `aws`
//! backend and reqwest — so rustls cannot auto-select a provider and tonic's
//! client TLS panics on first use. Each binary has to pin one, and it has to be
//! the same one, which is why the choice lives here rather than in either tree.
use rustls::crypto::aws_lc_rs;
/// Pin rustls's process-wide default provider to aws-lc-rs, matching the
/// volume server's TLS config. Idempotent: the first call wins and every
/// later one is a no-op, so callers do not have to coordinate.
pub fn install_default_crypto_provider() {
let _ = aws_lc_rs::default_provider().install_default();
}
#[cfg(test)]
mod tests {
use super::install_default_crypto_provider;
#[test]
fn installing_is_idempotent_and_leaves_a_default_behind() {
install_default_crypto_provider();
install_default_crypto_provider();
assert!(rustls::crypto::CryptoProvider::get_default().is_some());
}
}
+121 -127
View File
@@ -503,7 +503,7 @@ dependencies = [
"rustls-pki-types",
"tokio",
"tokio-rustls",
"tower 0.5.3",
"tower",
"tracing",
]
@@ -628,13 +628,13 @@ dependencies = [
[[package]]
name = "axum"
version = "0.7.9"
version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "edca88bc138befd0323b20752846e6587272d3b03b0343c8ea28a6f819e6e71f"
checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90"
dependencies = [
"async-trait",
"axum-core",
"bytes",
"form_urlencoded",
"futures-util",
"http 1.4.0",
"http-body 1.0.1",
@@ -648,14 +648,13 @@ dependencies = [
"multer",
"percent-encoding",
"pin-project-lite",
"rustversion",
"serde",
"serde_core",
"serde_json",
"serde_path_to_error",
"serde_urlencoded",
"sync_wrapper",
"tokio",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -663,19 +662,17 @@ dependencies = [
[[package]]
name = "axum-core"
version = "0.4.5"
version = "0.5.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "09f2bd6146b97ae3359fa0cc6d6b376d9539582c7b4220f041a33ec24c226199"
checksum = "08c78f31d7b1291f7ee735c1c6780ccde7785daae9a9206026862dab7d8792d1"
dependencies = [
"async-trait",
"bytes",
"futures-util",
"futures-core",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"mime",
"pin-project-lite",
"rustversion",
"sync_wrapper",
"tower-layer",
"tower-service",
@@ -1654,19 +1651,13 @@ dependencies = [
"futures-core",
"futures-sink",
"http 1.4.0",
"indexmap 2.13.1",
"indexmap",
"slab",
"tokio",
"tokio-util",
"tracing",
]
[[package]]
name = "hashbrown"
version = "0.12.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a9ee70c43aaf417c914396645a0fa852624801b24ebb7ae78fe8272889ac888"
[[package]]
name = "hashbrown"
version = "0.14.5"
@@ -1860,7 +1851,7 @@ dependencies = [
"libc",
"percent-encoding",
"pin-project-lite",
"socket2 0.6.3",
"socket2",
"tokio",
"tower-service",
"tracing",
@@ -2027,16 +2018,6 @@ dependencies = [
"quick-error",
]
[[package]]
name = "indexmap"
version = "1.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bd070e393353796e801d209ad339e89596eb4c8d430d18ede6a1cced8fafbd99"
dependencies = [
"autocfg",
"hashbrown 0.12.3",
]
[[package]]
name = "indexmap"
version = "2.13.1"
@@ -2250,9 +2231,9 @@ dependencies = [
[[package]]
name = "matchit"
version = "0.7.3"
version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0e7465ac9959cc2b1404e8e2367b43684a6d13790fe23056cc8c6c5a6b7bcb94"
checksum = "47e1ffaa40ddd1f3ed91f717a33c8c0ee23fff369e3aa8772b9605cc1d22f4c3"
[[package]]
name = "md-5"
@@ -2619,17 +2600,18 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4c5cc86750666a3ed20bdaf5ca2a0344f9c67674cae0515bec2da16fbaa47db"
dependencies = [
"fixedbitset 0.4.2",
"indexmap 2.13.1",
"indexmap",
]
[[package]]
name = "petgraph"
version = "0.7.1"
version = "0.8.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3672b37090dbd86368a4145bc067582552b29c27377cad4e0a306c97f9bd7772"
checksum = "8701b58ea97060d5e5b155d383a69952a60943f0e6dfe30b04c287beb0b27455"
dependencies = [
"fixedbitset 0.5.7",
"indexmap 2.13.1",
"hashbrown 0.15.5",
"indexmap",
]
[[package]]
@@ -2836,12 +2818,12 @@ dependencies = [
[[package]]
name = "prost"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2796faa41db3ec313a31f7624d9286acf277b52de526150b7e69f3debf891ee5"
checksum = "528ac67416ff8646872a3c02cad9cc4ee5dc9f9540c9b10771855c95cb2e5ae1"
dependencies = [
"bytes",
"prost-derive 0.13.5",
"prost-derive 0.14.4",
]
[[package]]
@@ -2867,19 +2849,20 @@ dependencies = [
[[package]]
name = "prost-build"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042"
dependencies = [
"heck",
"itertools 0.14.0",
"log",
"multimap",
"once_cell",
"petgraph 0.7.1",
"petgraph 0.8.3",
"prettyplease",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"pulldown-cmark",
"pulldown-cmark-to-cmark",
"regex",
"syn",
"tempfile",
@@ -2900,9 +2883,9 @@ dependencies = [
[[package]]
name = "prost-derive"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf"
dependencies = [
"anyhow",
"itertools 0.14.0",
@@ -2922,11 +2905,11 @@ dependencies = [
[[package]]
name = "prost-types"
version = "0.13.5"
version = "0.14.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "52c2c1bf36ddb1a1c396b3601a3cec27c2462e45f07c386894ec3ccf5332bd16"
checksum = "f94967dc7688f3054c7fac87473ffae4cc4c3904800e2d9f5b857246d8963b0a"
dependencies = [
"prost 0.13.5",
"prost 0.14.4",
]
[[package]]
@@ -2993,6 +2976,26 @@ version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.11.0",
"memchr",
"unicase",
]
[[package]]
name = "pulldown-cmark-to-cmark"
version = "22.0.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab1ad36992cead65f02aa399a373a42730922f1525d988172634fdefdecb8a60"
dependencies = [
"pulldown-cmark",
]
[[package]]
name = "pxfm"
version = "0.1.28"
@@ -3018,7 +3021,7 @@ dependencies = [
"quinn-udp",
"rustc-hash",
"rustls",
"socket2 0.6.3",
"socket2",
"thiserror 2.0.18",
"tokio",
"tracing",
@@ -3056,7 +3059,7 @@ dependencies = [
"cfg_aliases",
"libc",
"once_cell",
"socket2 0.6.3",
"socket2",
"tracing",
"windows-sys 0.60.2",
]
@@ -3262,8 +3265,8 @@ dependencies = [
"tokio",
"tokio-rustls",
"tokio-util",
"tower 0.5.3",
"tower-http 0.6.8",
"tower",
"tower-http",
"tower-service",
"url",
"wasm-bindgen",
@@ -3484,6 +3487,13 @@ version = "1.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "94143f37725109f92c262ed2cf5e59bce7498c01bcc1502d7b9afe439a4e9f49"
[[package]]
name = "seaweed-common"
version = "0.1.0"
dependencies = [
"rustls",
]
[[package]]
name = "sec1"
version = "0.3.0"
@@ -3719,16 +3729,6 @@ version = "1.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1b6b67fb9a61334225b5b790716f609cd58395f895b3fe8b328786812a40bc3b"
[[package]]
name = "socket2"
version = "0.5.10"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e22376abed350d73dd1cd119b57ffccad95b4e585a7cda43e286245ce23c0678"
dependencies = [
"libc",
"windows-sys 0.52.0",
]
[[package]]
name = "socket2"
version = "0.6.3"
@@ -3990,7 +3990,7 @@ dependencies = [
"parking_lot 0.12.5",
"pin-project-lite",
"signal-hook-registry",
"socket2 0.6.3",
"socket2",
"tokio-macros",
"windows-sys 0.61.2",
]
@@ -4077,7 +4077,7 @@ version = "0.22.27"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
dependencies = [
"indexmap 2.13.1",
"indexmap",
"serde",
"serde_spanned",
"toml_datetime",
@@ -4093,11 +4093,10 @@ checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
[[package]]
name = "tonic"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "877c5b330756d856ffcc4553ab34a5684481ade925ecc54bcd1bf02b1d0d4d52"
checksum = "ac2a5518c70fa84342385732db33fb3f44bc4cc748936eb5833d2df34d6445ef"
dependencies = [
"async-stream",
"async-trait",
"axum",
"base64",
@@ -4111,13 +4110,12 @@ dependencies = [
"hyper-util",
"percent-encoding",
"pin-project",
"prost 0.13.5",
"rustls-pemfile",
"socket2 0.5.10",
"socket2",
"sync_wrapper",
"tokio",
"tokio-rustls",
"tokio-stream",
"tower 0.4.13",
"tower",
"tower-layer",
"tower-service",
"tracing",
@@ -4125,49 +4123,55 @@ dependencies = [
[[package]]
name = "tonic-build"
version = "0.12.3"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9557ce109ea773b399c9b9e5dca39294110b74f1f342cb347a80d1fce8c26a11"
checksum = "c68f61875ac5293cf72e6c8cf0158086428c82c37229e98c840878f1706b0322"
dependencies = [
"prettyplease",
"proc-macro2",
"prost-build 0.13.5",
"prost-types 0.13.5",
"quote",
"syn",
]
[[package]]
name = "tonic-reflection"
version = "0.12.3"
name = "tonic-prost"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "878d81f52e7fcfd80026b7fdb6a9b578b3c3653ba987f87f0dce4b64043cba27"
checksum = "50849f68853be452acf590cde0b146665b8d507b3b8af17261df47e02c209ea0"
dependencies = [
"prost 0.13.5",
"prost-types 0.13.5",
"tokio",
"tokio-stream",
"bytes",
"prost 0.14.4",
"tonic",
]
[[package]]
name = "tower"
version = "0.4.13"
name = "tonic-prost-build"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8fa9be0de6cf49e536ce1851f987bd21a43b771b09473c3549a6c853db37c1c"
checksum = "654e5643eff75d7f8c99197ce1440ed19a3474eada74c12bbac488b2cafdae27"
dependencies = [
"futures-core",
"futures-util",
"indexmap 1.9.3",
"pin-project",
"pin-project-lite",
"rand 0.8.7",
"slab",
"prettyplease",
"proc-macro2",
"prost-build 0.14.4",
"prost-types 0.14.4",
"quote",
"syn",
"tempfile",
"tonic-build",
]
[[package]]
name = "tonic-reflection"
version = "0.14.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "acccd136a4bf19810a1fde9c74edc6129b42a66b44d0c1c8aaa67aeb49a146a7"
dependencies = [
"prost 0.14.4",
"prost-types 0.14.4",
"tokio",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
"tokio-stream",
"tonic",
"tonic-prost",
]
[[package]]
@@ -4178,26 +4182,12 @@ checksum = "ebe5ef63511595f1344e2d5cfa636d973292adc0eec1f0ad45fae9f0851ab1d4"
dependencies = [
"futures-core",
"futures-util",
"indexmap",
"pin-project-lite",
"slab",
"sync_wrapper",
"tokio",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
name = "tower-http"
version = "0.5.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1e9cd434a998747dd2c4276bc96ee2e0c7a2eadf3cae88e52be55a05fa9053f5"
dependencies = [
"bitflags 2.11.0",
"bytes",
"http 1.4.0",
"http-body 1.0.1",
"http-body-util",
"pin-project-lite",
"tokio-util",
"tower-layer",
"tower-service",
"tracing",
@@ -4216,9 +4206,10 @@ dependencies = [
"http-body 1.0.1",
"iri-string",
"pin-project-lite",
"tower 0.5.3",
"tower",
"tower-layer",
"tower-service",
"tracing",
]
[[package]]
@@ -4495,7 +4486,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bb0e353e6a2fbdc176932bbaab493762eb1255a7900fe0fea1a2f96c296cc909"
dependencies = [
"anyhow",
"indexmap 2.13.1",
"indexmap",
"wasm-encoder",
"wasmparser",
]
@@ -4521,7 +4512,7 @@ checksum = "47b807c72e1bac69382b3a6fb3dbe8ea4c0ed87ff5629b8685ae6b9a611028fe"
dependencies = [
"bitflags 2.11.0",
"hashbrown 0.15.5",
"indexmap 2.13.1",
"indexmap",
"semver",
]
@@ -4564,6 +4555,7 @@ dependencies = [
"aws-config",
"aws-credential-types",
"aws-sdk-s3",
"aws-smithy-runtime-api",
"aws-types",
"axum",
"base64",
@@ -4582,7 +4574,6 @@ dependencies = [
"image",
"jsonwebtoken",
"kamadak-exif",
"lazy_static",
"libc",
"md-5",
"memmap2",
@@ -4591,8 +4582,8 @@ dependencies = [
"parking_lot 0.12.5",
"pprof",
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"prost 0.14.4",
"prost-types 0.14.4",
"protoc-bin-vendored",
"rand 0.10.2",
"redb",
@@ -4601,6 +4592,7 @@ dependencies = [
"rustls",
"rustls-pemfile",
"rusty-leveldb",
"seaweed-common",
"serde",
"serde_json",
"serde_urlencoded",
@@ -4613,13 +4605,15 @@ dependencies = [
"tokio-stream",
"toml",
"tonic",
"tonic-build",
"tonic-prost",
"tonic-prost-build",
"tonic-reflection",
"tower 0.4.13",
"tower-http 0.5.2",
"tower",
"tower-http",
"tracing",
"tracing-subscriber",
"uuid",
"windows-sys 0.61.2",
"x509-parser",
"xxhash-rust",
]
@@ -4966,7 +4960,7 @@ checksum = "b7c566e0f4b284dd6561c786d9cb0142da491f46a9fbed79ea69cdad5db17f21"
dependencies = [
"anyhow",
"heck",
"indexmap 2.13.1",
"indexmap",
"prettyplease",
"syn",
"wasm-metadata",
@@ -4997,7 +4991,7 @@ checksum = "9d66ea20e9553b30172b5e831994e35fbde2d165325bec84fc43dbf6f4eb9cb2"
dependencies = [
"anyhow",
"bitflags 2.11.0",
"indexmap 2.13.1",
"indexmap",
"log",
"serde",
"serde_derive",
@@ -5016,7 +5010,7 @@ checksum = "ecc8ac4bc1dc3381b7f59c34f00b67e18f910c2c0f50015669dde7def656a736"
dependencies = [
"anyhow",
"id-arena",
"indexmap 2.13.1",
"indexmap",
"log",
"semver",
"serde",
+33 -10
View File
@@ -1,7 +1,10 @@
[package]
name = "weed-volume"
version = "0.1.0"
edition = "2021"
edition = "2024"
# The edition needs 1.85; the dependency tree needs more. Verified with
# `cargo +1.91.1 check --all-targets` (1.90 fails on the AWS SDK).
rust-version = "1.91.1"
description = "SeaweedFS Volume Server — Rust implementation"
[lib]
@@ -20,32 +23,44 @@ default = ["5bytes"]
# Pulls redb's experimental_cursor (and therefore experimental-api-5).
redb-experimental-cursor = ["redb/experimental_cursor"]
[lints.clippy]
# Protobuf message literals keep `..Default::default()` on purpose: it is
# what lets a proto gain a field without touching every constructor.
needless_update = "allow"
# Every `unsafe` block states its precondition, right above the block.
undocumented_unsafe_blocks = "warn"
[dependencies]
# Helpers the Rust plugin workers (seaweed-worker) need as well. A path
# dependency because the two trees are separate cargo workspaces with no
# common root manifest.
seaweed-common = { path = "../seaweed-common" }
# Async runtime
tokio = { version = "1", features = ["full"] }
tokio-stream = { version = "0.1", features = ["net"] }
tokio-io-timeout = "1"
# gRPC + protobuf
tonic = { version = "0.12", features = ["tls"] }
tonic-reflection = "0.12"
prost = "0.13"
prost-types = "0.13"
tonic = { version = "0.14", features = ["tls-aws-lc"] }
tonic-prost = "0.14"
tonic-reflection = "0.14"
prost = "0.14"
prost-types = "0.14"
# HTTP server
axum = { version = "0.7", features = ["multipart"] }
axum = { version = "0.8", features = ["multipart"] }
http-body = "1"
hyper = { version = "1", features = ["full"] }
hyper-util = { version = "0.1", features = ["tokio", "service", "server-auto", "http1", "http2"] }
tower = "0.4"
tower-http = { version = "0.5", features = ["cors", "trace"] }
tower = { version = "0.5", features = ["util"] }
tower-http = { version = "0.6", features = ["cors", "trace"] }
# CLI
clap = { version = "4", features = ["derive"] }
# Metrics
prometheus = { version = "0.13", default-features = false, features = ["process"] }
lazy_static = "1"
# JWT
jsonwebtoken = { version = "10", features = ["rust_crypto"] }
@@ -135,11 +150,19 @@ aws-types = "1"
[target.'cfg(unix)'.dependencies]
pprof = { version = "0.15", features = ["prost-codec"] }
# GetDiskFreeSpaceExW for per-path disk capacity on Windows (0.61.2 already
# in the tree via tempfile/mio, so this unifies rather than adding a version).
[target.'cfg(windows)'.dependencies]
windows-sys = { version = "0.61", features = ["Win32_Storage_FileSystem"] }
[dev-dependencies]
tempfile = "3"
# Already a transitive dependency of aws-sdk-s3 at a single locked version;
# needed directly only for the canned HttpClient in remote_storage::s3 tests.
aws-smithy-runtime-api = "1"
[build-dependencies]
tonic-build = "0.12"
tonic-prost-build = "0.14"
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry.
+4 -1
View File
@@ -4,7 +4,10 @@ A drop-in replacement for the [SeaweedFS](https://github.com/seaweedfs/seaweedfs
## Building
Requires Rust 1.75+ (2021 edition).
Requires Rust 1.91.1+ (2024 edition), matching `rust-version` in `Cargo.toml`.
The patch release matters: 1.91.0 does not build. The edition itself only needs
1.85; the higher floor comes from the dependency tree — chiefly the AWS SDK — so
it moves with those crates. CI builds on the latest stable.
```bash
cd seaweed-volume
+7 -2
View File
@@ -3,11 +3,16 @@ fn main() -> Result<(), Box<dyn std::error::Error>> {
// one, so the build needs no package manager and always sees the same
// version. An explicit PROTOC still wins, for packagers supplying their own.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
// SAFETY: a build script's main runs single-threaded before anything
// else in this process, so no other thread can be reading the
// environment concurrently.
unsafe {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
tonic_build::configure()
tonic_prost_build::configure()
.build_server(true)
.build_client(true)
// filer.proto uses proto3 optional, which protoc rejects without this
+7
View File
@@ -168,6 +168,7 @@ message VacuumVolumeCheckRequest {
}
message VacuumVolumeCheckResponse {
double garbage_ratio = 1;
bool disk_space_low = 4; // the volume is read-only solely because its disk is low on space — a cause compaction itself reclaims
}
message VacuumVolumeCompactRequest {
@@ -236,6 +237,7 @@ message VolumeIncrementalCopyResponse {
message VolumeMountRequest {
uint32 volume_id = 1;
optional string collection = 2;
}
message VolumeMountResponse {
}
@@ -258,6 +260,10 @@ message VolumeDeleteRequest {
// when true, do not remove the cloud-tier object backing the volume.
// used for moves where another server is taking over the same .vif.
bool keep_remote_data = 3;
// when true, delete only if every needle is deleted: the volume held
// data once but nothing is live anymore. Passing either check,
// only_empty or this one, is enough to delete.
bool only_garbage = 4;
}
message VolumeDeleteResponse {
}
@@ -470,6 +476,7 @@ message VolumeEcShardsDeleteRequest {
repeated uint32 shard_ids = 3;
bool full_teardown = 4; // pre-encode cleanup: wipe every EC artifact + generation for this volume, not just shard_ids
int64 encode_ts_ns = 5; // full_teardown generation fence: delete only a disk whose .vif generation is strictly OLDER than this; preserve same-or-newer, generation 0, and an unreadable .vif. 0 => wipe-all (shell pre-encode / pre-upgrade)
uint32 delete_generations_older_than = 6; // post-commit cleanup: delete only staged <base>.*.v<N> artifacts with N strictly below this; 0 disables
}
message VolumeEcShardsDeleteResponse {
bool full_teardown_done = 1; // set by a new server that performed full_teardown; absent from an old server lets the caller detect the silent no-op
File diff suppressed because it is too large Load Diff
+1
View File
@@ -1,5 +1,6 @@
pub mod config;
pub mod images;
pub mod malloc_tuning;
pub mod metrics;
pub mod remote_storage;
pub mod security;
+56 -35
View File
@@ -6,19 +6,22 @@ use seaweed_volume::config::{self, VolumeServerConfig};
use seaweed_volume::metrics;
use seaweed_volume::pb::volume_server_pb::volume_server_server::VolumeServerServer;
use seaweed_volume::security::tls::{
build_rustls_server_config, build_rustls_server_config_with_grpc_client_auth,
install_default_crypto_provider, GrpcClientAuthPolicy, TlsPolicy,
GrpcClientAuthPolicy, TlsPolicy, build_rustls_server_config,
build_rustls_server_config_with_grpc_client_auth, install_default_crypto_provider,
};
use seaweed_volume::security::{Guard, SigningKey};
#[cfg(unix)]
use seaweed_volume::server::debug::build_debug_router;
use seaweed_volume::server::grpc_client::load_outgoing_grpc_tls;
use seaweed_volume::server::grpc_client::{
GRPC_INITIAL_WINDOW_SIZE, GRPC_KEEPALIVE_INTERVAL, GRPC_KEEPALIVE_TIMEOUT,
GRPC_MAX_MESSAGE_SIZE, load_outgoing_grpc_tls,
};
use seaweed_volume::server::grpc_server::VolumeGrpcService;
#[cfg(unix)]
use seaweed_volume::server::profiling::CpuProfileSession;
use seaweed_volume::server::request_id::GrpcRequestIdLayer;
use seaweed_volume::server::volume_server::{
build_metrics_router, RuntimeMetricsConfig, VolumeServerState,
RuntimeMetricsConfig, VolumeServerState, build_metrics_router,
};
use seaweed_volume::server::write_queue::WriteQueue;
use seaweed_volume::storage::store::Store;
@@ -31,14 +34,19 @@ type CpuProfileParam = Option<CpuProfileSession>;
#[cfg(not(unix))]
type CpuProfileParam = Option<()>;
const GRPC_MAX_MESSAGE_SIZE: usize = 1 << 30;
const GRPC_KEEPALIVE_INTERVAL: std::time::Duration = std::time::Duration::from_secs(60);
const GRPC_KEEPALIVE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(20);
const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
// The two settings that only make sense for the inbound server. The rest of
// this server's HTTP/2 tuning — keepalive, window sizes, message size — is
// imported from `server::grpc_client` above, which is also what the outgoing
// clients dial with, so the two directions cannot drift apart.
const GRPC_MAX_HEADER_LIST_SIZE: u32 = 8 * 1024 * 1024;
const GRPC_MAX_CONCURRENT_STREAMS: u32 = 1000;
fn main() {
// Before anything allocates: stop glibc from training its mmap threshold
// upward on our large EC buffers and turning them into heap it never
// returns. See seaweed_volume::malloc_tuning for the measurements.
let malloc_tuning = seaweed_volume::malloc_tuning::pin_mmap_threshold();
install_default_crypto_provider();
// Initialize tracing
@@ -65,6 +73,19 @@ fn main() {
"SeaweedFS Volume Server (Rust) v{}",
seaweed_volume::version::full_version()
);
match malloc_tuning {
seaweed_volume::malloc_tuning::MallocTuning::Pinned(bytes) => {
info!("pinned glibc M_MMAP_THRESHOLD to {} bytes", bytes)
}
seaweed_volume::malloc_tuning::MallocTuning::DeferredToEnv => info!(
"an allocator mmap-threshold override ({}) is set; leaving glibc's mmap threshold to the environment",
seaweed_volume::malloc_tuning::MMAP_THRESHOLD_ENV
),
seaweed_volume::malloc_tuning::MallocTuning::Failed => {
warn!("mallopt(M_MMAP_THRESHOLD) failed; large freed buffers may stay resident")
}
seaweed_volume::malloc_tuning::MallocTuning::NotApplicable => {}
}
// Register Prometheus metrics
metrics::register_metrics();
@@ -325,9 +346,6 @@ async fn run(
pre_stop_seconds: config.pre_stop_seconds,
volume_state_notify: tokio::sync::Notify::new(),
write_queue: std::sync::OnceLock::new(),
s3_tier_registry: std::sync::RwLock::new(
seaweed_volume::remote_storage::s3_tier::S3TierRegistry::new(),
),
read_mode: config.read_mode,
allow_untrusted_remote_endpoints: config.allow_untrusted_remote_endpoints,
master_url,
@@ -653,8 +671,7 @@ async fn run(
})
.await
} else {
let incoming =
tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
let incoming = tokio_stream::wrappers::TcpListenerStream::new(grpc_listener);
info!("gRPC server listening on {}", grpc_local_addr);
build_grpc_server_builder()
.layer(GrpcRequestIdLayer)
@@ -1040,15 +1057,17 @@ mod tests {
#[test]
fn test_grpc_server_tls_returns_none_when_files_are_missing() {
assert!(build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
"/missing/server.crt",
"/missing/server.key",
"/missing/ca.crt",
&TlsPolicy::default(),
"",
&[],
)
.is_none()
);
}
#[test]
@@ -1070,19 +1089,21 @@ mod tests {
"-----BEGIN CERTIFICATE-----\nZmFrZQ==\n-----END CERTIFICATE-----\n",
);
assert!(build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none());
assert!(
build_grpc_server_tls_acceptor(
&cert,
&key,
&ca,
&TlsPolicy {
min_version: "TLS 1.0".to_string(),
max_version: "TLS 1.1".to_string(),
cipher_suites: String::new(),
},
"",
&[],
)
.is_none()
);
}
#[test]
+467
View File
@@ -0,0 +1,467 @@
//! Keep glibc from silently converting large short-lived buffers into heap the
//! process never gives back.
//!
//! glibc serves an allocation with `mmap` when it is at least
//! `M_MMAP_THRESHOLD` (128 KiB by default), and `munmap`s it on free, so the
//! pages go straight back to the OS. That threshold is **adaptive**: whenever a
//! block that came from `mmap` is freed, glibc raises the threshold to that
//! block's size — up to 32 MiB — on the theory that a workload repeatedly
//! allocating buffers of that size is better served from the heap.
//!
//! For a volume server that theory is wrong in a specific, expensive way. EC
//! reconstruction and needle reassembly allocate large, short-lived buffers.
//! The first few are mmap'd and freed, which trains the threshold upward; every
//! later buffer of that size is then carved out of the heap instead. Heap
//! memory is only returned to the OS from the top of the arena, so those pages
//! stay resident as anonymous memory for the life of the process. They are
//! still *reusable* — this is not a leak, and a repeat workload does not grow
//! the footprint further — but under a hard cgroup `MemoryMax` they are
//! indistinguishable from a leak, because anonymous pages cannot be reclaimed
//! under pressure the way page cache can. The retained footprint eats exactly
//! the headroom that a burst of maintenance work needs, and the process is
//! OOM-killed while most of its resident memory is free-but-unreturned.
//!
//! Measured on a 17-node cluster (EC 10+4, `--index=redb`), one node, two
//! identical `ec.scrub -mode full` rounds over 10912 EC files each, comparing
//! the same unit restarted with and without a pinned threshold:
//!
//! | | baseline | round 1 | round 2 | 60s idle |
//! |---|---|---|---|---|
//! | default (adaptive) | 10 MB | 84 MB | 88 MB | **88 MB** |
//! | pinned threshold | 10 MB | 13 MB | 14 MB | **14 MB** |
//!
//! 78 MB retained versus 4 MB for identical work. On that cluster's heavier
//! mixed scrub workloads the same effect reached ~600 MB of retained anonymous
//! memory per volume server, against a 3 GiB cap.
//!
//! Calling `mallopt(M_MMAP_THRESHOLD, ...)` sets the threshold *and* disables
//! the dynamic adjustment, which is the documented behaviour of setting it
//! explicitly. We pin it to glibc's own default rather than inventing a value:
//! the goal is to stop the adaptation, not to second-guess the default.
/// glibc's own default `M_MMAP_THRESHOLD`. Pinning to this value changes
/// nothing about which allocations use `mmap` on a freshly started process; it
/// only prevents the threshold from drifting upward later.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
const DEFAULT_MMAP_THRESHOLD: libc::c_int = 128 * 1024;
/// Legacy environment variable glibc reads for the same setting. If an operator
/// has set it, honour their value and do not override it.
pub const MMAP_THRESHOLD_ENV: &str = "MALLOC_MMAP_THRESHOLD_";
/// Modern glibc tunables environment variable. Operators may set the threshold
/// via `GLIBC_TUNABLES=glibc.malloc.mmap_threshold=...` instead of the legacy
/// variable; that override is honoured too.
pub const GLIBC_TUNABLES_ENV: &str = "GLIBC_TUNABLES";
/// The tunable name within `GLIBC_TUNABLES` that maps to `M_MMAP_THRESHOLD`.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
const MMAP_THRESHOLD_TUNABLE: &str = "glibc.malloc.mmap_threshold";
/// Outcome of the tuning attempt, so the caller can log it and tests can assert
/// on it without inspecting global allocator state.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum MallocTuning {
/// Threshold pinned to `DEFAULT_MMAP_THRESHOLD`; dynamic adjustment is off.
Pinned(i32),
/// An allocator override (`MALLOC_MMAP_THRESHOLD_` or
/// `GLIBC_TUNABLES=glibc.malloc.mmap_threshold=...`) was set, so the
/// operator's value wins.
DeferredToEnv,
/// `mallopt` reported failure. Not fatal — the server runs, it just keeps
/// glibc's adaptive behaviour.
Failed,
/// Not glibc, so there is no adaptive threshold to pin.
NotApplicable,
}
/// Pin glibc's mmap threshold unless the operator has set an allocator override.
/// Safe to call more than once; call it before serving traffic, since the point
/// is to prevent the threshold from being trained upward by early allocations.
pub fn pin_mmap_threshold() -> MallocTuning {
pin_mmap_threshold_inner()
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn pin_mmap_threshold_inner() -> MallocTuning {
if operator_mmap_threshold_override_active() {
return MallocTuning::DeferredToEnv;
}
// SAFETY: `mallopt` is a libc entry point that takes two ints and mutates
// only allocator-internal tunables. It has no preconditions and no effect
// on memory this process already owns.
let rc = unsafe { libc::mallopt(libc::M_MMAP_THRESHOLD, DEFAULT_MMAP_THRESHOLD) };
if rc == 1 {
MallocTuning::Pinned(DEFAULT_MMAP_THRESHOLD)
} else {
MallocTuning::Failed
}
}
#[cfg(not(all(target_os = "linux", target_env = "gnu")))]
fn pin_mmap_threshold_inner() -> MallocTuning {
MallocTuning::NotApplicable
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn operator_mmap_threshold_override_active() -> bool {
// MALLOC_MMAP_THRESHOLD_: glibc calls atoi(value) then mallopt, which
// always sets the threshold and disables dynamic adjustment — even for
// empty, negative, or non-numeric values (atoi returns 0). So any presence
// of the variable means the operator's override is in effect.
std::env::var_os(MMAP_THRESHOLD_ENV).is_some()
|| usable_glibc_tunable_threshold(std::env::var_os(GLIBC_TUNABLES_ENV))
}
/// Look for `glibc.malloc.mmap_threshold=<value>` among the colon-separated
/// tunables in `GLIBC_TUNABLES`. glibc's `parse_tunables_string` (elf/dl-tunables.c)
/// rejects the **entire** string (returns -1) if it reaches `\0` before finding
/// `=` in a name (last entry has no `=`), or if any entry's value contains a
/// duplicate `=`. When `parse_tunables_string` returns -1, `parse_tunables`
/// prints a warning and returns immediately without applying ANY tunable —
/// including ones already parsed into the tunables array. We match that by
/// returning `false` for the entire string on any of those conditions.
///
/// glibc parses tunable values with `_dl_strtoul`, which accepts decimal,
/// `0x` hex, `0` octal, an optional sign (negatives wrap to `unsigned long`),
/// and requires the entire value to be consumed; we match that with
/// `dl_strtoul_consumes_all`.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn usable_glibc_tunable_threshold(tunables: Option<std::ffi::OsString>) -> bool {
let s = match tunables.and_then(|v| v.into_string().ok()) {
Some(s) => s,
None => return false,
};
if s.is_empty() {
return false;
}
// Parse the string character-by-character, matching glibc's
// parse_tunables_string logic exactly. Using split(':') would lose the
// distinction between an entry terminated by ':' (skip) and one terminated
// by '\0' with no '=' (reject entire string).
let bytes = s.as_bytes();
let mut pos = 0;
let mut found_threshold = false;
loop {
// Find where the name ends ('=', ':', or end of string).
let name_start = pos;
while pos < bytes.len() && bytes[pos] != b'=' && bytes[pos] != b':' {
pos += 1;
}
// End of string before '=' → glibc returns -1 (reject entire string).
if pos >= bytes.len() {
return false;
}
// ':' before '=' → glibc skips this entry and continues.
if bytes[pos] == b':' {
pos += 1;
continue;
}
// Skip the '='.
let name_end = pos;
pos += 1;
// Find where the value ends ('=', ':', or end of string).
let val_start = pos;
while pos < bytes.len() && bytes[pos] != b'=' && bytes[pos] != b':' {
pos += 1;
}
// '=' in value → glibc returns -1 (reject entire string).
if pos < bytes.len() && bytes[pos] == b'=' {
return false;
}
let key = &s[name_start..name_end];
let val = &s[val_start..pos];
if key == MMAP_THRESHOLD_TUNABLE && dl_strtoul_consumes_all(val) {
found_threshold = true;
}
// End of string → done.
if pos >= bytes.len() {
break;
}
// Skip the ':'.
pos += 1;
}
found_threshold
}
/// Replicate glibc's `_dl_strtoul` (elf/dl-misc.c) just enough to determine
/// whether it would consume the entire string — which is what
/// `tunable_parse_num` checks (`endptr == strval + len`). Returns `true` if
/// glibc would accept the value and apply it.
///
/// `_dl_strtoul` skips leading spaces/tabs, accepts an optional `+`/`-` sign,
/// and parses `0x`-prefixed hex, `0`-prefixed octal, or plain decimal. A
/// negative result wraps to `unsigned long` (`-1` → `SIZE_MAX`). If no digit is
/// found after the sign, the end pointer stays at the current position — which
/// still counts as "consumed" when the string is empty or whitespace-only
/// (value 0). On overflow, `_dl_strtoul` stops at the overflowing digit (endptr
/// does not reach the end), so `tunable_parse_num` rejects the value.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn dl_strtoul_consumes_all(s: &str) -> bool {
let bytes = s.as_bytes();
let mut pos = 0;
// Skip leading whitespace (spaces and tabs, matching _dl_strtoul).
while pos < bytes.len() && (bytes[pos] == b' ' || bytes[pos] == b'\t') {
pos += 1;
}
// Optional sign.
if pos < bytes.len() && (bytes[pos] == b'-' || bytes[pos] == b'+') {
pos += 1;
}
// Must have at least one digit (0-9) to start parsing, unless we're already
// at the end (empty / whitespace-only / sign-only → value 0, consumed).
if pos >= bytes.len() {
return true;
}
if bytes[pos] < b'0' || bytes[pos] > b'9' {
return false;
}
// Determine base: 0x → hex, 0 → octal, else decimal. _dl_strtoul unconditionally
// advances past "0x"/"0X" when the first char is '0' and the next is 'x'/'X',
// even if no hex digit follows — in that case the digit loop breaks immediately,
// endptr reaches the end, and the value is 0.
let base: u32 = if bytes[pos] == b'0'
&& pos + 1 < bytes.len()
&& (bytes[pos + 1] == b'x' || bytes[pos + 1] == b'X')
{
pos += 2; // skip "0x"
16
} else if bytes[pos] == b'0' {
8
} else {
10
};
// Parse digits with overflow detection, matching _dl_strtoul's cutoff/cutlim
// logic. On overflow, _dl_strtoul sets endptr to the overflowing digit and
// returns UINT64_MAX — so the value is NOT fully consumed and
// tunable_parse_num rejects it.
let mut result: u64 = 0;
let cutoff = u64::MAX / base as u64;
let cutlim = u64::MAX % base as u64;
while pos < bytes.len() {
let b = bytes[pos];
let digval: u32 = match digit_value(b, base) {
Some(v) => v,
None => break,
};
if result > cutoff || (result == cutoff && digval as u64 > cutlim) {
// Overflow: _dl_strtoul stops here, endptr points at this digit.
return false;
}
result *= base as u64;
result += digval as u64;
pos += 1;
}
// The entire string must be consumed (matching tunable_parse_num's check).
pos == bytes.len()
}
/// Returns the numeric value of a digit byte in the given base, or `None` if
/// the byte is not a valid digit in that base.
#[cfg(all(target_os = "linux", target_env = "gnu"))]
fn digit_value(b: u8, base: u32) -> Option<u32> {
if (b'0'..=b'0' + (base - 1).min(9) as u8).contains(&b) {
return Some((b - b'0') as u32);
}
if base == 16 {
if (b'a'..=b'f').contains(&b) {
return Some((b - b'a' + 10) as u32);
}
if (b'A'..=b'F').contains(&b) {
return Some((b - b'A' + 10) as u32);
}
}
None
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn env_override_constants_match_glibc_names() {
// Verified against the real accessor rather than a copy of the name, so
// renaming the constant cannot silently break the override contract.
assert_eq!(MMAP_THRESHOLD_ENV, "MALLOC_MMAP_THRESHOLD_");
assert_eq!(GLIBC_TUNABLES_ENV, "GLIBC_TUNABLES");
}
#[test]
fn calling_twice_is_stable() {
// Startup paths get re-entered in tests and in `weed mini`; the second
// call must not report a different outcome from the first.
let first = pin_mmap_threshold();
let second = pin_mmap_threshold();
assert_eq!(first, second);
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn pins_threshold_on_glibc_when_no_override_is_set() {
// The env override is not set in the test process, so this exercises the
// mallopt path. If an override happens to be present, defer to it.
if operator_mmap_threshold_override_active() {
assert_eq!(pin_mmap_threshold(), MallocTuning::DeferredToEnv);
return;
}
assert_eq!(
pin_mmap_threshold(),
MallocTuning::Pinned(DEFAULT_MMAP_THRESHOLD),
"mallopt(M_MMAP_THRESHOLD) should succeed on glibc"
);
}
#[cfg(not(all(target_os = "linux", target_env = "gnu")))]
#[test]
fn is_a_noop_off_glibc() {
// No glibc adaptive threshold exists off glibc, so there is nothing to
// pin regardless of any environment variables that happen to be set.
assert_eq!(pin_mmap_threshold(), MallocTuning::NotApplicable);
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn dl_strtoul_consumes_all_matches_glibc_parser() {
// Decimal — any non-empty decimal integer is accepted, including
// negative (wraps to unsigned) and zero.
assert!(dl_strtoul_consumes_all("131072"));
assert!(dl_strtoul_consumes_all("0"));
assert!(dl_strtoul_consumes_all("-1"));
assert!(dl_strtoul_consumes_all("-131072"));
// Values above i64::MAX are valid for glibc's unsigned parser.
assert!(dl_strtoul_consumes_all("9223372036854775808"));
// Hex with 0x prefix.
assert!(dl_strtoul_consumes_all("0x20000"));
assert!(dl_strtoul_consumes_all("0X20000"));
assert!(dl_strtoul_consumes_all("0x0"));
// Octal with leading 0.
assert!(dl_strtoul_consumes_all("010"));
// Leading whitespace (spaces and tabs) is skipped.
assert!(dl_strtoul_consumes_all(" 131072"));
assert!(dl_strtoul_consumes_all("\t0x20000"));
// Empty and whitespace-only strings are accepted (value 0).
assert!(dl_strtoul_consumes_all(""));
assert!(dl_strtoul_consumes_all(" "));
assert!(dl_strtoul_consumes_all("\t"));
// Sign-only strings are accepted: _dl_strtoul skips the sign, finds no
// digit, sets endptr to the position after the sign (== end of string),
// and returns 0. tunable_parse_num sees endptr == strval + len → true.
assert!(dl_strtoul_consumes_all("-"));
assert!(dl_strtoul_consumes_all("+"));
// Trailing garbage is rejected — _dl_strtoul stops at the first
// non-digit and tunable_parse_num requires the entire string consumed.
assert!(!dl_strtoul_consumes_all("131072abc"));
// In hex mode, a-f are digits, so "0x20000abc" is a valid hex number.
// Use a non-hex character like 'g' to test trailing garbage in hex.
assert!(!dl_strtoul_consumes_all("0x20000g"));
assert!(!dl_strtoul_consumes_all("128K"));
// Non-numeric strings are rejected.
assert!(!dl_strtoul_consumes_all("abc"));
// "0x" with no hex digits: _dl_strtoul advances past "0x", the digit loop
// breaks immediately (no hex digit), endptr reaches the end, value is 0.
// tunable_parse_num accepts it.
assert!(dl_strtoul_consumes_all("0x"));
assert!(dl_strtoul_consumes_all("0X"));
// Overflow: _dl_strtoul stops at the overflowing digit (endptr points
// there, not at the end), so tunable_parse_num rejects the value.
assert!(!dl_strtoul_consumes_all("18446744073709551616")); // u64::MAX + 1
assert!(!dl_strtoul_consumes_all("99999999999999999999")); // 20 nines
assert!(!dl_strtoul_consumes_all("0x10000000000000000")); // 2^64
// u64::MAX itself is accepted: the last digit (5) equals cutlim (=5),
// so the overflow check (digval > cutlim) is false.
assert!(dl_strtoul_consumes_all("18446744073709551615")); // u64::MAX
}
#[cfg(all(target_os = "linux", target_env = "gnu"))]
#[test]
fn usable_glibc_tunable_threshold_detects_mmap_threshold() {
// Decimal, hex, octal, negative, and zero values are all accepted by
// glibc's _dl_strtoul and cause the threshold to be pinned.
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=131072".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=0x20000".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=0".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=-1".into()
)));
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=9223372036854775808".into()
)));
// u64::MAX is accepted by _dl_strtoul (last digit == cutlim, no overflow).
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=18446744073709551615".into()
)));
// Appears alongside other tunables.
assert!(usable_glibc_tunable_threshold(Some(
"glibc.cpu.x=1:glibc.malloc.mmap_threshold=131072".into()
)));
// Leading ':' is accepted — glibc skips the empty entry and continues.
assert!(usable_glibc_tunable_threshold(Some(
":glibc.malloc.mmap_threshold=131072".into()
)));
// Empty value is accepted by _dl_strtoul (value 0).
assert!(usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=".into()
)));
// Non-numeric values are rejected by _dl_strtoul.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=abc".into()
)));
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=128K".into()
)));
// A malformed sibling entry (duplicate '=') makes glibc reject the
// entire string, so we must not accept the threshold entry either.
// This applies regardless of whether the threshold is before or after
// the malformed entry — parse_tunables_string returns -1, and
// parse_tunables discards all tunables without applying any.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.check=2=2:glibc.malloc.mmap_threshold=131072".into()
)));
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:glibc.malloc.check=2=2".into()
)));
// A trailing entry with no '=' makes glibc reject the entire string
// (parse_tunables_string hits '\0' before '=' and returns -1).
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:glibc.cpu.x".into()
)));
// A trailing ':' makes glibc reject the entire string (the empty entry
// after ':' hits '\0' before '=' and returns -1).
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.malloc.mmap_threshold=262144:".into()
)));
// Unrelated tunables do not count.
assert!(!usable_glibc_tunable_threshold(Some(
"glibc.cpu.x=1".into()
)));
assert!(!usable_glibc_tunable_threshold(None));
}
}
+238 -118
View File
@@ -3,10 +3,10 @@
//! Mirrors the Go SeaweedFS volume server metrics.
use prometheus::{
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounterVec, IntGauge, IntGaugeVec,
Opts, Registry, TextEncoder,
self, Encoder, GaugeVec, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge,
IntGaugeVec, Opts, Registry, TextEncoder,
};
use std::sync::Once;
use std::sync::{LazyLock, Once};
use crate::version;
@@ -16,203 +16,320 @@ pub struct PushGatewayConfig {
pub interval_seconds: u32,
}
lazy_static::lazy_static! {
pub static ref REGISTRY: Registry = Registry::new();
pub static REGISTRY: LazyLock<Registry> = LazyLock::new(Registry::new);
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
// ---- Request metrics (Go: VolumeServerRequestCounter, VolumeServerRequestHistogram) ----
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static ref REQUEST_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_request_total", "Volume server requests"),
/// Request counter with labels `type` (HTTP method) and `code` (HTTP status).
pub static REQUEST_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_request_total",
"Volume server requests",
),
&["type", "code"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Request duration histogram with label `type` (HTTP method).
pub static ref REQUEST_DURATION: HistogramVec = HistogramVec::new(
/// Request duration histogram with label `type` (HTTP method).
pub static REQUEST_DURATION: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_request_seconds",
"Volume server request duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
// ---- Handler counters (Go: VolumeServerHandlerCounter) ----
/// Handler-level operation counter with label `type`.
pub static ref HANDLER_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_handler_total", "Volume server handler counters"),
/// Handler-level operation counter with label `type`.
pub static HANDLER_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_handler_total",
"Volume server handler counters",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
// ---- Vacuuming metrics (Go: VolumeServerVacuuming*) ----
/// Vacuuming compact counter with label `success` (true/false).
pub static ref VACUUMING_COMPACT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_compact_count", "Counter of volume vacuuming Compact counter"),
/// Vacuuming compact counter with label `success` (true/false).
pub static VACUUMING_COMPACT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_compact_count",
"Counter of volume vacuuming Compact counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming commit counter with label `success` (true/false).
pub static ref VACUUMING_COMMIT_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_volumeServer_vacuuming_commit_count", "Counter of volume vacuuming commit counter"),
/// Vacuuming commit counter with label `success` (true/false).
pub static VACUUMING_COMMIT_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_vacuuming_commit_count",
"Counter of volume vacuuming commit counter",
),
&["success"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static ref VACUUMING_HISTOGRAM: HistogramVec = HistogramVec::new(
/// Vacuuming duration histogram with label `type` (compact/commit).
pub static VACUUMING_HISTOGRAM: LazyLock<HistogramVec> = LazyLock::new(|| {
HistogramVec::new(
HistogramOpts::new(
"SeaweedFS_volumeServer_vacuuming_seconds",
"Volume vacuuming duration in seconds",
).buckets(exponential_buckets(0.0001, 2.0, 24)),
)
.buckets(exponential_buckets(0.0001, 2.0, 24)),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
// ---- Volume gauges (Go: VolumeServerVolumeGauge, VolumeServerReadOnlyVolumeGauge) ----
/// Volumes per collection and type (volume/ec_shards).
pub static ref VOLUME_GAUGE: GaugeVec = GaugeVec::new(
/// Volumes per collection and type (volume/ec_shards).
pub static VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_volumes", "Number of volumes"),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Read-only volumes per collection and type.
pub static ref READ_ONLY_VOLUME_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_read_only_volumes", "Number of read-only volumes."),
/// Read-only volumes per collection and type.
pub static READ_ONLY_VOLUME_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_read_only_volumes",
"Number of read-only volumes.",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Maximum number of volumes this server can hold.
pub static ref MAX_VOLUMES: IntGauge = IntGauge::new(
/// Maximum number of volumes this server can hold.
pub static MAX_VOLUMES: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_max_volumes",
"Maximum number of volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
// ---- Disk size gauges (Go: VolumeServerDiskSizeGauge) ----
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static ref DISK_SIZE_GAUGE: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_total_disk_size", "Actual disk size used by volumes"),
/// Actual disk size used by volumes per collection and type (normal/deleted_bytes/ec).
pub static DISK_SIZE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_total_disk_size",
"Actual disk size used by volumes",
),
&["collection", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
// ---- Resource gauges (Go: VolumeServerResourceGauge) ----
/// Disk resource usage per directory and type (all/used/free/avail).
pub static ref RESOURCE_GAUGE: GaugeVec = GaugeVec::new(
/// Disk resource usage per directory and type (all/used/free/avail).
pub static RESOURCE_GAUGE: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_volumeServer_resource", "Server resource usage"),
&["name", "type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
// ---- In-flight gauges (Go: VolumeServerInFlightRequestsGauge, InFlightDownload/UploadSize) ----
/// In-flight requests per HTTP method.
pub static ref INFLIGHT_REQUESTS_GAUGE: IntGaugeVec = IntGaugeVec::new(
Opts::new("SeaweedFS_volumeServer_in_flight_requests", "Current number of in-flight requests being handled by volume server."),
/// In-flight requests per HTTP method.
pub static INFLIGHT_REQUESTS_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_in_flight_requests",
"Current number of in-flight requests being handled by volume server.",
),
&["type"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent download limit in bytes.
pub static ref CONCURRENT_DOWNLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent download limit in bytes.
pub static CONCURRENT_DOWNLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_download_limit",
"Limit for total concurrent download size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Concurrent upload limit in bytes.
pub static ref CONCURRENT_UPLOAD_LIMIT: IntGauge = IntGauge::new(
/// Concurrent upload limit in bytes.
pub static CONCURRENT_UPLOAD_LIMIT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_concurrent_upload_limit",
"Limit for total concurrent upload size in bytes",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight download bytes.
pub static ref INFLIGHT_DOWNLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight download bytes.
pub static INFLIGHT_DOWNLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_download_size",
"In flight total download size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current in-flight upload bytes.
pub static ref INFLIGHT_UPLOAD_SIZE: IntGauge = IntGauge::new(
/// Current in-flight upload bytes.
pub static INFLIGHT_UPLOAD_SIZE: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"SeaweedFS_volumeServer_in_flight_upload_size",
"In flight total upload size.",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static ref UPLOAD_ERROR_COUNTER: IntCounterVec = IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created");
/// Upload error counter by HTTP status code. Code "0" = transport error (no response).
pub static UPLOAD_ERROR_COUNTER: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new("SeaweedFS_upload_error_total",
"Counter of upload errors by HTTP status code. Code 0 means transport error (no response received)."),
&["code"],
).expect("metric can be created")
});
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
// ---- Scrubbing metrics (Go: VolumeServerScrub*) ----
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static ref SCRUB_LAST_TIME_SECONDS: GaugeVec = GaugeVec::new(
/// Last scrub execution time, as seconds since UNIX epoch, with label `mode`.
pub static SCRUB_LAST_TIME_SECONDS: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_last_time_seconds",
"Last scrub execution time, as seconds since UNIX epoch.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_VOLUME_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall volumes with issues detected during scrubbing, with label `mode`.
pub static SCRUB_VOLUME_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_volume_failures",
"Counter of overall volumes with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static ref SCRUB_SHARD_FAILURES: IntCounterVec = IntCounterVec::new(
/// Counter of overall EC shards with issues detected during scrubbing, with label `mode`.
pub static SCRUB_SHARD_FAILURES: LazyLock<IntCounterVec> = LazyLock::new(|| {
IntCounterVec::new(
Opts::new(
"SeaweedFS_volumeServer_scrub_shard_failures",
"Counter of overall EC shards with issues detected during scrubbing.",
),
&["mode"],
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Legacy aliases for backward compat with existing code ----
/// Counter of storage read/write EIO errors on volumes and EC shards.
/// Mirrors Go's VolumeServerStorageIoErrorCounter.
pub static STORAGE_IO_ERROR_COUNTER: LazyLock<IntCounter> = LazyLock::new(|| {
IntCounter::new(
"SeaweedFS_volumeServer_storage_io_error_total",
"Counter of storage read/write EIO errors on volumes and EC shards.",
)
.expect("metric can be created")
});
/// Total number of volumes on this server (flat gauge).
pub static ref VOLUMES_TOTAL: IntGauge = IntGauge::new(
"volume_server_volumes_total",
"Total number of volumes",
).expect("metric can be created");
/// Number of volumes quarantined due to storage IO errors.
/// Mirrors Go's VolumeServerIoQuarantineGauge.
pub static IO_QUARANTINE_GAUGE: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new(
"SeaweedFS_volumeServer_io_quarantine",
"Number of volumes or EC shards quarantined due to storage IO errors.",
),
&["kind"],
)
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static ref DISK_SIZE_BYTES: IntGaugeVec = IntGaugeVec::new(
// ---- Legacy aliases for backward compat with existing code ----
/// Total number of volumes on this server (flat gauge).
pub static VOLUMES_TOTAL: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new("volume_server_volumes_total", "Total number of volumes")
.expect("metric can be created")
});
/// Disk size in bytes per directory.
pub static DISK_SIZE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_size_bytes", "Disk size in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Disk free bytes per directory.
pub static ref DISK_FREE_BYTES: IntGaugeVec = IntGaugeVec::new(
/// Disk free bytes per directory.
pub static DISK_FREE_BYTES: LazyLock<IntGaugeVec> = LazyLock::new(|| {
IntGaugeVec::new(
Opts::new("volume_server_disk_free_bytes", "Disk free space in bytes"),
&["dir"],
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Current number of in-flight requests (flat gauge).
pub static ref INFLIGHT_REQUESTS: IntGauge = IntGauge::new(
/// Current number of in-flight requests (flat gauge).
pub static INFLIGHT_REQUESTS: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_inflight_requests",
"Current number of in-flight requests",
).expect("metric can be created");
)
.expect("metric can be created")
});
/// Total number of files stored across all volumes.
pub static ref VOLUME_FILE_COUNT: IntGauge = IntGauge::new(
/// Total number of files stored across all volumes.
pub static VOLUME_FILE_COUNT: LazyLock<IntGauge> = LazyLock::new(|| {
IntGauge::new(
"volume_server_volume_file_count",
"Total number of files stored across all volumes",
).expect("metric can be created");
)
.expect("metric can be created")
});
// ---- Build info (Go: BuildInfo) ----
// ---- Build info (Go: BuildInfo) ----
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static ref BUILD_INFO: GaugeVec = GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created");
}
/// Build information gauge, always set to 1. Matches Go:
/// Namespace="SeaweedFS", Subsystem="build", Name="info",
/// labels: version, commit, sizelimit, goos, goarch.
pub static BUILD_INFO: LazyLock<GaugeVec> = LazyLock::new(|| {
GaugeVec::new(
Opts::new("SeaweedFS_build_info", "A metric with a constant '1' value labeled by version, commit, sizelimit, goos, and goarch from which SeaweedFS was built."),
&["version", "commit", "sizelimit", "goos", "goarch"],
).expect("metric can be created")
});
/// Generate exponential bucket boundaries for histograms.
fn exponential_buckets(start: f64, factor: f64, count: usize) -> Vec<f64> {
@@ -232,6 +349,7 @@ pub const DOWNLOAD_LIMIT_COND: &str = "downloadLimitCondition";
pub const UPLOAD_LIMIT_COND: &str = "uploadLimitCondition";
pub const READ_PROXY_REQ: &str = "readProxyRequest";
pub const READ_REDIRECT_REQ: &str = "readRedirectRequest";
pub const READ_DELETED_NEEDLE: &str = "readDeletedNeedle";
pub const EMPTY_READ_PROXY_LOC: &str = "emptyReadProxyLocaction";
pub const FAILED_READ_PROXY_REQ: &str = "failedReadProxyRequest";
@@ -283,6 +401,8 @@ pub fn register_metrics() {
Box::new(SCRUB_LAST_TIME_SECONDS.clone()),
Box::new(SCRUB_VOLUME_FAILURES.clone()),
Box::new(SCRUB_SHARD_FAILURES.clone()),
Box::new(STORAGE_IO_ERROR_COUNTER.clone()),
Box::new(IO_QUARANTINE_GAUGE.clone()),
// Legacy metrics
Box::new(VOLUMES_TOTAL.clone()),
Box::new(DISK_SIZE_BYTES.clone()),
@@ -358,10 +478,8 @@ fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
type_value = Some(label.get_value().to_string());
}
}
if matches_collection {
if let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
if matches_collection && let Some(ref tv) = type_value {
let _ = gauge.remove_label_values(&[collection, tv]);
}
}
}
@@ -408,7 +526,7 @@ pub async fn push_metrics_once(
#[cfg(test)]
mod tests {
use super::*;
use axum::{routing::put, Router};
use axum::{Router, routing::put};
use std::sync::{Arc, Mutex};
#[test]
@@ -480,7 +598,9 @@ mod tests {
register_metrics();
VOLUME_GAUGE.with_label_values(&["pics", "volume"]).set(2.0);
VOLUME_GAUGE.with_label_values(&["pics", "ec_shards"]).set(3.0);
VOLUME_GAUGE
.with_label_values(&["pics", "ec_shards"])
.set(3.0);
READ_ONLY_VOLUME_GAUGE
.with_label_values(&["pics", "volume"])
.set(1.0);
@@ -119,7 +119,11 @@ pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
/// reachable for callers whose target legitimately sits on an internal network
/// (peer volume servers), while still blocking loopback, link-local (IMDS) and
/// unspecified. Mirrors Go's `checkBlockedIPPolicy`.
pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool) -> Result<(), String> {
pub fn check_blocked_ip_policy(
endpoint: &str,
ip: IpAddr,
allow_private: bool,
) -> Result<(), String> {
// Normalize IPv4-mapped IPv6 (`::ffff:a.b.c.d`) to its IPv4 form so the
// IPv4 deny rules apply. The OS routes these to the embedded IPv4 address,
// so without this `::ffff:127.0.0.1` / `::ffff:169.254.169.254` would slip
@@ -173,10 +177,10 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if let IpAddr::V6(v6) = ip
&& let Some(v4) = embedded_transition_ipv4(v6)
{
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
Ok(())
}
@@ -214,9 +218,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
// Authority is everything up to the first '/', '?', or '#'.
let after = &trimmed[scheme_end + 3..];
let authority_end = after
.find(|c| c == '/' || c == '?' || c == '#')
.unwrap_or(after.len());
let authority_end = after.find(['/', '?', '#']).unwrap_or(after.len());
let authority = &after[..authority_end];
// Strip optional userinfo ("user:pass@").
@@ -233,7 +235,7 @@ fn precheck_endpoint(endpoint: &str) -> Result<HostCheck, String> {
return Err(format!(
"remote endpoint {:?} has a malformed IPv6 host",
endpoint
))
));
}
}
} else {
@@ -309,7 +311,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
return Err("replica target is empty".to_string());
}
if trimmed.contains("://") || trimmed.contains(['/', '?', '#', '@', '\\']) {
return Err(format!("replica target {:?} must be a bare host:port", target));
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
// Require an explicit host:port, handling `[IPv6]:port`. A bracketless IPv6
@@ -318,12 +323,22 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let host = if let Some(rest) = trimmed.strip_prefix('[') {
match rest.split_once(']') {
Some((h, port)) if port.starts_with(':') && port.len() > 1 => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
} else {
match trimmed.rsplit_once(':') {
Some((h, port)) if !port.is_empty() && !h.contains(':') => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
_ => {
return Err(format!(
"replica target {:?} must be a bare host:port",
target
));
}
}
};
@@ -342,7 +357,10 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let addrs = resolve_host(host).await?;
if addrs.is_empty() {
return Err(format!("resolve replica target host {:?}: no addresses", host));
return Err(format!(
"resolve replica target host {:?}: no addresses",
host
));
}
for ip in addrs {
check_blocked_ip_policy(target, ip, true)?;
@@ -350,6 +368,56 @@ pub async fn validate_replica_target(target: &str) -> Result<(), String> {
Ok(())
}
/// Resolve `host`, re-apply the replica deny list (private peers allowed) to
/// every resolved address, and connect to the first one that passes -- the
/// connect-time twin of [`validate_replica_target`], so a hostname whose DNS
/// answer flips to a blocked address after the up-front check is still refused.
/// Mirrors Go's `guardedDialerPolicy` with allowPrivate=true.
pub async fn guarded_tcp_connect(
host: &str,
port: u16,
endpoint: &str,
) -> std::io::Result<tokio::net::TcpStream> {
use std::io::{Error, ErrorKind};
let denied = |e: String| Error::new(ErrorKind::PermissionDenied, e);
if is_blocked_imds_host(&host.to_ascii_lowercase()) {
return Err(denied(format!(
"remote endpoint {:?} targets instance metadata service",
endpoint
)));
}
if let Ok(ip) = host.parse::<IpAddr>() {
check_blocked_ip_policy(endpoint, ip, true).map_err(denied)?;
return tokio::net::TcpStream::connect((ip, port)).await;
}
let lookup = tokio::net::lookup_host((host.to_string(), port));
let addrs = tokio::time::timeout(std::time::Duration::from_secs(2), lookup)
.await
.map_err(|_| {
Error::new(
ErrorKind::TimedOut,
format!("resolve remote endpoint host {:?}: timed out", host),
)
})??;
let mut first_block_err: Option<String> = None;
for addr in addrs {
if let Err(e) = check_blocked_ip_policy(endpoint, addr.ip(), true) {
if first_block_err.is_none() {
first_block_err = Some(e);
}
continue;
}
return tokio::net::TcpStream::connect(addr).await;
}
Err(denied(first_block_err.unwrap_or_else(|| {
format!("resolve remote endpoint host {:?}: no addresses", host)
})))
}
#[cfg(test)]
mod tests {
use super::*;
@@ -380,22 +448,30 @@ mod tests {
#[test]
fn rejects_empty_and_bad_scheme() {
assert!(precheck_endpoint("").unwrap_err().contains("empty"));
assert!(precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https"));
assert!(precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https"));
assert!(
precheck_endpoint("ftp://example.com/")
.unwrap_err()
.contains("http or https")
);
assert!(
precheck_endpoint("example.com/")
.unwrap_err()
.contains("http or https")
);
}
#[test]
fn rejects_imds_hostnames() {
assert!(precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service"));
assert!(precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service"));
assert!(
precheck_endpoint("http://metadata.google.internal/")
.unwrap_err()
.contains("metadata service")
);
assert!(
precheck_endpoint("http://metadata/")
.unwrap_err()
.contains("metadata service")
);
}
#[test]
@@ -415,27 +491,41 @@ mod tests {
#[test]
fn check_blocked_ip_matches_resolved_categories() {
// Mirror Go's "host resolves to X" cases at the address level.
assert!(check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local"));
assert!(check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private"));
assert!(check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT"));
assert!(check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("169.254.10.20"))
.unwrap_err()
.contains("link-local")
);
assert!(
check_blocked_ip("e", ip("10.1.2.3"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("172.20.0.5"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("192.168.1.1"))
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("100.64.0.42"))
.unwrap_err()
.contains("CGNAT")
);
assert!(
check_blocked_ip("e", ip("fc00::1"))
.unwrap_err()
.contains("private")
);
assert!(check_blocked_ip("e", ip("52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("2606:4700:4700::1111")).is_ok());
}
@@ -478,33 +568,45 @@ mod tests {
assert!(check_blocked_ip("e", ip("2001::f7f7:f7f7")).is_ok());
assert!(check_blocked_ip("e", ip("::808:808")).is_ok());
// Bracketed transition literal via the full endpoint path.
assert!(precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata"));
assert!(
precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata")
);
}
#[test]
fn rejects_ipv4_mapped_ipv6() {
// IPv4-mapped IPv6 must be unmapped so the IPv4 rules catch it.
assert!(check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private"));
assert!(
check_blocked_ip("e", ip("::ffff:127.0.0.1"))
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("::ffff:169.254.169.254"))
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip("e", ip("::ffff:10.0.0.1"))
.unwrap_err()
.contains("private")
);
// A mapped public address still passes, and genuine IPv6 loopback is
// still caught by the V6 path.
assert!(check_blocked_ip("e", ip("::ffff:52.216.10.10")).is_ok());
assert!(check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback"));
assert!(
check_blocked_ip("e", ip("::1"))
.unwrap_err()
.contains("loopback")
);
// Bracketed mapped literal via the full endpoint path.
assert!(precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback"));
assert!(
precheck_endpoint("http://[::ffff:127.0.0.1]/")
.unwrap_err()
.contains("loopback")
);
}
#[test]
@@ -514,60 +616,86 @@ mod tests {
assert!(check_blocked_ip_policy("e", ip("192.168.1.5"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("100.64.0.42"), true).is_ok());
// Loopback / IMDS / unspecified stay blocked even when private is allowed.
assert!(check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified"));
assert!(
check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified")
);
}
#[tokio::test]
async fn validate_replica_target_rejects_and_allows() {
// A path plus a trailing ?a= would otherwise swallow ?type=replicate.
assert!(validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("")
.await
.unwrap_err()
.contains("empty"));
assert!(
validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port")
);
assert!(
validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback")
);
assert!(
validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata")
);
assert!(
validate_replica_target("")
.await
.unwrap_err()
.contains("empty")
);
// Legitimate peer volume servers on private networks pass.
assert!(validate_replica_target("10.0.0.7:8080").await.is_ok());
assert!(validate_replica_target("192.168.1.5:8080").await.is_ok());
+1 -1
View File
@@ -7,7 +7,7 @@ pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::{validate_remote_endpoint, validate_replica_target};
pub use endpoint_guard::{guarded_tcp_connect, validate_remote_endpoint, validate_replica_target};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
+221 -23
View File
@@ -2,9 +2,10 @@
//!
//! Works with AWS S3, MinIO, SeaweedFS S3, and all S3-compatible providers.
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::Client;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::error::{DisplayErrorContext, SdkError};
use aws_sdk_s3::primitives::ByteStream;
use super::{RemoteEntry, RemoteStorageClient, RemoteStorageError};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
@@ -25,6 +26,23 @@ impl S3RemoteStorageClient {
endpoint: &str,
force_path_style: bool,
) -> Self {
let client = Client::from_conf(
Self::config_builder(access_key, secret_key, region, endpoint, force_path_style)
.build(),
);
S3RemoteStorageClient { client, conf }
}
/// Build the SDK config for the given credentials and endpoint. Split out so
/// tests can attach a canned HTTP client before building the [`Client`].
fn config_builder(
access_key: &str,
secret_key: &str,
region: &str,
endpoint: &str,
force_path_style: bool,
) -> aws_sdk_s3::config::Builder {
let region = if region.is_empty() {
"us-east-1"
} else {
@@ -49,9 +67,7 @@ impl S3RemoteStorageClient {
s3_config = s3_config.endpoint_url(endpoint);
}
let client = Client::from_conf(s3_config.build());
S3RemoteStorageClient { client, conf }
s3_config
}
}
@@ -75,13 +91,14 @@ impl RemoteStorageClient for S3RemoteStorageClient {
req = req.range(format!("bytes={}-", offset));
}
let resp = req.send().await.map_err(|e| {
let msg = format!("{}", e);
if msg.contains("NoSuchKey") || msg.contains("404") {
let resp = req.send().await.map_err(|e| match e {
// Go compares `aerr.Code()` to NoSuchKey on GET
// (s3_storage_client.go:436): a bare 404 maps to "NotFound"
// and stays a generic error, as it does here.
SdkError::ServiceError(ref se) if se.err().is_no_such_key() => {
RemoteStorageError::ObjectNotFound(format!("{}/{}", loc.bucket, key))
} else {
RemoteStorageError::Other(format!("s3 get object: {}", e))
}
e => RemoteStorageError::Other(format!("s3 get object: {}", DisplayErrorContext(&e))),
})?;
let data = resp
@@ -108,7 +125,9 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.body(ByteStream::from(data.to_vec()))
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 put object: {}", e)))?;
.map_err(|e| {
RemoteStorageError::Other(format!("s3 put object: {}", DisplayErrorContext(&e)))
})?;
Ok(RemoteEntry {
size: data.len() as i64,
@@ -134,13 +153,18 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.key(key)
.send()
.await
.map_err(|e| {
let msg = format!("{}", e);
if msg.contains("404") || msg.contains("NotFound") {
.map_err(|e| match e {
// Go checks only the raw HTTP status on HEAD
// (s3_storage_client.go:373): a HEAD response carries no
// error body, so a 404 is not-found whatever code the SDK
// assigns, and a non-404 is not.
SdkError::ServiceError(ref se) if se.raw().status().as_u16() == 404 => {
RemoteStorageError::ObjectNotFound(format!("{}/{}", loc.bucket, key))
} else {
RemoteStorageError::Other(format!("s3 head object: {}", e))
}
e => RemoteStorageError::Other(format!(
"s3 head object: {}",
DisplayErrorContext(&e)
)),
})?;
Ok(RemoteEntry {
@@ -160,18 +184,17 @@ impl RemoteStorageClient for S3RemoteStorageClient {
.key(key)
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 delete object: {}", e)))?;
.map_err(|e| {
RemoteStorageError::Other(format!("s3 delete object: {}", DisplayErrorContext(&e)))
})?;
Ok(())
}
async fn list_buckets(&self) -> Result<Vec<String>, RemoteStorageError> {
let resp = self
.client
.list_buckets()
.send()
.await
.map_err(|e| RemoteStorageError::Other(format!("s3 list buckets: {}", e)))?;
let resp = self.client.list_buckets().send().await.map_err(|e| {
RemoteStorageError::Other(format!("s3 list buckets: {}", DisplayErrorContext(&e)))
})?;
Ok(resp
.buckets()
@@ -184,3 +207,178 @@ impl RemoteStorageClient for S3RemoteStorageClient {
&self.conf
}
}
#[cfg(test)]
pub(crate) mod tests {
use super::*;
use aws_sdk_s3::config::http::{HttpRequest, HttpResponse};
use aws_sdk_s3::config::retry::RetryConfig;
use aws_sdk_s3::config::{HttpClient, RuntimeComponents};
use aws_sdk_s3::primitives::SdkBody;
use aws_smithy_runtime_api::client::http::{
HttpConnector, HttpConnectorFuture, HttpConnectorSettings, SharedHttpConnector,
};
use aws_smithy_runtime_api::http::StatusCode;
/// An SDK HTTP client that answers every request with one canned response,
/// so the error-mapping paths can be exercised without a network or a
/// running S3 server.
#[derive(Debug, Clone)]
pub(crate) struct CannedResponse {
pub(crate) status: u16,
pub(crate) body: &'static str,
}
impl HttpConnector for CannedResponse {
fn call(&self, _request: HttpRequest) -> HttpConnectorFuture {
let status = StatusCode::try_from(self.status).expect("valid HTTP status");
HttpConnectorFuture::ready(Ok(HttpResponse::new(status, SdkBody::from(self.body))))
}
}
impl HttpClient for CannedResponse {
fn http_connector(
&self,
_settings: &HttpConnectorSettings,
_components: &RuntimeComponents,
) -> SharedHttpConnector {
SharedHttpConnector::new(self.clone())
}
}
fn client_with(status: u16, body: &'static str) -> S3RemoteStorageClient {
let config = S3RemoteStorageClient::config_builder(
"AKIATEST",
"secret",
"us-east-1",
"http://127.0.0.1:1",
true,
)
.http_client(CannedResponse { status, body })
.retry_config(RetryConfig::disabled())
.build();
S3RemoteStorageClient {
client: Client::from_conf(config),
conf: RemoteConf::default(),
}
}
fn location() -> RemoteStorageLocation {
RemoteStorageLocation {
name: "remote".to_string(),
bucket: "bucket".to_string(),
path: "/dir/missing".to_string(),
..Default::default()
}
}
pub(crate) const NO_SUCH_KEY: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>NoSuchKey</Code><Message>The specified key does not exist.</Message><Key>dir/missing</Key></Error>"#;
const NOT_FOUND_BODY: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>NotFound</Code><Message>Not Found</Message></Error>"#;
const ACCESS_DENIED: &str = r#"<?xml version="1.0" encoding="UTF-8"?>
<Error><Code>AccessDenied</Code><Message>Access Denied</Message></Error>"#;
#[tokio::test]
async fn get_no_such_key_is_object_not_found() {
let err = client_with(404, NO_SUCH_KEY)
.read_file(&location(), 0, 0)
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::ObjectNotFound(path) if path == "bucket/dir/missing"),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn get_bare_404_is_not_object_not_found() {
// Go compares codes, not statuses, on GET: a body-less 404 stays generic.
let err = client_with(404, "")
.read_file(&location(), 0, 0)
.await
.unwrap_err();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
}
#[tokio::test]
async fn head_404_is_object_not_found() {
let err = client_with(404, "")
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::ObjectNotFound(path) if path == "bucket/dir/missing"),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn head_404_with_foreign_error_body_is_object_not_found() {
// The raw status check makes a 404 not-found whatever body it carries.
let err = client_with(404, NO_SUCH_KEY)
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(err, RemoteStorageError::ObjectNotFound(_)),
"expected ObjectNotFound, got {err:?}"
);
}
#[tokio::test]
async fn head_not_found_code_on_a_non_404_status_is_not_object_not_found() {
// A NotFound body on a non-404 status stays an error, as in Go.
let err = client_with(400, NOT_FOUND_BODY)
.stat_file(&location())
.await
.unwrap_err();
assert!(
matches!(&err, RemoteStorageError::Other(msg) if msg.contains("NotFound")),
"expected Other naming the code, got {err:?}"
);
}
#[tokio::test]
async fn get_access_denied_keeps_service_error_code() {
let err = client_with(403, ACCESS_DENIED)
.read_file(&location(), 0, 0)
.await
.unwrap_err();
let msg = err.to_string();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
assert!(
msg.contains("AccessDenied"),
"message should carry the S3 error code, got: {msg}"
);
assert!(
!msg.ends_with("service error"),
"message should not be the bare SdkError Display, got: {msg}"
);
}
#[tokio::test]
async fn head_access_denied_keeps_service_error_code() {
let err = client_with(403, ACCESS_DENIED)
.stat_file(&location())
.await
.unwrap_err();
let msg = err.to_string();
assert!(
matches!(err, RemoteStorageError::Other(_)),
"expected Other, got {err:?}"
);
assert!(
msg.contains("AccessDenied"),
"message should carry the S3 error code, got: {msg}"
);
}
}
+410 -59
View File
@@ -7,15 +7,66 @@ use std::collections::HashMap;
use std::future::Future;
use std::sync::{Arc, OnceLock, RwLock};
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::types::{CompletedMultipartUpload, CompletedPart};
use aws_sdk_s3::Client;
use aws_sdk_s3::config::http::HttpResponse;
use aws_sdk_s3::config::{BehaviorVersion, Credentials, Region};
use aws_sdk_s3::error::{DisplayErrorContext, SdkError};
use aws_sdk_s3::operation::get_object::GetObjectError;
use aws_sdk_s3::operation::head_object::HeadObjectError;
use aws_sdk_s3::types::{CompletedMultipartUpload, CompletedPart};
use tokio::io::{AsyncReadExt, AsyncSeekExt, AsyncWriteExt};
use tokio::sync::Semaphore;
/// Concurrency limit for multipart upload/download (matches Go's s3manager).
const CONCURRENCY: usize = 5;
/// A tier transfer failure. The variant is what callers match on; the
/// message is the operator-facing text.
#[derive(Debug, thiserror::Error)]
pub enum TierError {
/// The remote object does not exist.
#[error("{0}")]
NotFound(String),
/// An S3 request or a local file operation failed.
#[error("{0}")]
Io(String),
/// The tier I/O runtime could not be built or dropped the task.
#[error("{0}")]
RuntimeUnavailable(String),
/// The progress callback asked to stop.
#[error("{0}")]
Aborted(String),
}
// Not-found rules as in remote_storage/s3.rs: HEAD by the raw 404 status,
// GET by the NoSuchKey code only.
fn head_object_error(key: &str, e: SdkError<HeadObjectError, HttpResponse>) -> TierError {
let message = format!("failed to head object {}: {}", key, DisplayErrorContext(&e));
match e {
SdkError::ServiceError(ref se) if se.raw().status().as_u16() == 404 => {
TierError::NotFound(message)
}
_ => TierError::Io(message),
}
}
fn get_object_error(
key: &str,
range: &str,
e: SdkError<GetObjectError, HttpResponse>,
) -> TierError {
let message = format!(
"failed to get object {} range {}: {}",
key,
range,
DisplayErrorContext(&e)
);
match e {
SdkError::ServiceError(ref se) if se.err().is_no_such_key() => TierError::NotFound(message),
_ => TierError::Io(message),
}
}
/// Configuration for an S3 tier backend.
#[derive(Debug, Clone)]
pub struct S3TierConfig {
@@ -89,7 +140,7 @@ impl S3TierBackend {
&self,
file_path: &str,
progress_fn: F,
) -> Result<(String, u64), String>
) -> Result<(String, u64), TierError>
where
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
@@ -97,7 +148,7 @@ impl S3TierBackend {
let metadata = tokio::fs::metadata(file_path)
.await
.map_err(|e| format!("failed to stat file {}: {}", file_path, e))?;
.map_err(|e| TierError::Io(format!("failed to stat file {}: {}", file_path, e)))?;
let file_size = metadata.len();
// Calculate part size: start at 64MB, scale up for very large files (matches Go)
@@ -119,11 +170,16 @@ impl S3TierBackend {
)
.send()
.await
.map_err(|e| format!("failed to create multipart upload: {}", e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to create multipart upload: {}",
DisplayErrorContext(&e)
))
})?;
let upload_id = create_resp
.upload_id()
.ok_or_else(|| "no upload_id in multipart upload response".to_string())?
.ok_or_else(|| TierError::Io("no upload_id in multipart upload response".to_string()))?
.to_string();
// Build list of (part_number, offset, size) for all parts
@@ -159,19 +215,21 @@ impl S3TierBackend {
let _permit = sem
.acquire()
.await
.map_err(|e| format!("semaphore error: {}", e))?;
.map_err(|e| TierError::Io(format!("semaphore error: {}", e)))?;
// Read this part's data from the file at the correct offset
let mut file = tokio::fs::File::open(&fp)
.await
.map_err(|e| format!("failed to open file {}: {}", fp, e))?;
.map_err(|e| TierError::Io(format!("failed to open file {}: {}", fp, e)))?;
file.seek(std::io::SeekFrom::Start(off))
.await
.map_err(|e| format!("failed to seek to offset {}: {}", off, e))?;
.map_err(|e| {
TierError::Io(format!("failed to seek to offset {}: {}", off, e))
})?;
let mut buf = vec![0u8; size];
file.read_exact(&mut buf)
.await
.map_err(|e| format!("failed to read file at offset {}: {}", off, e))?;
file.read_exact(&mut buf).await.map_err(|e| {
TierError::Io(format!("failed to read file at offset {}: {}", off, e))
})?;
let upload_part_resp = client
.upload_part()
@@ -183,7 +241,12 @@ impl S3TierBackend {
.send()
.await
.map_err(|e| {
format!("failed to upload part {} at offset {}: {}", pn, off, e)
TierError::Io(format!(
"failed to upload part {} at offset {}: {}",
pn,
off,
DisplayErrorContext(&e)
))
})?;
let e_tag = upload_part_resp.e_tag().unwrap_or_default().to_string();
@@ -202,9 +265,9 @@ impl S3TierBackend {
};
(guard.1)(uploaded as i64, pct)
};
progress_result?;
progress_result.map_err(TierError::Aborted)?;
Ok::<_, String>(
Ok::<_, TierError>(
CompletedPart::builder()
.e_tag(e_tag)
.part_number(pn)
@@ -219,7 +282,7 @@ impl S3TierBackend {
for handle in handles {
let part = handle
.await
.map_err(|e| format!("upload task panicked: {}", e))??;
.map_err(|e| TierError::Io(format!("upload task panicked: {}", e)))??;
completed_parts.push(part);
}
@@ -236,9 +299,14 @@ impl S3TierBackend {
.multipart_upload(completed_upload)
.send()
.await
.map_err(|e| format!("failed to complete multipart upload: {}", e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to complete multipart upload: {}",
DisplayErrorContext(&e)
))
})?;
Ok::<(), String>(())
Ok::<(), TierError>(())
}
.await;
@@ -281,7 +349,7 @@ impl S3TierBackend {
dest_path: &str,
key: &str,
progress_fn: F,
) -> Result<u64, String>
) -> Result<u64, TierError>
where
F: FnMut(i64, f32) -> Result<(), String> + Send + Sync + 'static,
{
@@ -293,7 +361,7 @@ impl S3TierBackend {
.key(key)
.send()
.await
.map_err(|e| format!("failed to head object {}: {}", key, e))?;
.map_err(|e| head_object_error(key, e))?;
let file_size = head_resp.content_length().unwrap_or(0) as u64;
@@ -305,10 +373,12 @@ impl S3TierBackend {
.truncate(true)
.open(dest_path)
.await
.map_err(|e| format!("failed to open dest file {}: {}", dest_path, e))?;
.map_err(|e| {
TierError::Io(format!("failed to open dest file {}: {}", dest_path, e))
})?;
file.set_len(file_size)
.await
.map_err(|e| format!("failed to set file length: {}", e))?;
.map_err(|e| TierError::Io(format!("failed to set file length: {}", e)))?;
}
let part_size: u64 = 64 * 1024 * 1024;
@@ -344,7 +414,7 @@ impl S3TierBackend {
let _permit = sem
.acquire()
.await
.map_err(|e| format!("semaphore error: {}", e))?;
.map_err(|e| TierError::Io(format!("semaphore error: {}", e)))?;
let end = off + size - 1;
let range = format!("bytes={}-{}", off, end);
@@ -356,13 +426,13 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(&key, &range, e))?;
let body = get_resp
.body
.collect()
.await
.map_err(|e| format!("failed to read body: {}", e))?;
.map_err(|e| TierError::Io(format!("failed to read body: {}", e)))?;
let bytes = body.into_bytes();
// Write at the correct offset (like Go's WriteAt)
@@ -370,13 +440,17 @@ impl S3TierBackend {
.write(true)
.open(&dp)
.await
.map_err(|e| format!("failed to open dest file {}: {}", dp, e))?;
.map_err(|e| {
TierError::Io(format!("failed to open dest file {}: {}", dp, e))
})?;
file.seek(std::io::SeekFrom::Start(off))
.await
.map_err(|e| format!("failed to seek to offset {}: {}", off, e))?;
.map_err(|e| {
TierError::Io(format!("failed to seek to offset {}: {}", off, e))
})?;
file.write_all(&bytes)
.await
.map_err(|e| format!("failed to write to {}: {}", dp, e))?;
.map_err(|e| TierError::Io(format!("failed to write to {}: {}", dp, e)))?;
// Report progress. The lock is released before the result is
// propagated so an aborting callback cannot poison the mutex
@@ -392,9 +466,9 @@ impl S3TierBackend {
};
(guard.1)(downloaded as i64, pct)
};
progress_result?;
progress_result.map_err(TierError::Aborted)?;
Ok::<_, String>(())
Ok::<_, TierError>(())
}));
}
@@ -402,7 +476,7 @@ impl S3TierBackend {
for handle in handles {
handle
.await
.map_err(|e| format!("download task panicked: {}", e))??;
.map_err(|e| TierError::Io(format!("download task panicked: {}", e)))??;
}
// fsync the file so its content is durable before the caller trims the .vif
@@ -411,16 +485,21 @@ impl S3TierBackend {
.write(true)
.open(dest_path)
.await
.map_err(|e| format!("failed to open {} for fsync: {}", dest_path, e))?;
.map_err(|e| TierError::Io(format!("failed to open {} for fsync: {}", dest_path, e)))?;
synced
.sync_all()
.await
.map_err(|e| format!("failed to fsync {}: {}", dest_path, e))?;
.map_err(|e| TierError::Io(format!("failed to fsync {}: {}", dest_path, e)))?;
Ok(file_size)
}
pub async fn read_range(&self, key: &str, offset: u64, size: usize) -> Result<Vec<u8>, String> {
pub async fn read_range(
&self,
key: &str,
offset: u64,
size: usize,
) -> Result<Vec<u8>, TierError> {
let end = offset + (size as u64).saturating_sub(1);
let range = format!("bytes={}-{}", offset, end);
let resp = self
@@ -431,29 +510,35 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(key, &range, e))?;
let body = resp
.body
.collect()
.await
.map_err(|e| format!("failed to read object {} body: {}", key, e))?;
.map_err(|e| TierError::Io(format!("failed to read object {} body: {}", key, e)))?;
Ok(body.into_bytes().to_vec())
}
/// Delete a file from S3.
pub async fn delete_file(&self, key: &str) -> Result<(), String> {
pub async fn delete_file(&self, key: &str) -> Result<(), TierError> {
self.client
.delete_object()
.bucket(&self.bucket)
.key(key)
.send()
.await
.map_err(|e| format!("failed to delete object {}: {}", key, e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to delete object {}: {}",
key,
DisplayErrorContext(&e)
))
})?;
Ok(())
}
pub fn delete_file_blocking(&self, key: &str) -> Result<(), String> {
pub fn delete_file_blocking(&self, key: &str) -> Result<(), TierError> {
let client = self.client.clone();
let bucket = self.bucket.clone();
let key = key.to_string();
@@ -464,7 +549,13 @@ impl S3TierBackend {
.key(&key)
.send()
.await
.map_err(|e| format!("failed to delete object {}: {}", key, e))?;
.map_err(|e| {
TierError::Io(format!(
"failed to delete object {}: {}",
key,
DisplayErrorContext(&e)
))
})?;
Ok(())
})
}
@@ -474,7 +565,7 @@ impl S3TierBackend {
key: &str,
offset: u64,
size: usize,
) -> Result<Vec<u8>, String> {
) -> Result<Vec<u8>, TierError> {
let client = self.client.clone();
let bucket = self.bucket.clone();
let key = key.to_string();
@@ -488,13 +579,12 @@ impl S3TierBackend {
.range(&range)
.send()
.await
.map_err(|e| format!("failed to get object {} range {}: {}", key, range, e))?;
.map_err(|e| get_object_error(&key, &range, e))?;
let body = resp
.body
.collect()
.await
.map_err(|e| format!("failed to read object {} body: {}", key, e))?;
let body =
resp.body.collect().await.map_err(|e| {
TierError::Io(format!("failed to read object {} body: {}", key, e))
})?;
Ok(body.into_bytes().to_vec())
})
}
@@ -555,18 +645,279 @@ pub fn global_s3_tier_registry() -> &'static RwLock<S3TierRegistry> {
GLOBAL_S3_TIER_REGISTRY.get_or_init(|| RwLock::new(S3TierRegistry::new()))
}
fn block_on_tier_future<F, T>(future: F) -> Result<T, String>
where
F: Future<Output = Result<T, String>> + Send + 'static,
T: Send + 'static,
{
std::thread::spawn(move || {
let runtime = tokio::runtime::Builder::new_current_thread()
/// The one process-wide runtime for tiered-S3 I/O issued from synchronous
/// storage code. A per-call runtime tore down the SDK's pooled connections
/// after every 64 KiB chunk, re-dialing TLS per read; a long-lived runtime
/// keeps the pool warm.
///
/// Built on first use. A build failure is returned, not cached or panicked:
/// callers sit inside `Volume::destroy` and needle reads, whose own error
/// paths must run, and a later call may succeed.
static TIER_RUNTIME: std::sync::Mutex<Option<tokio::runtime::Runtime>> =
std::sync::Mutex::new(None);
fn tier_handle() -> Result<tokio::runtime::Handle, TierError> {
let mut slot = TIER_RUNTIME
.lock()
.unwrap_or_else(|poisoned| poisoned.into_inner());
if slot.is_none() {
let runtime = tokio::runtime::Builder::new_multi_thread()
.worker_threads(2)
.thread_name("tier-io")
.enable_all()
.build()
.map_err(|e| format!("failed to build tokio runtime: {}", e))?;
runtime.block_on(future)
})
.join()
.map_err(|_| "tier runtime thread panicked".to_string())?
.map_err(|e| {
TierError::RuntimeUnavailable(format!(
"failed to build the tier I/O tokio runtime: {}",
e
))
})?;
*slot = Some(runtime);
}
Ok(slot.as_ref().expect("just initialised").handle().clone())
}
/// Run `future` on the tier runtime and block the calling thread until it
/// finishes. The caller may be a worker of *another* tokio runtime, so this
/// waits on a channel rather than `Handle::block_on`, which panics when
/// called from inside any runtime context.
fn block_on_tier_future<F, T>(future: F) -> Result<T, TierError>
where
F: Future<Output = Result<T, TierError>> + Send + 'static,
T: Send + 'static,
{
let handle = tier_handle()?;
let task = handle.spawn(future);
let (tx, rx) = std::sync::mpsc::sync_channel(1);
handle.spawn(async move {
// The receiver only goes away if the caller was unwound; nothing to
// report then.
let _ = tx.send(task.await);
});
match rx.recv() {
Ok(Ok(result)) => result,
Ok(Err(join_error)) => Err(describe_join_error(join_error)),
Err(_) => Err(TierError::RuntimeUnavailable(
"tier I/O runtime dropped the task before it finished".to_string(),
)),
}
}
/// Turn a `JoinError` into a message that keeps the panic payload, so an
/// SDK panic surfaces as "boom" rather than a fixed "thread panicked".
fn describe_join_error(join_error: tokio::task::JoinError) -> TierError {
if join_error.is_panic() {
let payload = join_error.into_panic();
let message = if let Some(s) = payload.downcast_ref::<&str>() {
(*s).to_string()
} else if let Some(s) = payload.downcast_ref::<String>() {
s.clone()
} else {
"non-string panic payload".to_string()
};
TierError::Io(format!("tier I/O task panicked: {}", message))
} else {
TierError::RuntimeUnavailable(format!("tier I/O task failed: {}", join_error))
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::remote_storage::s3::tests::{CannedResponse, NO_SUCH_KEY};
use std::collections::HashSet;
use tokio::runtime::Handle;
fn probe() -> Result<(tokio::runtime::Id, Option<String>), TierError> {
block_on_tier_future(async {
Ok((
Handle::current().id(),
std::thread::current().name().map(str::to_string),
))
})
}
#[test]
fn block_on_tier_future_reuses_one_runtime() {
let (first_runtime, first_thread) = probe().expect("first call");
let (second_runtime, second_thread) = probe().expect("second call");
assert_eq!(
first_runtime, second_runtime,
"each call must run on the same long-lived tier runtime"
);
assert_eq!(first_thread.as_deref(), Some("tier-io"));
assert_eq!(second_thread.as_deref(), Some("tier-io"));
let mut runtimes = HashSet::new();
for _ in 0..20 {
let (id, _) = probe().expect("probe");
runtimes.insert(id);
}
assert_eq!(runtimes.len(), 1);
}
#[test]
fn block_on_tier_future_returns_the_value_and_the_error() {
assert_eq!(block_on_tier_future(async { Ok(7u32) }).unwrap(), 7);
let err = block_on_tier_future::<_, u32>(async { Err(TierError::NotFound("nope".into())) })
.unwrap_err();
assert!(
matches!(&err, TierError::NotFound(m) if m == "nope"),
"{err:?}"
);
}
#[test]
fn block_on_tier_future_works_from_a_std_thread() {
let (id, _) = std::thread::spawn(probe)
.join()
.expect("probe thread")
.expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn block_on_tier_future_works_from_spawn_blocking() {
let (id, _) = tokio::task::spawn_blocking(probe)
.await
.expect("spawn_blocking")
.expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
// Called straight from another runtime's async context: the case that
// would panic with `Handle::block_on` ("Cannot start a runtime from
// within a runtime").
#[tokio::test]
async fn block_on_tier_future_works_from_a_current_thread_runtime() {
let (id, _) = probe().expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn block_on_tier_future_works_from_a_multi_thread_runtime_worker() {
let (id, _) = probe().expect("probe");
assert_eq!(id, tier_handle().expect("tier runtime").id());
assert_ne!(id, Handle::current().id());
}
#[test]
fn block_on_tier_future_reports_the_panic_payload() {
let err = block_on_tier_future::<_, ()>(async {
if std::hint::black_box(true) {
panic!("boom {}", 42);
}
Ok(())
})
.expect_err("a panicking future must be an error");
assert!(matches!(err, TierError::Io(_)), "got: {err:?}");
assert!(err.to_string().contains("boom 42"), "got: {err}");
assert!(err.to_string().contains("panicked"), "got: {err}");
}
#[test]
fn block_on_tier_future_reports_a_str_panic_payload() {
let err = block_on_tier_future::<_, ()>(async {
if std::hint::black_box(true) {
panic!("static boom");
}
Ok(())
})
.expect_err("a panicking future must be an error");
assert!(err.to_string().contains("static boom"), "got: {err}");
}
fn backend_answering(status: u16, body: &'static str) -> S3TierBackend {
let config = aws_sdk_s3::Config::builder()
.behavior_version(BehaviorVersion::latest())
.region(Region::new("us-east-1"))
.credentials_provider(Credentials::new("AKIATEST", "secret", None, None, "test"))
.endpoint_url("http://127.0.0.1:1")
.force_path_style(true)
.http_client(CannedResponse { status, body })
.retry_config(aws_sdk_s3::config::retry::RetryConfig::disabled())
.build();
S3TierBackend {
client: Client::from_conf(config),
bucket: "bucket".to_string(),
storage_class: "STANDARD".to_string(),
}
}
#[tokio::test]
async fn download_head_404_is_not_found() {
let tmp = tempfile::tempdir().unwrap();
let dest = tmp.path().join("1.dat");
let err = backend_answering(404, "")
.download_file(dest.to_str().unwrap(), "missing", |_, _| Ok(()))
.await
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
assert!(
err.to_string()
.starts_with("failed to head object missing: "),
"{err}"
);
}
#[tokio::test]
async fn download_head_403_is_io() {
let tmp = tempfile::tempdir().unwrap();
let dest = tmp.path().join("1.dat");
let err = backend_answering(403, "")
.download_file(dest.to_str().unwrap(), "denied", |_, _| Ok(()))
.await
.unwrap_err();
assert!(matches!(err, TierError::Io(_)), "{err:?}");
}
#[tokio::test]
async fn read_range_no_such_key_is_not_found() {
let err = backend_answering(404, NO_SUCH_KEY)
.read_range("missing", 0, 8)
.await
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
assert!(
err.to_string()
.starts_with("failed to get object missing range bytes=0-7: "),
"{err}"
);
}
#[tokio::test]
async fn read_range_bare_404_is_io() {
// As in Go, GET is not-found by the NoSuchKey code, not the status.
let err = backend_answering(404, "")
.read_range("missing", 0, 8)
.await
.unwrap_err();
assert!(matches!(err, TierError::Io(_)), "{err:?}");
}
#[test]
fn read_range_blocking_no_such_key_is_not_found() {
let err = backend_answering(404, NO_SUCH_KEY)
.read_range_blocking("missing", 0, 8)
.unwrap_err();
assert!(matches!(err, TierError::NotFound(_)), "{err:?}");
}
#[test]
fn backend_name_to_type_id_splits_on_dot() {
assert_eq!(
backend_name_to_type_id("s3"),
("s3".to_string(), "default".to_string())
);
assert_eq!(
backend_name_to_type_id("s3.eu"),
("s3".to_string(), "eu".to_string())
);
assert_eq!(
backend_name_to_type_id("s3.a.b"),
(String::new(), String::new())
);
}
}
+10 -8
View File
@@ -10,7 +10,7 @@ use std::collections::HashSet;
use std::net::IpAddr;
use std::time::{SystemTime, UNIX_EPOCH};
use jsonwebtoken::{decode, encode, Algorithm, DecodingKey, EncodingKey, Header, Validation};
use jsonwebtoken::{Algorithm, DecodingKey, EncodingKey, Header, Validation, decode, encode};
use serde::{Deserialize, Serialize};
// ============================================================================
@@ -297,10 +297,10 @@ impl Guard {
/// Extract host from "host:port" or "[::1]:port" format.
fn extract_host(addr: &str) -> String {
// Handle IPv6 with brackets
if addr.starts_with('[') {
if let Some(end) = addr.find(']') {
return addr[1..end].to_string();
}
if addr.starts_with('[')
&& let Some(end) = addr.find(']')
{
return addr[1..end].to_string();
}
// Handle host:port
if let Some(pos) = addr.rfind(':') {
@@ -481,9 +481,11 @@ mod tests {
let token = gen_jwt(&key, 3600, "3,01637037d6").unwrap();
// Correct file ID
assert!(guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok());
assert!(
guard
.check_jwt_for_file(Some(&token), "3,01637037d6", true)
.is_ok()
);
// Wrong file ID
let err = guard.check_jwt_for_file(Some(&token), "4,deadbeef", true);
+8 -7
View File
@@ -3,12 +3,12 @@ use std::fmt;
use std::sync::Arc;
use rustls::client::danger::HandshakeSignatureValid;
use rustls::crypto::aws_lc_rs;
use rustls::crypto::CryptoProvider;
use rustls::crypto::aws_lc_rs;
use rustls::pki_types::UnixTime;
use rustls::pki_types::{CertificateDer, PrivateKeyDer};
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::server::WebPkiClientVerifier;
use rustls::server::danger::{ClientCertVerified, ClientCertVerifier};
use rustls::{
CipherSuite, DigitallySignedStruct, DistinguishedName, RootCertStore, ServerConfig,
SignatureScheme, SupportedCipherSuite, SupportedProtocolVersion,
@@ -120,10 +120,11 @@ impl ClientCertVerifier for CommonNameVerifier {
// aws-lc-rs and ring both get linked transitively, so rustls can't auto-select
// a provider and tonic's client TLS panics on first use. Pin the default to
// aws-lc-rs, matching the server config. Idempotent.
pub fn install_default_crypto_provider() {
let _ = aws_lc_rs::default_provider().install_default();
}
// aws-lc-rs, matching the server config. Idempotent. The body lives in
// seaweed-common so this binary and the Rust plugin workers cannot end up
// installing different providers; re-exported here so callers keep their
// import path.
pub use seaweed_common::tls::install_default_crypto_provider;
pub fn build_rustls_server_config(
cert_path: &str,
@@ -376,7 +377,7 @@ fn go_tls_version_for_supported(version: &SupportedProtocolVersion) -> GoTlsVers
#[cfg(test)]
mod tests {
use super::{build_supported_versions, common_name_is_allowed, parse_cipher_suites, TlsPolicy};
use super::{TlsPolicy, build_supported_versions, common_name_is_allowed, parse_cipher_suites};
use rustls::crypto::aws_lc_rs;
use std::collections::HashSet;
+2 -2
View File
@@ -1,9 +1,9 @@
use axum::Router;
use axum::body::Body;
use axum::extract::Query;
use axum::http::{header, StatusCode};
use axum::http::{StatusCode, header};
use axum::response::{IntoResponse, Response};
use axum::routing::{any, get};
use axum::Router;
use pprof::protos::Message;
use serde::Deserialize;
+352 -45
View File
@@ -1,16 +1,42 @@
//! Construction of the volume server's *outgoing* gRPC clients: TLS material,
//! endpoint tuning, dial bounds, and the three client constructors every call
//! site goes through.
//!
//! The keepalive, window-size and message-size constants below are shared with
//! the *inbound* server built in `main.rs`, which imports them from here rather
//! than declaring its own. Changing one therefore changes both directions at
//! once, which is deliberate: a volume server talks to its peers with the same
//! HTTP/2 settings it offers them.
use std::error::Error;
use std::fmt;
use std::time::Duration;
use hyper::http::Uri;
use tonic::service::interceptor::InterceptedService;
use tonic::transport::{Certificate, Channel, ClientTlsConfig, Endpoint, Identity};
use tonic::{Request, Status};
use crate::config::VolumeServerConfig;
use crate::pb::filer_pb::seaweed_filer_client::SeaweedFilerClient;
use crate::pb::master_pb::seaweed_client::SeaweedClient;
use crate::pb::volume_server_pb::volume_server_client::VolumeServerClient;
use crate::server::request_id::outgoing_request_id_interceptor;
pub const GRPC_MAX_MESSAGE_SIZE: usize = 1 << 30;
const GRPC_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(60);
const GRPC_KEEPALIVE_TIMEOUT: Duration = Duration::from_secs(20);
const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
pub const GRPC_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(60);
pub const GRPC_KEEPALIVE_TIMEOUT: Duration = Duration::from_secs(20);
pub const GRPC_INITIAL_WINDOW_SIZE: u32 = 16 * 1024 * 1024;
/// Bound on the TCP connect of every outgoing dial. `build_grpc_endpoint` is
/// private and `connect_channel` is the only way out of this module, so every
/// call site picks this up whether it thinks about timeouts or not.
///
/// It bounds the TCP handshake only — tonic hands it to
/// `HttpConnector::set_connect_timeout`. A peer that completes the handshake
/// and then stalls in the TLS or HTTP/2 exchange is not covered; callers that
/// need that bound wrap the whole dial (see `connect_ping_target`).
const GRPC_CONNECT_TIMEOUT: Duration = Duration::from_secs(5);
#[derive(Clone, Debug)]
pub struct OutgoingGrpcTlsConfig {
@@ -40,7 +66,9 @@ pub fn load_outgoing_grpc_tls(
(&config.grpc_client_cert_file, &config.grpc_client_key_file)
} else {
if !config.grpc_client_cert_file.is_empty() || !config.grpc_client_key_file.is_empty() {
tracing::warn!("grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key");
tracing::warn!(
"grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key"
);
}
(&config.grpc_cert_file, &config.grpc_key_file)
};
@@ -79,7 +107,7 @@ pub fn grpc_endpoint_uri(grpc_host_port: &str, tls: Option<&OutgoingGrpcTlsConfi
format!("{}://{}", scheme, grpc_host_port)
}
pub fn build_grpc_endpoint(
fn build_grpc_endpoint(
grpc_host_port: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
) -> Result<Endpoint, GrpcClientError> {
@@ -115,6 +143,184 @@ pub fn build_grpc_endpoint(
Ok(endpoint)
}
/// Connect `endpoint` through a connector that re-validates every resolved
/// address at connect time (Go's `guardedDialerPolicy` mirror), pinning a
/// validated copy/tail source against DNS rebinding. `allow_untrusted`
/// preserves the plain connect for operators that opted out.
pub async fn connect_guarded(
endpoint: Endpoint,
target: &str,
allow_untrusted: bool,
) -> Result<Channel, GrpcClientError> {
if allow_untrusted {
return endpoint
.connect()
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)));
}
let target_owned = target.to_string();
let connector = tower::service_fn(move |uri: Uri| {
let target = target_owned.clone();
async move {
let host = uri.host().unwrap_or_default().to_string();
let port = uri.port_u16().unwrap_or(80);
crate::remote_storage::guarded_tcp_connect(&host, port, &target)
.await
.map(hyper_util::rt::TokioIo::new)
}
});
endpoint
.connect_with_connector(connector)
.await
.map_err(|e| GrpcClientError(format!("connect {} failed: {}", target, e)))
}
/// How a dial is bounded.
///
/// `connect_timeout` is handed to the TCP connector. `request_timeout` becomes
/// [`Endpoint::timeout`], which tonic installs as a `GrpcTimeout` layer in
/// front of *every* request the resulting channel carries — it is not a
/// property of one call.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct GrpcDialOptions {
/// Bound on establishing the connection to the peer.
pub connect_timeout: Duration,
/// Deadline applied to each RPC on the channel, or `None` to leave them
/// unbounded.
pub request_timeout: Option<Duration>,
}
impl GrpcDialOptions {
/// A short request/response call: connect within 5 s, answer within 10 s.
pub fn unary() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: Some(Duration::from_secs(10)),
}
}
/// A call the peer may take a while to answer: connect within 5 s, answer
/// within 30 s.
pub fn long() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: Some(Duration::from_secs(30)),
}
}
/// A bounded connect with no deadline on the RPCs themselves.
///
/// `request_timeout` must stay `None` here. [`Endpoint::timeout`] is not a
/// transfer budget: tonic layers it as a `GrpcTimeout` around the
/// response future, which resolves when the server's *first response
/// headers* arrive, so it bounds how long the peer may take to start
/// answering — per request, for every request the channel carries. A 10 s
/// value picked to suit one short call would therefore also be the header
/// deadline for the `VolumeCopy` that shares the dial, and a busy source
/// that takes longer than that to open its file would lose the whole copy.
/// `VolumeCopy`, `VolumeTailSender` and `VolumeEcShardsCopy` have never
/// carried one.
pub fn stream() -> Self {
Self {
connect_timeout: GRPC_CONNECT_TIMEOUT,
request_timeout: None,
}
}
}
/// Dial a peer and return a connected channel.
///
/// The error carries only the transport failure: every caller already wraps it
/// with the address and the operation it was attempting.
pub async fn connect_channel(
grpc_host_port: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
opts: GrpcDialOptions,
) -> Result<Channel, GrpcClientError> {
let mut endpoint =
build_grpc_endpoint(grpc_host_port, tls)?.connect_timeout(opts.connect_timeout);
if let Some(request_timeout) = opts.request_timeout {
endpoint = endpoint.timeout(request_timeout);
}
endpoint
.connect()
.await
.map_err(|e| GrpcClientError(e.to_string()))
}
/// Dial a copy/tail source and return a connected channel, re-validating every
/// resolved address at connect time.
///
/// The guarded equivalent of [`connect_channel`]: same `opts` bounds, but the
/// dial goes through [`connect_guarded`] so a source address that passed
/// validation cannot be re-pointed by DNS between the check and the connect.
/// The bounds are applied to the endpoint *before* delegating, so the
/// `allow_untrusted` opt-out is timed too.
///
/// `target` is the caller-facing source address (the unparsed
/// `"ip:port.grpcPort"` form), which is what the guard pins against; the error
/// carries only the transport failure, as every caller already wraps it with
/// the address and the operation it was attempting.
pub async fn connect_channel_guarded(
grpc_host_port: &str,
target: &str,
tls: Option<&OutgoingGrpcTlsConfig>,
opts: GrpcDialOptions,
allow_untrusted: bool,
) -> Result<Channel, GrpcClientError> {
let mut endpoint =
build_grpc_endpoint(grpc_host_port, tls)?.connect_timeout(opts.connect_timeout);
if let Some(request_timeout) = opts.request_timeout {
endpoint = endpoint.timeout(request_timeout);
}
connect_guarded(endpoint, target, allow_untrusted).await
}
/// The outgoing request-id interceptor as a concrete type, so the client
/// aliases below can name it.
pub type RequestIdInterceptor = fn(Request<()>) -> Result<Request<()>, Status>;
/// A volume-server client with the request-id interceptor attached.
pub type VolumeServerGrpcClient =
VolumeServerClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// A master client with the request-id interceptor attached.
pub type MasterGrpcClient = SeaweedClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// A filer client with the request-id interceptor attached.
pub type FilerGrpcClient = SeaweedFilerClient<InterceptedService<Channel, RequestIdInterceptor>>;
/// Wrap a connected channel in a volume-server client that forwards the
/// current request id and lifts both message-size limits.
pub fn volume_server_client(channel: Channel) -> VolumeServerGrpcClient {
VolumeServerClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Wrap a connected channel in a master client that forwards the current
/// request id and lifts both message-size limits.
pub fn master_client(channel: Channel) -> MasterGrpcClient {
SeaweedClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Wrap a connected channel in a filer client that forwards the current
/// request id and lifts both message-size limits.
pub fn filer_client(channel: Channel) -> FilerGrpcClient {
SeaweedFilerClient::with_interceptor(
channel,
outgoing_request_id_interceptor as RequestIdInterceptor,
)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE)
}
/// Parse a SeaweedFS server address (`"ip:port.grpcPort"` or
/// `"ip:port"`) into the `host:grpcPort` form `build_grpc_endpoint`
/// expects. With the trailing `.grpcPort` segment, that segment IS
@@ -124,53 +330,27 @@ pub fn build_grpc_endpoint(
/// Shared between `grpc_server.rs` and the distributed-EC-read path
/// in `store_ec.rs` — keep this as the single source of truth so the
/// HTTP↔gRPC port translation can't drift between callers.
///
/// The rule itself lives in `seaweed_common::address`, which the Rust
/// plugin workers share; this wrapper only flattens the typed error
/// back to the `String` its callers already handle. Unbracketed IPv6
/// literals come back bracketed, which this copy used to get wrong.
pub fn parse_grpc_address(source: &str) -> Result<String, String> {
let colon_idx = source
.rfind(':')
.ok_or_else(|| format!("cannot parse address: {}", source))?;
let host = &source[..colon_idx];
let port_part = &source[colon_idx + 1..];
if let Some(dot_idx) = port_part.rfind('.') {
// Format: "ip:port.grpcPort". Validate BOTH ports as u16
// so a malformed HTTP port (e.g. `host:abc.18080`) is
// rejected here rather than tripping a downstream
// `build_grpc_endpoint` URI parse failure with a less
// useful error.
let http_port = &port_part[..dot_idx];
let grpc_port = &port_part[dot_idx + 1..];
http_port
.parse::<u16>()
.map_err(|e| format!("invalid http port {:?}: {}", http_port, e))?;
grpc_port
.parse::<u16>()
.map_err(|e| format!("invalid grpc port {:?}: {}", grpc_port, e))?;
return Ok(format!("{}:{}", host, grpc_port));
}
// Format: "ip:port" → grpc = port + 10000. Reject inputs whose
// implicit grpc port would overflow the TCP port range (e.g.
// `host:60000` produces 70000 — invalid). Without this check
// the cast silently wraps and the endpoint call later fails
// with an opaque connection error.
let port: u16 = port_part
.parse()
.map_err(|e| format!("invalid port {:?}: {}", port_part, e))?;
let grpc_port = port as u32 + 10000;
if grpc_port > u16::MAX as u32 {
return Err(format!(
"implicit grpc port out of range: {} + 10000 = {}",
port, grpc_port
));
}
Ok(format!("{}:{}", host, grpc_port))
seaweed_common::address::to_grpc_address(source).map_err(|e| e.to_string())
}
#[cfg(test)]
mod tests {
use super::{build_grpc_endpoint, grpc_endpoint_uri, load_outgoing_grpc_tls};
use super::{
GrpcDialOptions, build_grpc_endpoint, connect_channel, grpc_endpoint_uri,
load_outgoing_grpc_tls, volume_server_client,
};
use crate::config::{NeedleMapKind, ReadMode, VolumeServerConfig};
use crate::pb::volume_server_pb;
use crate::security::tls::TlsPolicy;
use crate::server::request_id::scope_request_id;
use std::sync::{Arc, Mutex};
use std::time::Duration;
const TEST_CERT_PEM: &str = "-----BEGIN CERTIFICATE-----\nMIIBPDCB76ADAgECAhRuRPQgeAu43BT/M7EfAWSdapVdYDAFBgMrZXAwFDESMBAG\nA1UEAwwJbG9jYWxob3N0MB4XDTI2MDcwNTE2MTUyOVoXDTM2MDcwMjE2MTUyOVow\nFDESMBAGA1UEAwwJbG9jYWxob3N0MCowBQYDK2VwAyEAr/3bNIFI+8V32oCiY6y+\nXRFmZpdNQ2g//VtRkT+nQg+jUzBRMB0GA1UdDgQWBBTsy9tLf1zPiXCQfgci6zNi\ndEzRSjAfBgNVHSMEGDAWgBTsy9tLf1zPiXCQfgci6zNidEzRSjAPBgNVHRMBAf8E\nBTADAQH/MAUGAytlcANBAIvsdw0IbvOBBkb9cd7BfMJfIP9pQQrAL03pCRWJFnFh\nSysaLVgFXI4T078IiaM874oO+iB+5vNbWEpc7CkGow4=\n-----END CERTIFICATE-----\n";
const TEST_KEY_PEM: &str = "-----BEGIN PRIVATE KEY-----\nMC4CAQAwBQYDK2VwBCIEIHbyn71Kk+Y7KT3sBctit7uZpErpoH6qDbFj6P8qGaZH\n-----END PRIVATE KEY-----\n";
@@ -365,4 +545,131 @@ mod tests {
let err = parse_grpc_address("hostname").unwrap_err();
assert!(err.contains("cannot parse"), "{}", err);
}
#[test]
fn test_parse_grpc_address_brackets_ipv6_literals() {
use super::parse_grpc_address;
// This used to come back as `::1:29333`, which is not a valid
// authority: `build_grpc_endpoint` reads the last colon as the port
// separator and rejects the rest.
assert_eq!(parse_grpc_address("::1:19333").unwrap(), "[::1]:29333");
assert_eq!(parse_grpc_address("::1:9333.19333").unwrap(), "[::1]:19333");
// Already bracketed, so it is left alone.
assert_eq!(parse_grpc_address("[::1]:9333").unwrap(), "[::1]:19333");
}
#[test]
fn test_build_grpc_endpoint_accepts_an_ipv6_master_address() {
use super::parse_grpc_address;
let endpoint = build_grpc_endpoint(&parse_grpc_address("::1:9333").unwrap(), None).unwrap();
assert_eq!(endpoint.uri().port_u16(), Some(19333));
}
/// A minimal HTTP/2 server that records the gRPC request headers it is
/// sent and answers every call with a trailers-only `unimplemented`. It is
/// enough to prove what a helper-built client puts on the wire, without
/// standing up the whole `VolumeServer` service behind a tonic server.
async fn serve_header_capture() -> (u16, Arc<Mutex<Option<String>>>) {
use hyper::service::service_fn;
use hyper_util::rt::{TokioExecutor, TokioIo};
let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await.unwrap();
let port = listener.local_addr().unwrap().port();
let seen: Arc<Mutex<Option<String>>> = Arc::new(Mutex::new(None));
let captured = Arc::clone(&seen);
tokio::spawn(async move {
while let Ok((stream, _)) = listener.accept().await {
let captured = Arc::clone(&captured);
tokio::spawn(async move {
let _ = hyper::server::conn::http2::Builder::new(TokioExecutor::new())
.serve_connection(
TokioIo::new(stream),
service_fn(move |req: hyper::Request<hyper::body::Incoming>| {
let captured = Arc::clone(&captured);
async move {
let value = req
.headers()
.get("x-amz-request-id")
.and_then(|v| v.to_str().ok())
.map(str::to_string);
*captured.lock().unwrap() = value;
Ok::<_, std::convert::Infallible>(
hyper::http::Response::builder()
.status(200)
.header("content-type", "application/grpc")
.header("grpc-status", "12")
.body(tonic::body::Body::empty())
.unwrap(),
)
}
}),
)
.await;
});
}
});
(port, seen)
}
#[tokio::test]
async fn test_helper_built_client_sends_the_scoped_request_id() {
let (port, seen) = serve_header_capture().await;
let channel = connect_channel(
&format!("127.0.0.1:{}", port),
None,
GrpcDialOptions::unary(),
)
.await
.expect("dial the header-capturing server");
let mut client = volume_server_client(channel);
// The interceptor has a request id to forward only inside a scope, so
// the call has to run inside one for this to test anything.
let _ = scope_request_id("REQUEST-ID-ON-THE-WIRE".to_string(), async move {
client
.ping(volume_server_pb::PingRequest {
target: String::new(),
target_type: String::new(),
})
.await
})
.await;
assert_eq!(
seen.lock().unwrap().as_deref(),
Some("REQUEST-ID-ON-THE-WIRE"),
"a client built by volume_server_client must carry the outgoing request id"
);
}
#[test]
fn test_dial_presets_match_the_call_sites_they_replace() {
assert_eq!(
GrpcDialOptions::unary().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::unary().request_timeout,
Some(Duration::from_secs(10))
);
assert_eq!(
GrpcDialOptions::long().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::long().request_timeout,
Some(Duration::from_secs(30))
);
assert_eq!(
GrpcDialOptions::stream().connect_timeout,
Duration::from_secs(5)
);
assert_eq!(
GrpcDialOptions::stream().request_timeout,
None,
"a streaming dial must not put a per-request deadline on the channel"
);
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+13 -9
View File
@@ -36,16 +36,20 @@ pub fn collect_mem_status() -> volume_server_pb::MemStatus {
#[cfg(target_os = "linux")]
fn get_system_memory_linux() -> Option<(u64, u64)> {
unsafe {
let mut info: libc::sysinfo = std::mem::zeroed();
if libc::sysinfo(&mut info) == 0 {
let unit = info.mem_unit as u64;
let total = info.totalram as u64 * unit;
let free = info.freeram as u64 * unit;
return Some((total, free));
}
// SAFETY: `libc::sysinfo` is plain data — integers and trailing padding,
// no pointers and no restricted niches — so the all-zero value is a valid
// one for the kernel to overwrite.
let mut info: libc::sysinfo = unsafe { std::mem::zeroed() };
// SAFETY: `&mut info` is a live, aligned, exclusive pointer to a
// `sysinfo` that the kernel only writes through, and its fields are read
// below only after the call reports success.
if unsafe { libc::sysinfo(&mut info) } != 0 {
return None;
}
None
let unit = info.mem_unit as u64;
let total = info.totalram as u64 * unit;
let free = info.freeram as u64 * unit;
Some((total, free))
}
#[cfg(target_os = "linux")]
+113
View File
@@ -1,3 +1,8 @@
use tonic::Status;
use crate::remote_storage::s3_tier::TierError;
use crate::storage::volume::VolumeError;
#[cfg(unix)]
pub mod debug;
pub mod grpc_client;
@@ -13,3 +18,111 @@ pub mod store_ec;
pub mod ui;
pub mod volume_server;
pub mod write_queue;
/// Map a storage error onto the gRPC code that describes it.
impl From<VolumeError> for Status {
fn from(err: VolumeError) -> Self {
let message = err.to_string();
match err {
VolumeError::NotFound
| VolumeError::VolumeNotFound(_)
| VolumeError::Tier(TierError::NotFound(_)) => Status::not_found(message),
VolumeError::ReadOnly(_) | VolumeError::NotEmpty => {
Status::failed_precondition(message)
}
VolumeError::InsufficientSpace { .. } => Status::resource_exhausted(message),
VolumeError::AlreadyExists => Status::already_exists(message),
_ => Status::internal(message),
}
}
}
/// Same mapping, with the RPC's own context prefixed (`compact volume 7: ...`).
pub fn status_with_context(context: &str, err: VolumeError) -> Status {
let status = Status::from(err);
Status::new(status.code(), format!("{context}: {}", status.message()))
}
/// Render a configured disk directory as an absolute path for display, so the
/// status JSON and the UI show the same thing for a relative `-dir`. Falls
/// back to the configured spelling when the current directory cannot be read.
pub(crate) fn absolute_display_path(path: &str) -> String {
let p = std::path::Path::new(path);
if p.is_absolute() {
return path.to_string();
}
std::env::current_dir()
.map(|cwd| cwd.join(p).to_string_lossy().to_string())
.unwrap_or_else(|_| path.to_string())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::storage::types::VolumeId;
#[test]
fn test_volume_error_maps_to_grpc_code() {
use tonic::Code;
let code = |e: VolumeError| Status::from(e).code();
assert_eq!(
code(VolumeError::VolumeNotFound(VolumeId(7))),
Code::NotFound
);
assert_eq!(code(VolumeError::NotFound), Code::NotFound);
assert_eq!(
code(VolumeError::ReadOnly(VolumeId(7))),
Code::FailedPrecondition
);
assert_eq!(
VolumeError::ReadOnly(VolumeId(7)).to_string(),
"volume 7 is read only"
);
assert_eq!(
code(VolumeError::InsufficientSpace {
vid: VolumeId(7),
required: 2,
free: 1,
}),
Code::ResourceExhausted
);
assert_eq!(code(VolumeError::AlreadyExists), Code::AlreadyExists);
assert_eq!(code(VolumeError::NotInitialized), Code::Internal);
assert_eq!(
code(TierError::NotFound("gone".into()).into()),
Code::NotFound
);
for tier in [
TierError::Io("io".into()),
TierError::RuntimeUnavailable("rt".into()),
TierError::Aborted("bye".into()),
] {
assert_eq!(code(tier.into()), Code::Internal);
}
let status = status_with_context(
"backend s3.default copy file /data/1.dat",
TierError::NotFound("failed to head object k: NotFound".into()).into(),
);
assert_eq!(status.code(), Code::NotFound);
assert_eq!(
status.message(),
"backend s3.default copy file /data/1.dat: failed to head object k: NotFound"
);
let status = status_with_context(
"compact volume 7",
VolumeError::InsufficientSpace {
vid: VolumeId(7),
required: 2,
free: 1,
},
);
assert_eq!(status.code(), Code::ResourceExhausted);
assert_eq!(
status.message(),
"compact volume 7: not enough free space: required 2, free 1"
);
}
}
+3 -3
View File
@@ -29,11 +29,11 @@ impl<S> Layer<S> for GrpcRequestIdLayer {
impl<S, B> Service<http::Request<B>> for GrpcRequestIdService<S>
where
S: Service<http::Request<B>, Response = http::Response<tonic::body::BoxBody>> + Send + 'static,
S: Service<http::Request<B>, Response = http::Response<tonic::body::Body>> + Send + 'static,
S::Future: Send + 'static,
B: Send + 'static,
{
type Response = http::Response<tonic::body::BoxBody>;
type Response = http::Response<tonic::body::Body>;
type Error = S::Error;
type Future = Pin<Box<dyn Future<Output = Result<Self::Response, Self::Error>> + Send>>;
@@ -57,7 +57,7 @@ where
let future = self.inner.call(request);
Box::pin(async move {
let mut response: http::Response<tonic::body::BoxBody> =
let mut response: http::Response<tonic::body::Body> =
scope_request_id(request_id.clone(), future).await?;
if let Ok(value) = HeaderValue::from_str(&request_id) {
response.headers_mut().insert("x-amz-request-id", value);
File diff suppressed because it is too large Load Diff
+7 -17
View File
@@ -1,5 +1,6 @@
use std::fmt::Write as _;
use crate::server::absolute_display_path;
use crate::server::server_stats;
use crate::server::volume_server::VolumeServerState;
use crate::storage::store::Store;
@@ -423,13 +424,12 @@ fn collect_ui_data(
shard_id: shard.shard_id,
size: shard_size,
});
if created_at == "-" {
if let Ok(metadata) = std::fs::metadata(shard.file_name()) {
if let Ok(modified) = metadata.modified() {
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
if created_at == "-"
&& let Ok(metadata) = std::fs::metadata(shard.file_name())
&& let Ok(modified) = metadata.modified()
{
let ts: chrono::DateTime<chrono::Local> = modified.into();
created_at = ts.format("%Y-%m-%d %H:%M").to_string();
}
}
let preferred_size = ec_volume.dat_file_size.max(0) as u64;
@@ -451,16 +451,6 @@ fn collect_ui_data(
(disk_rows, volumes, remote_volumes, ec_volumes)
}
fn absolute_display_path(path: &str) -> String {
let p = std::path::Path::new(path);
if p.is_absolute() {
return path.to_string();
}
std::env::current_dir()
.map(|cwd| cwd.join(p).to_string_lossy().to_string())
.unwrap_or_else(|_| path.to_string())
}
fn join_i64(values: &[i64]) -> String {
values
.iter()
+24 -27
View File
@@ -14,12 +14,12 @@ use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32, Ordering};
use std::sync::{Arc, RwLock};
use axum::{
extract::{connect_info::ConnectInfo, Request, State},
http::{header, HeaderValue, Method, StatusCode},
Router,
extract::{Request, State, connect_info::ConnectInfo},
http::{HeaderValue, Method, StatusCode, header},
middleware::{self, Next},
response::{IntoResponse, Response},
routing::{any, get},
Router,
};
use crate::config::ReadMode;
@@ -73,8 +73,6 @@ pub struct VolumeServerState {
pub volume_state_notify: tokio::sync::Notify,
/// Optional batched write queue for improved throughput under load.
pub write_queue: std::sync::OnceLock<WriteQueue>,
/// Registry of S3 tier backends for tiered storage operations.
pub s3_tier_registry: std::sync::RwLock<crate::remote_storage::s3_tier::S3TierRegistry>,
/// Read mode: local, proxy, or redirect for non-local volumes.
pub read_mode: ReadMode,
/// If true, FetchAndWriteNeedle skips remote S3 endpoint validation,
@@ -200,9 +198,7 @@ pub fn to_http_address(addr: &str) -> std::borrow::Cow<'_, str> {
// rather than being silently rewritten. Mirrors the validation already
// done in `to_grpc_address` for the inverse direction.
if let (Ok(_), Ok(_)) = (http_port.parse::<u16>(), grpc_port.parse::<u16>()) {
return std::borrow::Cow::Owned(
addr[..ports_sep_index + 1 + dot_idx].to_string(),
);
return std::borrow::Cow::Owned(addr[..ports_sep_index + 1 + dot_idx].to_string());
}
}
std::borrow::Cow::Borrowed(addr)
@@ -312,16 +308,15 @@ async fn admin_store_handler(state: State<Arc<VolumeServerState>>, request: Requ
)
}
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -358,16 +353,15 @@ async fn public_store_handler(state: State<Arc<VolumeServerState>>, request: Req
}
_ => StatusCode::OK.into_response(),
};
if method == Method::GET {
if let Some(response_bytes) = response
if method == Method::GET
&& let Some(response_bytes) = response
.headers()
.get(header::CONTENT_LENGTH)
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse::<i64>().ok())
.filter(|value| *value > 0)
{
super::server_stats::record_bytes_out(response_bytes);
}
{
super::server_stats::record_bytes_out(response_bytes);
}
super::server_stats::record_request_close();
crate::metrics::INFLIGHT_REQUESTS_GAUGE
@@ -432,13 +426,13 @@ pub fn build_admin_router_with_ui(state: Arc<VolumeServerState>, ui_enabled: boo
.route("/healthz", get(handlers::healthz_handler))
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(admin_store_handler))
.route("/:path", any(admin_store_handler))
.route("/:vid/:fid", any(admin_store_handler))
.route("/:vid/:fid/:filename", any(admin_store_handler))
.route("/{path}", any(admin_store_handler))
.route("/{vid}/{fid}", any(admin_store_handler))
.route("/{vid}/{fid}/{filename}", any(admin_store_handler))
.fallback(admin_store_handler);
if ui_enabled {
// Note: /stats/* endpoints are commented out in Go's volume_server.go (L130-134).
@@ -455,13 +449,13 @@ pub fn build_public_router(state: Arc<VolumeServerState>) -> Router {
Router::new()
.route("/favicon.ico", get(handlers::favicon_handler))
.route(
"/seaweedfsstatic/*path",
"/seaweedfsstatic/{*path}",
get(handlers::static_asset_handler),
)
.route("/", any(public_store_handler))
.route("/:path", any(public_store_handler))
.route("/:vid/:fid", any(public_store_handler))
.route("/:vid/:fid/:filename", any(public_store_handler))
.route("/{path}", any(public_store_handler))
.route("/{vid}/{fid}", any(public_store_handler))
.route("/{vid}/{fid}/{filename}", any(public_store_handler))
.fallback(public_store_handler)
.layer(middleware::from_fn(common_headers_middleware))
.with_state(state)
@@ -516,7 +510,10 @@ mod tests {
// "host:abc.def"), and silently rewriting it would just hide the bug.
assert_eq!(to_http_address("host:abc.def"), "host:abc.def");
assert_eq!(to_http_address("host:9333.notaport"), "host:9333.notaport");
assert_eq!(to_http_address("host:notaport.19333"), "host:notaport.19333");
assert_eq!(
to_http_address("host:notaport.19333"),
"host:notaport.19333"
);
// Out-of-range ports must not be silently truncated either.
assert_eq!(to_http_address("host:99999.19333"), "host:99999.19333");
}
+66 -8
View File
@@ -3,8 +3,8 @@
//! Instead of each upload handler directly calling `write_needle`, writes are
//! submitted to a queue. A background worker drains the queue in batches (up to
//! 128 entries), groups them by volume ID, and processes them together under a
//! single store lock. Requests that asked for `fsync` are flushed by
//! `write_needle` itself, one flush per durable write.
//! single store lock. Durable writes to a volume share their .dat and .idx
//! flushes (see `Volume::write_needles_grouped`).
use std::sync::Arc;
@@ -159,8 +159,12 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
let mut store = state.store.write().unwrap();
for (vid, entries) in groups {
for (mut needle, fsync, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle, fsync);
let (mut writes, senders): (Vec<_>, Vec<_>) = entries
.into_iter()
.map(|(needle, fsync, response_tx)| ((needle, fsync), response_tx))
.unzip();
let results = store.write_volume_needles(vid, &mut writes);
for (response_tx, result) in senders.into_iter().zip(results) {
// Send result back; ignore error if receiver dropped.
let _ = response_tx.send(result);
}
@@ -178,8 +182,8 @@ mod tests {
use crate::server::volume_server::RuntimeMetricsConfig;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::store::Store;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
use std::sync::RwLock;
use std::sync::atomic::{AtomicBool, AtomicI64, AtomicU32};
let store = Store::new(NeedleMapKind::InMemory);
let guard = Guard::new(&[], SigningKey(vec![]), 0, SigningKey(vec![]), 0);
@@ -207,9 +211,6 @@ mod tests {
pre_stop_seconds: 0,
volume_state_notify: tokio::sync::Notify::new(),
write_queue: std::sync::OnceLock::new(),
s3_tier_registry: std::sync::RwLock::new(
crate::remote_storage::s3_tier::S3TierRegistry::new(),
),
read_mode: crate::config::ReadMode::Local,
allow_untrusted_remote_endpoints: false,
master_url: String::new(),
@@ -315,6 +316,63 @@ mod tests {
}
}
/// The queue hands a volume's batch to the grouped path, so ten durable
/// writes cost one .dat sync and one .idx sync, not ten of each.
#[test]
fn test_process_batch_group_commits_fsync_writes() {
use crate::config::MinFreeSpace;
use crate::storage::types::{DiskType, NeedleId};
use crate::storage::volume::VolumeSpec;
let tmp = tempfile::TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let state = make_test_state();
{
let mut store = state.store.write().unwrap();
store
.add_location(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
store
.add_volume(VolumeId(1), DiskType::HardDrive, &VolumeSpec::default())
.unwrap();
}
let mut receivers = Vec::new();
let batch = (1..=10u64)
.map(|id| {
let (response_tx, response_rx) = oneshot::channel();
receivers.push(response_rx);
WriteRequest {
volume_id: VolumeId(1),
needle: Needle {
id: NeedleId(id),
cookie: 0x1111.into(),
data: vec![id as u8; 8],
data_size: 8,
..Needle::default()
},
fsync: true,
response_tx,
}
})
.collect();
process_batch(state.clone(), batch);
for mut rx in receivers {
assert!(matches!(rx.try_recv().unwrap(), Ok((_, _, false))));
}
let store = state.store.read().unwrap();
let (_, vol) = store.find_volume(VolumeId(1)).unwrap();
assert_eq!(vol.sync_counts_for_test(), (1, 1));
}
#[tokio::test]
async fn test_write_queue_dropped_sender() {
// When the queue is dropped, subsequent submits should fail gracefully.
+331 -197
View File
@@ -15,15 +15,15 @@ use tracing::warn;
use crate::config::MinFreeSpace;
use crate::storage::erasure_coding::ec_bitrot::remove_bitrot_sidecars;
use crate::storage::erasure_coding::ec_shard::{
EcVolumeShard, DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
DATA_SHARDS_COUNT, ERASURE_CODING_LARGE_BLOCK_SIZE, ERASURE_CODING_SMALL_BLOCK_SIZE,
EcVolumeShard, ShardId,
};
use crate::storage::erasure_coding::ec_volume::EcVolume;
use crate::storage::erasure_coding::ec_volume::{EcVolume, is_usable_ecx_file};
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::super_block::{ReplicaPlacement, SUPER_BLOCK_SIZE};
use crate::storage::super_block::SUPER_BLOCK_SIZE;
use crate::storage::types::*;
use crate::storage::volume::{
remove_volume_files, volume_file_name, VifVolumeInfo, Volume, VolumeError,
VifVolumeInfo, Volume, VolumeError, VolumeSpec, remove_volume_files, volume_file_name,
};
/// A single disk location managing volumes in one directory.
@@ -131,10 +131,10 @@ impl DiskLocation {
for entry in entries {
let entry = entry?;
let name = entry.file_name().into_string().unwrap_or_default();
if let Some((collection, vid)) = parse_volume_filename(&name) {
if seen.insert((collection.clone(), vid)) {
dat_files.push((collection, vid));
}
if let Some((collection, vid)) = parse_volume_filename(&name)
&& seen.insert((collection.clone(), vid))
{
dat_files.push((collection, vid));
}
}
@@ -205,7 +205,6 @@ impl DiskLocation {
continue;
}
// Load existing data only; never create a phantom `.dat`. A lone
// `.vif`/`.idx` (e.g. an EC sidecar whose `.ecx` is on a sibling
// disk) would otherwise have Volume::new write an 8-byte stub that
@@ -280,30 +279,33 @@ impl DiskLocation {
let opened = Mutex::new(Vec::with_capacity(to_load.len()));
std::thread::scope(|scope| {
for _ in 0..workers {
scope.spawn(|| loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
match Volume::new(
&self.directory,
&self.idx_directory,
collection,
*vid,
needle_map_kind,
None, // replica placement read from superblock
None, // TTL read from superblock
0, // no preallocate on load
Version::current(),
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
scope.spawn(|| {
loop {
let i = next.fetch_add(1, Ordering::Relaxed);
let Some((vid, collections)) = to_load.get(i) else {
return;
};
for collection in collections {
// Replica placement and TTL are read back from the
// superblock, and a load never preallocates.
match Volume::new(
&self.directory,
&self.idx_directory,
*vid,
needle_map_kind,
&VolumeSpec {
collection,
..Default::default()
},
) {
Ok(mut v) => {
v.location_disk_space_low = self.is_disk_space_low.clone();
opened.lock().unwrap().push((collection.clone(), *vid, v));
break;
}
Err(e) => {
warn!(volume_id = vid.0, error = %e, "failed to load volume");
}
}
}
}
@@ -327,10 +329,10 @@ impl DiskLocation {
.strip_suffix(".cpc")
.or_else(|| name.strip_suffix(".cpd"))
.or_else(|| name.strip_suffix(".cpx"));
if let Some(stem) = stem {
if let Some(key) = parse_collection_volume_id(stem) {
pending.insert(key);
}
if let Some(stem) = stem
&& let Some(key) = parse_collection_volume_id(stem)
{
pending.insert(key);
}
}
}
@@ -374,8 +376,10 @@ impl DiskLocation {
let mut expected_shard_size: Option<i64> = None;
let dat_exists = match fs::metadata(&dat_path) {
Ok(meta) if meta.len() > SUPER_BLOCK_SIZE as u64 => {
expected_shard_size =
Some(calculate_expected_shard_size(meta.len() as i64, data_shards));
expected_shard_size = Some(calculate_expected_shard_size(
meta.len() as i64,
data_shards,
));
true
}
Ok(_) => false,
@@ -399,7 +403,13 @@ impl DiskLocation {
if size != prev {
// Inconsistent sizes signal corruption or mixed
// generations; not trusted for deletion -> keep.
warn!(volume_id = vid.0, shard = i, size, expected = prev, "EC shard size mismatch; keeping shards");
warn!(
volume_id = vid.0,
shard = i,
size,
expected = prev,
"EC shard size mismatch; keeping shards"
);
return true;
}
} else {
@@ -426,11 +436,16 @@ impl DiskLocation {
if shard_count == 0 {
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size) {
if actual < expected {
warn!(volume_id = vid.0, actual, expected, "shards smaller than the .dat's full encode; reclaiming the complete .dat");
return false;
}
if let (Some(actual), Some(expected)) = (actual_shard_size, expected_shard_size)
&& actual < expected
{
warn!(
volume_id = vid.0,
actual,
expected,
"shards smaller than the .dat's full encode; reclaiming the complete .dat"
);
return false;
}
true
}
@@ -464,6 +479,13 @@ impl DiskLocation {
if self.idx_directory != self.directory {
remove_bitrot_sidecars(&idx_base)?;
}
// Staged 2PC generations (<base>.ecNN.v<N>, versioned .ecx/.ecj/.vif)
// belong to this volume's EC state too; leaving them orphans the files.
crate::storage::erasure_coding::ec_shard::remove_ec_generation_files(&base, 0)?;
if self.idx_directory != self.directory {
crate::storage::erasure_coding::ec_shard::remove_ec_generation_files(&idx_base, 0)?;
}
Ok(())
}
@@ -510,10 +532,10 @@ impl DiskLocation {
pub(crate) fn ec_generation_ts_ns(&self, collection: &str, vid: VolumeId) -> Option<i64> {
for dir in [&self.directory, &self.idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Ok(s) = fs::read_to_string(&vif) {
if let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s) {
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if let Ok(s) = fs::read_to_string(&vif)
&& let Ok(vi) = serde_json::from_str::<VifVolumeInfo>(&s)
{
return Some(vi.ec_shard_config.map(|c| c.encode_ts_ns).unwrap_or(0));
}
if self.directory == self.idx_directory {
break;
@@ -545,27 +567,19 @@ impl DiskLocation {
pub fn create_volume(
&mut self,
vid: VolumeId,
collection: &str,
needle_map_kind: NeedleMapKind,
replica_placement: Option<ReplicaPlacement>,
ttl: Option<crate::storage::needle::ttl::TTL>,
preallocate: u64,
version: Version,
spec: &VolumeSpec<'_>,
) -> Result<(), VolumeError> {
let mut v = Volume::new(
&self.directory,
&self.idx_directory,
collection,
vid,
needle_map_kind,
replica_placement,
ttl,
preallocate,
version,
spec,
)?;
v.location_disk_space_low = self.is_disk_space_low.clone();
crate::metrics::VOLUME_GAUGE
.with_label_values(&[collection, "volume"])
.with_label_values(&[spec.collection, "volume"])
.inc();
self.volumes.insert(vid, v);
Ok(())
@@ -591,13 +605,20 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
only_empty: bool,
only_garbage: bool,
keep_remote_data: bool,
) -> Result<(), VolumeError> {
// Refuse before removing: a refused destroy must leave it mounted.
if let Some(v) = self.volumes.get(&vid)
&& v.is_compacting()
{
return Err(v.compacting_error());
}
if let Some(mut v) = self.volumes.remove(&vid) {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&v.collection, "volume"])
.dec();
v.destroy(only_empty, keep_remote_data)?;
v.destroy(only_empty, only_garbage, keep_remote_data)?;
Ok(())
} else {
Err(VolumeError::NotFound)
@@ -618,7 +639,7 @@ impl DiskLocation {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&v.collection, "volume"])
.dec();
if let Err(e) = v.destroy(false, false) {
if let Err(e) = v.destroy(false, false, false) {
warn!(volume_id = vid.0, error = %e, "delete collection: failed to destroy volume");
}
}
@@ -669,8 +690,7 @@ impl DiskLocation {
pub fn free_volume_count(&self) -> i32 {
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let max = self.max_volume_count.load(Ordering::Relaxed);
let free_count = (max as i64 - self.volumes.len() as i64)
* DATA_SHARDS_COUNT as i64
let free_count = (max as i64 - self.volumes.len() as i64) * DATA_SHARDS_COUNT as i64
- self.ec_shard_count() as i64;
let effective_free = free_count / DATA_SHARDS_COUNT as i64;
if effective_free > 0 {
@@ -773,22 +793,18 @@ impl DiskLocation {
/// Mirrors `DiskLocation.HasEcxFileOnDisk` in
/// `weed/storage/disk_location_ec.go`. Skips entries that are
/// directories so a stray dir named `<collection>_<vid>.ecx` doesn't
/// register as a present index file.
/// register as a present index file. A 0-byte `.ecx` is a corrupt stub
/// left by a failed EC distribute copy; it must not steer placement
/// toward this disk, so it counts as absent (Go requires `Size() > 0`).
pub fn has_ecx_file_on_disk(&self, collection: &str, vid: VolumeId) -> bool {
let idx_base = volume_file_name(&self.idx_directory, collection, vid);
let idx_path = format!("{}.ecx", idx_base);
if let Ok(meta) = fs::metadata(&idx_path) {
if !meta.is_dir() {
return true;
}
if is_usable_ecx_file(&format!("{}.ecx", idx_base)) {
return true;
}
if self.idx_directory != self.directory {
let data_base = volume_file_name(&self.directory, collection, vid);
let data_path = format!("{}.ecx", data_base);
if let Ok(meta) = fs::metadata(&data_path) {
if !meta.is_dir() {
return true;
}
if is_usable_ecx_file(&format!("{}.ecx", data_base)) {
return true;
}
}
false
@@ -799,6 +815,20 @@ impl DiskLocation {
self.ec_volumes.remove(&vid)
}
/// Drop the in-memory EC volume for vid and close its descriptors without
/// deleting files, so a following unlink frees the inodes instead of
/// leaving open fds serving the old bytes. Mirrors Go's unloadEcVolume.
pub fn unload_ec_volume(&mut self, vid: VolumeId) {
if let Some(mut ec_vol) = self.ec_volumes.remove(&vid) {
for _ in 0..ec_vol.shard_count() {
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&ec_vol.collection, "ec_shards"])
.dec();
}
ec_vol.close();
}
}
/// Mount EC shards for a volume on this location.
///
/// `source_disk_type` is the source volume's disk type carried on the
@@ -811,7 +841,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
source_disk_type: &str,
) -> Result<(), VolumeError> {
let idx_dir = self.idx_directory.clone();
@@ -833,7 +863,7 @@ impl DiskLocation {
&mut self,
vid: VolumeId,
collection: &str,
shard_ids: &[u32],
shard_ids: &[ShardId],
idx_dir: &str,
source_disk_type: &str,
) -> Result<(), VolumeError> {
@@ -845,14 +875,10 @@ impl DiskLocation {
// propagate the error to the caller.
let created = !self.ec_volumes.contains_key(&vid);
if created {
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid)
.map_err(VolumeError::Io)?;
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid).map_err(VolumeError::Io)?;
self.ec_volumes.insert(vid, ec_vol);
}
let ec_vol = self
.ec_volumes
.get_mut(&vid)
.expect("just inserted above");
let ec_vol = self.ec_volumes.get_mut(&vid).expect("just inserted above");
// When the orchestrator supplied a source disk type on the Mount
// RPC, override the EC volume's disk type so heartbeats report
// under the source volume's disk type (#9423). When the caller
@@ -871,10 +897,10 @@ impl DiskLocation {
// keep the existing registration (mirrors Go's AddEcVolumeShard
// added=false) — re-adding would replace a serving fd and bump
// the ec_shards gauge without growing the mounted count.
if ec_vol.has_shard(shard_id as u8) {
if ec_vol.has_shard(shard_id) {
continue;
}
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id as u8);
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id);
shard.disk_type = ec_vol.disk_type.clone();
if let Err(e) = ec_vol.add_shard(shard) {
// The shard was dropped (its descriptors closed) inside the
@@ -902,14 +928,14 @@ impl DiskLocation {
/// caller passes a shard that lives on a sibling disk
/// (cross-disk reconcile makes that the common case for the same
/// `vid` after reconciliation).
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[u32]) {
pub fn unmount_ec_shards(&mut self, vid: VolumeId, shard_ids: &[ShardId]) {
if let Some(ec_vol) = self.ec_volumes.get_mut(&vid) {
let collection = ec_vol.collection.clone();
for &shard_id in shard_ids {
if !ec_vol.has_shard(shard_id as u8) {
if !ec_vol.has_shard(shard_id) {
continue;
}
ec_vol.remove_shard(shard_id as u8);
let _ = ec_vol.remove_shard(shard_id);
crate::metrics::VOLUME_GAUGE
.with_label_values(&[&collection, "ec_shards"])
.dec();
@@ -969,7 +995,7 @@ impl DiskLocation {
}
entries.sort();
let mut same_volume_shards: Vec<(String, u32)> = Vec::new(); // (filename, shard_id)
let mut same_volume_shards: Vec<(String, ShardId)> = Vec::new(); // (filename, shard_id)
let mut prev_vid: Option<VolumeId> = None;
let mut prev_collection: String = String::new();
@@ -1034,7 +1060,12 @@ impl DiskLocation {
/// Validate + mount a (collection, vid) group when its `.ecx` is
/// found. Mirrors `handleFoundEcxFile` in
/// `weed/storage/disk_location_ec.go`.
fn handle_found_ecx_file(&mut self, shards: &[(String, u32)], collection: &str, vid: VolumeId) {
fn handle_found_ecx_file(
&mut self,
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) {
let base = volume_file_name(&self.directory, collection, vid);
let dat_path = format!("{}.dat", base);
let dat_exists = check_dat_file_exists(&dat_path);
@@ -1048,7 +1079,7 @@ impl DiskLocation {
return;
}
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if let Err(e) = self.mount_ec_shards(vid, collection, &shard_ids, "") {
// A mount failure (corrupt/locked .ecx, EMFILE, transient I/O) is
// not proof the shards are disposable -- validate_ec_volume already
@@ -1057,8 +1088,7 @@ impl DiskLocation {
// delete on a load error.
warn!(
volume_id = vid.0,
"Failed to load EC shards: {}; keeping files for retry",
e,
"Failed to load EC shards: {}; keeping files for retry", e,
);
self.unmount_ec_shards(vid, &shard_ids);
}
@@ -1071,7 +1101,7 @@ impl DiskLocation {
/// distributed-EC shards waiting for cross-disk reconciliation.
fn check_orphaned_shards(
&self,
shards: &[(String, u32)],
shards: &[(String, ShardId)],
collection: &str,
vid: VolumeId,
) -> bool {
@@ -1107,7 +1137,7 @@ impl DiskLocation {
/// Close all volumes.
pub fn close(&mut self) {
for (_, v) in self.volumes.iter_mut() {
for v in self.volumes.values_mut() {
v.close();
}
self.volumes.clear();
@@ -1127,20 +1157,60 @@ pub fn get_disk_stats(path: &str) -> (u64, u64) {
Ok(p) => p,
Err(_) => return (0, 0),
};
unsafe {
let mut stat: libc::statvfs = std::mem::zeroed();
if libc::statvfs(c_path.as_ptr(), &mut stat) == 0 {
let all = stat.f_blocks as u64 * stat.f_frsize as u64;
let free = stat.f_bavail as u64 * stat.f_frsize as u64;
return (all, free);
}
// SAFETY: `libc::statvfs` is plain data — integers and reserved
// padding, no pointers and no restricted niches — so the all-zero
// value is a valid one for the call to overwrite.
let mut stat: libc::statvfs = unsafe { std::mem::zeroed() };
// SAFETY: `c_path` is a live NUL-terminated `CString` that outlives
// the call, and `&mut stat` is a live, aligned, exclusive pointer the
// kernel only writes through; the fields are read below only after
// the call reports success.
if unsafe { libc::statvfs(c_path.as_ptr(), &mut stat) } != 0 {
return (0, 0);
}
(0, 0)
let all = stat.f_blocks as u64 * stat.f_frsize as u64;
let free = stat.f_bavail as u64 * stat.f_frsize as u64;
(all, free)
}
#[cfg(not(unix))]
#[cfg(windows)]
{
let _ = path;
(0, 0)
use std::os::windows::ffi::OsStrExt;
// Canonicalize so symlinks, `.`/`..` segments, and relative paths
// resolve to the real location before querying. `\\?\`-prefixed
// extended-length paths and UNC (`\\?\UNC\...`) are passed through
// untouched: GetDiskFreeSpaceExW accepts them as-is.
let canonical = match std::fs::canonicalize(path) {
Ok(p) => p,
Err(_) => return (0, 0),
};
// UTF-16 with trailing NUL for the Win32 wide-string call.
let mut wide: Vec<u16> = canonical.as_os_str().encode_wide().collect();
// UNC directory names must end in a backslash for GetDiskFreeSpaceExW.
if !wide.ends_with(&[0x5C]) {
wide.push(0x5C);
}
wide.push(0);
// SAFETY: `wide` is NUL-terminated; the out-params are valid u64
// writes; the call has no other preconditions.
unsafe {
let mut free_available: u64 = 0;
let mut total: u64 = 0;
let ok = windows_sys::Win32::Storage::FileSystem::GetDiskFreeSpaceExW(
wide.as_ptr(),
&mut free_available,
&mut total,
std::ptr::null_mut(),
);
if ok == 0 {
return (0, 0);
}
return (total, free_available);
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("get_disk_stats is implemented for unix and windows only");
}
}
@@ -1176,7 +1246,12 @@ fn rm_if_present(path: String) -> io::Result<()> {
}
}
fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &str, vid: VolumeId) -> usize {
fn ec_data_shards_from_vif(
directory: &str,
idx_directory: &str,
collection: &str,
vid: VolumeId,
) -> usize {
for dir in [directory, idx_directory] {
let vif = format!("{}.vif", volume_file_name(dir, collection, vid));
if let Some(ds) = fs::read_to_string(&vif)
@@ -1184,10 +1259,9 @@ fn ec_data_shards_from_vif(directory: &str, idx_directory: &str, collection: &st
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
.and_then(|vi| vi.ec_shard_config)
.map(|c| c.data_shards as usize)
&& ds > 0
{
if ds > 0 {
return ds;
}
return ds;
}
if directory == idx_directory {
break;
@@ -1223,7 +1297,7 @@ fn parse_collection_volume_id(base: &str) -> Option<(String, VolumeId)> {
/// `pub(crate)` re-export of [`parse_ec_shard_extension`] for the
/// cross-disk reconcile in `store_ec_reconcile.rs`.
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<ShardId> {
parse_ec_shard_extension(ext)
}
@@ -1237,7 +1311,7 @@ pub(crate) fn is_ec_shard_extension(ext: &str) -> Option<u32> {
/// shardId > 255` guard. The 3-digit form (`.ec100`–`.ec255`) is
/// retained so the parser can still recognise shards from custom
/// 32+ ratios that fit in a u8 even though OSS only ships 10+4.
fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
fn parse_ec_shard_extension(ext: &str) -> Option<ShardId> {
let rest = ext.strip_prefix(".ec")?;
if rest.len() < 2 || rest.len() > 3 {
return None;
@@ -1246,7 +1320,7 @@ fn parse_ec_shard_extension(ext: &str) -> Option<u32> {
if id > 255 {
return None;
}
Some(id)
ShardId::try_from(id).ok()
}
/// Robust check that a `.dat` with actual data exists. An empty `.dat`
@@ -1265,7 +1339,7 @@ fn check_dat_file_exists(path: &str) -> bool {
/// True when a `.vif` references remote-tier files: a remote-only volume
/// that has no local `.dat` but must still load via the remote path,
/// rather than be skipped as a lone EC sidecar.
fn vif_references_remote_file(vif_path: &str) -> bool {
pub(crate) fn vif_references_remote_file(vif_path: &str) -> bool {
fs::read_to_string(vif_path)
.ok()
.and_then(|s| serde_json::from_str::<VifVolumeInfo>(&s).ok())
@@ -1308,7 +1382,10 @@ fn remove_empty_ec_dat_stub(volume_name: &str, idx_name: &str, vid: VolumeId) ->
return false;
}
warn!(volume_id = vid.0, "removing leftover empty .dat stub for EC volume");
warn!(
volume_id = vid.0,
"removing leftover empty .dat stub for EC volume"
);
let _ = fs::remove_file(&dat_path);
let _ = fs::remove_file(format!("{}.idx", idx_name));
true
@@ -1331,6 +1408,17 @@ mod tests {
use super::*;
use tempfile::TempDir;
/// get_disk_stats must report real capacity for a real path on every
/// platform (Windows included) — consumers treat total==0 as "unknown"
/// and leave available_space at 0, which breaks volume assignment.
#[test]
fn test_get_disk_stats_reports_capacity_for_real_path() {
let tmp = TempDir::new().unwrap();
let (total, free) = get_disk_stats(tmp.path().to_str().unwrap());
assert!(total > 0, "expected total>0, got {total}");
assert!(free > 0, "expected free>0, got {free}");
}
/// When `-dir.idx` is configured the EC `.vif` may live in the idx
/// directory; the sweep must look there too, not only the data dir.
#[test]
@@ -1353,7 +1441,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&vif).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&vif).unwrap(),
)
.unwrap();
assert!(
remove_empty_ec_dat_stub(&vbase, &ibase, VolumeId(42)),
@@ -1369,16 +1461,30 @@ mod tests {
fn test_validate_ec_volume_partial_dat_next_to_full_shards_keeps() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(70));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let full = calculate_expected_shard_size(30 * 1024 * 1024, ds);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(full as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(full as u64)
.unwrap();
}
// Partial .dat: bigger than a superblock so it is not swept as a stub,
// but smaller than what these shards encode.
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(5 * 1024 * 1024).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(5 * 1024 * 1024)
.unwrap();
assert!(
loc.validate_ec_volume("", VolumeId(70)),
"full-size shards beside a smaller (stale/partial) .dat must be kept",
@@ -1392,15 +1498,29 @@ mod tests {
fn test_validate_ec_volume_interrupted_encode_reclaims() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let loc = DiskLocation::new(dir, dir, 10, DiskType::HardDrive, MinFreeSpace::Percent(1.0), Vec::new()).unwrap();
let loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
let base = volume_file_name(dir, "", VolumeId(71));
let ds = crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
let dat_size = 30 * 1024 * 1024i64;
std::fs::File::create(format!("{}.dat", base)).unwrap().set_len(dat_size as u64).unwrap();
std::fs::File::create(format!("{}.dat", base))
.unwrap()
.set_len(dat_size as u64)
.unwrap();
let partial = calculate_expected_shard_size(dat_size, ds) / 3;
assert!(partial > 0);
for i in 0..ds {
std::fs::File::create(format!("{}.ec{:02}", base, i)).unwrap().set_len(partial as u64).unwrap();
std::fs::File::create(format!("{}.ec{:02}", base, i))
.unwrap()
.set_len(partial as u64)
.unwrap();
}
assert!(
!loc.validate_ec_volume("", VolumeId(71)),
@@ -1445,7 +1565,11 @@ mod tests {
}),
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
// A .vif with no EC config reads as generation 0 (recovered/pre-upgrade live volume).
@@ -1454,12 +1578,20 @@ mod tests {
version: 3,
..Default::default()
};
std::fs::write(format!("{}.vif", dbase), serde_json::to_string(&no_cfg).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", dbase),
serde_json::to_string(&no_cfg).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(0));
// idx-dir fallback: only the idx dir holds the .vif.
std::fs::remove_file(format!("{}.vif", dbase)).unwrap();
std::fs::write(format!("{}.vif", ibase), serde_json::to_string(&with_gen).unwrap()).unwrap();
std::fs::write(
format!("{}.vif", ibase),
serde_json::to_string(&with_gen).unwrap(),
)
.unwrap();
assert_eq!(loc.ec_generation_ts_ns("", vid), Some(4242));
}
@@ -1499,16 +1631,8 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 1);
assert!(loc.find_volume(VolumeId(1)).is_some());
@@ -1532,24 +1656,15 @@ mod tests {
Vec::new(),
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(
VolumeId(2),
"test",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "test",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1592,12 +1707,11 @@ mod tests {
.unwrap();
loc.create_volume(
VolumeId(9),
"good",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "good",
..Default::default()
},
)
.unwrap();
loc.close();
@@ -1641,29 +1755,13 @@ mod tests {
)
.unwrap();
loc.create_volume(
VolumeId(1),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(
VolumeId(2),
"",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
loc.create_volume(VolumeId(1), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
loc.create_volume(VolumeId(2), NeedleMapKind::InMemory, &VolumeSpec::default())
.unwrap();
assert_eq!(loc.volumes_len(), 2);
loc.delete_volume(VolumeId(1), false, false).unwrap();
loc.delete_volume(VolumeId(1), false, false, false).unwrap();
assert_eq!(loc.volumes_len(), 1);
assert!(loc.find_volume(VolumeId(1)).is_none());
}
@@ -1684,32 +1782,29 @@ mod tests {
loc.create_volume(
VolumeId(1),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(2),
"pics",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "pics",
..Default::default()
},
)
.unwrap();
loc.create_volume(
VolumeId(3),
"docs",
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec {
collection: "docs",
..Default::default()
},
)
.unwrap();
assert_eq!(loc.volumes_len(), 3);
@@ -1719,6 +1814,34 @@ mod tests {
assert!(loc.find_volume(VolumeId(3)).is_some());
}
/// A 0-byte `.ecx` is the stub a failed EC distribute copy leaves behind.
/// Go's HasEcxFileOnDisk requires Size() > 0 so the stub cannot pin
/// placement to a disk that has no usable index.
#[test]
fn test_has_ecx_file_on_disk_ignores_zero_byte_stub() {
let tmp = TempDir::new().unwrap();
let data = tmp.path().join("data");
let idx = tmp.path().join("idx");
fs::create_dir_all(&data).unwrap();
fs::create_dir_all(&idx).unwrap();
let loc = DiskLocation::new(
data.to_str().unwrap(),
idx.to_str().unwrap(),
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
fs::write(idx.join("pics_7.ecx"), b"").unwrap();
assert!(!loc.has_ecx_file_on_disk("pics", VolumeId(7)));
// A real index in the data dir still counts, stub or no stub.
fs::write(data.join("pics_7.ecx"), [0u8; 16]).unwrap();
assert!(loc.has_ecx_file_on_disk("pics", VolumeId(7)));
}
#[test]
fn test_disk_location_delete_collection_removes_ec_volumes() {
let tmp = TempDir::new().unwrap();
@@ -1735,6 +1858,8 @@ mod tests {
let shard_path = format!("{}/pics_7.ec00", dir);
std::fs::write(&shard_path, b"ec-shard").unwrap();
// An EC volume needs its .ecx to mount.
std::fs::write(format!("{}/pics_7.ecx", dir), [0u8; 16]).unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "").unwrap();
assert!(loc.has_ec_volume(VolumeId(7)));
@@ -1772,7 +1897,9 @@ mod tests {
// mount_ec_shards with source_disk_type="ssd" — simulating the
// VolumeEcShardsMount RPC path.
std::fs::write(format!("{}/pics_7.ec00", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd").unwrap();
std::fs::write(format!("{}/pics_7.ecx", dir), [0u8; 16]).unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[0], "ssd")
.unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume mounted");
assert_eq!(
@@ -1789,7 +1916,9 @@ mod tests {
std::fs::write(format!("{}/pics_7.ec01", dir), b"ec-shard").unwrap();
loc.mount_ec_shards(VolumeId(7), "pics", &[1], "").unwrap();
{
let ec_vol = loc.find_ec_volume(VolumeId(7)).expect("ec volume still mounted");
let ec_vol = loc
.find_ec_volume(VolumeId(7))
.expect("ec volume still mounted");
assert_eq!(
ec_vol.disk_type,
DiskType::Ssd,
@@ -1868,10 +1997,12 @@ mod tests {
// A collection name unique to this test: the gauge is process-global
// and sibling tests running in parallel touch other labels.
std::fs::write(format!("{}/dupmount_11.ec00", dir), b"shard bytes").unwrap();
std::fs::write(format!("{}/dupmount_11.ecx", dir), [0u8; 16]).unwrap();
let gauge = crate::metrics::VOLUME_GAUGE.with_label_values(&["dupmount", "ec_shards"]);
let before = gauge.get();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "").unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.expect("a duplicate mount must succeed as a no-op");
@@ -1947,8 +2078,11 @@ mod tests {
let path = format!("{}/{}_{}.ec{:02}", dir, collection, vid.0, sid);
std::fs::write(&path, b"shard data nonempty").unwrap();
}
std::fs::write(format!("{}/{}_{}.ecx", dir, collection, vid.0), vec![0u8; 20])
.unwrap();
std::fs::write(
format!("{}/{}_{}.ecx", dir, collection, vid.0),
vec![0u8; 20],
)
.unwrap();
std::fs::write(format!("{}/{}_{}.ecj", dir, collection, vid.0), b"").unwrap();
std::fs::write(
format!("{}/{}_{}.vif", dir, collection, vid.0),
@@ -30,6 +30,7 @@ use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums, EcShardConfig,
};
use crate::storage::erasure_coding::ec_shard::MAX_SHARD_COUNT;
use crate::storage::io::read_exact_at;
use crate::storage::needle::crc::CRC;
/// Canonical extension for the checksum sidecar. Generation 0 (legacy/fresh
@@ -164,16 +165,16 @@ pub fn remove_bitrot_sidecars(base: &str) -> io::Result<()> {
};
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res {
if first_err.is_none() {
first_err = Some(e);
}
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
record(rm(format!("{}{}", base, BITROT_SIDECAR_EXT).into()));
let path = Path::new(base);
if let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) {
let prefix = format!("{}{}.v", fname.to_string_lossy(), BITROT_SIDECAR_EXT);
let prefix = format!("{}{}.v", fname.display(), BITROT_SIDECAR_EXT);
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries.flatten() {
@@ -203,7 +204,7 @@ pub fn new_encode_uuid() -> Vec<u8> {
/// Reports whether `block_size` is a power of two in [1 MiB, MAX_BITROT_BLOCK_SIZE].
pub fn is_pow2_multiple_of_1mib(block_size: u32) -> bool {
block_size >= (1 << 20) && block_size <= MAX_BITROT_BLOCK_SIZE && block_size.count_ones() == 1
((1 << 20)..=MAX_BITROT_BLOCK_SIZE).contains(&block_size) && block_size.count_ones() == 1
}
/// Returns ceil(covered_size / block_size).
@@ -402,7 +403,7 @@ pub fn validate_manifest(
total
));
}
let mut seen = vec![false; MAX_SHARD_COUNT];
let mut seen = [false; MAX_SHARD_COUNT];
for s in &prot.shards {
if s.shard_id >= total as u32 {
return Err(format!(
@@ -505,7 +506,21 @@ pub fn verify_shard_file_blocks(
entry: &EcShardChecksums,
block_size: i64,
) -> io::Result<Vec<usize>> {
let f = File::open(path)?;
verify_shard_blocks(&File::open(path)?, entry, block_size)
}
/// Same verification against an ALREADY-OPEN shard handle.
///
/// Go's `ChecksumScrub` reads through `shard.ReadAt`, i.e. the handle the
/// EcVolumeShard already holds, so a concurrent teardown that unlinks the shard
/// cannot turn an intentional removal into a scrub read error. A scrub that
/// runs with the store lock released has to read the same way — see
/// `EcChecksumScrubPlan`.
pub fn verify_shard_blocks(
f: &File,
entry: &EcShardChecksums,
block_size: i64,
) -> io::Result<Vec<usize>> {
let file_size = f.metadata()?.len() as i64;
let want = unpack_u32_le(&entry.block_crc32c);
@@ -523,7 +538,7 @@ pub fn verify_shard_file_blocks(
break;
}
let to_read = to_read as usize;
read_full_at(&f, &mut buf[..to_read], offset as u64)?;
read_exact_at(f, &mut buf[..to_read], offset as u64)?;
if CRC::new(&buf[..to_read]).0 != *want_crc {
mismatched.push(i);
}
@@ -532,33 +547,6 @@ pub fn verify_shard_file_blocks(
Ok(mismatched)
}
/// Reads exactly `buf.len()` bytes from `f` at `offset`, erroring on early EOF.
fn read_full_at(f: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
let mut total = 0usize;
while total < buf.len() {
#[cfg(unix)]
let n = {
use std::os::unix::fs::FileExt;
f.read_at(&mut buf[total..], offset + total as u64)?
};
#[cfg(not(unix))]
let n = {
use std::io::{Read, Seek, SeekFrom};
let mut fc = f.try_clone()?;
fc.seek(SeekFrom::Start(offset + total as u64))?;
fc.read(&mut buf[total..])?
};
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"short read on shard block",
));
}
total += n;
}
Ok(())
}
/// Builds the `EcShardConfig` proto for the given layout. The bitrot sidecar
/// carries its own top-level encode_uuid, so the nested config leaves it empty.
pub fn ec_shard_config(data_shards: u32, parity_shards: u32, block_size: i64) -> EcShardConfig {
@@ -613,7 +601,10 @@ mod tests {
save_bitrot_sidecar(path, &prot).unwrap();
let bytes = std::fs::read(path).unwrap();
let hex: String = bytes.iter().map(|b| format!("{:02x}", b)).collect();
assert_eq!(hex, CANONICAL_HEX, "Rust .ecsum bytes drifted from the Go canonical form");
assert_eq!(
hex, CANONICAL_HEX,
"Rust .ecsum bytes drifted from the Go canonical form"
);
let _ = std::fs::remove_file(path);
}
@@ -648,7 +639,11 @@ mod tests {
format!("{}.ecsum.v1", base),
format!("{}.ecsum.v7", base),
] {
assert!(!std::path::Path::new(&p).exists(), "{} should be removed", p);
assert!(
!std::path::Path::new(&p).exists(),
"{} should be removed",
p
);
}
assert!(std::path::Path::new(&keep_shard).exists());
assert!(std::path::Path::new(&keep_other_vid).exists());
@@ -667,7 +662,9 @@ mod tests {
assert!(!is_pow2_multiple_of_1mib(1 << 19)); // 512 KiB, too small
assert!(!is_pow2_multiple_of_1mib(3 << 20)); // 3 MiB, not pow2
assert!(!is_pow2_multiple_of_1mib(128 * 1024 * 1024)); // pow2 but > MAX_BITROT_BLOCK_SIZE
assert!(!is_pow2_multiple_of_1mib(DEFAULT_BITROT_BLOCK_SIZE as u32 + 1));
assert!(!is_pow2_multiple_of_1mib(
DEFAULT_BITROT_BLOCK_SIZE as u32 + 1
));
}
#[test]
@@ -721,12 +718,7 @@ mod tests {
#[test]
fn test_save_load_roundtrip() {
let tmp = tempfile::TempDir::new().unwrap();
let path = tmp
.path()
.join("vol.ecsum")
.to_str()
.unwrap()
.to_string();
let path = tmp.path().join("vol.ecsum").to_str().unwrap().to_string();
let mut builder = ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64);
builder.write(b"hello world");
@@ -887,8 +879,7 @@ mod tests {
assert_eq!(resolve_status(&notfound, 0, 10, 4), BitrotStatus::Off);
// Integrity failure => Invalid.
let bad: Result<EcBitrotProtection, BitrotLoadError> =
Err(BitrotLoadError::BadMagic(0));
let bad: Result<EcBitrotProtection, BitrotLoadError> = Err(BitrotLoadError::BadMagic(0));
assert_eq!(resolve_status(&bad, 0, 10, 4), BitrotStatus::Invalid);
// Generation mismatch => Off.
@@ -3,10 +3,12 @@
//! Rebuilds the original .dat + .idx files from data shards (.ec00-.ec09)
//! and the sorted index (.ecx) + deletion journal (.ecj).
use std::collections::HashSet;
use std::fs::File;
use std::io::{self, Read, Write};
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::erasure_coding::ec_volume::read_ecj_ids;
use crate::storage::idx;
use crate::storage::needle::needle::get_actual_size;
use crate::storage::super_block::SUPER_BLOCK_SIZE;
@@ -20,18 +22,21 @@ use crate::storage::volume::{fsync_dir, volume_file_name};
/// `dir` is used both for reading `.ec00` and `.ecx`. For split-disk
/// reconciled volumes call [`find_dat_file_size_with_dirs`] instead.
pub fn find_dat_file_size(dir: &str, collection: &str, volume_id: VolumeId) -> io::Result<i64> {
find_dat_file_size_with_dirs(dir, dir, collection, volume_id)
let deleted = read_ecj_deletions(&[dir], collection, volume_id)?;
find_dat_file_size_with_dirs(dir, dir, collection, volume_id, &deleted)
}
/// Like [`find_dat_file_size`] but lets the caller pass separate dirs
/// for `.ec00` (the data shard) and `.ecx` (the sealed index). This
/// is the form needed when shards are split across data dirs and the
/// `.ecx` lives on a sibling disk's idx dir (#9252).
/// `.ecx` lives on a sibling disk's idx dir (#9252). Needles in `deleted`
/// count as deleted.
pub fn find_dat_file_size_with_dirs(
ec00_dir: &str,
ecx_dir: &str,
collection: &str,
volume_id: VolumeId,
deleted: &HashSet<NeedleId>,
) -> io::Result<i64> {
let ec00_base = volume_file_name(ec00_dir, collection, volume_id);
let ecx_base = volume_file_name(ecx_dir, collection, volume_id);
@@ -53,9 +58,9 @@ pub fn find_dat_file_size_with_dirs(
for i in 0..entry_count {
let start = i * NEEDLE_MAP_ENTRY_SIZE;
let (_, offset, size) =
let (key, offset, size) =
idx_entry_from_bytes(&ecx_data[start..start + NEEDLE_MAP_ENTRY_SIZE]);
if size.is_deleted() {
if size.is_deleted() || deleted.contains(&key) {
continue;
}
let entry_stop = offset.to_actual_offset() + get_actual_size(size, version);
@@ -67,72 +72,87 @@ pub fn find_dat_file_size_with_dirs(
Ok(dat_size)
}
/// Reconstruct a .dat file from EC data shards.
///
/// Reads from .ec00-.ec09 and writes a new .dat file. All data shards
/// must live in `dir`. For the cross-disk reconciled layout where
/// shards are split across multiple data dirs of the same node, use
/// [`write_dat_file_from_shards_with_dirs`] instead.
pub fn write_dat_file_from_shards(
dir: &str,
/// Whether the `.ecx` in `ecx_dir` indexes a needle deleted neither there nor
/// in `deleted`.
pub fn has_live_needles(
ecx_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let dirs: Vec<String> = (0..data_shards).map(|_| dir.to_string()).collect();
write_dat_file_from_shards_with_dirs(
dir,
collection,
volume_id,
dat_file_size,
encoded_dat_file_size,
data_shards,
&dirs,
large_block_size,
small_block_size,
)
deleted: &HashSet<NeedleId>,
) -> io::Result<bool> {
let ecx_base = volume_file_name(ecx_dir, collection, volume_id);
let ecx_data = std::fs::read(format!("{}.ecx", ecx_base))?;
let (entries, _) = ecx_data.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>();
Ok(entries.iter().any(|entry| {
let (key, _, size) = idx_entry_from_bytes(entry);
!size.is_deleted() && !deleted.contains(&key)
}))
}
/// Reconstruct a .dat file from EC data shards, taking the source
/// directory for each shard separately.
///
/// `dat_dir` is where the produced `.dat` is written. `shard_dirs[i]`
/// is the directory holding shard `i`. For the simple "all shards in
/// one dir" case both can be the same value.
/// Distinct needle ids journaled in the `.ecj` of any of `dirs`. Go folds the
/// journal into the `.ecx` (RebuildEcxFile) before a decode; reading it leaves
/// the sealed index untouched. Only NotFound means "no journal".
pub fn read_ecj_deletions(
dirs: &[&str],
collection: &str,
volume_id: VolumeId,
) -> io::Result<HashSet<NeedleId>> {
let mut ids = HashSet::new();
for (i, dir) in dirs.iter().enumerate() {
if dirs[..i].contains(dir) {
continue;
}
let path = format!("{}.ecj", volume_file_name(dir, collection, volume_id));
let file = match File::open(&path) {
Ok(file) => file,
Err(e) if e.kind() == io::ErrorKind::NotFound => continue,
Err(e) => return Err(e),
};
let len = file.metadata()?.len();
read_ecj_ids(&file, len, &mut ids)?;
}
Ok(ids)
}
/// What it takes to rebuild a volume's .dat from its EC data shards.
///
/// Mirrors Go's `WriteDatFile(baseFileName, datFileSize,
/// encodedDatFileSize, shardFileNames)` shape — Go passes per-shard
/// paths so a reconciled volume with shards split across disks of the
/// same volume server can still be decoded back to a regular .dat
/// (seaweedfs/seaweedfs#9252).
#[derive(Clone, Copy, Debug)]
pub struct DatRebuild<'a> {
/// Where the produced `.dat` is written.
pub dat_dir: &'a str,
pub collection: &'a str,
pub volume_id: VolumeId,
/// The number of bytes to write, i.e. the live data extent from
/// [`find_dat_file_size`].
pub dat_file_size: i64,
/// The .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Zero when the .vif does not
/// record the encode-time size; the layout is then inferred from the
/// shard size.
pub encoded_dat_file_size: i64,
pub data_shards: usize,
/// `shard_dirs[i]` is the directory holding shard `i`. `None` means every
/// data shard sits in `dat_dir`.
pub shard_dirs: Option<&'a [String]>,
/// The volume's shard block layout, e.g. `EcVolume::large_block_size()`
/// / `small_block_size()` from its .vif EC config.
pub large_block_size: usize,
pub small_block_size: usize,
}
/// Reconstruct a .dat file from EC data shards.
///
/// `dat_file_size` is the number of bytes to write, i.e. the live data
/// extent from [`find_dat_file_size`]. `encoded_dat_file_size` is the
/// .dat size at encode time, which fixed the shard block layout:
/// deletions can move the live extent below the large-block row
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Pass zero when the .vif does
/// not record the encode-time size to infer the layout from the shard
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
write_dat_file(
/// Reads from .ec00-.ec09 and writes a new .dat file, from one directory or
/// from the per-shard directories of a cross-disk reconciled volume.
pub fn write_dat_file_from_shards(spec: &DatRebuild<'_>) -> io::Result<()> {
let DatRebuild {
dat_dir,
collection,
volume_id,
@@ -142,21 +162,15 @@ pub fn write_dat_file_from_shards_with_dirs(
shard_dirs,
large_block_size,
small_block_size,
)
}
#[allow(clippy::too_many_arguments)]
fn write_dat_file(
dat_dir: &str,
collection: &str,
volume_id: VolumeId,
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
} = *spec;
let same_dir: Vec<String>;
let shard_dirs: &[String] = match shard_dirs {
Some(dirs) => dirs,
None => {
same_dir = vec![dat_dir.to_string(); data_shards];
&same_dir
}
};
if data_shards == 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
@@ -233,10 +247,10 @@ fn write_dat_file(
// Read large blocks
while encoded_remaining >= large_row_size && remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = large_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -255,10 +269,10 @@ fn write_dat_file(
// Read small blocks
while remaining > 0 {
for i in 0..data_shards {
for (i, shard) in shards[..data_shards].iter().enumerate() {
let to_write = small_block_size.min(remaining as usize);
let mut buf = vec![0u8; to_write];
let n = shards[i].read_at(&mut buf, shard_offset)?;
let n = shard.read_at(&mut buf, shard_offset)?;
if n != to_write {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
@@ -300,54 +314,61 @@ fn write_dat_file(
/// Write .idx file from .ecx index + .ecj deletion journal.
///
/// Copies sorted .ecx entries to .idx, then appends tombstones for
/// deleted needles from .ecj.
/// See [`write_idx_file_from_ec_index_with_dirs`]; everything lives in `dir`.
pub fn write_idx_file_from_ec_index(
dir: &str,
collection: &str,
volume_id: VolumeId,
) -> io::Result<()> {
let base = volume_file_name(dir, collection, volume_id);
let ecx_path = format!("{}.ecx", base);
let ecj_path = format!("{}.ecj", base);
let idx_path = format!("{}.idx", base);
let deleted = read_ecj_deletions(&[dir], collection, volume_id)?;
let dat_file_size = find_dat_file_size_with_dirs(dir, dir, collection, volume_id, &deleted)?;
write_idx_file_from_ec_index_with_dirs(dir, dir, collection, volume_id, &deleted, dat_file_size)
}
/// Write the `.idx` for a `.dat` decoded to `dat_file_size` bytes, from the
/// `.ecx` in `ecx_dir`, into `idx_dir`.
///
/// Copies the `.ecx` rows, then appends one tombstone per row whose needle is
/// in `deleted`. A deleted needle at or past `dat_file_size` was cut from the
/// `.dat`, so its row is dropped: a row pointing past the end of the `.dat`
/// makes the volume load read-only.
pub fn write_idx_file_from_ec_index_with_dirs(
ecx_dir: &str,
idx_dir: &str,
collection: &str,
volume_id: VolumeId,
deleted: &HashSet<NeedleId>,
dat_file_size: i64,
) -> io::Result<()> {
let ecx_path = format!("{}.ecx", volume_file_name(ecx_dir, collection, volume_id));
let idx_path = format!("{}.idx", volume_file_name(idx_dir, collection, volume_id));
// Write to a temp file and atomically rename into place, so a crash
// mid-write never leaves a partial .idx at the final name beside the
// source shards.
let tmp_path = format!("{}.tmp", idx_path);
let write_result = (|| -> io::Result<()> {
// Copy .ecx to the temp .idx
std::fs::copy(&ecx_path, &tmp_path)?;
// Append deletions from .ecj as tombstones. Read the journal directly
// and treat only NotFound as "no journal": Path::exists would also
// swallow a permission/IO error and silently skip deletions, which
// would resurrect deleted needles as live.
let mut idx_file = std::fs::OpenOptions::new()
.write(true)
.append(true)
.open(&tmp_path)?;
match std::fs::read(&ecj_path) {
Ok(ecj_data) => {
let count = ecj_data.len() / NEEDLE_ID_SIZE;
for i in 0..count {
let start = i * NEEDLE_ID_SIZE;
let needle_id = NeedleId::from_bytes(&ecj_data[start..start + NEEDLE_ID_SIZE]);
idx::write_index_entry(
&mut idx_file,
needle_id,
Offset::default(),
TOMBSTONE_FILE_SIZE,
)?;
}
let mut ecx_file = File::open(&ecx_path)?;
let mut idx_file = io::BufWriter::new(File::create(&tmp_path)?);
let mut tombstoned = Vec::new();
idx::walk_index_file(&mut ecx_file, 0, |key, offset, size| {
let is_deleted = size.is_deleted() || deleted.contains(&key);
if is_deleted && offset.to_actual_offset() >= dat_file_size {
return Ok(());
}
Err(e) if e.kind() == io::ErrorKind::NotFound => {}
Err(e) => return Err(e),
idx::write_index_entry(&mut idx_file, key, offset, size)?;
if !size.is_deleted() && deleted.contains(&key) {
tombstoned.push(key);
}
Ok(())
})?;
for key in tombstoned {
idx::write_index_entry(&mut idx_file, key, Offset::default(), TOMBSTONE_FILE_SIZE)?;
}
// fsync, rename, then fsync the dir so the decoded .idx is durable and
// atomically published before the caller deletes the source shards.
let idx_file = idx_file.into_inner().map_err(|e| e.into_error())?;
idx_file.sync_all()?;
drop(idx_file);
// Windows rename does not replace an existing file on every version;
@@ -372,7 +393,7 @@ mod tests {
use crate::storage::erasure_coding::ec_encoder;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -384,13 +405,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -430,16 +447,17 @@ mod tests {
std::fs::remove_file(format!("{}/1.idx", dir)).unwrap();
// Reconstruct from EC shards
write_dat_file_from_shards(
dir,
"",
VolumeId(1),
original_dat_size as i64,
original_dat_size as i64,
write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(1),
dat_file_size: original_dat_size as i64,
encoded_dat_file_size: original_dat_size as i64,
data_shards,
block_size as usize,
block_size as usize,
)
shard_dirs: None,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
})
.unwrap();
write_idx_file_from_ec_index(dir, "", VolumeId(1)).unwrap();
@@ -459,13 +477,9 @@ mod tests {
let v2 = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -485,29 +499,29 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
// No shard files exist, so de-striping must fail and publish nothing:
// neither the final .dat nor a partial .dat.tmp may remain.
let res = write_dat_file_from_shards(
dir,
"",
VolumeId(7),
100,
100,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
);
let res = write_dat_file_from_shards(&DatRebuild {
dat_dir: dir,
collection: "",
volume_id: VolumeId(7),
dat_file_size: 100,
encoded_dat_file_size: 100,
data_shards: 10,
shard_dirs: None,
large_block_size: ERASURE_CODING_LARGE_BLOCK_SIZE,
small_block_size: ERASURE_CODING_SMALL_BLOCK_SIZE,
});
assert!(res.is_err());
assert!(!std::path::Path::new(&format!("{}/7.dat", dir)).exists());
assert!(!std::path::Path::new(&format!("{}/7.dat.tmp", dir)).exists());
}
// Decoding when .vif does not record the encode-time size: the layout is
// inferred from the shard size, except when that is an exact large-block
// multiple and the live extent reaches the ambiguous region.
#[test]
fn test_write_dat_file_fallback_layout() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -545,11 +559,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -567,7 +583,17 @@ mod tests {
-> io::Result<Vec<u8>> {
let out = format!("{}/{}", dir, sub);
std::fs::create_dir_all(&out).unwrap();
write_dat_file(&out, "", VolumeId(1), live, encoded, 10, shard_dirs, LARGE, SMALL)?;
write_dat_file_from_shards(&DatRebuild {
dat_dir: &out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live,
encoded_dat_file_size: encoded,
data_shards: 10,
shard_dirs: Some(shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})?;
Ok(std::fs::read(format!("{}/1.dat", out)).unwrap())
};
@@ -581,14 +607,20 @@ mod tests {
// each shard exactly one large block, indistinguishable from one large row
let (dir, shard_dirs, _) = encode("ambig1", large_row_size - 1);
let err = decode_to(&dir, "out", large_row_size / 2, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
// two-row equivalent: decoding within the agreed prefix still works
let (dir, shard_dirs, original) = encode("ambig2", 2 * large_row_size - 1);
let decoded = decode_to(&dir, "outa", large_row_size, 0, &shard_dirs).unwrap();
assert_eq!(&original[..large_row_size as usize], &decoded[..]);
let err = decode_to(&dir, "outb", large_row_size + 1, 0, &shard_dirs).unwrap_err();
assert!(err.to_string().contains("does not identify the block layout"));
assert!(
err.to_string()
.contains("does not identify the block layout")
);
}
// Decoding after deletions moved the live extent below the large-block row
@@ -597,7 +629,7 @@ mod tests {
#[test]
fn test_write_dat_file_after_tail_deletion() {
use crate::storage::erasure_coding::ec_bitrot::{
ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use reed_solomon_erasure::galois_8::ReedSolomon;
@@ -637,11 +669,13 @@ mod tests {
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
ec_encoder::EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: SMALL,
large_block_size: LARGE,
small_block_size: SMALL,
},
)
.unwrap();
for shard in &mut shards {
@@ -657,17 +691,17 @@ mod tests {
std::fs::create_dir(&out_dir).unwrap();
let out = out_dir.to_str().unwrap();
let decode = |live_size: i64, encoded_size: i64| -> Vec<u8> {
write_dat_file(
out,
"",
VolumeId(1),
live_size,
encoded_size,
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: live_size,
encoded_dat_file_size: encoded_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.unwrap();
let path = format!("{}/1.dat", out);
let decoded = std::fs::read(&path).unwrap();
@@ -702,17 +736,58 @@ mod tests {
assert_ne!(&original[..(large_row_size / 2) as usize], &control[..]);
// the live extent can never exceed the encode-time size
assert!(write_dat_file(
out,
"",
VolumeId(1),
dat_size + 1,
dat_size,
data_shards,
&shard_dirs,
LARGE,
SMALL,
)
.is_err());
assert!(
write_dat_file_from_shards(&DatRebuild {
dat_dir: out,
collection: "",
volume_id: VolumeId(1),
dat_file_size: dat_size + 1,
encoded_dat_file_size: dat_size,
data_shards,
shard_dirs: Some(&shard_dirs),
large_block_size: LARGE,
small_block_size: SMALL,
})
.is_err()
);
}
/// A journal many chunks long that repeats a few ids reads back as those
/// ids, from each dir once, whatever its length.
#[test]
fn test_read_ecj_deletions_collects_distinct_ids_across_dirs() {
let tmp = TempDir::new().unwrap();
let data = tmp.path().join("data");
let idx = tmp.path().join("idx");
let missing = tmp.path().join("missing");
std::fs::create_dir_all(&data).unwrap();
std::fs::create_dir_all(&idx).unwrap();
let (data, idx, missing) = (
data.to_str().unwrap(),
idx.to_str().unwrap(),
missing.to_str().unwrap(),
);
let entry = |id: u64| {
let mut buf = [0u8; NEEDLE_ID_SIZE];
NeedleId(id).to_bytes(&mut buf);
buf
};
// Past two load chunks of three repeating ids, then an id only in the
// last chunk and a torn trailing record.
let mut ecj = Vec::new();
while ecj.len() <= 2 * (1 << 20) {
for id in [1, 2, 3] {
ecj.extend_from_slice(&entry(id));
}
}
ecj.extend_from_slice(&entry(7));
ecj.extend_from_slice(&entry(8)[..3]);
std::fs::write(format!("{idx}/1.ecj"), &ecj).unwrap();
std::fs::write(format!("{data}/1.ecj"), entry(9)).unwrap();
let ids = read_ecj_deletions(&[data, idx, idx, missing], "", VolumeId(1)).unwrap();
let expected: HashSet<NeedleId> = [1, 2, 3, 7, 9].into_iter().map(NeedleId).collect();
assert_eq!(ids, expected);
}
}
@@ -5,16 +5,12 @@
use std::fs::File;
use std::io;
#[cfg(not(unix))]
use std::io::{Read, Seek, SeekFrom};
use reed_solomon_erasure::galois_8::ReedSolomon;
use crate::pb::volume_server_pb::{
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums,
};
use crate::pb::volume_server_pb::{ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums};
use crate::storage::erasure_coding::ec_bitrot::{
self, ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
self, DEFAULT_BITROT_BLOCK_SIZE, ShardChecksumBuilder,
};
use crate::storage::erasure_coding::ec_shard::*;
use crate::storage::idx;
@@ -50,7 +46,7 @@ pub fn write_ec_files(
let dat_size = dat_file.metadata()?.len() as i64;
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
// Create shard files
let total_shards = data_shards + parity_shards;
@@ -77,11 +73,13 @@ pub fn write_ec_files(
&rs,
&mut shards,
&mut builders,
data_shards,
parity_shards,
ENCODE_BUFFER_SIZE,
block_size as usize,
block_size as usize,
EcEncodeLayout {
data_shards,
parity_shards,
buffer_size: ENCODE_BUFFER_SIZE,
large_block_size: block_size as usize,
small_block_size: block_size as usize,
},
)?;
// Close all shards
@@ -162,7 +160,7 @@ pub fn rebuild_ec_files(
}
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
@@ -175,7 +173,7 @@ pub fn rebuild_ec_files(
let mut shard_size = 0;
for (i, shard) in shards.iter_mut().enumerate() {
if !missing_shard_ids.contains(&(i as u32)) {
if let Ok(_) = shard.open() {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
@@ -185,7 +183,7 @@ pub fn rebuild_ec_files(
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
if let Ok(_) = alt.open() {
if alt.open().is_ok() {
let size = alt.file_size();
if size > shard_size {
shard_size = size;
@@ -251,12 +249,8 @@ pub fn rebuild_ec_files(
}
// Reconstruct missing shards
rs.reconstruct(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon reconstruct: {:?}", e),
)
})?;
rs.reconstruct(&mut buffers)
.map_err(|e| io::Error::other(format!("reed-solomon reconstruct: {:?}", e)))?;
// Write recovered data into the missing shards
for i in missing_shard_ids {
@@ -284,40 +278,63 @@ pub fn rebuild_ec_files(
/// FULL walk only reads live data-shard intervals, so on its own it can't catch
/// bitrot in a parity shard or an unwalked region. Move to mode 4 (CHECKSUM) and
/// drop it from mode 2 once the `.ecsum` subsystem lands.
///
/// `dirs` is indexed BY SHARD ID: each entry is the directory holding that
/// shard, or `None` when no disk mounts it. A reconciled volume's shards can be
/// split across disks, so a single directory cannot address them all.
pub fn verify_ec_shards(
dir: &str,
dirs: &[Option<String>],
collection: &str,
volume_id: VolumeId,
data_shards: usize,
parity_shards: usize,
) -> io::Result<(Vec<u32>, Vec<String>)> {
let rs = ReedSolomon::new(data_shards, parity_shards)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
.map_err(|e| io::Error::other(format!("reed-solomon init: {:?}", e)))?;
let total_shards = data_shards + parity_shards;
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
let mut shards: Vec<Option<EcVolumeShard>> = (0..total_shards)
.map(|i| {
dirs.get(i)
.and_then(|d| d.as_ref())
.map(|d| EcVolumeShard::new(d, collection, volume_id, i as u8))
})
.collect();
let mut shard_size = 0;
let mut broken_shards = std::collections::HashSet::new();
let mut details = Vec::new();
for (i, shard) in shards.iter_mut().enumerate() {
if let Ok(_) = shard.open() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
for (i, slot) in shards.iter_mut().enumerate() {
match slot.as_mut() {
// Not a match guard: a binding is immutable until the guard ends,
// and `open()` needs `&mut self`.
Some(shard) => {
if shard.open().is_ok() {
let size = shard.file_size();
if size > shard_size {
shard_size = size;
}
} else {
broken_shards.insert(i as u32);
details.push(format!("failed to open or missing shard {}", i));
}
}
None => {
broken_shards.insert(i as u32);
details.push(format!("shard {} is not mounted on any disk", i));
}
} else {
broken_shards.insert(i as u32);
details.push(format!("failed to open or missing shard {}", i));
}
}
if shard_size == 0 || broken_shards.len() >= parity_shards {
// Can't do much if we don't know the size or have too many missing
return Ok((broken_shards.into_iter().collect(), details));
// Can't do much if we don't know the size or have too many missing.
// Sort like the normal path below: a `HashSet` iteration order would
// make this return shard ids in an arbitrary order, and enough `None`
// entries in `dirs` now reach this branch for a caller to notice.
let mut broken_vec: Vec<u32> = broken_shards.into_iter().collect();
broken_vec.sort_unstable();
return Ok((broken_vec, details));
}
let block_size = ERASURE_CODING_SMALL_BLOCK_SIZE;
@@ -331,7 +348,17 @@ pub fn verify_ec_shards(
let mut read_failed = false;
for i in 0..total_shards {
if !broken_shards.contains(&(i as u32)) {
if let Err(e) = shards[i].read_at(&mut buffers[i], offset) {
// The `None` arm is defensive and unreachable: the open loop
// put every unmounted slot in `broken_shards`, which this
// branch already skipped. Kept because the `Option` forces
// some handling here, and an error is the only shape that
// cannot quietly feed an unread buffer into the parity
// comparison below. Nothing needs to cover it.
let read = match shards[i].as_mut() {
Some(shard) => shard.read_at(&mut buffers[i], offset),
None => Err(io::Error::new(io::ErrorKind::NotFound, "shard not mounted")),
};
if let Err(e) = read {
broken_shards.insert(i as u32);
details.push(format!("read error shard {}: {}", i, e));
read_failed = true;
@@ -345,27 +372,27 @@ pub fn verify_ec_shards(
if !read_failed {
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
if !is_valid {
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
if let Ok(is_valid) = rs.verify(&slice_ptrs)
&& !is_valid
{
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
// is corrupted without recalculating parities or syndromes, so we just
// log that this batch has corruption. Wait, we can test each parity shard!
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for i in data_shards..total_shards {
verify_buffers[i].fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
let mut verify_buffers = buffers.clone();
// Clear the parity parts
for buf in &mut verify_buffers[data_shards..total_shards] {
buf.fill(0);
}
if rs.encode(&mut verify_buffers).is_ok() {
for i in 0..total_shards {
if buffers[i] != verify_buffers[i] {
broken_shards.insert(i as u32);
details.push(format!(
"parity mismatch on shard {} at offset {}",
i, offset
));
}
}
}
@@ -377,7 +404,7 @@ pub fn verify_ec_shards(
}
// Close all shards
for shard in &mut shards {
for shard in shards.iter_mut().flatten() {
shard.close();
}
@@ -398,22 +425,23 @@ pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::R
// Read all idx entries
let mut idx_file = File::open(idx_path)?;
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
let mut last: std::collections::HashMap<NeedleId, (Offset, Size)> =
std::collections::HashMap::new();
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
entries.push((key, offset, size));
last.insert(key, (offset, size));
Ok(())
})?;
// Sort by NeedleId, then by actual offset so later entries come last
entries.sort_by_key(|&(key, offset, _)| (key, offset.to_actual_offset()));
// Remove duplicates (keep last/latest entry for each key).
// dedup_by_key keeps the first in each run, so we reverse first,
// dedup, then reverse back.
entries.reverse();
entries.dedup_by_key(|entry| entry.0);
entries.reverse();
let mut entries: Vec<(NeedleId, Offset, Size)> = last
.into_iter()
.filter_map(|(key, (offset, size))| {
if size.is_deleted() || offset.is_zero() {
None
} else {
Some((key, offset, size))
}
})
.collect();
entries.sort_by_key(|&(key, _o, _s)| key);
// Write sorted entries to .ecx
let mut ecx_file = File::create(ecx_path)?;
@@ -457,7 +485,7 @@ pub fn rebuild_ecx_file(
.collect();
for (i, shard) in shards.iter_mut().enumerate() {
if let Err(_) = shard.open() {
if shard.open().is_err() {
let mut found = false;
for &other_dir in additional_dirs {
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
@@ -474,7 +502,7 @@ pub fn rebuild_ecx_file(
}
return Err(io::Error::new(
io::ErrorKind::NotFound,
format!("cannot open data shard for ecx rebuild"),
"cannot open data shard for ecx rebuild".to_string(),
));
}
}
@@ -482,7 +510,7 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
let total_data_size = shard_size * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
@@ -505,7 +533,7 @@ pub fn rebuild_ecx_file(
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
(shard_size - 1).max(0)
};
// Read version from superblock (first byte of logical data)
@@ -554,7 +582,8 @@ pub fn rebuild_ecx_file(
}
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
let needle_id = NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let needle_id =
NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
let size = Size::from_bytes(&header_buf[COOKIE_SIZE + NEEDLE_ID_SIZE..header_size]);
// Validate: stop if we hit zero cookie+id (end of data)
@@ -607,7 +636,6 @@ pub fn rebuild_ecx_file(
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
@@ -674,30 +702,50 @@ fn read_from_data_shards(
/// the uniform block is.
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// Shape of one encode run: the Reed-Solomon split and the block sizes that
/// fix where every byte of the .dat lands in the shards. Mirrors Go's
/// `ECContext`. `buffer_size` must divide both block sizes.
#[derive(Clone, Copy, Debug)]
pub(crate) struct EcEncodeLayout {
pub(crate) data_shards: usize,
pub(crate) parity_shards: usize,
/// Bytes of each shard's block handled per sub-batch; bounds memory at
/// `total_shards * buffer_size` however large the blocks are.
pub(crate) buffer_size: usize,
pub(crate) large_block_size: usize,
pub(crate) small_block_size: usize,
}
/// Encode the .dat file data into shard files.
///
/// Uses a two-phase approach matching Go's ec_encoder.go:
/// 1. Process as many large blocks as possible
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
#[allow(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
dat_size: i64,
rs: &ReedSolomon,
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
buffer_size: usize,
large_block_size: usize,
small_block_size: usize,
layout: EcEncodeLayout,
) -> io::Result<()> {
let EcEncodeLayout {
data_shards,
parity_shards,
buffer_size,
large_block_size,
small_block_size,
} = layout;
let total_shards = data_shards + parity_shards;
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
.map(|_| vec![0u8; buffer_size])
.collect();
let mut buffers: Vec<Vec<u8>> = (0..total_shards).map(|_| vec![0u8; buffer_size]).collect();
let mut run = EncodeRun {
dat_file,
rs,
buffers: &mut buffers,
shards,
builders,
data_shards,
};
let mut remaining = dat_size;
let mut offset: u64 = 0;
@@ -706,16 +754,7 @@ pub(crate) fn encode_dat_file(
let large_row_size = large_block_size * data_shards;
while remaining >= large_row_size as i64 {
encode_data(
dat_file,
offset,
large_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, large_block_size)?;
offset += large_row_size as u64;
remaining -= large_row_size as i64;
}
@@ -725,16 +764,7 @@ pub(crate) fn encode_dat_file(
while remaining > 0 {
let to_process = remaining.min(small_row_size as i64);
encode_data(
dat_file,
offset,
small_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
)?;
run.encode_row(offset, small_block_size)?;
offset += to_process as u64;
remaining -= to_process;
}
@@ -742,108 +772,66 @@ pub(crate) fn encode_dat_file(
Ok(())
}
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
/// Everything one encode run streams through: the source .dat, the codec, a
/// buffer per shard, and the per-shard file and checksum sinks.
struct EncodeRun<'a> {
dat_file: &'a File,
rs: &'a ReedSolomon,
buffers: &'a mut [Vec<u8>],
shards: &'a mut [EcVolumeShard],
builders: &'a mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
encode_one_batch(
dat_file,
row_offset + (b * buffer_size) as u64,
block_size,
rs,
buffers,
shards,
builders,
data_shards,
)?;
}
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
fn encode_one_batch(
dat_file: &File,
offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
for i in 0..data_shards {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
impl EncodeRun<'_> {
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches
/// so arbitrarily large blocks never require block-sized allocations.
/// Mirrors Go's encodeData.
fn encode_row(&mut self, row_offset: u64, block_size: usize) -> io::Result<()> {
let buffer_size = self.buffers[0].len();
if !block_size.is_multiple_of(buffer_size) {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
)
})?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
for (i, buf) in buffers.iter().enumerate() {
shards[i].write_all(buf)?;
builders[i].write(buf);
}
Ok(())
}
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let mut n = 0;
while n < buf.len() {
#[cfg(unix)]
let r = {
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buf[n..], offset + n as u64)?
};
#[cfg(not(unix))]
let r = {
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(offset + n as u64))?;
f.read(&mut buf[n..])?
};
if r == 0 {
break;
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
self.encode_one_batch(row_offset + (b * buffer_size) as u64, block_size)?;
}
n += r;
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block
/// in this row. Mirrors Go's encodeDataOneBatch.
fn encode_one_batch(&mut self, offset: u64, block_size: usize) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the
// buffers are reused across batches, so the tail must be cleared
// explicitly.
for (i, buf) in self.buffers[..self.data_shards].iter_mut().enumerate() {
let read_offset = offset + (i * block_size) as u64;
let n = crate::storage::io::read_full_at(self.dat_file, buf, read_offset)?;
buf[n..].fill(0);
}
// Encode parity shards
self.rs
.encode(&mut *self.buffers)
.map_err(|e| io::Error::other(format!("reed-solomon encode: {:?}", e)))?;
// Write all shard buffers to files and feed the same bytes to each
// shard's bitrot checksum builder, keeping covered_size == on-disk
// length.
for (i, buf) in self.buffers.iter().enumerate() {
self.shards[i].write_all(buf)?;
self.builders[i].write(buf);
}
Ok(())
}
Ok(n)
}
#[cfg(test)]
@@ -851,7 +839,7 @@ mod tests {
use super::*;
use crate::storage::needle::needle::Needle;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use tempfile::TempDir;
#[test]
@@ -863,13 +851,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -914,13 +898,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=n {
@@ -993,13 +973,9 @@ mod tests {
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=20 {
@@ -1031,7 +1007,10 @@ mod tests {
let victim = format!("{}/1.ec03", dir);
let full = std::fs::metadata(&victim).unwrap().len();
assert!(full > 0, "encoded shard should be non-empty");
let f = std::fs::OpenOptions::new().write(true).open(&victim).unwrap();
let f = std::fs::OpenOptions::new()
.write(true)
.open(&victim)
.unwrap();
f.set_len(full / 2).unwrap();
drop(f);
@@ -1185,19 +1164,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_uniform_layout() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1227,7 +1202,10 @@ mod tests {
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
let rebuilt = std::fs::read(&ecx_path).unwrap();
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
assert_eq!(
canonical, rebuilt,
"rebuilt .ecx must match the encode-time .ecx"
);
}
// A truncated data shard must FAIL the .ecx rebuild, not publish the
@@ -1235,19 +1213,15 @@ mod tests {
#[test]
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(3),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for i in 1u64..=12 {
@@ -1345,13 +1319,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1429,13 +1399,9 @@ mod tests {
let mut v = Volume::new(
dat_dir,
idx_dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -1457,4 +1423,189 @@ mod tests {
"should fail when idx_dir doesn't contain .idx"
);
}
/// Write a real 10+4 encoded volume into `dir`.
///
/// Unlike `make_volume_with_needles` and `encode_sample_volume` this seeds
/// a caller-chosen directory, which is what a split-disk test needs: the
/// shards have to be scattered out of the directory they were encoded into.
fn seed_encoded_volume(dir: &str, vid: VolumeId) {
let mut v = Volume::new(
dir,
dir,
vid,
NeedleMapKind::InMemory,
&VolumeSpec::default(),
)
.unwrap();
for i in 1..=8 {
let data = format!("test data for needle {} with a bit more length", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
write_ec_files(dir, dir, "", vid, 10, 4).unwrap();
}
/// Shards split across two directories must all be found. Passing one dir
/// per shard is what lets a reconciled volume's parity be checked at all.
#[test]
fn test_verify_ec_shards_reads_shards_from_multiple_dirs() {
let tmp = TempDir::new().unwrap();
let src = tmp.path().join("src");
let d0 = tmp.path().join("d0");
let d1 = tmp.path().join("d1");
for d in [&src, &d0, &d1] {
std::fs::create_dir_all(d).unwrap();
}
let src_s = src.to_str().unwrap();
seed_encoded_volume(src_s, VolumeId(1));
// Move shards 0..=6 to d0 and 7..=13 to d1.
let mut dirs: Vec<Option<String>> = Vec::new();
for id in 0..14u8 {
let target = if id < 7 { &d0 } else { &d1 };
std::fs::rename(
format!("{}/1.ec{:02}", src_s, id),
format!("{}/1.ec{:02}", target.to_str().unwrap(), id),
)
.unwrap();
dirs.push(Some(target.to_str().unwrap().to_string()));
}
let (broken, details) = verify_ec_shards(&dirs, "", VolumeId(1), 10, 4).unwrap();
assert!(
broken.is_empty(),
"split-dir shards reported broken: {:?}",
details
);
}
/// A shard no disk holds is a missing shard, not a panic and not a silent
/// pass: it is REPORTED, by id, with a message that distinguishes "no disk
/// holds this shard" from "the disk holds it but it won't open".
///
/// Read the scope literally. This does NOT show that the mounted shards
/// verify clean. `dirs[5] = None` puts shard 5 in `broken_shards` before
/// the block loop starts, so every iteration takes the
/// `else { read_failed = true; }` arm and the Reed-Solomon comparison never
/// runs at all. `broken == vec![5]` therefore holds because the other 13
/// were never verified, not because they verified clean -- a parity check
/// over intact shards is what
/// `test_verify_ec_shards_reads_shards_from_multiple_dirs` and the
/// end-to-end split-disk FULL scrub establish.
#[test]
fn test_verify_ec_shards_treats_a_none_dir_as_missing() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
seed_encoded_volume(dir, VolumeId(1));
let mut dirs: Vec<Option<String>> = (0..14).map(|_| Some(dir.to_string())).collect();
dirs[5] = None;
let (broken, details) = verify_ec_shards(&dirs, "", VolumeId(1), 10, 4).unwrap();
assert_eq!(
broken,
vec![5],
"an unmounted shard must be reported, and only it: {:?}",
details
);
// "no disk holds this shard" and "the disk holds it but it won't open"
// are different operator problems, which is why they carry different
// messages. Asserting only the id would let one masquerade as the other.
assert!(
details.iter().any(|d| d.contains("not mounted")),
"an unmounted shard must be distinguished from an unopenable one, got {:?}",
details
);
}
#[test]
fn test_encode_drops_tombstone_last_wins() {
use crate::storage::idx;
use crate::storage::types::{NeedleId, Offset, Size, TOMBSTONE_FILE_SIZE};
let tmp = tempfile::TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let idx_path = format!("{}/t.idx", dir);
let ecx_path = format!("{}/t.ecx", dir);
let key = NeedleId(12345);
{
let mut f = std::fs::File::create(&idx_path).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
}
super::write_sorted_ecx_from_idx(&idx_path, &ecx_path).unwrap();
let mut found = false;
{
let mut f = std::fs::File::open(&ecx_path).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found = true;
}
Ok(())
})
.unwrap();
}
assert!(!found, "tombstoned key must not appear in .ecx");
let idx2 = format!("{}/t2.idx", dir);
let ecx2 = format!("{}/t2.ecx", dir);
{
let mut f = std::fs::File::create(&idx2).unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), TOMBSTONE_FILE_SIZE).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(2048), Size(200))
.unwrap();
}
super::write_sorted_ecx_from_idx(&idx2, &ecx2).unwrap();
let mut found2 = false;
{
let mut f = std::fs::File::open(&ecx2).unwrap();
idx::walk_index_file(&mut f, 0, |k, o, s| {
if k == key {
found2 = true;
assert_eq!(o.to_actual_offset(), 2048);
assert_eq!(s, Size(200));
}
Ok(())
})
.unwrap();
}
assert!(found2, "re-created key must appear live");
// Zero offset with non-negative size is also a deletion: Go
// readNeedleMap (`if !offset.IsZero() && !size.IsDeleted() { Set }
// else { Delete }`) and CompactNeedleMap::load_from_idx both treat
// it as deleted. Encode must drop it too, or the .ecx live-map
// mismatches replay.
let idx3 = format!("{}/t3.idx", dir);
let ecx3 = format!("{}/t3.ecx", dir);
{
let mut f = std::fs::File::create(&idx3).unwrap();
idx::write_index_entry(&mut f, key, Offset::from_actual_offset(1024), Size(100))
.unwrap();
idx::write_index_entry(&mut f, key, Offset::default(), Size(0)).unwrap();
}
super::write_sorted_ecx_from_idx(&idx3, &ecx3).unwrap();
let mut found3 = false;
{
let mut f = std::fs::File::open(&ecx3).unwrap();
idx::walk_index_file(&mut f, 0, |k, _o, _s| {
if k == key {
found3 = true;
}
Ok(())
})
.unwrap();
}
assert!(
!found3,
"zero-offset row must not appear in .ecx even with non-negative size"
);
}
}
@@ -16,6 +16,20 @@ pub const ERASURE_CODING_SMALL_BLOCK_SIZE: usize = 1024 * 1024; // 1MB
pub type ShardId = u8;
/// Validate a wire shard id. `ShardId` is `u8` but only 0..MAX_SHARD_COUNT are valid.
/// Rejects 256 (would truncate to 0 and delete .ec00) and 270 (would alias 14).
pub fn shard_id_try_from(v: u32) -> Result<ShardId, String> {
if v < MAX_SHARD_COUNT as u32 {
Ok(v as ShardId)
} else {
Err(format!(
"invalid shard id {} (max {})",
v,
MAX_SHARD_COUNT - 1
))
}
}
/// A single erasure-coded shard file.
pub struct EcVolumeShard {
pub volume_id: VolumeId,
@@ -73,28 +87,14 @@ impl EcVolumeShard {
Ok(())
}
/// Read data at a specific offset.
/// Read data at a specific offset, filling `buf` unless the shard ends first.
pub fn read_at(&self, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let file = self
.ecd_file
.as_ref()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(not(unix))]
{
use std::io::{Read, Seek, SeekFrom};
// File::read_at is unix-only; fall back to seek + read.
// We need a mutable reference for seek/read, so clone the handle.
let mut f = file.try_clone()?;
f.seek(SeekFrom::Start(offset))?;
f.read(buf)
}
crate::storage::io::read_full_at(file, buf, offset)
}
/// Write data to the shard file (appends).
@@ -102,7 +102,7 @@ impl EcVolumeShard {
let file = self
.ecd_file
.as_mut()
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "shard file not open"))?;
.ok_or_else(|| io::Error::other("shard file not open"))?;
file.write_all(data)?;
self.ecd_file_size += data.len() as i64;
Ok(())
@@ -112,6 +112,21 @@ impl EcVolumeShard {
self.ecd_file_size
}
/// A duplicate of the mounted shard handle, for a reader that has to
/// outlive the store guard.
///
/// This is the same descriptor `read_at` serves from, so it carries the
/// `O_NOATIME` from `open_volume_file` and keeps pointing at the shard
/// that was mounted, whatever later happens to the path. `dup` shares the
/// kernel file offset, which is why every read through it must be
/// positional (`read_at`), never seek-based.
pub fn try_clone_file(&self) -> io::Result<File> {
self.ecd_file
.as_ref()
.ok_or_else(|| io::Error::other("shard file not open"))?
.try_clone()
}
/// Protobuf descriptor for this shard. Mirrors Go's ToEcShardInfo.
pub fn to_ec_shard_info(&self) -> crate::pb::volume_server_pb::EcShardInfo {
crate::pb::volume_server_pb::EcShardInfo {
@@ -180,6 +195,104 @@ impl ShardBits {
}
}
/// Parses the generation of a 2PC-staged `<base>.v<N>` file: `None` means the
/// name is not a generation file of `base`.
pub fn ec_file_generation(name: &str, base: &str) -> Option<u32> {
let suffix = name.strip_prefix(&format!("{}.v", base))?;
match suffix.parse::<u32>() {
Ok(g) if g > 0 => Some(g),
_ => None,
}
}
/// Removes 2PC generation files staged under `base`:
/// `<base>.ecNN.v<N>`, `<base>.ecx.v<N>`, `<base>.ecj.v<N>`, `<base>.ecsum.v<N>`
/// and `<base>.vif.v<N>`. `generations_older_than == 0` removes every
/// generation; otherwise only generations strictly below it. Returns the
/// first real removal failure. Mirrors Go's `RemoveEcGenerationFiles`.
pub fn remove_ec_generation_files(base: &str, generations_older_than: u32) -> io::Result<()> {
let path = std::path::Path::new(base);
let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) else {
return Ok(());
};
let ec_prefix = format!("{}.ec", fname.to_string_lossy());
let vif_name = format!("{}.vif", fname.to_string_lossy());
let mut first_err: Option<io::Error> = None;
let mut record = |res: io::Result<()>| {
if let Err(e) = res
&& first_err.is_none()
{
first_err = Some(e);
}
};
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries {
let entry = match entry {
Ok(entry) => entry,
Err(e) => {
// A skipped entry means an incomplete sweep; report it
// instead of pretending the cleanup finished.
record(Err(e));
continue;
}
};
let name = entry.file_name().to_string_lossy().into_owned();
let Some((artifact, _)) = name.rsplit_once(".v") else {
continue;
};
if artifact != vif_name && !artifact.starts_with(&ec_prefix) {
continue;
}
let Some(generation) = ec_file_generation(&name, artifact) else {
continue;
};
if generations_older_than > 0 && generation >= generations_older_than {
continue;
}
record(match fs::remove_file(entry.path()) {
Err(e) if e.kind() != io::ErrorKind::NotFound => Err(e),
_ => Ok(()),
});
}
}
Err(e) if e.kind() != io::ErrorKind::NotFound => record(Err(e)),
Err(_) => {}
}
match first_err {
Some(e) => Err(e),
None => Ok(()),
}
}
/// Removes every staged generation `<shard_file>.v<N>` of one shard file.
/// Returns true when at least one generation file was removed.
pub fn remove_ec_shard_generations(shard_file: &str) -> io::Result<bool> {
let path = std::path::Path::new(shard_file);
let (Some(parent), Some(fname)) = (path.parent(), path.file_name()) else {
return Ok(false);
};
let fname = fname.to_string_lossy().into_owned();
let mut removed = false;
match fs::read_dir(parent) {
Ok(entries) => {
for entry in entries {
let entry = entry?;
let name = entry.file_name().to_string_lossy().into_owned();
if ec_file_generation(&name, &fname).is_some() {
match fs::remove_file(entry.path()) {
Err(e) if e.kind() != io::ErrorKind::NotFound => return Err(e),
_ => removed = true,
}
}
}
}
Err(e) if e.kind() != io::ErrorKind::NotFound => return Err(e),
Err(_) => {}
}
Ok(removed)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -236,4 +349,42 @@ mod tests {
let shard = EcVolumeShard::new("/data", "", VolumeId(7), 13);
assert_eq!(shard.file_name(), "/data/7.ec13");
}
#[test]
fn test_shard_id_try_from_u32_rejects_overflow() {
use super::{MAX_SHARD_COUNT, shard_id_try_from};
assert_eq!(shard_id_try_from(0).unwrap(), 0u8);
assert_eq!(shard_id_try_from(14).unwrap(), 14u8);
assert_eq!(shard_id_try_from(31).unwrap(), 31u8);
assert!(shard_id_try_from(32).is_err());
assert!(shard_id_try_from(256).is_err());
assert!(shard_id_try_from(270).is_err());
assert!(shard_id_try_from(u32::MAX).is_err());
assert_eq!(MAX_SHARD_COUNT, 32);
}
#[test]
fn test_shard_batch_validation_is_atomic_rejects_without_partial_prefix() {
use super::shard_id_try_from;
// The mount/unmount handlers pre-validate the ENTIRE req.shard_ids into
// a Vec<ShardId> BEFORE acquiring the write lock or mutating any EC
// state. This test pins the validation half of that contract at the
// unit level: a batch like [0, 32] must fail as a whole, so by
// construction no validated prefix (e.g. shard 0) is ever applied.
// The handler-level tests below assert the no-state-change half.
let batch = vec![0u32, 32u32];
let validated: Result<Vec<_>, _> =
batch.iter().map(|&sid| shard_id_try_from(sid)).collect();
assert!(
validated.is_err(),
"batch {:?} must be rejected as a whole",
batch
);
// A fully-valid batch still validates cleanly.
let ok: Result<Vec<_>, _> = [0u32, 1u32, 13u32]
.iter()
.map(|&sid| shard_id_try_from(sid))
.collect();
assert_eq!(ok.unwrap(), vec![0u8, 1u8, 13u8]);
}
}
File diff suppressed because it is too large Load Diff
@@ -11,7 +11,7 @@ pub mod ec_shard;
pub mod ec_volume;
pub use ec_shard::{
EcVolumeShard, ShardId, DATA_SHARDS_COUNT, MAX_SHARD_COUNT, MIN_TOTAL_DISKS,
PARITY_SHARDS_COUNT, TOTAL_SHARDS_COUNT,
DATA_SHARDS_COUNT, EcVolumeShard, MAX_SHARD_COUNT, MIN_TOTAL_DISKS, PARITY_SHARDS_COUNT,
ShardId, TOTAL_SHARDS_COUNT,
};
pub use ec_volume::EcVolume;
+136 -8
View File
@@ -21,12 +21,26 @@ where
let mut buf = vec![0u8; NEEDLE_MAP_ENTRY_SIZE * ROWS_TO_READ];
loop {
let count = match reader.read(&mut buf) {
Ok(0) => return Ok(()),
Ok(n) => n,
Err(ref e) if e.kind() == io::ErrorKind::UnexpectedEof => return Ok(()),
Err(e) => return Err(e),
};
// Fill the batch before decoding: `read` may return a count that is
// not a multiple of the entry size, and a split entry would misalign
// every later row. Go is immune: `ReadAt` fills or errors.
let mut count = 0;
let mut eof = false;
while count < buf.len() {
match reader.read(&mut buf[count..]) {
Ok(0) => {
eof = true;
break;
}
Ok(n) => count += n,
Err(ref e) if e.kind() == io::ErrorKind::Interrupted => continue,
Err(ref e) if e.kind() == io::ErrorKind::UnexpectedEof => {
eof = true;
break;
}
Err(e) => return Err(e),
}
}
let mut i = 0;
while i + NEEDLE_MAP_ENTRY_SIZE <= count {
@@ -34,6 +48,11 @@ where
f(key, offset, size)?;
i += NEEDLE_MAP_ENTRY_SIZE;
}
// A trailing partial entry at EOF is ignored, as Go does on `io.EOF`.
if eof {
return Ok(());
}
}
}
@@ -57,7 +76,7 @@ pub fn check_index_file<R: Read + Seek>(
errs.push(format!("walk index file: {}", e));
}
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3 .0.cmp(&b.3 .0)));
entries.sort_by(|a, b| a.2.cmp(&b.2).then(a.3.0.cmp(&b.3.0)));
// Offset-0 logical tombstones (remote-tier deletes) occupy no physical extent,
// so they cannot overlap anything — exclude them from the overlap check. They
@@ -177,6 +196,111 @@ mod tests {
data
}
/// Reader that hands back at most `chunk` bytes per `read`. 7 is coprime
/// with the 17-byte entry size, so nearly every read ends mid-entry. With
/// `interrupts`, every other call fails with `ErrorKind::Interrupted`.
struct ShortReader {
inner: Cursor<Vec<u8>>,
chunk: usize,
interrupts: bool,
interrupt_next: bool,
}
impl ShortReader {
fn new(data: Vec<u8>, interrupts: bool) -> Self {
ShortReader {
inner: Cursor::new(data),
chunk: 7,
interrupts,
interrupt_next: false,
}
}
}
impl Read for ShortReader {
fn read(&mut self, buf: &mut [u8]) -> io::Result<usize> {
if self.interrupt_next {
self.interrupt_next = false;
return Err(io::Error::from(io::ErrorKind::Interrupted));
}
self.interrupt_next = self.interrupts;
let n = buf.len().min(self.chunk);
self.inner.read(&mut buf[..n])
}
}
impl Seek for ShortReader {
fn seek(&mut self, pos: SeekFrom) -> io::Result<u64> {
self.inner.seek(pos)
}
}
fn walk_all<R: Read + Seek>(reader: &mut R, start_from: u64) -> Vec<(NeedleId, i64, Size)> {
let mut collected = Vec::new();
walk_index_file(reader, start_from, |key, offset, size| {
collected.push((key, offset.to_actual_offset(), size));
Ok(())
})
.unwrap();
collected
}
/// More than one ROWS_TO_READ batch, so the walk crosses a buffer refill.
fn many_entries() -> Vec<(NeedleId, Offset, Size)> {
(0..(ROWS_TO_READ as u64 * 2 + 37))
.map(|i| {
(
NeedleId(i * 7 + 1),
Offset::from_actual_offset(i as i64 * 128),
Size(i as i32 + 1),
)
})
.collect()
}
#[test]
fn test_walk_index_file_short_reads_keep_alignment() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
assert_eq!(expected.len(), ROWS_TO_READ * 2 + 37);
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_walk_index_file_retries_interrupted_reads() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
let mut short = ShortReader::new(data, true);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_walk_index_file_short_reads_start_from() {
let data = idx_bytes(&many_entries());
let expected = walk_all(&mut Cursor::new(data.clone()), 0);
let start = ROWS_TO_READ as u64 + 5;
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, start), expected[start as usize..]);
}
#[test]
fn test_walk_index_file_ignores_trailing_partial_entry() {
// A torn final entry is dropped without an error, as Go does on io.EOF.
let entries = many_entries();
let mut data = idx_bytes(&entries);
data.extend_from_slice(&[0xAB; NEEDLE_MAP_ENTRY_SIZE - 1]);
let expected = walk_all(&mut Cursor::new(idx_bytes(&entries)), 0);
assert_eq!(walk_all(&mut Cursor::new(data.clone()), 0), expected);
let mut short = ShortReader::new(data, false);
assert_eq!(walk_all(&mut short, 0), expected);
}
#[test]
fn test_check_index_file_clean() {
let data = idx_bytes(&[
@@ -213,7 +337,11 @@ mod tests {
let size = data.len() as i64;
let (count, errs) = check_index_file(&mut Cursor::new(data), size, Version(3));
assert_eq!(count, 2, "tombstone row is still counted: {:?}", errs);
assert!(errs.is_empty(), "offset-0 tombstone must not overlap: {:?}", errs);
assert!(
errs.is_empty(),
"offset-0 tombstone must not overlap: {:?}",
errs
);
}
#[test]
+208
View File
@@ -0,0 +1,208 @@
//! Positional file reads.
//!
//! Every read here is "these bytes at this offset", never "the next bytes".
//! The handles are shared — `.dat` and `.idx` descriptors are borrowed from
//! [`file_pool`](super::needle_map::file_pool), a mounted EC shard's handle is
//! duplicated into a scrub plan — so no caller may rely on a file position.
//!
//! On unix that is `pread(2)` through `std::os::unix::fs::FileExt`. On Windows
//! it is `seek_read`, which passes the offset through `OVERLAPPED`, so the read
//! itself is independent of the current cursor.
//!
//! What these helpers replace is `try_clone()` + `seek()` + `read()`. A
//! duplicated handle shares one kernel file offset with the original, so that
//! sequence is two syscalls against state another thread can move in between:
//! the seek positions the offset, a concurrent reader or an append moves it,
//! and the read returns bytes from somewhere else entirely. `seek_read` carries
//! its own offset in a single call, so there is no window.
//!
//! `seek_read` does still advance the cursor as a side effect — Windows updates
//! the file pointer even for an `OVERLAPPED` read — which nothing here relies
//! on. A caller that genuinely needs a private position must open the file
//! again rather than duplicate a handle; see `Volume::dat_scan_plan` in
//! [`storage::volume`](super::volume).
use std::fs::File;
use std::io;
/// Reads exactly `buf.len()` bytes from `file` starting at `offset`.
///
/// Fails with [`io::ErrorKind::UnexpectedEof`] if the file ends first.
pub(crate) fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)?;
}
#[cfg(windows)]
{
if read_full_at(file, buf, offset)? < buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
Ok(())
}
/// Reads up to `buf.len()` bytes from `file` starting at `offset`, returning
/// how many were read.
///
/// A short read — including `0` at or past end of file — is not an error; use
/// [`read_exact_at`] when the whole buffer must be filled.
pub(crate) fn read_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
file.seek_read(buf, offset)
}
#[cfg(not(any(unix, windows)))]
{
compile_error!("Platform not supported: only unix and windows are supported");
}
}
/// Reads into `buf` at `offset` until it is full or the file ends, retrying
/// interrupted reads; returns how many bytes were read.
///
/// Unlike [`read_at`], a count below `buf.len()` always means end of file.
pub(crate) fn read_full_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
fill_at(|b, at| read_at(file, b, at), buf, offset)
}
fn fill_at(
mut read: impl FnMut(&mut [u8], u64) -> io::Result<usize>,
buf: &mut [u8],
offset: u64,
) -> io::Result<usize> {
let mut filled = 0;
while filled < buf.len() {
match read(&mut buf[filled..], offset + filled as u64) {
Ok(0) => break,
Ok(n) => filled += n,
Err(err) if err.kind() == io::ErrorKind::Interrupted => {}
Err(err) => return Err(err),
}
}
Ok(filled)
}
#[cfg(test)]
mod tests {
use super::{fill_at, read_at, read_exact_at, read_full_at};
use std::io::{ErrorKind, Write};
fn temp_file(bytes: &[u8]) -> tempfile::NamedTempFile {
let mut f = tempfile::NamedTempFile::new().expect("temp file");
f.write_all(bytes).expect("write");
f.flush().expect("flush");
f
}
#[test]
fn read_exact_at_fills_the_whole_buffer() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 10];
read_exact_at(f.as_file(), &mut buf, 0).expect("read");
assert_eq!(&buf, b"0123456789");
}
#[test]
fn read_exact_at_reads_from_the_offset() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 4];
read_exact_at(f.as_file(), &mut buf, 3).expect("read");
assert_eq!(&buf, b"3456");
// The helper is positional: a second read at a lower offset sees the
// bytes at that offset, not wherever the first read left a cursor.
let mut again = [0u8; 4];
read_exact_at(f.as_file(), &mut again, 1).expect("read");
assert_eq!(&again, b"1234");
}
#[test]
fn read_exact_at_short_file_is_unexpected_eof() {
let f = temp_file(b"0123");
let mut buf = [0u8; 8];
let err = read_exact_at(f.as_file(), &mut buf, 0).expect_err("short file");
assert_eq!(err.kind(), ErrorKind::UnexpectedEof);
}
#[test]
fn read_at_allows_a_short_read_at_eof() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 8];
let n = read_at(f.as_file(), &mut buf, 6).expect("read");
assert_eq!(n, 4);
assert_eq!(&buf[..n], b"6789");
// Entirely past the end is zero bytes, not an error.
let n = read_at(f.as_file(), &mut buf, 10).expect("read");
assert_eq!(n, 0);
}
/// A source that returns at most `chunk` bytes per call and fails with
/// `Interrupted` on its first call, like a network mount under a signal.
fn chunked(src: &[u8], chunk: usize) -> impl FnMut(&mut [u8], u64) -> std::io::Result<usize> {
let mut interrupted = false;
move |buf, at| {
if !interrupted {
interrupted = true;
return Err(ErrorKind::Interrupted.into());
}
let at = (at as usize).min(src.len());
let n = buf.len().min(chunk).min(src.len() - at);
buf[..n].copy_from_slice(&src[at..at + n]);
Ok(n)
}
}
#[test]
fn fill_at_fills_across_short_and_interrupted_reads() {
let src: Vec<u8> = (0..=255).collect();
let mut buf = [0u8; 100];
let n = fill_at(chunked(&src, 7), &mut buf, 50).expect("read");
assert_eq!(n, buf.len());
assert_eq!(&buf[..], &src[50..150]);
}
#[test]
fn fill_at_stops_at_end_of_source() {
let src: Vec<u8> = (0..=255).collect();
let mut buf = [0u8; 100];
let n = fill_at(chunked(&src, 7), &mut buf, 200).expect("read");
assert_eq!(n, 56);
assert_eq!(&buf[..n], &src[200..]);
}
#[test]
fn fill_at_propagates_other_errors() {
let mut buf = [0u8; 8];
let err = fill_at(|_, _| Err(ErrorKind::PermissionDenied.into()), &mut buf, 0)
.expect_err("error");
assert_eq!(err.kind(), ErrorKind::PermissionDenied);
}
#[test]
fn read_full_at_returns_the_short_count_only_at_eof() {
let f = temp_file(b"0123456789");
let mut buf = [0u8; 8];
assert_eq!(read_full_at(f.as_file(), &mut buf, 0).expect("read"), 8);
assert_eq!(&buf, b"01234567");
assert_eq!(read_full_at(f.as_file(), &mut buf, 6).expect("read"), 4);
assert_eq!(&buf[..4], b"6789");
assert_eq!(read_full_at(f.as_file(), &mut buf, 10).expect("read"), 0);
}
}
+354
View File
@@ -0,0 +1,354 @@
//! Consecutive storage-media error tracking shared by `Volume` and
//! `EcVolume`. Mirrors Go's `weed/storage/io_error.go`.
use std::io;
use std::sync::Mutex;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
/// Consecutive storage-media errors allowed before the volume is quarantined.
pub(crate) const IO_ERROR_TOLERANCE: i32 = 3;
/// Returns true for I/O errors that indicate faulty storage media, not
/// transient/network failures. On Unix this is EIO; on Windows it covers
/// ERROR_CRC and ERROR_IO_DEVICE, which the kernel returns for failing disks.
pub(crate) fn is_storage_io_error(e: &io::Error) -> bool {
#[cfg(unix)]
{
e.raw_os_error() == Some(libc::EIO)
}
#[cfg(windows)]
{
const ERROR_CRC: i32 = 23;
const ERROR_IO_DEVICE: i32 = 1117;
return e.raw_os_error() == Some(ERROR_CRC) || e.raw_os_error() == Some(ERROR_IO_DEVICE);
}
#[cfg(not(any(unix, windows)))]
{
false
}
}
/// Consecutive storage-media error state for one volume. `quarantined` is
/// sticky: once set it survives later successful I/O and is lifted only by
/// `reset_io_error_state`.
#[derive(Default)]
pub(crate) struct IoErrorTracker {
last: Mutex<Option<String>>,
/// The consecutive error count in the low 32 bits and, in the high 32,
/// how many times it has been cleared. They share one word so that
/// `record_success_at` updates both in one step: reads record their
/// outcomes here without the volume's write lock.
streak: AtomicU64,
quarantined: AtomicBool,
}
const STREAK_COUNT_BITS: u64 = 0xffff_ffff;
fn streak_count(streak: u64) -> i32 {
(streak & STREAK_COUNT_BITS) as i32
}
/// `streak` with its count cleared and one more clear on record.
fn streak_cleared(streak: u64) -> u64 {
(streak >> 32).wrapping_add(1) << 32
}
/// A point in the error streak, taken where a write landed whose success
/// is only recorded later. See `IoErrorTracker::record_success_at`.
#[derive(Clone, Copy)]
pub(crate) struct StreakMark(u64);
impl IoErrorTracker {
/// `Some(e)` records a failure, `None` a success. Only storage-media
/// failures count; every other outcome clears the count and last error.
pub(crate) fn check_read_write_error(&self, err: Option<&io::Error>) {
if let Some(e) = err
&& is_storage_io_error(e)
{
self.streak.fetch_add(1, Ordering::Relaxed);
if let Ok(mut guard) = self.last.lock() {
*guard = Some(e.to_string());
}
crate::metrics::STORAGE_IO_ERROR_COUNTER.inc();
return;
}
self.clear_count();
self.clear_last();
}
fn clear_count(&self) {
self.update_streak(|streak| Some(streak_cleared(streak)));
}
fn clear_last(&self) {
if let Ok(mut guard) = self.last.lock()
&& guard.is_some()
{
*guard = None;
}
}
/// Apply `f` to the streak atomically; `None` leaves it as it is.
/// Returns the streak `f` produced, if any.
fn update_streak(&self, mut f: impl FnMut(u64) -> Option<u64>) -> Option<u64> {
let mut updated = None;
let _ = self
.streak
.fetch_update(Ordering::Relaxed, Ordering::Relaxed, |streak| {
updated = f(streak);
updated
});
updated
}
pub(crate) fn mark(&self) -> StreakMark {
StreakMark(self.streak.load(Ordering::Relaxed))
}
/// Record a success as if it had come at `mark`: the errors counted
/// before the mark are cleared and the ones counted since still stand,
/// as they would had each outcome been recorded in order. A streak
/// cleared since the mark is left as it is.
pub(crate) fn record_success_at(&self, mark: StreakMark) {
let before = streak_count(mark.0);
let updated = self.update_streak(|streak| {
if streak >> 32 != mark.0 >> 32 {
return None;
}
if streak_count(streak) <= before {
return Some(streak_cleared(streak));
}
Some(streak - before as u64)
});
// The last error stays when one counted since the mark is left.
if updated.is_some_and(|streak| streak_count(streak) == 0) {
self.clear_last();
}
}
/// The last recorded error, the consecutive count, and the quarantine flag.
pub(crate) fn get_io_error_state(&self) -> (Option<String>, i32, bool) {
let err = self.last.lock().ok().and_then(|g| g.clone());
let count = self.count();
let quarantined = self.quarantined.load(Ordering::Relaxed);
(err, count, quarantined)
}
fn count(&self) -> i32 {
streak_count(self.streak.load(Ordering::Relaxed))
}
pub(crate) fn should_quarantine(&self) -> bool {
self.quarantined.load(Ordering::Relaxed) || self.count() >= IO_ERROR_TOLERANCE
}
pub(crate) fn mark_io_quarantined(&self) {
self.quarantined.store(true, Ordering::Relaxed);
}
pub(crate) fn reset_io_error_state(&self) {
self.clear_count();
self.quarantined.store(false, Ordering::Relaxed);
if let Ok(mut guard) = self.last.lock() {
*guard = None;
}
}
#[cfg(test)]
pub(crate) fn set_last_io_error_for_test(&self, err: Option<&str>) {
if let Ok(mut guard) = self.last.lock() {
*guard = err.map(|value| value.to_string());
}
if err.is_some() {
self.update_streak(|streak| {
Some((streak & !STREAK_COUNT_BITS) | IO_ERROR_TOLERANCE as u64)
});
} else {
self.clear_count();
}
}
}
// The tracker only reacts to errors `is_storage_io_error` recognises, which is
// nothing at all on a platform that is neither Unix nor Windows.
#[cfg(all(test, any(unix, windows)))]
mod tests {
use super::*;
/// An OS error the platform reports for failing storage media.
#[cfg(unix)]
fn media_error() -> io::Error {
io::Error::from_raw_os_error(libc::EIO)
}
/// An OS error the platform reports for failing storage media.
#[cfg(windows)]
fn media_error() -> io::Error {
const ERROR_IO_DEVICE: i32 = 1117;
io::Error::from_raw_os_error(ERROR_IO_DEVICE)
}
#[test]
fn check_read_write_error_counts_consecutive_media_errors() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let (last, count, quarantined) = tracker.get_io_error_state();
assert_eq!(last, Some(media_error().to_string()));
assert_eq!(count, 2);
assert!(!quarantined);
}
#[test]
fn success_clears_the_count_and_the_last_error() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn non_media_error_clears_the_count() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&io::Error::new(
io::ErrorKind::NotFound,
"no such file",
)));
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn should_quarantine_only_once_the_tolerance_is_reached() {
let tracker = IoErrorTracker::default();
for _ in 1..IO_ERROR_TOLERANCE {
tracker.check_read_write_error(Some(&media_error()));
assert!(!tracker.should_quarantine());
}
tracker.check_read_write_error(Some(&media_error()));
assert!(tracker.should_quarantine());
}
#[test]
fn quarantine_survives_later_successful_io() {
let tracker = IoErrorTracker::default();
tracker.mark_io_quarantined();
tracker.check_read_write_error(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, true));
assert!(tracker.should_quarantine());
}
#[test]
fn success_at_a_mark_keeps_only_the_errors_after_it() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.check_read_write_error(Some(&media_error()));
tracker.record_success_at(mark);
assert_eq!(
tracker.get_io_error_state(),
(Some(media_error().to_string()), 1, false)
);
}
#[test]
fn success_at_a_mark_with_nothing_after_it_clears_the_streak() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.record_success_at(mark);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
#[test]
fn success_at_a_mark_leaves_a_streak_cleared_since() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
tracker.check_read_write_error(None);
tracker.check_read_write_error(Some(&media_error()));
tracker.record_success_at(mark);
assert_eq!(
tracker.get_io_error_state(),
(Some(media_error().to_string()), 1, false)
);
}
/// Reads update the tracker without the volume's write lock, so a
/// success replayed at a mark must not lose the errors they record
/// while it runs.
#[test]
fn success_at_a_mark_keeps_concurrent_errors() {
use std::sync::{Arc, Barrier};
const READERS: i32 = 4;
const ERRORS: i32 = 200;
for _ in 0..500 {
let tracker = Arc::new(IoErrorTracker::default());
tracker.check_read_write_error(Some(&media_error()));
tracker.check_read_write_error(Some(&media_error()));
let mark = tracker.mark();
let start = Arc::new(Barrier::new(READERS as usize + 1));
let readers: Vec<_> = (0..READERS)
.map(|_| {
let (tracker, start) = (tracker.clone(), start.clone());
std::thread::spawn(move || {
start.wait();
for _ in 0..ERRORS {
tracker.check_read_write_error(Some(&media_error()));
}
})
})
.collect();
start.wait();
while tracker.get_io_error_state().1 < 2 + READERS * ERRORS / 2 {
std::hint::spin_loop();
}
tracker.record_success_at(mark);
for reader in readers {
reader.join().unwrap();
}
// The two errors before the mark are cleared; every error the
// readers recorded after it stands.
assert_eq!(tracker.get_io_error_state().1, READERS * ERRORS);
}
}
#[test]
fn reset_io_error_state_lifts_the_quarantine() {
let tracker = IoErrorTracker::default();
tracker.check_read_write_error(Some(&media_error()));
tracker.mark_io_quarantined();
tracker.reset_io_error_state();
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
assert!(!tracker.should_quarantine());
}
#[test]
fn test_helper_arms_a_sustained_error() {
let tracker = IoErrorTracker::default();
tracker.set_last_io_error_for_test(Some("input/output error"));
assert!(tracker.should_quarantine());
assert_eq!(
tracker.get_io_error_state(),
(
Some("input/output error".to_string()),
IO_ERROR_TOLERANCE,
false
)
);
tracker.set_last_io_error_for_test(None);
assert_eq!(tracker.get_io_error_state(), (None, 0, false));
}
}
+2
View File
@@ -1,6 +1,8 @@
pub mod disk_location;
pub mod erasure_coding;
pub mod idx;
pub(crate) mod io;
pub(crate) mod io_error;
pub mod needle;
pub mod needle_map;
pub mod store;
+4 -2
View File
@@ -21,7 +21,7 @@ impl CRC {
/// Legacy `.Value()` function — deprecated in Go but needed for backward compat check.
/// Formula: (crc >> 15 | crc << 17) + 0xa282ead8
pub fn legacy_value(&self) -> u32 {
(self.0 >> 15 | self.0 << 17).wrapping_add(0xa282ead8)
self.0.rotate_right(15).wrapping_add(0xa282ead8)
}
}
@@ -67,7 +67,9 @@ mod tests {
fn test_crc_legacy_value() {
let crc = CRC(0x12345678);
let v = crc.legacy_value();
let expected = (0x12345678u32 >> 15 | 0x12345678u32 << 17).wrapping_add(0xa282ead8);
// (0x12345678 >> 15 | 0x12345678 << 17) + 0xa282ead8, worked out by hand so
// the test checks the rotate rather than restating it.
let expected = 0x4f730f40_u32;
assert_eq!(v, expected);
}
}
+4
View File
@@ -1,4 +1,8 @@
pub mod crc;
#[expect(
clippy::module_inception,
reason = "needle/needle.rs mirrors the Go package layout"
)]
pub mod needle;
pub mod ttl;
+120 -33
View File
@@ -198,8 +198,8 @@ impl Needle {
/// the data payload from disk at all, matching Go's `ReadNeedleMeta`.
pub fn read_paged_meta(
&mut self,
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
header_bytes: &[u8], // first 20 bytes: NEEDLE_HEADER_SIZE + DATA_SIZE_SIZE
meta_bytes: &[u8], // tail: non-data body metadata + checksum + timestamp + padding
offset: i64,
expected_size: Size,
version: Version,
@@ -560,7 +560,7 @@ impl Needle {
// Padding to 8-byte alignment
let padding = padding_length(self.size, version).0 as usize;
buf.extend(std::iter::repeat(0u8).take(padding));
buf.extend(std::iter::repeat_n(0u8, padding));
buf
}
@@ -581,23 +581,19 @@ impl Needle {
// ============================================================================
/// Compute padding to align needle to NEEDLE_PADDING_SIZE (8 bytes).
///
/// The sum is formed in i64: a size read from a corrupt header can sit near
/// `i32::MAX`, and adding the header, checksum and timestamp widths to it in
/// i32 would overflow (a panic with overflow checks, a wrapped padding
/// without). The result is at most NEEDLE_PADDING_SIZE, so it fits `Size`.
pub fn padding_length(needle_size: Size, version: Version) -> Size {
if version == VERSION_3 {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32
+ needle_size.0
+ NEEDLE_CHECKSUM_SIZE as i32
+ TIMESTAMP_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
let fixed = if version == VERSION_3 {
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE + TIMESTAMP_SIZE
} else {
Size(
NEEDLE_PADDING_SIZE as i32
- ((NEEDLE_HEADER_SIZE as i32 + needle_size.0 + NEEDLE_CHECKSUM_SIZE as i32)
% NEEDLE_PADDING_SIZE as i32),
)
}
NEEDLE_HEADER_SIZE + NEEDLE_CHECKSUM_SIZE
};
let unpadded = fixed as i64 + needle_size.0 as i64;
Size((NEEDLE_PADDING_SIZE as i64 - unpadded % NEEDLE_PADDING_SIZE as i64) as i32)
}
/// Body length = Size + Checksum + [Timestamp] + Padding.
@@ -619,6 +615,30 @@ pub fn get_actual_size(size: Size, version: Version) -> i64 {
NEEDLE_HEADER_SIZE as i64 + needle_body_length(size, version)
}
/// Validate a wire-supplied needle body size before any `as usize` cast.
/// Rejects negative/deleted sizes and bodies larger than the gRPC max message.
/// Size(0) is allowed: empty/anomalous entries and tombstones read as size 0
/// (actual_size = header+checksum+pad > 0, safe alloc, no wrap).
/// Transport cap only: storage paths must NOT use this cap — see volume.rs
/// guards (a >1GiB stored needle from a high-limit cluster must remain
/// readable/compaction-safe). Keep `get_actual_size` unchanged (it
/// intentionally returns negative for deleted index entries).
pub fn validate_wire_size(size: Size) -> Result<(), String> {
if size.0 < 0 {
return Err(format!("invalid needle size {}", size.0));
}
// Keep in sync with canonical `GRPC_MAX_MESSAGE_SIZE` in server/grpc_client.rs:10
// (duplicated here to avoid a storage->server import and prevent drift).
const WIRE_MAX_NEEDLE_SIZE: i32 = 1 << 30;
if size.0 > WIRE_MAX_NEEDLE_SIZE {
return Err(format!(
"needle size {} exceeds max {}",
size.0, WIRE_MAX_NEEDLE_SIZE
));
}
Ok(())
}
/// Read 5 bytes as a u64 (big-endian, zero-padded high bytes).
fn bytes_to_u64_5(bytes: &[u8]) -> u64 {
assert!(bytes.len() >= 5);
@@ -729,6 +749,14 @@ pub fn parse_needle_id_cookie(s: &str) -> Result<(NeedleId, Cookie), String> {
(s, None)
};
// Every length check and the split below are in BYTES, so a multi-byte
// character would let `split` land inside one and panic the slice. Hex is
// ASCII by definition; reject anything else up front, as Go's ParseUint
// does a step later.
if !hex_part.is_ascii() {
return Err("KeyHash must be ASCII hex.".to_string());
}
// Go: len(key_hash_string) <= CookieSize*2 => error (must be > 8 hex chars)
if hex_part.len() <= COOKIE_SIZE * 2 {
return Err("KeyHash is too short.".to_string());
@@ -770,7 +798,9 @@ pub fn parse_needle_id_cookie(s: &str) -> Result<(NeedleId, Cookie), String> {
#[derive(Debug, thiserror::Error)]
pub enum NeedleError {
#[error("size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}")]
#[error(
"size mismatch at offset {offset}: found id={id} size={found:?}, expected size={expected:?}"
)]
SizeMismatch {
offset: i64,
id: NeedleId,
@@ -806,6 +836,30 @@ pub enum NeedleError {
mod tests {
use super::*;
/// A fid whose hex part carries multi-byte UTF-8 must be rejected, not
/// panic. `split` is a byte offset into `hex_part`; before the ASCII guard
/// `&hex_part[..split]` could land inside a character. `GET /3,ééééa` is
/// nine bytes, so it passes the length checks and splits at byte 1 —
/// halfway through the first `é`. Go's `ParseUint` just errors.
#[test]
fn parse_needle_id_cookie_rejects_non_ascii_instead_of_panicking() {
for s in ["ééééa", "ééééaaaaa", "0123456é9abc", "ééééa_1"] {
assert!(
parse_needle_id_cookie(s).is_err(),
"non-ASCII fid {:?} must be an error",
s
);
}
}
/// The ASCII guard must not change any accepted input.
#[test]
fn parse_needle_id_cookie_still_accepts_ascii_hex() {
let (id, cookie) = parse_needle_id_cookie("01637037d6").unwrap();
assert_eq!(id, NeedleId(0x01));
assert_eq!(cookie, Cookie(0x637037d6));
}
#[test]
fn test_parse_header() {
let mut buf = [0u8; NEEDLE_HEADER_SIZE];
@@ -824,11 +878,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v3() {
let mut n = Needle::default();
n.cookie = Cookie(42);
n.id = NeedleId(100);
n.data = b"hello world".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(42),
id: NeedleId(100),
data: b"hello world".to_vec(),
flags: 0,
..Needle::default()
};
n.set_has_name();
n.name = b"test.txt".to_vec();
n.name_size = 8;
@@ -867,11 +923,13 @@ mod tests {
#[test]
fn test_needle_write_read_round_trip_v2() {
let mut n = Needle::default();
n.cookie = Cookie(77);
n.id = NeedleId(200);
n.data = b"data v2".to_vec();
n.flags = 0;
let mut n = Needle {
cookie: Cookie(77),
id: NeedleId(200),
data: b"data v2".to_vec(),
flags: 0,
..Needle::default()
};
let bytes = n.write_bytes(VERSION_2);
let expected_size = get_actual_size(n.size, VERSION_2);
@@ -886,10 +944,12 @@ mod tests {
#[test]
fn test_read_bytes_meta_only_handles_tombstone_v3() {
let mut tombstone = Needle::default();
tombstone.cookie = Cookie(0x1234abcd);
tombstone.id = NeedleId(300);
tombstone.append_at_ns = 999_999;
let mut tombstone = Needle {
cookie: Cookie(0x1234abcd),
id: NeedleId(300),
append_at_ns: 999_999,
..Needle::default()
};
let bytes = tombstone.write_bytes(VERSION_3);
@@ -917,6 +977,21 @@ mod tests {
}
}
#[test]
fn padding_length_does_not_overflow_on_a_corrupt_size() {
// A header read from a corrupt or truncated file can carry any i32
// size. The scanners bound it against the bytes left before sizing a
// buffer, but on a volume with more than 2 GiB left a size near
// i32::MAX passes that bound, so the padding arithmetic itself must
// not overflow. Overflow checks are on in test builds, so an i32 sum
// here would panic rather than wrap.
for version in [VERSION_2, VERSION_3] {
let padding = padding_length(Size(i32::MAX), version).0 as i64;
assert!((1..=NEEDLE_PADDING_SIZE as i64).contains(&padding));
assert_eq!(get_actual_size(Size(i32::MAX), version) % 8, 0);
}
}
#[test]
fn test_file_id_parse() {
let fid = FileId::parse("3,01637037d6").unwrap();
@@ -961,4 +1036,16 @@ mod tests {
assert_eq!(fid.key, NeedleId(0x123));
assert_eq!(fid.cookie, Cookie(0));
}
#[test]
fn test_validate_wire_size_boundaries() {
assert!(validate_wire_size(Size(-100)).is_err());
assert!(validate_wire_size(Size(-1)).is_err());
assert!(validate_wire_size(Size(0)).is_ok());
assert!(validate_wire_size(Size(1024)).is_ok());
assert!(validate_wire_size(Size(1)).is_ok());
assert!(validate_wire_size(Size(1 << 30)).is_ok());
assert!(validate_wire_size(Size((1 << 30) + 1)).is_err());
assert!(validate_wire_size(Size(i32::MAX)).is_err());
}
}
+122 -25
View File
@@ -80,8 +80,14 @@ impl TTL {
if s.is_empty() {
return Ok(TTL::EMPTY);
}
// The unit is read as the last BYTE and the count as everything before
// it, so a trailing multi-byte character would split inside itself and
// panic. A TTL is digits plus a one-letter unit; reject the rest.
if !s.is_ascii() {
return Err(format!("invalid TTL {:?}: must be ASCII", s));
}
let last_byte = s.as_bytes()[s.len() - 1];
let (num_str, unit_byte) = if last_byte >= b'0' && last_byte <= b'9' {
let (num_str, unit_byte) = if last_byte.is_ascii_digit() {
// All digits — default to minutes (matching Go)
(s, b'm')
} else {
@@ -144,40 +150,73 @@ fn fit_ttl_count(count: u32, unit: u8) -> TTL {
const MINUTE_SECS: u64 = 60;
// First pass: try exact fits from largest to smallest
if seconds % YEAR_SECS == 0 && seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
if seconds.is_multiple_of(YEAR_SECS) && seconds / YEAR_SECS < 256 {
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
if seconds % MONTH_SECS == 0 && seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
if seconds.is_multiple_of(MONTH_SECS) && seconds / MONTH_SECS < 256 {
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds % WEEK_SECS == 0 && seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
if seconds.is_multiple_of(WEEK_SECS) && seconds / WEEK_SECS < 256 {
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds % DAY_SECS == 0 && seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
if seconds.is_multiple_of(DAY_SECS) && seconds / DAY_SECS < 256 {
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds % HOUR_SECS == 0 && seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
if seconds.is_multiple_of(HOUR_SECS) && seconds / HOUR_SECS < 256 {
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
// Minutes: truncating division
if seconds / MINUTE_SECS < 256 {
return TTL { count: (seconds / MINUTE_SECS) as u8, unit: TTL_UNIT_MINUTE };
return TTL {
count: (seconds / MINUTE_SECS) as u8,
unit: TTL_UNIT_MINUTE,
};
}
// Second pass: truncating division from smallest to largest
if seconds / HOUR_SECS < 256 {
return TTL { count: (seconds / HOUR_SECS) as u8, unit: TTL_UNIT_HOUR };
return TTL {
count: (seconds / HOUR_SECS) as u8,
unit: TTL_UNIT_HOUR,
};
}
if seconds / DAY_SECS < 256 {
return TTL { count: (seconds / DAY_SECS) as u8, unit: TTL_UNIT_DAY };
return TTL {
count: (seconds / DAY_SECS) as u8,
unit: TTL_UNIT_DAY,
};
}
if seconds / WEEK_SECS < 256 {
return TTL { count: (seconds / WEEK_SECS) as u8, unit: TTL_UNIT_WEEK };
return TTL {
count: (seconds / WEEK_SECS) as u8,
unit: TTL_UNIT_WEEK,
};
}
if seconds / MONTH_SECS < 256 {
return TTL { count: (seconds / MONTH_SECS) as u8, unit: TTL_UNIT_MONTH };
return TTL {
count: (seconds / MONTH_SECS) as u8,
unit: TTL_UNIT_MONTH,
};
}
if seconds / YEAR_SECS < 256 {
return TTL { count: (seconds / YEAR_SECS) as u8, unit: TTL_UNIT_YEAR };
return TTL {
count: (seconds / YEAR_SECS) as u8,
unit: TTL_UNIT_YEAR,
};
}
TTL::EMPTY
}
@@ -207,6 +246,16 @@ impl fmt::Display for TTL {
mod tests {
use super::*;
/// `?ttl=5%C3%A9` must be an error, not a panic. The unit is taken as the
/// last *byte*, so a trailing multi-byte character made `&s[..s.len()-1]`
/// split inside it.
#[test]
fn ttl_read_rejects_non_ascii_instead_of_panicking() {
for s in ["5é", "é", "3🦀", "12é"] {
assert!(TTL::read(s).is_err(), "non-ASCII TTL {:?} must error", s);
}
}
#[test]
fn test_ttl_parse() {
let ttl = TTL::read("3m").unwrap();
@@ -225,7 +274,13 @@ mod tests {
// 24h normalizes to 1d via fitTtlCount
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl.to_seconds(), 86400);
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -271,12 +326,24 @@ mod tests {
fn test_ttl_overflow_normalizes() {
// Go's ReadTTL calls fitTtlCount: 300m = 18000s = 5h (exact fit)
let ttl = TTL::read("300m").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_HOUR
}
);
// 256h = 921600s. Doesn't fit in hours (256 >= 256), doesn't fit exact in days.
// Second pass: 921600/86400 = 10 (truncated) < 256 -> 10d
let ttl = TTL::read("256h").unwrap();
assert_eq!(ttl, TTL { count: 10, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 10,
unit: TTL_UNIT_DAY
}
);
}
#[test]
@@ -284,19 +351,49 @@ mod tests {
// Go's ReadTTL calls fitTtlCount which normalizes to coarsest unit.
// 120m -> 2h, 7d -> 1w, 24h -> 1d.
let ttl = TTL::read("120m").unwrap();
assert_eq!(ttl, TTL { count: 2, unit: TTL_UNIT_HOUR });
assert_eq!(
ttl,
TTL {
count: 2,
unit: TTL_UNIT_HOUR
}
);
let ttl = TTL::read("7d").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_WEEK });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_WEEK
}
);
let ttl = TTL::read("24h").unwrap();
assert_eq!(ttl, TTL { count: 1, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 1,
unit: TTL_UNIT_DAY
}
);
// Values that don't simplify stay as-is
let ttl = TTL::read("5d").unwrap();
assert_eq!(ttl, TTL { count: 5, unit: TTL_UNIT_DAY });
assert_eq!(
ttl,
TTL {
count: 5,
unit: TTL_UNIT_DAY
}
);
let ttl = TTL::read("3m").unwrap();
assert_eq!(ttl, TTL { count: 3, unit: TTL_UNIT_MINUTE });
assert_eq!(
ttl,
TTL {
count: 3,
unit: TTL_UNIT_MINUTE
}
);
}
}
+239 -213
View File
@@ -97,12 +97,13 @@ impl NeedleMapMetric {
self.file_byte_count
.fetch_add(new_size.0 as u64, Ordering::Relaxed);
// Go: if oldSize > 0 && oldSize.IsValid() { LogDeletionCounter(oldSize) }
if let Some(old_val) = old {
if old_val.size.0 > 0 && old_val.size.is_valid() {
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
if let Some(old_val) = old
&& old_val.size.0 > 0
&& old_val.size.is_valid()
{
self.deletion_count.fetch_add(1, Ordering::Relaxed);
self.deletion_byte_count
.fetch_add(old_val.size.0 as u64, Ordering::Relaxed);
}
}
@@ -225,6 +226,12 @@ pub struct CompactNeedleMap {
idx_file_offset: u64,
}
impl Default for CompactNeedleMap {
fn default() -> Self {
Self::new()
}
}
impl CompactNeedleMap {
/// Create a new empty in-memory map.
pub fn new() -> Self {
@@ -240,6 +247,8 @@ impl CompactNeedleMap {
pub fn load_from_idx<R: Read + Seek>(reader: &mut R, version: Version) -> io::Result<Self> {
let mut nm = CompactNeedleMap::new();
idx::walk_index_file(reader, 0, |key, offset, size| {
// A read-only load attaches no writer, so this is its only size.
nm.idx_file_offset += NEEDLE_MAP_ENTRY_SIZE as u64;
nm.metric.maybe_set_max_needle_end(offset, size, version);
if offset.is_zero() || size.is_deleted() {
nm.delete_from_map(key);
@@ -410,9 +419,9 @@ impl CompactNeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
{
self.map.ascending_visit(f)
}
@@ -465,9 +474,9 @@ impl RedbNeedleMap {
/// loses at most the writes since the last checkpoint from redb, and
/// the next load replays them from .idx.
fn begin_write_no_fsync(db: &Database) -> io::Result<redb::WriteTransaction> {
let mut txn = db.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = db
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
let _ = txn.set_durability(Durability::None);
Ok(txn)
}
@@ -501,7 +510,7 @@ impl RedbNeedleMap {
pub fn checkpoint(&mut self, sync_idx: bool) -> io::Result<()> {
let txn = self.begin_checkpoint(sync_idx)?;
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
self.writes_since_checkpoint = 0;
Ok(())
}
@@ -516,17 +525,17 @@ impl RedbNeedleMap {
if sync_idx {
self.sync()?;
}
let mut txn = self.db_or_err()?.begin_write().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb begin_write: {}", e))
})?;
let mut txn = self
.db_or_err()?
.begin_write()
.map_err(|e| io::Error::other(format!("redb begin_write: {}", e)))?;
txn.set_quick_repair(true);
if self.idx_file.is_some() {
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, self.idx_file_offset).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, self.idx_file_offset)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
Ok(txn)
}
@@ -538,22 +547,20 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.create(db_path)
.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb create error: {}", e))
})?;
.map_err(|e| io::Error::other(format!("redb create error: {}", e)))?;
// Ensure tables exist
let txn = Self::begin_write_no_fsync(&db)?;
{
let _table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let _meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table meta: {}", e))
})?;
let _table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let _meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table meta: {}", e)))?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
Ok(RedbNeedleMap {
db: Some(db),
@@ -572,16 +579,14 @@ impl RedbNeedleMap {
fn save_idx_size_meta(&self, idx_size: u64) -> io::Result<()> {
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
{
let mut meta = txn.open_table(META_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e))
})?;
meta.insert(META_IDX_SIZE, idx_size).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert meta: {}", e))
})?;
let mut meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
meta.insert(META_IDX_SIZE, idx_size)
.map_err(|e| io::Error::other(format!("redb insert meta: {}", e)))?;
}
txn.commit().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb commit meta: {}", e))
})?;
txn.commit()
.map_err(|e| io::Error::other(format!("redb commit meta: {}", e)))?;
Ok(())
}
@@ -590,22 +595,26 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let meta = txn
.open_table(META_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open meta: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open meta: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `meta`; bind the match so the
// temporary Result is dropped before `meta`.
let result = match meta.get(META_IDX_SIZE) {
// ReadableTable::get guard borrows `meta`; edition 2024 drops the tail
// expression's temporaries before `meta`, so no extra binding is needed.
match meta.get(META_IDX_SIZE) {
Ok(Some(guard)) => Ok(Some(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get meta: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get meta: {}", e))),
}
}
/// Test-only read of META `idx_size` through the live handle. See
/// [`test_support::live_meta_idx_size`] for why durability tests use
/// this instead of copying the open `.rdb`.
#[cfg(test)]
pub(crate) fn live_meta_idx_size(&self) -> Option<u64> {
self.read_idx_size_meta().unwrap()
}
/// Load from an .idx file, reusing an existing .rdb if it is consistent.
@@ -648,7 +657,7 @@ impl RedbNeedleMap {
let db = Database::builder()
.set_cache_size(cache_bytes)
.open(db_path)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open: {}", e)))?;
let mut nm = RedbNeedleMap {
db: Some(db),
@@ -663,14 +672,11 @@ impl RedbNeedleMap {
let stored_idx_size = nm
.read_idx_size_meta()?
.ok_or_else(|| io::Error::new(io::ErrorKind::Other, "no idx_size in redb meta"))?;
.ok_or_else(|| io::Error::other("no idx_size in redb meta"))?;
if stored_idx_size > idx_size {
// .idx shrank — corrupted or truncated, need full rebuild
return Err(io::Error::new(
io::ErrorKind::Other,
"idx file smaller than stored size",
));
return Err(io::Error::other("idx file smaller than stored size"));
}
// Counters come from the whole .idx history, never from the table,
@@ -683,40 +689,37 @@ impl RedbNeedleMap {
let start_entry = stored_idx_size / NEEDLE_MAP_ENTRY_SIZE as u64;
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
idx::walk_index_file(reader, start_entry, |key, offset, size| {
let key_u64: u64 = key.into();
if offset.is_zero() || size.is_deleted() {
// Delete: store a tombstone (negative size, original
// offset) over a live value; already deleted is a no-op.
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64) {
if old.size.is_valid() {
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
)
})?;
}
if let Ok(Some(old)) = nm.get_via_table(&table, key_u64)
&& old.size.is_valid()
{
let deleted_nv = NeedleValue {
offset: old.offset,
size: Size(-(old.size.0)),
};
let packed = pack_needle_value(&deleted_nv);
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
} else {
let packed = pack_needle_value(&NeedleValue { offset, size });
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
Ok(())
})?;
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
}
@@ -734,10 +737,7 @@ impl RedbNeedleMap {
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
@@ -790,13 +790,13 @@ impl RedbNeedleMap {
let txn = Self::begin_write_no_fsync(nm.db.as_ref().unwrap())?;
{
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
if !unlinked {
table.retain(|_, _| false).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb retain: {}", e))
})?;
table
.retain(|_, _| false)
.map_err(|e| io::Error::other(format!("redb retain: {}", e)))?;
}
#[cfg(not(feature = "redb-experimental-cursor"))]
@@ -804,30 +804,33 @@ impl RedbNeedleMap {
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
table.insert(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb insert: {}", e))
})?;
table
.insert(key_u64, packed.as_slice())
.map_err(|e| io::Error::other(format!("redb insert: {}", e)))?;
}
}
#[cfg(feature = "redb-experimental-cursor")]
{
let mut cursor = table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
let mut cursor =
table
.upper_bound_mut(Bound::<u64>::Unbounded)
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb upper_bound_mut: {}", e),
)
})?;
for (key, nv) in &entries {
let key_u64: u64 = (*key).into();
let packed = pack_needle_value(nv);
cursor.insert_before(key_u64, packed.as_slice()).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
cursor
.insert_before(key_u64, packed.as_slice())
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("redb insert_before: {}", e),
)
})?;
}
cursor.close().map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb cursor close: {}", e))
@@ -835,7 +838,7 @@ impl RedbNeedleMap {
}
}
txn.commit()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb commit: {}", e)))?;
nm.save_idx_size_meta(idx_size)?;
Ok(())
@@ -901,23 +904,16 @@ impl RedbNeedleMap {
Ok(t) => t,
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb open_table: {}", e),
));
return Err(io::Error::other(format!("redb open_table: {}", e)));
}
};
let result = match table.insert(key_u64, packed.as_slice()) {
match table.insert(key_u64, packed.as_slice()) {
Ok(prev) => prev.and_then(|g| packed_to_needle_value(g.value())),
Err(e) => {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
};
result
}
};
match txn.commit() {
Ok(()) => old,
@@ -925,8 +921,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -935,12 +930,9 @@ impl RedbNeedleMap {
// visible and redb refuses further writes. Keep
// the .idx row (do NOT truncate) and reopen from
// .idx to repair redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after put commit error failed: {}",
reopen_err
);
tracing::warn!("redb reopen after put commit error failed: {}", reopen_err);
}
return Err(err);
}
@@ -968,39 +960,32 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
// experimental-api-5 drops inherent ReadOnlyTable::get ('static guard).
// ReadableTable::get guard borrows `table`; bind the match so the
// temporary Result is dropped before `table`.
let result = match table.get(key_u64) {
// ReadableTable::get guard borrows `table`; edition 2024 drops the tail
// expression's temporaries before `table`, so no extra binding is needed.
match table.get(key_u64) {
Ok(Some(guard)) => Ok(packed_to_needle_value(guard.value())),
Ok(None) => Ok(None),
Err(e) => Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
)),
};
result
Err(e) => Err(io::Error::other(format!("redb get: {}", e))),
}
}
/// Mark a needle as deleted. Appends tombstone to .idx file, negates size in redb.
pub fn delete(&mut self, key: NeedleId, offset: Offset) -> io::Result<Option<Size>> {
let key_u64: u64 = key.into();
let txn = Self::begin_write_no_fsync(self.db_or_err()?)?;
let mut table = txn.open_table(NEEDLE_TABLE).map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e))
})?;
let mut table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let old = match table.get(key_u64) {
Ok(Some(guard)) => packed_to_needle_value(guard.value()),
Ok(None) => None,
Err(e) => {
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb get: {}", e),
));
return Err(io::Error::other(format!("redb get: {}", e)));
}
};
let Some(old) = old.filter(|nv| nv.size.is_valid()) else {
@@ -1021,10 +1006,7 @@ impl RedbNeedleMap {
drop(table);
if let Err(e) = insert_res {
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
format!("redb insert: {}", e),
));
return Err(io::Error::other(format!("redb insert: {}", e)));
}
match txn.commit() {
Ok(()) => {}
@@ -1032,8 +1014,7 @@ impl RedbNeedleMap {
// Transaction rolled back, database still usable:
// truncate the orphan .idx row.
self.truncate_idx_to_offset();
return Err(io::Error::new(
io::ErrorKind::Other,
return Err(io::Error::other(
"redb commit: Transaction was poisoned by a panic",
));
}
@@ -1042,7 +1023,7 @@ impl RedbNeedleMap {
// and redb refuses further writes. Keep the .idx row
// (do NOT truncate) and reopen from .idx to repair
// redb's internal state.
let err = io::Error::new(io::ErrorKind::Other, format!("redb commit: {}", e));
let err = io::Error::other(format!("redb commit: {}", e));
if let Err(reopen_err) = self.reopen_from_idx() {
tracing::warn!(
"redb reopen after delete commit error failed: {}",
@@ -1105,10 +1086,10 @@ impl RedbNeedleMap {
/// after the orphan, `idx_file_offset` advances past it, and a later
/// checkpoint records an offset that makes the reload skip the orphan.
fn truncate_idx_to_offset(&mut self) {
if let Some(ref mut idx_file) = self.idx_file {
if let Err(e) = idx_file.truncate_to(self.idx_file_offset) {
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
if let Some(ref mut idx_file) = self.idx_file
&& let Err(e) = idx_file.truncate_to(self.idx_file_offset)
{
tracing::warn!("failed to truncate orphan .idx row: {}", e);
}
}
@@ -1146,18 +1127,12 @@ impl RedbNeedleMap {
let read_file = std::fs::OpenOptions::new()
.read(true)
.open(&idx_path)
.map_err(|e| {
io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e))
})?;
.map_err(|e| io::Error::other(format!("reopen: open .idx {}: {}", idx_path, e)))?;
let actual_idx_size = read_file.metadata()?.len();
let mut reader = io::BufReader::new(read_file);
let reopened = Self::load_from_idx(
&self.rdb_path,
&mut reader,
self.version,
self.cache_bytes,
)?;
let reopened =
Self::load_from_idx(&self.rdb_path, &mut reader, self.version, self.cache_bytes)?;
// Preserve the append writer and the paths/version/cache; adopt the
// repaired database, metrics, and idx_file_offset from the reload.
@@ -1198,10 +1173,10 @@ impl RedbNeedleMap {
let txn = self
.db_or_err()?
.begin_read()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb begin_read: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb begin_read: {}", e)))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb open_table: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb open_table: {}", e)))?;
let mut file = std::fs::OpenOptions::new()
.write(true)
@@ -1212,18 +1187,17 @@ impl RedbNeedleMap {
// redb iterates in key order (u64 ascending)
let iter = table
.iter()
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("redb iter: {}", e)))?;
.map_err(|e| io::Error::other(format!("redb iter: {}", e)))?;
for entry in iter {
let (key_guard, val_guard) = entry.map_err(|e| {
io::Error::new(io::ErrorKind::Other, format!("redb iter next: {}", e))
})?;
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb iter next: {}", e)))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if let Some(nv) = packed_to_needle_value(bytes) {
if nv.size.is_valid() {
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
if let Some(nv) = packed_to_needle_value(bytes)
&& nv.size.is_valid()
{
idx::write_index_entry(&mut file, NeedleId(key_u64), nv.offset, nv.size)?;
}
}
file.sync_all()?;
@@ -1231,9 +1205,10 @@ impl RedbNeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, mut f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, mut f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
let txn = self
.db_or_err()
@@ -1413,6 +1388,18 @@ impl NeedleMap {
}
}
/// Skew the live file count away from what the `.idx` holds, so tests
/// can build a volume whose reported count disagrees with a reload.
#[cfg(test)]
pub(crate) fn add_file_count_for_test(&self, delta: i64) {
let metric = match self {
NeedleMap::InMemory(nm) => &nm.metric,
NeedleMap::Redb(nm) => &nm.metric,
NeedleMap::SortedFile(_) => panic!("sorted-file needle maps are read-only"),
};
metric.file_count.fetch_add(delta, Ordering::Relaxed);
}
/// Largest (offset + actual size) seen during the load walk; 0 if the
/// map is empty. Used at volume load to detect .idx entries that
/// reference past the end of .dat (issue #8928) without a second scan.
@@ -1471,9 +1458,10 @@ impl NeedleMap {
}
/// Visit all entries in ascending order by needle ID.
pub fn ascending_visit<F>(&self, f: F) -> Result<(), String>
pub fn ascending_visit<F, E>(&self, f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
match self {
NeedleMap::InMemory(nm) => nm.ascending_visit(f),
@@ -1494,7 +1482,7 @@ impl NeedleMap {
// The visitor never fails, so neither can this.
let _ = nm.ascending_visit(|id, nv| {
entries.push((id, *nv));
Ok(())
Ok::<(), std::convert::Infallible>(())
});
Ok(entries)
}
@@ -1512,22 +1500,30 @@ impl NeedleMap {
pub(crate) mod test_support {
use super::*;
/// The `.idx` size recorded in the durable state of the `.rdb` at
/// `rdb_path`, read from a copy taken while the map may still be open:
/// exactly what a crash would leave behind. `None` when nothing durable
/// has been recorded yet.
pub(crate) fn durable_idx_size(rdb_path: &Path) -> Option<u64> {
let copy = rdb_path.with_extension("crash-copy.rdb");
std::fs::copy(rdb_path, &copy).unwrap();
let db = Database::open(&copy).unwrap();
let txn = db.begin_read().unwrap();
let meta = txn.open_table(META_TABLE).ok()?;
let size = meta.get(META_IDX_SIZE).unwrap().map(|g| g.value());
drop(meta);
drop(txn);
drop(db);
let _ = std::fs::remove_file(&copy);
size
/// The `.idx` size in the map's META table, read through the live
/// handle.
///
/// The load path records the `.idx` size with `Durability::None`, and
/// every `put`/`delete` also commits non-durably, so before the first
/// checkpoint this is the load-time value (`Some(0)` for a fresh map) —
/// NOT the crash-durable `None` a copy of the open `.rdb` would show.
/// A live read is the only portable observation: redb 4.2.0 takes an
/// exclusive whole-file lock, which is advisory on Unix but mandatory
/// on Windows, so copying the open `.rdb` fails there with OS error 33.
///
/// It still pins the property under test: the only *durable* META
/// writer is `checkpoint`, so any value other than the load-time one
/// proves a checkpoint recorded progress — and the post-checkpoint
/// value equals the durable one, because checkpoints commit with
/// `Durability::Immediate`. What is lost vs the old copy: strict crash
/// fidelity — a hard crash pre-checkpoint would leave META absent
/// rather than `Some(0)` (loader-equivalent outcomes: full rebuild vs
/// replay-from-0, both correct). A clean close+reopen cannot recover
/// that distinction either: dropping the `Database` flushes pending
/// non-durable commits, so a reopened handle reads `Some(0)` just like
/// the live one.
pub(crate) fn live_meta_idx_size(nm: &RedbNeedleMap) -> Option<u64> {
nm.live_meta_idx_size()
}
}
@@ -1682,6 +1678,7 @@ mod tests {
.read(true)
.write(true)
.create(true)
.truncate(false)
.open(&idx_path)
.unwrap();
let idx_size = idx_file.metadata().unwrap().len();
@@ -1857,7 +1854,7 @@ mod tests {
let mut live = 0u64;
nm.ascending_visit(|_, _| {
live += 1;
Ok(())
Ok::<(), String>(())
})
.unwrap();
assert_eq!(live, N - 1);
@@ -2070,7 +2067,7 @@ mod tests {
let mut visited = Vec::new();
nm.ascending_visit(|id, nv| {
visited.push((id, nv.size));
Ok(())
Ok::<(), String>(())
})
.unwrap();
@@ -2168,8 +2165,14 @@ mod tests {
// server opens one redb database per volume, so the process-wide
// ceiling is roughly (volumes x budget).
assert_eq!(NeedleMapKind::Redb.redb_cache_bytes(), 4 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbMedium.redb_cache_bytes(), 8 * 1024 * 1024);
assert_eq!(NeedleMapKind::RedbLarge.redb_cache_bytes(), 16 * 1024 * 1024);
assert_eq!(
NeedleMapKind::RedbMedium.redb_cache_bytes(),
8 * 1024 * 1024
);
assert_eq!(
NeedleMapKind::RedbLarge.redb_cache_bytes(),
16 * 1024 * 1024
);
}
#[test]
@@ -2198,7 +2201,7 @@ mod tests {
#[test]
fn test_redb_checkpoint_is_explicit_and_due_every_interval() {
use test_support::durable_idx_size;
use test_support::live_meta_idx_size;
// Every non-durable redb commit leaves bookkeeping behind until a
// durable one clears it, so a writable map asks for a checkpoint on
@@ -2208,8 +2211,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, _idx_path) = open_writable_redb(dir.path());
for i in 1..EXPECTED_INTERVAL {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
assert!(!nm.checkpoint_due(), "due after only {i} writes");
}
nm.put(
@@ -2219,20 +2226,29 @@ mod tests {
)
.unwrap();
assert!(nm.checkpoint_due());
assert_eq!(durable_idx_size(&db_path), None, "put() must not commit durably");
// No checkpoint taken yet: META still holds the load-time .idx size.
// put() only commits non-durably, so the live value is unchanged.
assert_eq!(
live_meta_idx_size(&nm),
Some(0),
"put() must not record checkpoint progress"
);
nm.checkpoint(true).unwrap();
assert!(!nm.checkpoint_due());
assert_eq!(
durable_idx_size(&db_path),
live_meta_idx_size(&nm),
Some(EXPECTED_INTERVAL * NEEDLE_MAP_ENTRY_SIZE as u64),
"checkpoint records how much of the .idx the table reflects"
);
// Snapshot the .rdb while the map is still open: what a crash leaves.
// Everything is durable after the checkpoint, so the map is closed
// first and the snapshot sees the same bytes on every platform.
// (Copying while open fails on Windows, where redb's file lock is
// mandatory: what a crash leaves.)
drop(nm);
let crash_copy = dir.path().join("crash.rdb");
std::fs::copy(&db_path, &crash_copy).unwrap();
drop(nm);
let db = Database::open(&crash_copy).unwrap();
let txn = db.begin_read().unwrap();
let table = txn.open_table(NEEDLE_TABLE).unwrap();
@@ -2247,8 +2263,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
nm.close();
drop(nm);
@@ -2272,8 +2292,12 @@ mod tests {
let dir = tempfile::tempdir().unwrap();
let (mut nm, db_path, idx_path) = open_writable_redb(dir.path());
for i in 1..=5u64 {
nm.put(NeedleId(i), Offset::from_actual_offset((i * 8) as i64), Size(1))
.unwrap();
nm.put(
NeedleId(i),
Offset::from_actual_offset((i * 8) as i64),
Size(1),
)
.unwrap();
}
// Drop without close(): redb makes the table durable on drop, but the
// recorded .idx size stays at its load-time value (0), so the reload
@@ -2344,12 +2368,14 @@ mod tests {
reloaded.deleted_count(),
reloaded.deleted_size(),
);
assert_eq!(
after, live,
"close_first={close_first} rebuild={rebuild}"
);
assert_eq!(after, live, "close_first={close_first} rebuild={rebuild}");
assert_eq!(reloaded.get(NeedleId(1)).unwrap().unwrap().size, Size(200));
assert!(reloaded.get(NeedleId(2)).unwrap().map_or(true, |v| v.size.is_deleted()));
assert!(
reloaded
.get(NeedleId(2))
.unwrap()
.is_none_or(|v| v.size.is_deleted())
);
}
}
}
@@ -31,7 +31,7 @@ struct CompactEntry {
}
impl CompactEntry {
fn to_needle_value(&self) -> NeedleValue {
fn to_needle_value(self) -> NeedleValue {
NeedleValue {
offset: Offset::from_bytes(&self.offset),
size: self.size,
@@ -205,6 +205,18 @@ mod tests {
use std::os::unix::fs::FileExt;
borrowed.read_exact_at(&mut buf, 0).unwrap();
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = 0;
while filled < buf.len() {
let n = borrowed.seek_read(&mut buf[filled..], at).unwrap();
assert!(n != 0, "unexpected EOF in seek_read");
filled += n;
at += n as u64;
}
}
assert_eq!(&buf, b"first");
}
@@ -132,9 +132,7 @@ mod tests {
let mut seen = SeenKeys::new(10_000, FALSE_POSITIVE_RATE);
// Fresh keys may occasionally collide (that is the false-positive
// rate), but only rarely.
let fresh_reported_seen = (0..10_000u64)
.filter(|&key| seen.test_and_add(key))
.count();
let fresh_reported_seen = (0..10_000u64).filter(|&key| seen.test_and_add(key)).count();
assert!(
fresh_reported_seen < 50,
"fresh keys reported seen: {fresh_reported_seen}"
@@ -18,6 +18,7 @@ use std::sync::{Mutex, RwLock};
use super::file_pool::pooled_index_files;
use crate::storage::idx;
use crate::storage::io::read_exact_at;
use crate::storage::needle_map::{CompactNeedleMap, NeedleMapMetric, NeedleValue};
use crate::storage::types::*;
@@ -133,9 +134,7 @@ impl SortedFileNeedleMap {
}
let file = pooled_index_files()
.borrow(&self.db_file_name, false)
.map_err(|e| {
io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e))
})?;
.map_err(|e| io::Error::new(e.kind(), format!("open {}: {}", self.db_file_name, e)))?;
match search_sorted_index(&file, self.db_file_size, key)? {
Some((_, offset, size)) => Ok(Some(NeedleValue { offset, size })),
None => Ok(None),
@@ -226,10 +225,7 @@ impl SortedFileNeedleMap {
.fail_sdx_mark
.load(std::sync::atomic::Ordering::Relaxed)
{
return Err(io::Error::new(
io::ErrorKind::Other,
"injected .sdx mark failure",
));
return Err(io::Error::other("injected .sdx mark failure"));
}
let mut buf = [0u8; SIZE_SIZE];
TOMBSTONE_FILE_SIZE.to_bytes(&mut buf);
@@ -309,7 +305,7 @@ impl SortedFileNeedleMap {
let rows = rows_per_read.min(entry_count - done) as usize;
let bytes = &mut block[..rows * NEEDLE_MAP_ENTRY_SIZE];
read_exact_at(&file, bytes, done * NEEDLE_MAP_ENTRY_SIZE as u64)?;
for entry in bytes.chunks_exact(NEEDLE_MAP_ENTRY_SIZE) {
for entry in bytes.as_chunks::<NEEDLE_MAP_ENTRY_SIZE>().0 {
let (key, offset, size) = idx_entry_from_bytes(entry);
if !size.is_valid() || pending.contains_key(&key) {
continue; // deleted in place, or still awaiting that mark
@@ -321,9 +317,11 @@ impl SortedFileNeedleMap {
Ok(())
}
pub fn ascending_visit<F>(&self, mut f: F) -> Result<(), String>
/// Visit all live entries in ascending order by needle ID.
pub fn ascending_visit<F, E>(&self, mut f: F) -> Result<(), E>
where
F: FnMut(NeedleId, &NeedleValue) -> Result<(), String>,
F: FnMut(NeedleId, &NeedleValue) -> Result<(), E>,
E: From<String>,
{
let mut visit_error = None;
self.visit_live_entries(|id, nv| {
@@ -333,7 +331,7 @@ impl SortedFileNeedleMap {
}
Ok(())
})
.map_err(|e| visit_error.take().unwrap_or_else(|| e.to_string()))
.map_err(|e| visit_error.take().unwrap_or_else(|| E::from(e.to_string())))
}
pub fn iter_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
@@ -525,32 +523,6 @@ fn search_sorted_index(
Ok(None)
}
fn read_exact_at(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
file.read_exact_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;
let mut filled = 0;
let mut at = offset;
while filled < buf.len() {
let n = file.seek_read(&mut buf[filled..], at)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"unexpected EOF in seek_read",
));
}
filled += n;
at += n as u64;
}
Ok(())
}
}
fn write_at(file: &File, buf: &[u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
@@ -703,10 +675,11 @@ mod tests {
// without a reload — the same contract Go's Get has, where callers
// check size.is_deleted().
assert!(m.get(NeedleId(2)).unwrap().unwrap().size.is_deleted());
assert!(m
.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none());
assert!(
m.delete(NeedleId(2), Offset::from_actual_offset(16))
.unwrap()
.is_none()
);
assert!(!m.get(NeedleId(1)).unwrap().unwrap().size.is_deleted());
}
@@ -1002,7 +975,8 @@ mod tests {
// The retry is a no-op: no second tombstone, no double counting.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
assert_eq!(m.deleted_count(), deleted_before + 2);
@@ -1032,7 +1006,8 @@ mod tests {
);
// And a retry must not append a second tombstone for it.
assert_eq!(
m.delete(NeedleId(1), Offset::from_actual_offset(8)).unwrap(),
m.delete(NeedleId(1), Offset::from_actual_offset(8))
.unwrap(),
None
);
}
@@ -1062,7 +1037,7 @@ mod tests {
let mut visited = Vec::new();
m.ascending_visit(|id, _| {
visited.push(id);
Ok(())
Ok::<(), String>(())
})
.unwrap();
assert_eq!(visited, vec![NeedleId(2)]);
File diff suppressed because it is too large Load Diff
+37 -4
View File
@@ -8,7 +8,7 @@ use std::path::Path;
use tracing::{info, warn};
use crate::storage::disk_location::{parse_collection_volume_id_pub, DiskLocation};
use crate::storage::disk_location::{DiskLocation, parse_collection_volume_id_pub};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -131,6 +131,12 @@ impl Store {
let Some(base) = name.strip_suffix(".ecx") else {
continue;
};
// A 0-byte .ecx is a corrupt stub from a failed copy, not a
// credible owner — skip it so the scan keeps looking for a
// real index on a sibling disk (Go's indexEcxOwners).
if !ent.metadata().is_ok_and(|m| m.len() > 0) {
continue;
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
@@ -286,9 +292,7 @@ fn collect_shard_disk_volumes(loc: &DiskLocation) -> HashMap<EcKey, Vec<String>>
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
out.entry(EcKey { collection, vid })
.or_default()
.push(name);
out.entry(EcKey { collection, vid }).or_default().push(name);
}
out
}
@@ -419,4 +423,33 @@ mod tests {
let post = fs::read(dir0.join(format!("{}_{}.ecx", collection, vid))).unwrap();
assert_eq!(post, ecx_local, "mirror overwrote dir0's existing .ecx");
}
/// The mirror shares Go's indexEcxOwners, which skips a 0-byte `.ecx`:
/// a stub must not be chosen as the source to mirror from.
#[test]
fn mirror_owner_index_skips_zero_byte_ecx() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
fs::create_dir_all(&dir0).unwrap();
fs::create_dir_all(&dir1).unwrap();
let collection = "video-recordings";
let vid = 4123u32;
plant_ecx(&dir0, collection, vid, b"");
plant_ecx(&dir1, collection, vid, &[0xA1u8; 20]);
let mut store = Store::new(NeedleMapKind::InMemory);
add_loc(&mut store, &dir0);
add_loc(&mut store, &dir1);
let owners = store.index_ecx_owners_for_mirror();
let owner = owners
.get(&EcKey {
collection: collection.to_string(),
vid: VolumeId(vid),
})
.expect("the valid .ecx on disk 1 must be indexed");
assert_eq!(owner.location, 1);
}
}
+284 -90
View File
@@ -21,7 +21,7 @@ use std::fs;
use tracing::{error, info, warn};
use crate::storage::disk_location::{is_ec_shard_extension, parse_collection_volume_id_pub};
use crate::storage::erasure_coding::ec_shard::DATA_SHARDS_COUNT;
use crate::storage::erasure_coding::ec_shard::{DATA_SHARDS_COUNT, ShardId};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
@@ -37,6 +37,7 @@ pub(crate) struct EcVolumeMissingIndex {
pub data_dir: String,
}
#[cfg(test)]
pub(crate) fn ec_local_ecx_path(dir: &str, collection: &str, vid: VolumeId) -> String {
if collection.is_empty() {
format!("{}/{}.ecx", dir, vid.0)
@@ -80,6 +81,11 @@ struct EcxOwnerInfo {
idx_dir: String,
}
/// One unit of reconcile work: the disk holding orphan shards, the volume
/// they belong to, the shard files, the `.ecx` owner, and whether the
/// mirror already installed sidecars locally (`use_local_idx`).
type OrphanShardLoad = (usize, EcKey, Vec<(String, ShardId)>, EcxOwnerInfo, bool);
impl Store {
/// Run cross-disk orphan-shard reconciliation. Should be called
/// after every DiskLocation has finished its per-disk EC scan.
@@ -98,7 +104,7 @@ impl Store {
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
let mut to_load: Vec<OrphanShardLoad> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -112,14 +118,11 @@ impl Store {
);
continue;
};
let local_ecx = ec_local_ecx_path(&loc.idx_directory, &key.collection, key.vid);
let local_ecx_in_data = ec_local_ecx_path(&loc.directory, &key.collection, key.vid);
let use_local_idx = std::path::Path::new(&local_ecx).exists()
|| std::path::Path::new(&local_ecx_in_data).exists();
// A 0-byte local stub is not a mirrored index (Go gates this fast
// path on HasEcxFileOnDisk); mount against the owner instead.
let use_local_idx = loc.has_ecx_file_on_disk(&key.collection, key.vid);
if !use_local_idx
&& owner.location == loc_idx
&& owner.idx_dir == loc.idx_directory
if !use_local_idx && owner.location == loc_idx && owner.idx_dir == loc.idx_directory
{
// Same-disk no-op: load_all_ec_shards already
// tried and logged the failure.
@@ -132,7 +135,7 @@ impl Store {
for (loc_idx, key, shards, owner, use_local_idx) in to_load {
let shard_names: Vec<&str> = shards.iter().map(|(n, _)| n.as_str()).collect();
let loc_dir = self.locations[loc_idx].directory.clone();
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let shard_ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
if use_local_idx {
info!(
@@ -293,10 +296,10 @@ impl Store {
// may be sole copies of a distributed volume.
let mut node_wide_bits = ev.shard_bits().0;
for other in &self.locations {
if let Some(other_ev) = other.find_ec_volume(*vid) {
if other_ev.collection == ev.collection {
node_wide_bits |= other_ev.shard_bits().0;
}
if let Some(other_ev) = other.find_ec_volume(*vid)
&& other_ev.collection == ev.collection
{
node_wide_bits |= other_ev.shard_bits().0;
}
}
let node_wide = node_wide_bits.count_ones() as usize;
@@ -419,6 +422,12 @@ impl Store {
let Some(base) = name.strip_suffix(".ecx") else {
continue;
};
// A 0-byte .ecx is a corrupt stub from a failed copy, not a
// credible owner — skip it so the scan keeps looking for a
// real index on a sibling disk (Go's indexEcxOwners).
if !ent.metadata().is_ok_and(|m| m.len() > 0) {
continue;
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
@@ -474,13 +483,13 @@ impl Store {
/// Unlike `reconcile_ec_shards_across_disks` it needs no sibling disk, so a
/// single-disk store recovers once its index has been fetched from a peer.
fn load_orphan_ec_shards_with_local_index(&mut self) {
let mut work: Vec<(usize, EcKey, Vec<u32>)> = Vec::new();
let mut work: Vec<(usize, EcKey, Vec<ShardId>)> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
for (key, shards) in collect_orphan_ec_shards(loc, loc_idx) {
if !loc.has_ecx_file_on_disk(&key.collection, key.vid) {
continue;
}
let ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let ids: Vec<ShardId> = shards.iter().map(|(_, sid)| *sid).collect();
work.push((loc_idx, key, ids));
}
}
@@ -499,6 +508,53 @@ impl Store {
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, ShardId)>> {
let mut orphans: HashMap<EcKey, Vec<(String, ShardId)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid)
&& ecv.has_shard(shard_id)
{
continue;
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
#[cfg(test)]
mod tests {
use super::*;
@@ -541,7 +597,13 @@ mod tests {
std::fs::write(&p, b"shard data nonempty").unwrap();
}
fn write_index_files(idx_dir: &str, collection: &str, vid: u32, data_shards: u32, parity_shards: u32) {
fn write_index_files(
idx_dir: &str,
collection: &str,
vid: u32,
data_shards: u32,
parity_shards: u32,
) {
// Minimal sealed .ecx (the loader only opens the file; it
// doesn't parse it during placement).
std::fs::write(
@@ -583,6 +645,34 @@ mod tests {
.unwrap();
}
/// A 0-byte `.ecx` is not a credible owner (Go's indexEcxOwners skips
/// it): picking the stub would hide the valid index on the sibling disk.
#[test]
fn test_index_ecx_owners_skips_zero_byte_stub() {
let (store, _tmp) = make_test_store(2, None);
let d0 = store.locations[0].directory.clone();
let d1 = store.locations[1].directory.clone();
std::fs::write(ec_local_ecx_path(&d0, "pics", VolumeId(7)), b"").unwrap();
write_index_files(&d1, "pics", 7, 10, 4);
let owners = store.index_ecx_owners();
let owner = owners
.get(&EcKey {
collection: "pics".to_string(),
vid: VolumeId(7),
})
.expect("the valid .ecx on disk 1 must be indexed");
assert_eq!(owner.location, 1);
assert_eq!(owner.idx_dir, d1);
// A stub with no real index anywhere owns nothing.
std::fs::write(ec_local_ecx_path(&d0, "pics", VolumeId(8)), b"").unwrap();
assert!(!store.index_ecx_owners().contains_key(&EcKey {
collection: "pics".to_string(),
vid: VolumeId(8),
}));
}
/// An empty `.dat` (<= a superblock, i.e. zero needles) for an EC volume
/// is a leftover stub from the pre-fix loader. It must be swept on startup,
/// not loaded as a phantom empty volume. With the same vid's stub on two
@@ -895,15 +985,74 @@ mod tests {
// dir1 owns the .ecx and so already has shard 1 mounted via
// its own load_all_ec_shards.
let ev1 = store.locations[1].find_ec_volume(VolumeId(vid));
assert!(ev1.is_some(), "baseline broken: dir1 should have mounted shard 1");
assert!(
ev1.is_some(),
"baseline broken: dir1 should have mounted shard 1"
);
// dir0's shards must be reconciled across to its own
// ec_volumes map, pointing at dir1's idx dir.
let ev0 = store.locations[0]
.find_ec_volume(VolumeId(vid))
.expect("dir0 should now have an EcVolume after reconcile");
assert!(ev0.has_shard(0), "shard 0 missing from dir0 after reconcile");
assert!(ev0.has_shard(12), "shard 12 missing from dir0 after reconcile");
assert!(
ev0.has_shard(0),
"shard 0 missing from dir0 after reconcile"
);
assert!(
ev0.has_shard(12),
"shard 12 missing from dir0 after reconcile"
);
}
/// dir0 holds orphan shards next to a 0-byte `.ecx` stub from a failed
/// copy; the real index is on dir1. The stub must not count as a
/// locally-mirrored index (Go gates that fast path on HasEcxFileOnDisk),
/// or the shards get registered against an empty index.
#[test]
fn test_reconcile_ignores_zero_byte_local_ecx_stub() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
std::fs::create_dir_all(&dir0).unwrap();
std::fs::create_dir_all(&dir1).unwrap();
let collection = "grafana-loki";
let vid = 1094u32;
write_shard(dir0.to_str().unwrap(), collection, vid, 0);
write_shard(dir1.to_str().unwrap(), collection, vid, 1);
write_index_files(dir1.to_str().unwrap(), collection, vid, 10, 4);
let mut store = Store::new(NeedleMapKind::InMemory);
for dir in [&dir0, &dir1] {
store
.add_location(
dir.to_str().unwrap(),
dir.to_str().unwrap(),
100,
DiskType::HardDrive,
MinFreeSpace::Percent(0.0),
Vec::new(),
)
.unwrap();
}
// Plant the stub after the startup scan so only the reconcile decision
// is under test, then drop dir0's mount and reconcile again.
store.locations[0].remove_ec_volume(VolumeId(vid));
std::fs::write(
ec_local_ecx_path(dir0.to_str().unwrap(), collection, VolumeId(vid)),
b"",
)
.unwrap();
store.reconcile_ec_shards_across_disks();
let ev0 = store.locations[0]
.find_ec_volume(VolumeId(vid))
.expect("dir0's shard must be mounted against the owner's index");
assert!(ev0.has_shard(0));
assert_eq!(ev0.ecx_actual_dir(), dir1.to_str().unwrap());
}
/// PR 9244 review case: idx_directory is configured but the
@@ -1012,7 +1161,13 @@ mod tests {
assert!(store.locations[0].find_ec_volume(VolumeId(vid)).is_none());
// Shard files must still exist on disk for operator recovery.
for sid in [0u8, 12u8] {
let p = format!("{}/{}_{}.ec{:02}", dir0.to_str().unwrap(), collection, vid, sid);
let p = format!(
"{}/{}_{}.ec{:02}",
dir0.to_str().unwrap(),
collection,
vid,
sid
);
assert!(
std::path::Path::new(&p).exists(),
"orphan shard {} was destroyed",
@@ -1077,10 +1232,12 @@ mod tests {
assert!(ev1.has_shard(6), "dir1 shard missing");
// Nothing left to recover.
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
#[test]
@@ -1112,10 +1269,12 @@ mod tests {
.unwrap();
}
assert!(store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid)));
assert!(
store
.collect_ec_volumes_missing_index()
.iter()
.all(|m| m.vid != VolumeId(vid))
);
}
/// Helper: build a 2-disk store where reconcile produces the
@@ -1184,6 +1343,87 @@ mod tests {
assert!(!std::ptr::eq(ev0, ev1));
}
/// `find_ec_volume` returns only disk 0's runtime, which is what hides
/// sibling-disk shards from every scrub mode. The plural lookup must
/// return one runtime per disk holding the vid, in location order.
#[test]
fn test_find_all_ec_volumes_returns_every_disk() {
let (store, _tmp) = build_split_disk_store(7010);
let vid = VolumeId(7010);
let all = store.find_all_ec_volumes(vid);
assert_eq!(
all.len(),
2,
"expected one EcVolume per disk holding the vid"
);
// Disk 0 carries shards 0 and 12; disk 1 carries shard 1.
assert!(all[0].has_shard(0));
assert!(all[0].has_shard(12));
assert!(all[1].has_shard(1));
// The singular lookup sees only the first — the bug being fixed.
let first = store.find_ec_volume(vid).unwrap();
assert!(std::ptr::eq(first, all[0]));
// A vid nobody mounts yields an empty vec, not a panic.
assert!(store.find_all_ec_volumes(VolumeId(9999)).is_empty());
}
/// End-to-end: with the vid mounted on two disks, a scrub driven through
/// the Store must reach BOTH disks' shards. Before the aggregation fix
/// `find_ec_volume` returned disk 0 and disk 1's shard 1 was never read.
#[test]
fn test_scrub_plans_reach_every_disk_through_the_store() {
use crate::storage::erasure_coding::ec_volume::{
EcChecksumScrubPlan, EcLocalScrubPlan, merge_ec_runtimes,
};
let (store, _tmp) = build_split_disk_store(7030);
let vid = VolumeId(7030);
let runtimes = store.find_all_ec_volumes(vid);
assert_eq!(runtimes.len(), 2);
// Reachability is the invariant, so assert on the resolved slots rather
// than on scrub message text: shards 0 and 12 live on disk 0, shard 1 on
// disk 1. The old first-match lookup could never see shard 1.
let merged = merge_ec_runtimes(&runtimes).expect("two runtimes merge");
assert!(merged.slots[0].is_some(), "disk 0's shard 0 unreachable");
assert!(merged.slots[12].is_some(), "disk 0's shard 12 unreachable");
assert!(
merged.slots[1].is_some(),
"disk 1's shard 1 unreachable — the bug"
);
assert!(
merged.skipped.is_empty(),
"same generation: {:?}",
merged.skipped
);
// Shard 1 is owned by the sibling runtime, not the anchor.
let (owner, _) = merged.slots[1].unwrap();
assert!(std::ptr::eq(owner, runtimes[1]));
// Both plans build over the union rather than over disk 0 alone.
assert!(EcChecksumScrubPlan::for_volumes(&runtimes).is_some());
assert!(EcLocalScrubPlan::for_volumes(&runtimes).is_some());
// ...and `is_some()` is a real question: `for_volumes` has exactly one
// `None` (the vanished-volume case), so without this the two lines above
// would hold for any input at all.
assert!(EcChecksumScrubPlan::for_volumes(&[]).is_none());
assert!(EcLocalScrubPlan::for_volumes(&[]).is_none());
// Regression guard: a single-runtime view still sees only its own disk,
// which is exactly what made aggregation necessary.
let disk0 = merge_ec_runtimes(&[runtimes[0]]).unwrap();
assert!(
disk0.slots.get(1).copied().flatten().is_none(),
"disk 0's runtime must not see the sibling's shard"
);
}
/// `Store::unmount_ec_shards` used to return after the first
/// location with the vid, so a request to unmount a shard that
/// lives on a sibling disk became a silent no-op. After the fix,
@@ -1236,7 +1476,7 @@ mod tests {
let vid = VolumeId(7004);
let collection = "grafana-loki";
store.delete_ec_shards(vid, collection, &[1]);
store.delete_ec_shards(vid, collection, &[1]).unwrap();
// Shard 1 file is gone on disk 1.
let p1 = format!(
@@ -1267,17 +1507,22 @@ mod tests {
let (_ev, dirs) = store.collect_ec_shard_dirs(vid, max_shards).unwrap();
// Shards 0 and 12 → disk 0's directory.
assert_eq!(dirs[0].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(dirs[12].as_deref(), Some(store.locations[0].directory.as_str()));
assert_eq!(
dirs[0].as_deref(),
Some(store.locations[0].directory.as_str())
);
assert_eq!(
dirs[12].as_deref(),
Some(store.locations[0].directory.as_str())
);
// Shard 1 → disk 1's directory.
assert_eq!(dirs[1].as_deref(), Some(store.locations[1].directory.as_str()));
assert_eq!(
dirs[1].as_deref(),
Some(store.locations[1].directory.as_str())
);
// Unmounted shards → None.
for sid in [2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13] {
assert_eq!(
dirs[sid], None,
"shard {} unexpectedly reported a dir",
sid,
);
assert_eq!(dirs[sid], None, "shard {} unexpectedly reported a dir", sid,);
}
}
@@ -1586,11 +1831,7 @@ mod tests {
vec![0u8; 20],
)
.unwrap();
std::fs::write(
ec_dir.join(format!("{}_{}.ecj", collection, vid)),
b"",
)
.unwrap();
std::fs::write(ec_dir.join(format!("{}_{}.ecj", collection, vid)), b"").unwrap();
let mut store = Store::new(NeedleMapKind::InMemory);
store
@@ -1689,50 +1930,3 @@ mod tests {
assert!(std::path::Path::new(&format!("{}.ecx", ec_base)).exists());
}
}
/// Walk a disk's data directory and return the `.ec??` shard files
/// that are present on disk but not yet registered in the location's
/// `ec_volumes` map. Keyed by (collection, vid) so callers can match
/// each group against its `.ecx`-owning disk in one lookup. Zero-byte
/// shard files are ignored — same shape as `load_all_ec_shards`.
fn collect_orphan_ec_shards(
loc: &crate::storage::disk_location::DiskLocation,
_loc_idx: usize,
) -> HashMap<EcKey, Vec<(String, u32)>> {
let mut orphans: HashMap<EcKey, Vec<(String, u32)>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return orphans;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
let Some(shard_id) = is_ec_shard_extension(ext) else {
continue;
};
// Ignore zero-byte shards. Use the DirEntry's metadata so we
// don't pay a second stat syscall per file beyond what
// read_dir already returned.
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
// Skip shards that are already registered to an EcVolume.
if let Some(ecv) = loc.find_ec_volume(vid) {
if ecv.has_shard(shard_id as u8) {
continue;
}
}
let key = EcKey { collection, vid };
orphans.entry(key).or_default().push((name, shard_id));
}
orphans
}
+29
View File
@@ -221,6 +221,35 @@ mod tests {
use super::*;
use crate::storage::types::*;
/// Multi-byte input must be an error, not a panic: `to_digit` on the
/// leading characters rejects it before `chars[2]` is ever indexed.
#[test]
fn replica_placement_rejects_non_ascii_instead_of_panicking() {
for s in ["é", "0é", "é0", "🦀", "ééé"] {
assert!(
ReplicaPlacement::from_string(s).is_err(),
"non-ASCII replication {:?} must error",
s
);
}
}
/// The ASCII guard must not change any accepted input, including the
/// zero-padding shorthands.
#[test]
fn replica_placement_still_accepts_ascii_shorthands() {
assert_eq!(
ReplicaPlacement::from_string("1").unwrap(),
ReplicaPlacement::from_string("001").unwrap()
);
assert_eq!(
ReplicaPlacement::from_string("01").unwrap(),
ReplicaPlacement::from_string("001").unwrap()
);
let rp = ReplicaPlacement::from_string("010").unwrap();
assert_eq!(rp.diff_rack_count, 1);
}
#[test]
fn test_super_block_round_trip() {
let sb = SuperBlock {
+4 -24
View File
@@ -155,7 +155,7 @@ impl Size {
return 0;
}
if self.0 < 0 {
return (self.0 * -1) as u32;
return -self.0 as u32;
}
self.0 as u32
}
@@ -284,8 +284,9 @@ impl fmt::Display for Offset {
// DiskType
// ============================================================================
#[derive(Debug, Clone, PartialEq, Eq, Hash)]
#[derive(Debug, Clone, PartialEq, Eq, Hash, Default)]
pub enum DiskType {
#[default]
HardDrive,
Ssd,
Custom(String),
@@ -319,12 +320,6 @@ impl fmt::Display for DiskType {
}
}
impl Default for DiskType {
fn default() -> Self {
DiskType::HardDrive
}
}
// ============================================================================
// VolumeId
// ============================================================================
@@ -397,7 +392,7 @@ impl From<u8> for Version {
///
/// Fields are split into request-side options (set by the caller) and response-side
/// flags (set during the read to communicate status back).
#[derive(Debug, Clone)]
#[derive(Debug, Clone, Default)]
pub struct ReadOption {
// -- request --
/// If true, allow reading needles that have been soft-deleted.
@@ -423,21 +418,6 @@ pub struct ReadOption {
pub read_buffer_size: i32,
}
impl Default for ReadOption {
fn default() -> Self {
ReadOption {
read_deleted: false,
attempt_meta_only: false,
must_meta_only: false,
is_meta_only: false,
volume_revision: 0,
is_out_of_range: false,
has_slow_read: false,
read_buffer_size: 0,
}
}
}
// ============================================================================
// NeedleMapEntry helpers (for .idx file)
// ============================================================================
File diff suppressed because it is too large Load Diff
@@ -10,7 +10,7 @@ use crate::storage::needle::Needle;
use crate::storage::super_block::SuperBlock;
use crate::storage::types::*;
use crate::storage::volume::{
fsync_dir, needle_disk_end, scan_volume_file, Volume, VolumeError, VolumeFileVisitor,
Volume, VolumeError, VolumeFileVisitor, fsync_dir, needle_disk_end, scan_volume_file,
};
/// Writes one .idx row per .dat record, in .dat append order, which is the
@@ -105,11 +105,11 @@ impl Volume {
#[cfg(test)]
mod tests {
use crate::storage::needle::crc::CRC;
use crate::storage::needle::Needle;
use crate::storage::needle::crc::CRC;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::*;
use crate::storage::volume::Volume;
use crate::storage::volume::{Volume, VolumeSpec};
use std::fs;
use std::path::Path;
use tempfile::TempDir;
@@ -145,13 +145,9 @@ mod tests {
let mut v = Volume::new(
data,
old_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
for id in 1..=3 {
@@ -167,13 +163,9 @@ mod tests {
let reopened = Volume::new(
data,
new_idx,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
@@ -207,13 +199,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -233,13 +221,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
@@ -261,13 +245,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -290,13 +270,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
@@ -318,13 +294,9 @@ mod tests {
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
v.write_needle(&mut needle(1), true, false).unwrap();
@@ -356,13 +328,9 @@ mod tests {
let reopened = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap();
drop(reopened);
+61 -16
View File
@@ -9,10 +9,10 @@ use std::path::Path;
use tracing::info;
use crate::storage::idx;
use crate::storage::needle::needle::needle_body_length;
use crate::storage::needle::Needle;
use crate::storage::needle::needle::needle_body_length;
use crate::storage::types::*;
use crate::storage::volume::{fsync_dir, Volume, VolumeError};
use crate::storage::volume::{Volume, VolumeError, fsync_dir};
/// Needles found in the head of .dat, keyed by id, plus the ids in .dat order.
type DatHeadNeedles = (HashMap<NeedleId, (Offset, Size)>, Vec<NeedleId>);
@@ -122,7 +122,19 @@ impl Volume {
} else {
found.remove(&id);
}
offset += NEEDLE_HEADER_SIZE as i64 + needle_body_length(size, version);
let record_size = NEEDLE_HEADER_SIZE as i64 + needle_body_length(size, version);
// A corrupt header can make the record length zero or negative;
// the scan cannot advance past it.
if record_size <= 0 {
return Err(VolumeError::Io(io::Error::new(
io::ErrorKind::InvalidData,
format!(
"corrupt needle header at offset {offset}: size {}, record length {record_size}",
size.0
),
)));
}
offset += record_size;
}
Ok((found, order))
@@ -215,20 +227,19 @@ mod tests {
use super::*;
use crate::storage::needle::crc::CRC;
use crate::storage::needle_map::NeedleMapKind;
use std::os::unix::fs::{FileExt, PermissionsExt};
use crate::storage::volume::VolumeSpec;
use std::io::{Seek, SeekFrom};
#[cfg(unix)]
use std::os::unix::fs::PermissionsExt;
use tempfile::TempDir;
fn open_volume(dir: &str) -> Volume {
Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
&VolumeSpec::default(),
)
.unwrap()
}
@@ -256,7 +267,7 @@ mod tests {
/// writes (key, offset 0, tombstone) rows over the front of .idx instead of
/// appending them.
fn clobber_idx_head(idx_path: &str, keys: &[u64]) {
let file = OpenOptions::new().write(true).open(idx_path).unwrap();
let mut file = OpenOptions::new().write(true).open(idx_path).unwrap();
for (i, key) in keys.iter().enumerate() {
let mut row = Vec::new();
idx::write_index_entry(
@@ -266,8 +277,13 @@ mod tests {
TOMBSTONE_FILE_SIZE,
)
.unwrap();
file.write_at(&row, (i * NEEDLE_MAP_ENTRY_SIZE) as u64)
// Positional write without Unix-only `FileExt::write_at`, so this
// helper (and the tests using it) also builds on Windows.
// Single-threaded test helper: no concurrent reader can move the
// offset between seek and write.
file.seek(SeekFrom::Start((i * NEEDLE_MAP_ENTRY_SIZE) as u64))
.unwrap();
file.write_all(&row).unwrap();
}
}
@@ -305,6 +321,7 @@ mod tests {
let size_before = idx_size(&idx_path);
// The rewrite replaces .idx wholesale, so it must not widen the mode.
#[cfg(unix)]
fs::set_permissions(&idx_path, fs::Permissions::from_mode(0o600)).unwrap();
// Deletes against needles 9..12 land on the front of .idx and take the
@@ -324,11 +341,14 @@ mod tests {
let want = size_before + 4 * NEEDLE_MAP_ENTRY_SIZE as u64;
assert_eq!(idx_size(&idx_path), want, "idx size after recovery");
assert_eq!(
fs::metadata(&idx_path).unwrap().permissions().mode() & 0o777,
0o600,
"idx mode after recovery"
);
#[cfg(unix)]
{
assert_eq!(
fs::metadata(&idx_path).unwrap().permissions().mode() & 0o777,
0o600,
"idx mode after recovery"
);
}
// The recovered rows go back in front, so .idx is in .dat append order
// again: the fingerprint is gone and the last row is still the .dat tail.
@@ -418,4 +438,29 @@ mod tests {
"needle 1 should have been recovered"
);
}
#[test]
fn test_scan_dat_head_fails_at_a_header_it_cannot_advance_past() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
write_test_volume(dir, 2);
let mut dat = OpenOptions::new()
.append(true)
.open(format!("{}/1.dat", dir))
.unwrap();
let mut corrupt = [0u8; NEEDLE_HEADER_SIZE];
NeedleId(99).to_bytes(&mut corrupt[4..12]);
Size(-100).to_bytes(&mut corrupt[12..16]);
dat.write_all(&corrupt).unwrap();
drop(dat);
let v = open_volume(dir);
let first = v.super_block.block_size() as i64;
let err = v.scan_dat_head(v.version(), first, 10).unwrap_err();
assert!(
matches!(&err, VolumeError::Io(e) if e.kind() == io::ErrorKind::InvalidData),
"expected a corrupt-data error, got {err:?}"
);
}
}

Some files were not shown because too many files have changed in this diff Show More