mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-01 04:05:54 +00:00
8fdcf69eb0ed733c2738168c745ebb2df7b44b69
15325
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8fdcf69eb0 |
s3api: report the stored checksum in GetObjectAttributes (#11529)
GetObjectAttributes accepted the Checksum attribute but never filled it in, as its comment said SeaweedFS did not store S3 checksums. PutObject and CompleteMultipartUpload store them now, and HeadObject returns them. Fill in Checksum from the same entry fields, with the ChecksumType and ChecksumCRC64NVME members the response did not have. Also run ceph/s3-tests' test_get_checksum_object_attributes in CI. |
||
|
|
988fc4f7ba |
s3api: do not store aws-chunked in an object's Content-Encoding (#11528)
* s3api: do not store aws-chunked in an object's Content-Encoding aws-chunked in Content-Encoding names the SigV4 streaming framing of the request body, which the gateway decodes on upload. PutObject and CreateMultipartUpload stored the header as sent, so an object uploaded with "gzip, aws-chunked" was served with that Content-Encoding, and one uploaded with "aws-chunked" alone was served as aws-chunked. S3 drops aws-chunked and keeps the other encodings. Also run ceph/s3-tests' test_object_content_encoding_aws_chunked in CI. * s3api: read every Content-Encoding field, and drop aws-chunked on copy A client can send aws-chunked and the object's own encoding as separate Content-Encoding fields. Only the first was read, so "aws-chunked" followed by "gzip" left the object without its gzip. Combine all the fields before dropping aws-chunked. CopyObject with the REPLACE directive stored the requested Content-Encoding as sent: drop aws-chunked there too. |
||
|
|
11e8c4c288 |
master: follow heartbeat read-only changes in the layout's replica flag (#11527)
A replica's read-only flag in the volume layout only moved on registration and on volume.mark. A change that arrived in the regular heartbeat updated the node's record, which the writable list follows, but not the layout flag, which the vacuum sweep reads. So the sweep kept trying volumes on a disk that had gone read-only while the server ran, and after a restart it skipped volumes that had since become writable again until the next restart (issue #11516). Apply the reported state to the flag for every changed volume. Only the flag: the writable list stays with EnsureCorrectWritables and its capacity guards. |
||
|
|
d3cd061c22 |
shell: say which read-only volumes volume.vacuum leaves alone (#11525)
* shell: say which read-only volumes volume.vacuum leaves alone volume.vacuum without -volumeId runs the same sweep as the automatic vacuum, which skips read-only volumes, and the master's response carries no result. An operator whose disk filled up runs the command, sees it return, and watches nothing change (issue #11516). Before issuing the request, list the read-only volumes whose garbage is at or above the threshold and point at -volumeId, which is the explicit path PR #9861 opened for them. The help text says the same. * shell: volume.vacuum hint survives a failed listing and looks at every replica Review follow-ups: a failed topology listing no longer stops a sweep without -volumeId, it only drops the hint; a volume counts as read-only when any replica is, with the garbage ratio taken from the replica that reports the most, which is what the sweep itself does; a converted index that reports deletes without sizes is listed rather than hidden; and the threshold is printed as given instead of rounded to two decimals. * shell: do not guess a garbage ratio for a converted index The master cannot compute one for a volume that reports deletes without their sizes, and a guess of 1 would send the operator to -volumeId for a volume the server may decline at that threshold. Leave it out and say so. |
||
|
|
895d49b55b |
s3/iam: manage roles through the IAM API, with an opt-in persistent role store (#11522)
* s3/iam: manage roles through the IAM API, with an opt-in persistent role store
Roles could only come from the IAM config file: the S3 server pinned the
role store to memory and the embedded IAM API had no role actions, so a
role could not be created, retrusted or revoked without editing the file
and restarting every gateway.
Role store
- Read the `roleStore` key (the IAMConfig field already existed). With an
IAM config file the default stays memory; with none it is the filer, as
for OIDC providers, so zero-config clusters keep runtime-created roles.
- Roles from the IAM config file never go into a persistent role store,
which outlives the file and may be shared by S3 servers with different
files. They are served from memory beneath the store, as OIDC providers
are: a stored role of the same name takes precedence, and deleting it
restores the file's. A config-file role cannot be changed or deleted
through the API (UnmodifiableEntity), and removing one from the file
removes it at the next start. An in-memory store holds them as records,
as before. They have no creation time, so CreateDate is omitted rather
than reporting when this server started. SetRoleStore installs a store
the same way, so a store set after startup keeps the config-file roles,
as SetOIDCProviderStore does for providers.
- Watch /etc/iam/roles and drop the cached role definitions on change. The
cached filer store otherwise serves a peer's stale role for up to its 5m
TTL, which keeps a revoked trust policy in force on the other gateways.
- Role stores wrap ErrRoleNotFound for a missing role; the filer store
used to report any failed lookup as "role not found". CreateRole proceeds
only on a confirmed absence, so an unreadable store cannot let it write
over an existing role.
IAM actions
- CreateRole, GetRole, ListRoles, DeleteRole, UpdateAssumeRolePolicy,
AttachRolePolicy, DetachRolePolicy, ListAttachedRolePolicies. The reads
are allowed in read-only mode.
- A role defined in the config file is reloaded from it at every start, so
changing or deleting it through the API is refused (UnmodifiableEntity)
rather than silently reverted.
- DeleteRole with policies attached is refused (DeleteConflict), as on AWS.
- Role names follow AWS's rules ([\w+=,.@-]{1,64}); a role is stored as
<name>.json in the filer, so this also keeps a name from leaving the role
store's directory. At most 10 managed policies per role (AWS's default
quota; MaxManagedPoliciesPerUser is 10 too), LimitExceeded beyond.
- DeletePolicy is refused (DeleteConflict) while a role attaches the
policy, as it already is for users and groups: roles attach policies by
name, so a policy created later under the deleted one's name would
otherwise take effect on the role.
- Role paths other than "/" and role tags are not stored, so they are
refused rather than dropped.
Role IDs and sessions
- Roles get a unique RoleId when first stored (random, AWS AROA form),
kept across updates; a config-file role gets a stable ID derived from its
name, since it is created again at every start.
- Sessions issued through AssumeRoleWithWebIdentity, AssumeRoleWithCredentials
and AssumeRole carry the role's ID (claim "rid"), and a request under a role
whose current ID differs is denied. Resolving a session's policies by role
name let a session outlive its role: once a role was deleted, a role later
created under the same name — with a different trust policy and different
policies — revived every unexpired session of the old one with the new
role's permissions. Sessions issued before this change carry no ID and are
unaffected until they expire.
Integration test (test/s3/iam, run with `make start-services`):
TestWebIdentityWithProviderAndRoleManagedThroughIAMAPI configures an OIDC
provider, a managed policy and a role entirely through the IAM API against a
JWKS served by the test, then checks the trusted subject gets credentials
scoped to the attached policy; another subject, a token signed by another
key, an unsigned token and a token for another audience are refused; and UpdateAssumeRolePolicy moves the
trust at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: bind every role session to its role, and change roles atomically
Review follow-ups.
Session binding
- The role-ID check ran only when a session carried no policy names, and
AssumeRole embeds the role's attached policies, so those sessions kept
their permissions after the role was deleted or recreated. The check
now runs for every session carrying a role ID, before policy selection.
- A named role that cannot be resolved at issuance gets no session,
instead of one with no role ID (which nothing binds).
- A config-file role's ID is derived from its name and trust policy, not
the name alone: a different role put in the file under the same name
gets a new ID, while an unchanged role keeps its sessions across restarts.
Role writes
- RoleStore gains UpdateRole, a read-modify-write that lands only if the
role is unchanged since the read, and otherwise re-reads and retries. The
filer store uses the filer's write conditions (IF_NOT_EXISTS for a new
role, IF_ENTRY_EQUAL otherwise). CreateRole, UpdateAssumeRolePolicy and
Attach/DetachRolePolicy all go through it, so two gateways no longer
overwrite each other's changes, a change racing a delete no longer
writes the role back, and of two concurrent creates one gets
EntityAlreadyExists.
- The filer store's ListRoles pages past 1,000 entries and fails on a
broken stream instead of returning what arrived, so DeletePolicy's
attachment check sees every role. ListRoles skips a role deleted between
listing and reading it.
- CreateRole validates first; a failed write is ServiceFailure, not
InvalidInput. Any Tags.* parameter is refused, not only the first key.
- ExecuteAction's skipPersist covers the S3ApiConfiguration only; the
comment now says so. Role and OIDC provider actions write their own stores.
Each fix has a test that fails without it. Against a real filer with two
gateways, concurrent AttachRolePolicy calls lost 1-4 of 8 attachments per
run before this change and none after.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: one role snapshot per decision; DeleteRole is atomic; watch a custom role store path
Review follow-ups.
- Authorization evaluates the policies of the role definition the session's
binding was checked against, instead of reading the role again: a role
replaced in between cannot lend a session its policies.
- AssumeRole and AssumeRoleWithLDAPIdentity issue the session from the
definition whose trust admits the caller (IAMManager.ResolveRoleForPrincipal),
and take its ID, duration cap and embedded policies from that same
definition. A role replaced after the caller's trust check by one that does
not trust the caller now yields AccessDenied, not a session bound to the
replacement.
- A RoleUpdate that returns nil deletes the role, on the same condition as a
write: the filer store deletes with ObjectTransaction on IF_ENTRY_EQUAL,
routed and locked like the conditional CreateEntry. DeleteRole decides
against the role it deletes, so a policy attached meanwhile on another
server is a DeleteConflict, and a delete never removes a role written
after its check.
- S3 servers watch the role store's configured basePath, not only
/etc/iam/roles, so a custom path also drops peers' cached roles on change.
Each has a test that fails without it. Live against a real filer: DeleteRole
refuses while a policy is attached and removes the entry once detached; all
test/s3/iam CI stages pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: state which roles DeletePolicy's attachment check can see
RolesAttachingPolicy sees the stored roles and this server's config-file
roles. A role defined only in another server's IAM config file is invisible
to it, so a config-file role that attaches a managed policy is protected
only on the servers whose file defines it. The doc comment now says so and
how to avoid it: keep such roles in every server's file, or attach only
config-file policies to config-file roles.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* iam: note that a role store set after startup is not watched for peer changes
S3 servers build their metadata watch list once, at startup, from the role
store installed then. SetRoleStore's doc now says that a filer-backed store
installed later with a different basePath is not watched, so peers' changes
to it reach this server's cached roles only when the cache expires.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
95e0b74fb6 |
s3/iam: retry a failed OIDC provider refresh until the store answers (#11521)
* s3/iam: retry a failed OIDC provider refresh until the store answers RefreshOIDCProvidersFromStore reports a failure and nothing retries it. Its callers can't: a metadata-subscription event reports each change once, so a refresh that found the filer unreachable on it (the filer restarting, say) left a peer's new provider untrusted, or a deleted one trusted, until some unrelated later change. The refresh after a local IAM API mutation has the same shape. Only the startup load retried. A failed refresh now retries in the background with the startup load's backoff until the store answers. At most one retry runs, however many refreshes fail meanwhile, and installing another store cancels it. The startup load uses the same path instead of its own. Seen on a SeaweedFS operator cluster whose filer restarted while an S3OIDCProvider was created: the gateway logged "OIDC provider refresh after /etc/iam/oidc-providers change failed: ... fail to dial". The operator's periodic re-apply happened to recover it; an IAM API client would not. * s3/iam: never retry or apply a superseded OIDC provider store, and never drop a failure during a retry Review of the retry (#11521) found two ways to lose the state it protects. A refresh of store A that failed as store B was installed could start a retry for A after B's install had cancelled retries. Nothing cancelled it, and when A answered it replaced B's providers in STS. The installed store now changes under the retry lock, a store that is no longer current gets no retry, and a snapshot of a replaced store is never handed to STS, even when the refresh listed it just before the swap. A refresh that failed while a retry ran was dropped by the at-most-one guard, though the retry might already have listed an older snapshot, so the change the failed refresh would have loaded stayed unloaded. The retry now runs once more after its success when a failure arrived meanwhile. Each has a test that fails without its guard. |
||
|
|
0978e7f833 |
vacuum: keep disk-full read-only volumes reclaimable (#11519)
* storage/topology: keep disk-full read-only volumes vacuumable The vacuum sweep skipped every read-only replica, so a volume that went read-only because its disk filled could never reclaim its garbage — the exact situation compaction exists for. The volume server now reports disk_space_low in VacuumVolumeCheckResponse, and the sweep skips a read-only replica only when the flag is clear. An explicit volumeId vacuum is unaffected: it already bypassed the read-only rule. The field takes number 4: 2 and 3 are downstream-allocated for tombstone retention, keeping the wire merge clean. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * storage: measure vacuum free space against live bytes The pre-compaction space check required the current .dat + .idx size free, which includes the garbage being reclaimed — on a nearly full disk that estimate can never fit, so the volume stayed garbage-bound forever. Measure against the estimated compacted output instead: superblock plus live index entries plus live content bytes, with the existing ten percent buffer unchanged. Mirrors the same check in the Rust volume server. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: count per-needle framing in the compacted-size estimate The live-bytes estimate covered each live needle's content and index entry but not its .dat framing (header, checksum, timestamp, padding — ~32 bytes on version 3). For small-needle volumes that is more than the 10% headroom, so a disk with space between the estimate and the real output still ran out mid-compaction. Rust side mirrors the same formula. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * storage: report disk_space_low only when it is the sole read-only cause Review feedback (ihnokim, greptile, devin): a volume read-only for low disk space AND an operator mark or I/O quarantine was still eligible for the automatic sweep, rewriting a copy meant to stay protected. The flag now reports only the benign sole-cause case in both servers. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * topology: fail closed when the read-only lookup misses in the sweep A heartbeat can drop the volume from the DataNode cache between the location-list copy and VacuumVolumeCheck; a lookup error previously skipped the read-only check entirely. Review feedback (coderabbit). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
38ce95d960 |
s3api: always write XML timestamps with three fractional digits (#11520)
CopyObject responses carried LastModified values such as
"2026-09-29T20:30:04.56Z": trailing zeros of the fractional seconds were
trimmed, and a whole-second value had no fraction at all. AWS S3 always
writes exactly three digits ("...04.560Z"), and clients that parse with a
fixed-width pattern reject anything else. minio-java 8.6.0
(yyyy-MM-dd'T'HH:mm:ss.SSS'Z') throws DateTimeParseException, so roughly
one CopyObject in ten fails on the client even though the copy succeeded.
Two causes:
- xsdDateTime marshalled with "2006-01-02T15:04:05.999999999", which
drops trailing zeros. It now writes UTC with ".000Z".
- CopyObjectResult.MarshalXML had a pointer receiver, but the handlers
pass the result by value, so encoding/xml never called it and fell back
to time.Time's RFC 3339 encoding. It now has a value receiver.
CopyPartResult had no custom marshaller at all; it now uses xsdDateTime.
Follow-up to #8394 / #8398, which truncated these timestamps to
milliseconds but kept the trimmed format.
|
||
|
|
9a454bbe2b | docs: regenerate star history chart | ||
|
|
757917f564 |
filer: evict remote-cached objects under storage pressure (#11515)
* filer: identify remote-mounted entries safe to drop under disk pressure ListEvictableRemoteEntries walks every mounted directory directly on the filer store (no lazy remote listing) and returns entries that hold local chunks fully synchronized with remote, ordered oldest-cached first. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: evict remote-cached chunks oldest-first and vacuum the garbage uncacheRemoteEntry applies the same transition remote.uncache does - cleared chunks plus a reset LastLocalSyncTsNs under the entry path lock - and evictRemoteCachedEntries serializes passes over all mounts until a byte target is met. Aged victims are preferred; a second pass accepts any synchronized cached entry when aged ones cannot cover the request, since a failed read is worse than a dropped hot object. Cleared chunks only become disk space after compaction, so reclaimRemoteCacheSpace pairs each pass with a rate-limited VacuumVolume call that also picks up orphaned partial fills. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: trigger remote cache eviction under storage pressure A periodic check (30s) reads disk usage from master topology and evicts remote-mounted cached chunks once any disk crosses -filer.remoteCacheEvictThreshold (default 0.9; 0 disables), with a vacuum pass to reclaim the tombstoned needles. The cold-read cache path also kicks the same reclaim when a fill fails on exhausted volumes - the request still falls back to streaming from the remote, but the cache stops being permanently wedged full. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: flush deletion queue before remote cache vacuum Vacuum ran immediately after eviction while evicted file IDs still sat in the asynchronous deletion queue, so compaction saw no garbage and the cache stayed wedged. Flush the queue synchronously first and shorten the vacuum cooldown so sustained pressure does not wait five minutes between reclaim passes. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: cover remote cache eviction under capacity pressure Unit tests pin the eligibility filter and oldest-first ordering; the integration test runs a constrained two-node setup that saturates the cache, verifies the oldest synced entry is evicted and vacuumed, and that a later read re-caches it. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: coalesce remote cache reclaim passes A failed cache fill used to queue behind any in-flight eviction, stacking full mount traversals during a write-failure storm. Skip the pass when one is already running; the caller falls back to streaming from remote regardless. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: stop the remote cache janitor on shutdown The eviction ticker kept running after Shutdown closed the metadata store and could traverse a closed store. Give the janitor a context cancelled from Shutdown and propagate it into its master RPCs and traversals. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: vacuum only tombstoned volumes and retry deferred passes VacuumVolume with no volume id swept every collection, compacting volumes unrelated to the cache fill that failed. Now the reclaim path collects the vids of file ids actually flushed from the deletion queue and compacts only those. Vids that land inside the vacuum cooldown stay in a pending set the janitor retries on each tick, so chunks evicted just after a sweep are not stranded until the next pressure event. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: count only pressured disks when evicting remote cache The janitor measured the largest excess on one disk but let bytes on healthy disks satisfy the reclaim target. Split the topology disk view per physical disk and count only chunk bytes whose volumes sit on an over-threshold disk; entries contributing nothing there are skipped. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: compare remote cache sync time at nanosecond precision Second-precision mtime comparisons let a local write in the same second as the last sync still qualify as evictable, discarding unsynced changes. Compare LastLocalSyncTsNs against full-precision mtime (mtime_ns round-trips through the entry codec), and apply the same fix to remote.uncache's inline check. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: invalidate remote sync stamp on local content change A local overwrite that keeps the remote entry's LastLocalSyncTsNs looks evictable even though the remote copy no longer matches, and some write paths stamp mtime at second precision so a timestamp comparison cannot catch it. UpdateEntry now clears the stamp when chunks change without a fresh stamp, leaving replicated updates authoritative. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: bound remote cache master rpcs and vacuum all evicted garbage VolumeList and VacuumVolume now run under a 30s context so a stalled master cannot wedge the eviction janitor. The targeted vacuum drops the garbage threshold so volumes with under 10% deleted bytes still compact. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: tolerate straggler fills in remote cache eviction test Detached fills from the concurrent wave keep racing the final checks: live chunks legitimately fill both volumes, and a re-cached object can be evicted again before its commit is observed. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: start remote cache eviction loop after filer init The janitor's first tick dereferences fs.filer; starting the goroutine before NewFiler assigns it could panic when startup exceeds an interval. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: keep remote cache vacuum intent across retries Evicted entries now record their chunk volumes for vacuum directly, so the intent survives whoever consumes the shared deletion queue first. A pending volume keeps several vacuum attempts so tombstones that land late are still compacted, and the janitor retries pending volumes under the reclaim mutex instead of flushing unrelated deletes every tick. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: bound each remote cache vacuum request independently A shared 30s deadline across pending volumes let one slow compaction cancel the rest. Each VacuumVolume now gets its own context, and pending volumes keep more attempts since the master reports request acceptance rather than compaction. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: tighten remote cache reclamation bound Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: treat chunk timestamp changes as content changes chunksEqual now also compares ModifiedTsNs so an update that rewrites a chunk record still invalidates the remote sync stamp. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: run remote cache queue flush under the reclaim context BatchDelete for flushed file ids now uses the caller's context instead of context.Background(), so a reclaim pass bounded by shutdown or timeout stops its deletes too. Other callers keep their existing behavior. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: scope remote cache vacuum to evicted volumes The flush no longer feeds the shared deletion queue's ids into the pending set — only evicted chunks' volumes are tracked, so ordinary deletions no longer pick up repeated vacuum attempts. The flush also runs under a shutdown-immune bounded context and is skipped when no volume is pending. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: retry remote cache vacuums even after unmount Pending volumes were only retried while a remote mount existed; removing the last mount skipped every later pass and left evicted bytes allocated. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: reclaim partial cache fills that run out of capacity A fill that fails midway queues its written chunks for deletion, but when no entries remain evictable the reclaim pass found no pending volumes and skipped the flush and vacuum entirely, leaving the partial garbage to the slow periodic vacuum while the disk stayed full. Mark the failed fill's chunk volumes pending so the pass tombstones and compacts them even when nothing was evicted. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a033ab76e5 | helm: add base dns names to admin cert (#11518) | ||
|
|
5fcf12c080 |
docker: publish release and latest images for linux/ppc64le and linux/s390x (#11517)
Add ppc64le and s390x to the normal and large_disk release image matrices, next to arm/v7 and 386, and to the manual container_latest fallback so that rebuilding `latest` does not drop them from the manifest. The Go binary is cross-compiled without cgo on the build host, and Dockerfile.go_build already falls back to empty Rust volume server and worker placeholders on platforms without pre-built Rust binaries. Only the small final-stage RUN steps run emulated, so these two platforms use QEMU on ubuntu-latest like arm/v7. Refs #5558 |
||
|
|
5da137233d |
s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart (#11510)
* s3/iam: persist IAM-managed OIDC providers in the filer, and trust them after a restart
The S3 server's IAM config loader never read the documented
`oidcProviderStore` key, so the OIDC provider store was always in memory:
a provider created with CreateOpenIDConnectProvider lived in one gateway's
process, was lost on restart, and was never seen by peers. The
/etc/iam/oidc-providers metadata subscription refreshed from that empty
in-memory store.
- Read `oidcProviderStore` and pass it to the IAM manager. With an IAM
config file the default stays memory. With no config file (zero-config
IAM, as `weed filer -s3` and operator-managed clusters run) it defaults
to the filer: there is nothing static to shadow, and providers created at
runtime otherwise vanish on restart.
- With a store that outlives the process, load the STS runtime view from it
at startup, so providers created on an earlier boot or on a peer are
trusted without waiting for the next mutation.
- If the store cannot be read at startup (a filer not up yet), the load is
retried in the background with backoff until it succeeds: the metadata
subscription reports only later changes, so providers already stored would
otherwise stay unknown to STS until one of them changed.
- Mark records mirrored from STS.Providers as `source: static-config`, and
at startup delete such records whose provider has left the config, so
removing a provider from the config file still revokes it. Records created
through the IAM API are never pruned.
- The filer store reported every failed lookup, an unreachable filer
included, as ErrOIDCProviderNotFound, which CreateOIDCProvider reads as
"free to create". Only a confirmed absence is now not-found.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: keep config-file OIDC providers out of a persistent store
Review of the previous commit found that mirroring the IAM config file's
providers into a persistent store, and pruning them when they leave the
file, breaks as soon as S3 servers share a filer:
- a server prunes stored config-file providers its own file does not list,
including ones a peer's file still defines (a zero-config server prunes
them all);
- mirroring overwrites an API-created provider with the same ARN and marks
it config-owned, so a later prune deletes it;
- a failed mirror write or a failed prune leaves a stale record trusted;
- a mirrored record is loaded into STS at startup as an IAM-managed provider
and shadows the config-file provider, dropping the settings a record does
not carry (jwksUri, roleMapping, policyClaim, ...).
A persistent store now never receives the config file's providers. STS keeps
serving them from its static configuration, as it always has; the IAM API
lists and returns them from memory, refuses to change or delete them
(UnmodifiableEntity; change them in the file) and to create another provider
with their ARN (EntityAlreadyExists). The store holds only providers created
through the IAM API, and those are what startup loads into STS. There is
nothing to prune, so the source marker is gone. An in-memory store keeps its
behaviour: the config file's providers are records in it, as before.
buildOIDCProviderFromRecord also carries PolicyClaim and
AllowedPrincipalTagKeys now; they were dropped whenever an API-created
provider was loaded into STS.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: send UnmodifiableEntity as a 400, not an internal error
The IAM API's error writer had no case for UnmodifiableEntity, which the
previous commit returns for a change to a config-file provider, so it went
out as a 500 ServiceFailure that clients retry. AWS sends it as a 400.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* s3/iam: document stored-over-config precedence, drop invented CreateDate, cancel superseded retries
Follow-ups from review of
|
||
|
|
62d4f9152a |
iam: evaluate trust policies deny-by-default (#11513)
* iam: evaluate trust policies deny-by-default EvaluateTrustPolicy seeded its result with the engine's DefaultEffect, so a non-matching trust-policy statement set still resolved to Allow when the IAM config sets policy.defaultEffect=Allow. A caller holding a validly signed token from a registered provider could then assume a role its trust policy does not admit. Trust policies now start from implicit deny, matching AWS semantics and the pre-d751623 behavior of evaluateTrustPolicy; DefaultEffect still governs identity-policy evaluation. Upgrade note: deployments on defaultEffect=Allow whose trust policies do not match their callers will see those assumptions refused. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: cover trust policy implicit deny under DefaultEffect=Allow Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
a901c1a5e2 |
filer: compare IF_ENTRY_EQUAL chunks by fid, not file_id (#11514)
The stored entry came through FindEntry, which restores chunk file ids from their fid form, while an expected entry built from a metadata-log event still carries the serialized form (file_id moved into fid). The proto.Equal saw file_id "" against the restored id and refused every stamp, so remote.sync re-uploaded each entry and the RemoteEntry stamp never landed. Clone both sides and run BeforeEntrySerialization before comparing, so chunks match on their fid and the file_id spelling is ignored; the stored entry and the request's ExpectedEntry are left untouched. Generated with [Devin](https://devin.ai) Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
67b19b2db5 |
build(deps): bump com.fasterxml.jackson.core:jackson-databind from 2.22.1 to 2.22.2 in /test/java/spark (#11511)
build(deps): bump com.fasterxml.jackson.core:jackson-databind Bumps [com.fasterxml.jackson.core:jackson-databind](https://github.com/FasterXML/jackson-databind) from 2.22.1 to 2.22.2. - [Commits](https://github.com/FasterXML/jackson-databind/compare/jackson-databind-2.22.1...jackson-databind-2.22.2) --- updated-dependencies: - dependency-name: com.fasterxml.jackson.core:jackson-databind dependency-version: 2.22.2 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c54f1fe69c |
build(deps): bump github/codeql-action from 4.38.1 to 4.38.2 (#11509)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.1 to 4.38.2. - [Release notes](https://github.com/github/codeql-action/releases) - [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/github/codeql-action/compare/v4.38.1...v4.38.2) --- updated-dependencies: - dependency-name: github/codeql-action dependency-version: 4.38.2 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
818ff69270 |
build(deps): bump cloud.google.com/go/kms from 1.34.0 to 1.35.0 (#11508)
--- updated-dependencies: - dependency-name: cloud.google.com/go/kms dependency-version: 1.35.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
db89ed198b |
build(deps): bump github.com/twmb/avro from 1.8.0 to 1.9.0 (#11507)
Bumps [github.com/twmb/avro](https://github.com/twmb/avro) from 1.8.0 to 1.9.0. - [Commits](https://github.com/twmb/avro/compare/v1.8.0...v1.9.0) --- updated-dependencies: - dependency-name: github.com/twmb/avro dependency-version: 1.9.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f7b5ca826b |
build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity from 1.14.0 to 1.14.1 (#11506)
build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/azidentity Bumps [github.com/Azure/azure-sdk-for-go/sdk/azidentity](https://github.com/Azure/azure-sdk-for-go) from 1.14.0 to 1.14.1. - [Release notes](https://github.com/Azure/azure-sdk-for-go/releases) - [Commits](https://github.com/Azure/azure-sdk-for-go/compare/sdk/azcore/v1.14.0...sdk/azidentity/v1.14.1) --- updated-dependencies: - dependency-name: github.com/Azure/azure-sdk-for-go/sdk/azidentity dependency-version: 1.14.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
878d23b95a |
build(deps): bump github.com/dustin/go-humanize from 1.0.1 to 1.1.0 (#11505)
Bumps [github.com/dustin/go-humanize](https://github.com/dustin/go-humanize) from 1.0.1 to 1.1.0. - [Commits](https://github.com/dustin/go-humanize/compare/v1.0.1...v1.1.0) --- updated-dependencies: - dependency-name: github.com/dustin/go-humanize dependency-version: 1.1.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1e0b243aed |
build(deps): bump golang.org/x/oauth2 from 0.36.0 to 0.37.0 (#11504)
Bumps [golang.org/x/oauth2](https://github.com/golang/oauth2) from 0.36.0 to 0.37.0. - [Commits](https://github.com/golang/oauth2/compare/v0.36.0...v0.37.0) --- updated-dependencies: - dependency-name: golang.org/x/oauth2 dependency-version: 0.37.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
79e2abdfbf | docs: regenerate star history chart | ||
|
|
530be3e373 | 4.48 4.48 | ||
|
|
9b3b12c607 |
filer: pin-aware reader cache eviction and stream release (#11503)
* filer: synchronize stream pins and release them on transitions Guard chunkStream.cacher with the ReaderCache lock everywhere: mount sections share one ChunkReadAt across concurrent reads, and unsynchronized release could double-unpin. Reads served from the chunk cache now detach the stream's pin instead of retaining the previous chunk. Eviction prefers unpinned downloaders so a pinned buffer is not dropped mid-stream. A new ReleaseStream lets callers drop their pin without destroying the shared cache; S3 and WebDAV readers use it. lastChunkFid becomes atomic since concurrent mount reads can update it. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: keep eviction bounded when every downloader is pinned Both eviction paths still fall back to a pinned victim when no unpinned one exists, so abandoned stream pins cannot bypass the downloader limit or stall the memory budget. Budget eviction also rechecks the pin under the ReaderCache lock at removal time: a stream that pinned the selected victim in between keeps it mapped and the selection retries. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: restore budget bookkeeping when a victim gets pinned mid-eviction removeUnpinned losing the pin race left the victim out of the idle list while still holding its reservation, making it unevictable even as the pinned fallback. Push it back when the reservation is still live. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
3abdef3202 |
filer: keep shared chunk buffers pinned while another stream reads them (#11502)
The ReaderCache is shared by all streams of a process (every S3 GET, for instance), but a ChunkReadAt released chunks as if it owned them: - moving on to the next chunk called UnCache on the previous one, destroying the buffer even when other streams were still inside it; - since #11384 a buffer is dropped once any reader has consumed it to the end and no read call is in flight. Streams copy out in slices (256 KiB in the S3 gateway), so between two calls a slower stream is not attached and loses the buffer to a faster one. Either way the slower stream refetches the whole chunk from the volume servers. With many clients downloading the same popular object at once, each chunk is fetched over and over; in production we saw the S3 gateway pull ~10 Gbit/s from volume servers while serving ~1 Gbit/s to clients. A ChunkReadAt now pins the chunk it is positioned in. The pin is taken and released only under the ReaderCache lock, since concurrent ReadAt calls on one ChunkReadAt (as in mount) share it. It is released when the stream reads the chunk to its end, moves to another chunk (including one served from the chunk cache), or falls back to random reads. A buffer is dropped once no stream pins it and no read is in progress, if it was consumed or its last stream left it; a read still in flight when the stream leaves drops it on detach, as UnCache did via destroy. Eviction by slot limit and memory budget is unchanged. lastChunkFid is now guarded as well: concurrent ReadAt calls raced on it. Tests: two ChunkReadAt instances streaming one object in interleaved slices fetch each chunk exactly once (2-3 times before); leaving a chunk for a chunk-cache hit or while another read is in flight releases it; concurrent ReadAt calls on one ChunkReadAt leave no pins behind under -race. |
||
|
|
43fd5b8d82 |
volume: reclaim staged EC shard generations left by the 2PC switch (#11501)
* volume: remove staged EC generation files on teardown and shard delete The 2PC generation switch stages each run as <base>.ecNN.v<N> plus versioned .ecx/.ecj/.vif files. Nothing on the volume server removes them: isEcDataShardFile only recognises the exact .ecNN name, so the staged files are invisible to every bookkeeping pass, and even full_teardown's wipe-all path left them behind. Each re-encode therefore leaks a full shard set per shard-holding disk. RemoveEcGenerationFiles sweeps <base>.ec*.v<N> and <base>.vif.v<N>, optionally keeping generations at or above a threshold; teardown and the reconcile wipe remove every generation, and a per-shard delete removes that shard's staged generations too. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: delete staged EC generations older than N via VolumeEcShardsDelete After a 2PC generation switch commits, the superseded generation's <base>.*.v<N> files sit on disk with no cleanup path: teardown removes everything, and a per-shard delete only touches the named shards, so the executor had no RPC that reclaims just the staged leftovers. delete_generations_older_than removes staged generation files strictly below the threshold on every disk. Versioned files are never mounted, so nothing is unloaded first; the committed generation and the canonical files are preserved. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: mirror staged EC generation cleanup Parity with the Go volume server: remove_ec_generation_files sweeps <base>.ec*.v<N> and <base>.vif.v<N> staged by the 2PC switch, called by remove_ec_volume_files (which covers both teardown paths) and the new delete_generations_older_than request field; delete_ec_shards removes a shard's staged generations along with the canonical file. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: match staged generation filenames literally filepath.Glob interprets metacharacters in the collection part of the base name, so a collection like a[bc] could match another volume's staged files (or miss its own). Scan the directory and compare names literally instead, mirroring the Rust read_dir implementation. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: report generation-sweep errors and drop the store lock first - snapshot the location base names under the read lock and run the filesystem sweep after dropping it, so a slow disk cannot stall the store; - record per-entry read_dir errors in remove_ec_generation_files and propagate them from remove_ec_shard_generations instead of flatten() skipping them; - warn when a staged-shard generation fails to delete rather than reporting success with files left behind. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume: fail shard delete when the staged-generation listing fails A transient ReadDir failure fell back to removing canonical shard names only: staged .v<N> files survived while the RPC still reported success, leaving the leak invisible to retrying callers. ENOENT still means the disk simply has no such directory; other listing errors now propagate. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: propagate staged-generation removal failures delete_ec_shards logged remove_ec_shard_generations errors and the RPC returned success while staged .v<N> files remained, diverging from the Go handler which surfaces the failure. The sweep keeps processing the remaining shards, retains the first error, and volume_ec_shards_delete maps it to Status::internal so callers can retry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: notify state change even when the shard sweep errors delete_ec_shards already deletes and unmounts the shards before returning a staged-generation failure, so returning early skipped volume_state_notify and the master kept routing to them until the next heartbeat. Notify before propagating the error. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
150a69fe11 |
master: make volume capacity reservation timeout configurable (#11426) (#11497)
* master: make volume capacity reservation timeout configurable (#11426) * master: expire reservations on reads, fix int timeout units - AvailableSpaceForReservation now expires reservations too: a node that is full of reservations is filtered out before TryReserveCapacity can clean them, which stranded expired capacity indefinitely. - Drop TryReserveCapacityWithTimeout: a per-call timeout lets one caller expire another's live reservations, and the Node interface stays stable for implementations outside this tree. - parseReservationTimeout no longer routes integer values through GetDuration, which read them as nanoseconds; bare numbers are seconds. The 5m fallback is now the shared DefaultReservationTimeout. --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
4fec65d949 |
filer: demote client-cancelled directory listing log from error (#11495) (#11496)
* filer: demote client-cancelled directory listing log from error (#11495) * filer: quote path in canceled listing log --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
f564918685 | docs: regenerate star history chart | ||
|
|
f9289f0570 |
s3: do not promote ?prefix into the object for non-List actions (#11494)
* s3: do not promote ?prefix into the object for non-List actions authRequestWithAuthType mapped an empty object to the prefix parameter for every action, so PUT /bucket?versioning&prefix=x authorized as Write:bucket/x. An object-scoped grant (Write:bucket/*) could then change bucket versioning, lifecycle, cors, and object-lock configuration, and the promoted object also made ResolveS3Action report s3:PutObject to attached IAM policies. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: treat GET ?uploads as a bucket listing for authorization Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: resolve the listing action through the bucket-level object resolveS3AuthTarget fed the promoted prefix to ResolveS3Action, so a bucket-level ?uploads request resolved as s3:GetObject on the prefix ARN in the admin explicit-deny check. Resolve both action and resource against the object the bucket listing actually scopes. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: resolve the listing action through the bucket-level object in AuthorizeAction Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: drop the unreachable object-level uploads case from the resolver test Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4303b3aa4c |
s3: keep a listing's start position inside the requested prefix (#11493)
* s3: a list marker that sorts past the prefix leaves nothing to list AWS scopes a listing to keys under Prefix; StartAfter, Marker and continuation tokens only reposition inside that range. A marker that diverges from the prefix at a larger byte is after every key the prefix can match, so the page is empty. normalizePrefixMarker used to keep such a marker as the walk cutoff at the bucket root, where the walk descends into the marker's own directory and returns keys the prefix never names. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: keep the listing variant's action when a prefix is promoted to object authRequestWithAuthType promotes ?prefix= into the object argument for the legacy CanDo path. ResolveS3Action treats a non-empty object as object-level, so a bucket-level ?versions or ?uploads request carrying a prefix missed its specific action and fell back to the base List action: an s3:ListBucket grant then covered s3:ListBucketVersions, and an explicit Deny on the specific action was skipped on the same path. Resolve the action against the same bucket-level object the resource ARN already uses. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: treat GET ?uploads as a bucket listing for authorization Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * Update weed/s3api/auth_credentials.go Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
a0ee7ba314 |
s3: ignore empty intermediate directories in bucketHasUserObjects (#11490) (#11491)
* s3: ignore empty intermediate directories in bucketHasUserObjects (#11490) * s3: keep nested reserved-named dirs from hiding user objects Reserved folders (.uploads, *.versions) are internal only at the bucket root; deeper entries with those names are user key prefixes and must be walked. Also treat a missing subdirectory as empty via isFilerNotFound (list errors cross gRPC as status errors, not the sentinel), let names containing backslashes count as objects, and walk iteratively so empty chains deeper than the old scan depth no longer report non-empty. * s3: treat reserved-named directories as internal at every level Object listing interprets .uploads and *.versions directories as internal storage wherever they appear, so walking them during the emptiness check would report invisible version remnants as user objects and block deletion. A reserved name on a file still counts, matching listing which only special-cases directories. * s3: count explicit directory objects under reserved names A directory object created by PutObject (MIME or prefix-object marker set) is user data even when named .uploads or *.versions; only a plain directory with a reserved name is internal storage. --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
a976b21010 |
s3: require dedicated object-lock permissions for x-amz-object-lock-* headers (#11492)
* s3: require dedicated object-lock permissions for x-amz-object-lock-* headers PutObject, CreateMultipartUpload, and PostPolicy honor the retention and legal-hold headers after only the route's s3:PutObject check, so a write-only principal could pin a version under COMPLIANCE retention that nobody can remove before its retain-until date. On AWS these headers require s3:PutObjectRetention / s3:PutObjectLegalHold. validateObjectLockHeaders is the shared funnel for all four call sites; it now authorizes the corresponding dedicated action when each header is present. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: record the verified POST-policy signer as the request identity The handler authenticated the form policy signature but stored only the signer's name, so downstream authorization (the object-lock header check) re-authenticated the form-signed request as anonymous and evaluated the wrong principal. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
e57f8c4d87 |
volume server: read GET/HEAD needles off the store lock, and only once (#11487)
The GET/HEAD handler read the needle synchronously on the tokio worker while holding store.read(): first a stream-info read that loaded the whole record just to parse its meta, then, for every needle that was not streamed (small, compressed, chunk manifest, image ops), a second full read. For a tiered volume each read is an S3 GET under the store lock, and a writer queued behind it parks every other store reader. The regular-volume read now runs in spawn_blocking. Under the store guard it only resolves a NeedleReadPlan (index lookup, a freshly opened .dat handle or the remote backend, offset, size); the guard is dropped before any needle data I/O. No data-file lease is held across the read either, since a writer waits for one while holding the store write lock. The index size decides the read, as in Go's readNeedle: a HEAD, a ranged read or a needle above the stream threshold reads only its header and meta tail (ReadNeedleMeta) and hands off to StreamingBody or the range path; everything else is read in full once, with its checksum verified. A compressed or manifest needle found by the meta read is then read in full once. The range-from-source read also moves to spawn_blocking. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
67691a1eea |
volume server: split volume_copy into phases and type the delete-after-status gate (#11485)
volume_copy was one ~400-line handler, and the rule that an existing local replica is deleted only after the source's ReadVolumeFileStatus succeeded was held by statement order alone. The keep_remote_data=true that the pre-copy delete and the failed-copy rollback must share was kept in sync by a comment pointing from one to the other. The handler is now a ~60-line orchestrator over connect_to_copy_source, SourceVolumeStatus::fetch, delete_existing_replica, plan_copy_destination and a VolumeCopyJob whose run() drives preallocate_dat, transfer_files, finish_copied_files and mount_and_reply, with cleanup_failed_copy on error. delete_existing_replica takes a &SourceVolumeStatus, which only fetch can construct (private field in a child module), so the delete cannot be called before the status RPC. Both deletes go through delete_replica_keep_remote. Pure refactor: call order, status codes and messages, cancellation checks, throttling, progress reports and cleanup are unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
02353444ac |
fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts (#11483)
* fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts Port of the Go VolumeCopy hardening in #11238 and #11252. - Pick the destination disk before deleting the existing replica, counting the slot that replica holds as free. If no disk qualifies, the healthy replica is kept instead of being deleted. - Read the source's VolumeStatus before and after the copy. When both succeed and the counts did not change, the mounted replica's file and deleted counts must match; on mismatch it is unmounted and its files removed. A failed "before" read skips the check; a failed "after" read fails the copy. * fix(volume-rust): let a departing caller cancel VolumeCopy's post-copy status read Go reads the source's status after the copy with stream.Context(), so the call ends when the caller leaves. The Rust call had no such link: a source that stalled there held the copied, unmounted files after the caller was gone. Race it against the response channel, like the other blocking steps, so the usual error cleanup removes the partial copy. |
||
|
|
68944e83a3 |
volume: typed tier errors so a missing remote object answers NotFound (#11484)
remote_storage/s3_tier.rs returned Result<_, String> from every
transfer (upload_file, download_file, read_range[_blocking],
delete_file[_blocking]) and from the tier runtime helpers. The tier
move handlers could only wrap that in Status::internal, so a .dat whose
remote object is gone was indistinguishable from an I/O failure to
weed shell.
Add TierError { NotFound, Io, RuntimeUnavailable, Aborted }. Each
variant carries the existing message verbatim. NotFound follows the
rules remote_storage/s3.rs already uses: raw 404 status on HEAD,
NoSuchKey code on GET; a bare 404 on GET stays Io. A progress-callback
Err becomes Aborted. VolumeError gains a transparent Tier variant and
From<VolumeError> for Status maps Tier(NotFound) to NotFound; the tier
move handlers go through status_with_context, so their message text is
unchanged. Every other tier failure is still Internal.
The remote needle read path keeps io::Error::other, so its error kind
and vacuum's handling of it do not change.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
00310f6588 |
volume server: run the vacuum compaction copy without the store lock (#11482)
* volume server: run the vacuum compaction copy without the store lock VacuumVolumeCompact held the store write lock for the whole live-needle copy, including every progress blocking_send on the 16-deep stream. On a large volume that is minutes with every read, write and heartbeat on the node parked behind it, long enough for the master to unregister the node. Split compaction the way Go's CompactByIndex runs it. A short locked step claims the volume's compacting flag, records the makeup_diff watermark (index size and compaction revision) and opens fresh .dat/.idx handles. The copy then replays .idx up to the watermark and copies from those handles with the store lock released; writes that land meanwhile are replayed by makeup_diff at commit, as before. The flag is an Arc<AtomicBool> released when the job is dropped, so every exit path clears it. Because the flag is now visible to other callers, the operations that would pull the files out from under the copy refuse while it is set: unmount (and VolumeConfigure, which unmounts and remounts), delete (checked before the volume is removed from the map, which a refused destroy used to leave unmounted), cleanup, and index relocation. A second compact and a commit stay no-ops, as in Go. The pre-copy fsync is dropped: the copy reads its own handles through the page cache and .cpd/.cpx are fsynced before commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: keep a read-only in-memory index's size for the compaction copy The unlocked copy replays .idx up to index_file_size(). A read-only volume whose .sdx could not be built loads its index into memory without a writer, so that size stayed 0: the copy came out empty and the commit replaced the volume with it. CompactNeedleMap::load_from_idx now records the rows it loaded, which is also what Go's IndexFileSize reports for a read-only index. The copy's index replay now stops reading at the recorded size instead of walking rows appended since, which makeup_diff replays anyway. Adds tests for compacting a read-only volume on both the sorted index and the in-memory fallback, and for VolumeConfigure stopping when the unmount is refused during a copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a vacuum copy as soon as its client is gone The progress callback only noticed a closed response stream when a report was due, every 128 MiB. With the copy now running outside the store lock, a copy nobody waits for keeps the volume marked compacting and so keeps refusing unmount, delete and cleanup until that next report. Check the stream on every callback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5c9c424a84 |
volume server: stream ReadAllNeedles without holding the store lock (#11481)
* volume server: stream ReadAllNeedles without holding the store lock read_all_needles held store.read() while Volume::read_all_needles read every live needle of the volume into a Vec, and kept holding it through the whole blocking_send loop. Memory grew with the volume, and a slow client parked the scan in a send with the guard held; needle writes and the heartbeat take store.write() on a writer-preferring lock, so the node stopped serving until the client caught up. Take a DatScanPlan (fresh .dat open, end bound) under a short guard and walk it with the guard released, sending one needle at a time. Each record is checked against the live needle map under a brief read guard, as the scan reaches it, and only a live record is parsed, so a damaged stale copy does not fail the stream. Records appended while a pass ran are walked by a follow-up plan, so a needle overwritten during the scan is streamed once, as its new copy. A vacuum commit or re-create of the volume during the scan fails the stream, since the map's offsets no longer describe the pinned file; the plan carries the volume instance and compaction revision for that check. DatScanPlan::scan_records yields records unparsed; scan keeps its behaviour on top of it. Volume::read_all_needles has no caller left and is removed; its tests move to the RPC. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: check ReadAllNeedles liveness only once the send can proceed The per-record liveness check ran before blocking_send, so a scan parked on a full channel held a record it had already judged live. An overwrite landing during that park left the old copy in the stream, and the continuation over appended records then streamed the new copy as well. Reserve channel space first, then take the store read guard, check the record against the needle map and enqueue it through the permit before releasing the guard. The wait for space still happens without the lock; the record is parsed before the guard is taken, and its parse error only counts if the record turns out to be live. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5218e68554 |
volume server: collect heartbeats under the store read lock, off the runtime (#11480)
* volume server: collect heartbeats under the store read lock, off the runtime Every pulse tick, options change and volume-state notification took store.write() for a whole heartbeat pass, directly on the async heartbeat task. The pass fstats every volume's .dat twice and hashes its report, so on a server with many volumes it held the store exclusively for the whole scan: reads and writes stalled, and with the writer-preferring RwLock a pending pass parked every new reader too. The pass only needs to mutate the store for a few rare actions: removing expired EC volumes, deleting expired volumes past their removal delay, and setting no-write on IO-quarantined volumes. It now runs under store.read(), records those as (disk, volume id) actions, and applies them afterwards under a short store.write() that is only taken when there is something to do. Each action re-checks its target under the write lock, so a volume written to, replaced or removed in between is left alone. Expired EC volumes are still removed before the volume pass, as before, because the EC shard count feeds the disk-space-low max volume count. Every pass runs on the blocking pool via spawn_blocking. The heartbeat message is unchanged for the same store state. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: take has_no_ec_shards with the heartbeat's volume list The heartbeat pass took has_no_ec_shards from the EC phase's read lock, then built the volume list under a second one. An EC shard mounted in between went out as "no EC shards" beside a volume list taken after the mount, and the master clears a server's EC registrations on that flag. has_no_ec_shards is now computed under the same read lock as the volume list, with the EC phase's filter: not expired, not quarantined, at least one shard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
00cffa028c | docs: regenerate star history chart | ||
|
|
a261f90e18 |
vacuum: let the sweep release volumes that stay empty and quiet (#11477)
* vacuum: let the sweep release volumes that stay empty and quiet Vacuuming reclaims bytes but not slots: a fully emptied volume stays registered to its collection forever, and since growth is gated only on slot count a store at 99% free disk can still refuse writes to other collections (#11429). volume.deleteEmpty exists but is manual-only. With -vacuumDeleteEmptyAfterSeconds (or master.vacuumDeleteEmptyAfterSeconds under weed server/mini; default 0, off) the automatic sweep now deletes replica copies that have stayed empty and quiet for that long, the same rule volume.deleteEmpty applies on demand: remote-backed copies are skipped, and every delete carries the volume server's onlyEmpty / onlyGarbage guards so a copy written since the last report is refused rather than removed. Copies that still hold data or were written recently stay; only a volume whose every copy is deleted leaves the sweep's work map, sparing a compaction of bytes that are all deleted. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: harden empty-volume sweep against partial and racing deletes Review follow-up on #11477: - delete a volume only when every replica copy is a verifiable empty-and-quiet candidate; deleting the empty copy of a volume whose sibling holds live files would silently cut its replica count (greptile P1). - drain the volume out of the writable list before deleting, the same drain the compact pass uses, so PickForWrite stops assigning it and pending writes settle (devin). - bound the VolumeDelete RPC so one stalled server cannot hold the vacuum lock indefinitely (greptile P1, reusing allocateVolumeTimeout). The vid2location panic scenario raised in review does not exist: VolumeLocationList methods are nil-receiver safe and a missing vid just fails enoughCopies, so a partially deleted volume skips compaction instead of crashing the sweep. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: unregister deleted empty replicas and prune the sweep list A successful VolumeDelete only updates the volume server; the master still tracked the replica and kept it in the sweep's location list for the compaction pass (coderabbit on #11477). Unregister the replica right after its delete succeeds and drop it from the sweep copy, so a partially deleted volume only compacts copies that still exist. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: pin deleting volumes out of the writable list across heartbeats Review follow-up on #11477 (greptile): DrainAndRemoveFromWritable only removed the volume once; a heartbeat landing between the drain and the replica deletes re-evaluated writability and re-added it, so a client write could reach a replica whose siblings were already gone and leave the volume under-replicated when the last copy refused its onlyEmpty delete. MarkDeleting records the vid in deletingVolumes — checked inside setVolumeWritable so heartbeat, capacity-recovery, and admin re-add paths all hold it out — and UnmarkDeleting releases it once the sweep finishes the copy pass. A partially deleted volume's surviving replicas then return to writable through the normal heartbeat path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: restore writability when a sweep delete survives Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
be29f44d87 |
s3: record requester identity before the authz verdict (#11479)
* s3: record requester identity before the authz verdict for audit Identity was only stored in request context on the success branch, so denied requests reached WriteErrorResponse without requester attribution and audit entries had empty requester/requester_arn/requester_identity. Authentication failures still resolve no identity, so unauthenticated denials stay unattributed. Fixes #11474 Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: keep the resolved identity through authz denial in Auth Review follow-up on #11479 (devin): authRequest discarded the identity on every error, so a request that authenticated fine but failed the action check still reached handleAuthResult with no identity and the deny path could not audit a requester. Auth now calls authRequestWithAuthType directly, the same entry AuthPostPolicy uses, so the resolved identity reaches the error writer; a failed authN still resolves no identity and stays unattributed. The regression test now signs a denied request end to end through iam.Auth. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2864bc0fe8 |
s3: honor configured session bounds on AssumeRole and LDAP identity (#11478)
* sts: export CalculateSessionDuration Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: honor configured session bounds on AssumeRole and LDAP identity prepareSTSCredentials hardcoded a one-hour session when the caller omitted DurationSeconds, so sts.tokenDuration was ignored and sts.maxSessionLength only clamped explicit requests: asking for 3600s against a 20m ceiling was rejected while omitting the parameter was granted a full hour (#11473). The two affected handlers now use the same default-then-cap calculation as AssumeRoleWithWebIdentity. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: keep MaxSessionDuration through role store copies copyRoleDefinition rebuilt RoleDefinition field by field and dropped MaxSessionDuration, so memory-backed role stores silently discarded the per-role session bound on every write and read (devin on #11478). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * sts: apply per-role MaxSessionDuration to resolved session durations Review follow-up on #11478 (devin): the role bound only ever applied to explicit DurationSeconds values — an omitted duration resolved to the configured default and sailed past a shorter role max on every assume path. - capDurationByRole now resolves min(requested||tokenDuration, roleMax), so AssumeRoleWithWebIdentity and AssumeRoleWithCredentials cap defaults the same way they cap explicit values. - prepareSTSCredentials caps the calculated duration at the named role's MaxSessionDuration, covering the AssumeRole and LDAP handlers; self-assumption has no role definition to consult. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: keep MaxSessionDuration through the cached role store genericCopyRoleDefinition drops MaxSessionDuration the same way copyRoleDefinition did, so the cached filer role store reads back a zero maximum and every downstream duration cap is skipped (greptile on #11478). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * sts: only materialize defaults that pass session duration validation Review follow-up on #11478 (greptile): materializing an omitted DurationSeconds into an explicit value could exceed the service's own input bound (a configured tokenDuration above maxSessionLength) and turn a previously working request into a validation error. capDurationByRole now leaves nil anything the service can resolve better itself, clamps a tightened default at maxSessionLengthSeconds, and floors a role bound below 900s to the tightest issuable value. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ab95d58b7c |
s3: keep dedicated object-lock actions pinned during action resolution (#11475)
* s3: keep dedicated object-lock actions pinned during action resolution A coarse action that already names a dedicated operation (governance bypass, retention, legal hold, bucket object-lock config) now resolves to itself before request shape is consulted. Previously a synthetic DELETE ?versionId authorization request re-resolved to s3:DeleteObjectVersion, so the bypass check was satisfied by the delete-version grant alone; with the pin it evaluates s3:BypassGovernanceRetention as intended. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: cover pinned object-lock actions against competing query params Locks in the resolution for every dedicated action in the pin set, incl. the retention and legal-hold shapes carrying versionId. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2f641a63d6 |
filer: honor is_moved only from ring member connections (#11456)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks updateLocalEntry records the RemoteEntry stamp after an upload by writing the event's entry back with UpdateEntry. The filer deletes every stored chunk absent from an updated entry, so when the file was rewritten while its upload was in flight (or the event is a replay), the stale snapshot deletes the rewrite's chunks: the entry then points at the new fid with no needle behind it, and the rewrite's own upload fails and is skipped as superseded. The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's chunk fids, evaluated by the filer under the path lock. A refused stamp means the filer moved past this event; the superseding event follows in the log and stamps the current entry, so the refusal is logged and skipped like a superseded upload. Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount, filer.remote.sync; hold the remote (docker pause) so one upload stays in flight, rewrite the file through the filer, unpause. Before: the entry's chunk is 404 on every volume server. After: the stale stamp is refused, the rewrite's chunk stays live and reads back after a vacuum. * filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a rewrite that touched inline content or metadata alone still compared equal and the stale snapshot overwrote the live entry. The new clause compares the whole stored entry against the event's entry under the same path lock. * filer: route conditional UpdateEntry to the entry's owner filer Two filers locking the same path locally could still pass a stale condition on the non-owner while the owner's entry had moved on. When a condition or expected_extended precondition is set, forward the request to the entry's owner the same way conditional CreateEntry does, with is_moved bounding the hop. * filer: compare IF_ENTRY_EQUAL against the normalized expected entry FindEntry grows FileSize to the chunk extent, so a raw event entry with FileSize still zero failed the condition on an unchanged file and the stamp was skipped, letting a replay upload the object again. * filer.remote.sync: classify refused stamps by gRPC status only A FailedPrecondition substring in an unrelated error would have been swallowed as a skipped stamp; status.FromError already unwraps. * remote sync: keep the event entry intact for IF_ENTRY_EQUAL * filer: honor is_moved only from ring member connections is_moved is caller-controlled, so a request could set it to skip owner routing and run a conditional check under a non-owner's lock. Verify the marker against the peer's connection address and the lock ring members; an unverified marker is ignored and the request routes like a fresh one. * filer: refuse unverifiable is_moved at a non-owner, cache ring IPs Follow-up fixes from review on the is_moved provenance check: - checkMovedMarker replaces "ignore and re-forward" for markers that did not arrive on a ring member's connection. Re-forwarding a claimed hop could cycle while rings disagree; instead the request is refused with FailedPrecondition unless this filer is the key's owner, in which case applying locally is correct anyway. - ringMemberIPs caches resolved member addresses per ring membership so hostname-advertising deployments do not pay a DNS lookup per forwarded request; failed lookups are not cached so a DNS blip self-heals. - DistributedUnlock no longer dereferences the nil response of a failed next-hop RPC. * filer: refuse unverifiable is_moved with PermissionDenied, not FailedPrecondition A routing refusal is different in kind from a write-condition mismatch: remote sync treats FailedPrecondition as a stale stamp and skips it, so reusing that code let a routing failure pass as synced. Owner checks now also run before the peer-IP lookup so the common accept path does no DNS. * filer: expire resolved ring member IPs after 5 minutes A member's hostname can re-resolve to a new IP while its ring address stays unchanged; caching forever would reject its genuine forwards until a membership change or restart. * filer: deduplicate concurrent ring member DNS lookups At cache expiry, parallel forwarded requests would each resolve every member hostname serially; singleflight collapses them into one lookup per ring membership. * filer: detach the shared ring lookup from the caller's context The singleflight winner's ctx is cancelled when its request ends; the shared result would then be an incomplete member list and genuine forwards denied. The lookup now runs on a detached context with its own deadline so a canceled caller cannot poison it. * filer: resolve ring member hostnames in parallel The shared lookup gave every member one serial budget, so a few slow resolutions could leave later members out of the cached list and reject their genuine forwards. Each member now resolves concurrently under its own detached deadline. * filer: gather literal member IPs before spawning lookups A ring mixing IP literals and hostnames raced: the literal appends ran unlocked alongside the resolver goroutines' locked appends. Split into two passes so only hostname results share the mutex. --------- Co-authored-by: jsas <1351492+jsas@users.noreply.github.com> |
||
|
|
80fd3635d2 |
volume: skip TTL last-write scan when it cannot fit its budget (#11472)
* volume: skip TTL last-write scan when it cannot fit its budget * Update weed/storage/volume_checking.go Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
7129e1178e |
s3: evaluate bucket policy before ACL public-read for anonymous requests (#11471)
* s3: evaluate bucket policy before ACL public-read for anonymous requests AuthWithPublicRead granted anonymous access on a public-read ACL before consulting the bucket policy, so an explicit Deny (e.g. s3:ListBucket) was skipped for anonymous callers while still enforced for authenticated ones. Run the policy engine first: a matching Deny or Allow is honored, otherwise fall through to the ACL grant as before. * s3: defer object-level anonymous requests to the handler's policy recheck Evaluating the bucket policy with a nil entry at middleware time makes tag conditions like s3:ExistingObjectTag/<key> resolve against missing values, so a conditional Deny could wrongly block anonymous Get/Head on a public bucket whose handler recheck would permit it. Object requests now take the ACL grant and let Get/HeadObjectHandler re-evaluate with the fetched entry; only bucket-level requests (List, HeadBucket), which have no such recheck, are decided by the middleware policy verdict. Reading the bucket config first also refreshes the compiled policy on a cache miss, so a remotely deleted policy cannot leave a stale verdict in the engine for nonresident buckets. * s3: recheck bucket policy before serving directory objects handleDirectoryObjectRequest runs before the object handlers' policy recheck, so directory content on a public-read bucket was served to anonymous callers without any policy evaluation. Evaluate the policy with the directory entry, matching the recheck the file path performs. |
||
|
|
3c1e8ca7a8 |
volume: never finish serving a needle whose data fails its CRC (#11467)
* storage: hold back last chunk until CRC verifies on whole-needle reads Above PagedReadLimit the needle is streamed: headers and body go out before the checksum is computed, so a corrupted needle was served as 200 with bad bytes and readers could not fall back to a replica. The final chunk is now written only after the checksum verifies; on a mismatch the response ends short of Content-Length and the client sees a failed transfer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: verify needle checksum on streamed reads (parity) Mirror the Go fix: carry the needle checksum in NeedleStreamInfo and have StreamingBody accumulate the CRC and verify it before emitting the last frame; a mismatch ends the body with an error so the client sees the transfer fail rather than receiving corrupt bytes that look complete. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: abort the transfer when a committed stream fails A writeFn failure after any byte or WriteHeader call leaves the declared status and Content-Length already sent; http.Error's text then joins the body and can exactly fill the withheld tail of a corrupted needle read — the client sees a complete 200 instead of a failed transfer to retry. Track whether the response is committed (headers sent, or bytes buffered for the deferred flush) and panic with http.ErrAbortHandler instead of appending an error body; pre-commit failures keep the 500 path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: drop the response writer wrapper from the committed-response check Counting buffered writes is enough: with no bytes buffered the status and headers cannot have gone out, and the range branches commit via the explicit WriteHeader call before writeFn runs. The extra ResponseWriter wrapper added a new Write sink site that CodeQL flags. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |