mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-09-29 11:15:34 +00:00
67691a1eeae9b87726ff09c8ceb317afb90ac3dd
15290
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
67691a1eea |
volume server: split volume_copy into phases and type the delete-after-status gate (#11485)
volume_copy was one ~400-line handler, and the rule that an existing local replica is deleted only after the source's ReadVolumeFileStatus succeeded was held by statement order alone. The keep_remote_data=true that the pre-copy delete and the failed-copy rollback must share was kept in sync by a comment pointing from one to the other. The handler is now a ~60-line orchestrator over connect_to_copy_source, SourceVolumeStatus::fetch, delete_existing_replica, plan_copy_destination and a VolumeCopyJob whose run() drives preallocate_dat, transfer_files, finish_copied_files and mount_and_reply, with cleanup_failed_copy on error. delete_existing_replica takes a &SourceVolumeStatus, which only fetch can construct (private field in a child module), so the delete cannot be called before the status RPC. Both deletes go through delete_replica_keep_remote. Pure refactor: call order, status codes and messages, cancellation checks, throttling, progress reports and cleanup are unchanged. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
02353444ac |
fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts (#11483)
* fix(volume-rust): reserve a disk before replacing a replica in VolumeCopy, and check record counts Port of the Go VolumeCopy hardening in #11238 and #11252. - Pick the destination disk before deleting the existing replica, counting the slot that replica holds as free. If no disk qualifies, the healthy replica is kept instead of being deleted. - Read the source's VolumeStatus before and after the copy. When both succeed and the counts did not change, the mounted replica's file and deleted counts must match; on mismatch it is unmounted and its files removed. A failed "before" read skips the check; a failed "after" read fails the copy. * fix(volume-rust): let a departing caller cancel VolumeCopy's post-copy status read Go reads the source's status after the copy with stream.Context(), so the call ends when the caller leaves. The Rust call had no such link: a source that stalled there held the copied, unmounted files after the caller was gone. Race it against the response channel, like the other blocking steps, so the usual error cleanup removes the partial copy. |
||
|
|
68944e83a3 |
volume: typed tier errors so a missing remote object answers NotFound (#11484)
remote_storage/s3_tier.rs returned Result<_, String> from every
transfer (upload_file, download_file, read_range[_blocking],
delete_file[_blocking]) and from the tier runtime helpers. The tier
move handlers could only wrap that in Status::internal, so a .dat whose
remote object is gone was indistinguishable from an I/O failure to
weed shell.
Add TierError { NotFound, Io, RuntimeUnavailable, Aborted }. Each
variant carries the existing message verbatim. NotFound follows the
rules remote_storage/s3.rs already uses: raw 404 status on HEAD,
NoSuchKey code on GET; a bare 404 on GET stays Io. A progress-callback
Err becomes Aborted. VolumeError gains a transparent Tier variant and
From<VolumeError> for Status maps Tier(NotFound) to NotFound; the tier
move handlers go through status_with_context, so their message text is
unchanged. Every other tier failure is still Internal.
The remote needle read path keeps io::Error::other, so its error kind
and vacuum's handling of it do not change.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
||
|
|
00310f6588 |
volume server: run the vacuum compaction copy without the store lock (#11482)
* volume server: run the vacuum compaction copy without the store lock VacuumVolumeCompact held the store write lock for the whole live-needle copy, including every progress blocking_send on the 16-deep stream. On a large volume that is minutes with every read, write and heartbeat on the node parked behind it, long enough for the master to unregister the node. Split compaction the way Go's CompactByIndex runs it. A short locked step claims the volume's compacting flag, records the makeup_diff watermark (index size and compaction revision) and opens fresh .dat/.idx handles. The copy then replays .idx up to the watermark and copies from those handles with the store lock released; writes that land meanwhile are replayed by makeup_diff at commit, as before. The flag is an Arc<AtomicBool> released when the job is dropped, so every exit path clears it. Because the flag is now visible to other callers, the operations that would pull the files out from under the copy refuse while it is set: unmount (and VolumeConfigure, which unmounts and remounts), delete (checked before the volume is removed from the map, which a refused destroy used to leave unmounted), cleanup, and index relocation. A second compact and a commit stay no-ops, as in Go. The pre-copy fsync is dropped: the copy reads its own handles through the page cache and .cpd/.cpx are fsynced before commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume: keep a read-only in-memory index's size for the compaction copy The unlocked copy replays .idx up to index_file_size(). A read-only volume whose .sdx could not be built loads its index into memory without a writer, so that size stayed 0: the copy came out empty and the commit replaced the volume with it. CompactNeedleMap::load_from_idx now records the rows it loaded, which is also what Go's IndexFileSize reports for a read-only index. The copy's index replay now stops reading at the recorded size instead of walking rows appended since, which makeup_diff replays anyway. Adds tests for compacting a read-only volume on both the sorted index and the in-memory fallback, and for VolumeConfigure stopping when the unmount is refused during a copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: stop a vacuum copy as soon as its client is gone The progress callback only noticed a closed response stream when a report was due, every 128 MiB. With the copy now running outside the store lock, a copy nobody waits for keeps the volume marked compacting and so keeps refusing unmount, delete and cleanup until that next report. Check the stream on every callback. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5c9c424a84 |
volume server: stream ReadAllNeedles without holding the store lock (#11481)
* volume server: stream ReadAllNeedles without holding the store lock read_all_needles held store.read() while Volume::read_all_needles read every live needle of the volume into a Vec, and kept holding it through the whole blocking_send loop. Memory grew with the volume, and a slow client parked the scan in a send with the guard held; needle writes and the heartbeat take store.write() on a writer-preferring lock, so the node stopped serving until the client caught up. Take a DatScanPlan (fresh .dat open, end bound) under a short guard and walk it with the guard released, sending one needle at a time. Each record is checked against the live needle map under a brief read guard, as the scan reaches it, and only a live record is parsed, so a damaged stale copy does not fail the stream. Records appended while a pass ran are walked by a follow-up plan, so a needle overwritten during the scan is streamed once, as its new copy. A vacuum commit or re-create of the volume during the scan fails the stream, since the map's offsets no longer describe the pinned file; the plan carries the volume instance and compaction revision for that check. DatScanPlan::scan_records yields records unparsed; scan keeps its behaviour on top of it. Volume::read_all_needles has no caller left and is removed; its tests move to the RPC. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: check ReadAllNeedles liveness only once the send can proceed The per-record liveness check ran before blocking_send, so a scan parked on a full channel held a record it had already judged live. An overwrite landing during that park left the old copy in the stream, and the continuation over appended records then streamed the new copy as well. Reserve channel space first, then take the store read guard, check the record against the needle map and enqueue it through the permit before releasing the guard. The wait for space still happens without the lock; the record is parsed before the guard is taken, and its parse error only counts if the record turns out to be live. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5218e68554 |
volume server: collect heartbeats under the store read lock, off the runtime (#11480)
* volume server: collect heartbeats under the store read lock, off the runtime Every pulse tick, options change and volume-state notification took store.write() for a whole heartbeat pass, directly on the async heartbeat task. The pass fstats every volume's .dat twice and hashes its report, so on a server with many volumes it held the store exclusively for the whole scan: reads and writes stalled, and with the writer-preferring RwLock a pending pass parked every new reader too. The pass only needs to mutate the store for a few rare actions: removing expired EC volumes, deleting expired volumes past their removal delay, and setting no-write on IO-quarantined volumes. It now runs under store.read(), records those as (disk, volume id) actions, and applies them afterwards under a short store.write() that is only taken when there is something to do. Each action re-checks its target under the write lock, so a volume written to, replaced or removed in between is left alone. Expired EC volumes are still removed before the volume pass, as before, because the EC shard count feeds the disk-space-low max volume count. Every pass runs on the blocking pool via spawn_blocking. The heartbeat message is unchanged for the same store state. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * volume server: take has_no_ec_shards with the heartbeat's volume list The heartbeat pass took has_no_ec_shards from the EC phase's read lock, then built the volume list under a second one. An EC shard mounted in between went out as "no EC shards" beside a volume list taken after the mount, and the master clears a server's EC registrations on that flag. has_no_ec_shards is now computed under the same read lock as the volume list, with the EC phase's filter: not expired, not quarantined, at least one shard. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
00cffa028c | docs: regenerate star history chart | ||
|
|
a261f90e18 |
vacuum: let the sweep release volumes that stay empty and quiet (#11477)
* vacuum: let the sweep release volumes that stay empty and quiet Vacuuming reclaims bytes but not slots: a fully emptied volume stays registered to its collection forever, and since growth is gated only on slot count a store at 99% free disk can still refuse writes to other collections (#11429). volume.deleteEmpty exists but is manual-only. With -vacuumDeleteEmptyAfterSeconds (or master.vacuumDeleteEmptyAfterSeconds under weed server/mini; default 0, off) the automatic sweep now deletes replica copies that have stayed empty and quiet for that long, the same rule volume.deleteEmpty applies on demand: remote-backed copies are skipped, and every delete carries the volume server's onlyEmpty / onlyGarbage guards so a copy written since the last report is refused rather than removed. Copies that still hold data or were written recently stay; only a volume whose every copy is deleted leaves the sweep's work map, sparing a compaction of bytes that are all deleted. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: harden empty-volume sweep against partial and racing deletes Review follow-up on #11477: - delete a volume only when every replica copy is a verifiable empty-and-quiet candidate; deleting the empty copy of a volume whose sibling holds live files would silently cut its replica count (greptile P1). - drain the volume out of the writable list before deleting, the same drain the compact pass uses, so PickForWrite stops assigning it and pending writes settle (devin). - bound the VolumeDelete RPC so one stalled server cannot hold the vacuum lock indefinitely (greptile P1, reusing allocateVolumeTimeout). The vid2location panic scenario raised in review does not exist: VolumeLocationList methods are nil-receiver safe and a missing vid just fails enoughCopies, so a partially deleted volume skips compaction instead of crashing the sweep. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: unregister deleted empty replicas and prune the sweep list A successful VolumeDelete only updates the volume server; the master still tracked the replica and kept it in the sweep's location list for the compaction pass (coderabbit on #11477). Unregister the replica right after its delete succeeds and drop it from the sweep copy, so a partially deleted volume only compacts copies that still exist. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: pin deleting volumes out of the writable list across heartbeats Review follow-up on #11477 (greptile): DrainAndRemoveFromWritable only removed the volume once; a heartbeat landing between the drain and the replica deletes re-evaluated writability and re-added it, so a client write could reach a replica whose siblings were already gone and leave the volume under-replicated when the last copy refused its onlyEmpty delete. MarkDeleting records the vid in deletingVolumes — checked inside setVolumeWritable so heartbeat, capacity-recovery, and admin re-add paths all hold it out — and UnmarkDeleting releases it once the sweep finishes the copy pass. A partially deleted volume's surviving replicas then return to writable through the normal heartbeat path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * vacuum: restore writability when a sweep delete survives Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
be29f44d87 |
s3: record requester identity before the authz verdict (#11479)
* s3: record requester identity before the authz verdict for audit Identity was only stored in request context on the success branch, so denied requests reached WriteErrorResponse without requester attribution and audit entries had empty requester/requester_arn/requester_identity. Authentication failures still resolve no identity, so unauthenticated denials stay unattributed. Fixes #11474 Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: keep the resolved identity through authz denial in Auth Review follow-up on #11479 (devin): authRequest discarded the identity on every error, so a request that authenticated fine but failed the action check still reached handleAuthResult with no identity and the deny path could not audit a requester. Auth now calls authRequestWithAuthType directly, the same entry AuthPostPolicy uses, so the resolved identity reaches the error writer; a failed authN still resolves no identity and stays unattributed. The regression test now signs a denied request end to end through iam.Auth. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2864bc0fe8 |
s3: honor configured session bounds on AssumeRole and LDAP identity (#11478)
* sts: export CalculateSessionDuration Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: honor configured session bounds on AssumeRole and LDAP identity prepareSTSCredentials hardcoded a one-hour session when the caller omitted DurationSeconds, so sts.tokenDuration was ignored and sts.maxSessionLength only clamped explicit requests: asking for 3600s against a 20m ceiling was rejected while omitting the parameter was granted a full hour (#11473). The two affected handlers now use the same default-then-cap calculation as AssumeRoleWithWebIdentity. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: keep MaxSessionDuration through role store copies copyRoleDefinition rebuilt RoleDefinition field by field and dropped MaxSessionDuration, so memory-backed role stores silently discarded the per-role session bound on every write and read (devin on #11478). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * sts: apply per-role MaxSessionDuration to resolved session durations Review follow-up on #11478 (devin): the role bound only ever applied to explicit DurationSeconds values — an omitted duration resolved to the configured default and sailed past a shorter role max on every assume path. - capDurationByRole now resolves min(requested||tokenDuration, roleMax), so AssumeRoleWithWebIdentity and AssumeRoleWithCredentials cap defaults the same way they cap explicit values. - prepareSTSCredentials caps the calculated duration at the named role's MaxSessionDuration, covering the AssumeRole and LDAP handlers; self-assumption has no role definition to consult. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iam: keep MaxSessionDuration through the cached role store genericCopyRoleDefinition drops MaxSessionDuration the same way copyRoleDefinition did, so the cached filer role store reads back a zero maximum and every downstream duration cap is skipped (greptile on #11478). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * sts: only materialize defaults that pass session duration validation Review follow-up on #11478 (greptile): materializing an omitted DurationSeconds into an explicit value could exceed the service's own input bound (a configured tokenDuration above maxSessionLength) and turn a previously working request into a validation error. capDurationByRole now leaves nil anything the service can resolve better itself, clamps a tightened default at maxSessionLengthSeconds, and floors a role bound below 900s to the tightest issuable value. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
ab95d58b7c |
s3: keep dedicated object-lock actions pinned during action resolution (#11475)
* s3: keep dedicated object-lock actions pinned during action resolution A coarse action that already names a dedicated operation (governance bypass, retention, legal hold, bucket object-lock config) now resolves to itself before request shape is consulted. Previously a synthetic DELETE ?versionId authorization request re-resolved to s3:DeleteObjectVersion, so the bypass check was satisfied by the delete-version grant alone; with the pin it evaluates s3:BypassGovernanceRetention as intended. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * s3: cover pinned object-lock actions against competing query params Locks in the resolution for every dedicated action in the pin set, incl. the retention and legal-hold shapes carrying versionId. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
2f641a63d6 |
filer: honor is_moved only from ring member connections (#11456)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks updateLocalEntry records the RemoteEntry stamp after an upload by writing the event's entry back with UpdateEntry. The filer deletes every stored chunk absent from an updated entry, so when the file was rewritten while its upload was in flight (or the event is a replay), the stale snapshot deletes the rewrite's chunks: the entry then points at the new fid with no needle behind it, and the rewrite's own upload fails and is skipped as superseded. The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's chunk fids, evaluated by the filer under the path lock. A refused stamp means the filer moved past this event; the superseding event follows in the log and stamps the current entry, so the refusal is logged and skipped like a superseded upload. Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount, filer.remote.sync; hold the remote (docker pause) so one upload stays in flight, rewrite the file through the filer, unpause. Before: the entry's chunk is 404 on every volume server. After: the stale stamp is refused, the rewrite's chunk stays live and reads back after a vacuum. * filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a rewrite that touched inline content or metadata alone still compared equal and the stale snapshot overwrote the live entry. The new clause compares the whole stored entry against the event's entry under the same path lock. * filer: route conditional UpdateEntry to the entry's owner filer Two filers locking the same path locally could still pass a stale condition on the non-owner while the owner's entry had moved on. When a condition or expected_extended precondition is set, forward the request to the entry's owner the same way conditional CreateEntry does, with is_moved bounding the hop. * filer: compare IF_ENTRY_EQUAL against the normalized expected entry FindEntry grows FileSize to the chunk extent, so a raw event entry with FileSize still zero failed the condition on an unchanged file and the stamp was skipped, letting a replay upload the object again. * filer.remote.sync: classify refused stamps by gRPC status only A FailedPrecondition substring in an unrelated error would have been swallowed as a skipped stamp; status.FromError already unwraps. * remote sync: keep the event entry intact for IF_ENTRY_EQUAL * filer: honor is_moved only from ring member connections is_moved is caller-controlled, so a request could set it to skip owner routing and run a conditional check under a non-owner's lock. Verify the marker against the peer's connection address and the lock ring members; an unverified marker is ignored and the request routes like a fresh one. * filer: refuse unverifiable is_moved at a non-owner, cache ring IPs Follow-up fixes from review on the is_moved provenance check: - checkMovedMarker replaces "ignore and re-forward" for markers that did not arrive on a ring member's connection. Re-forwarding a claimed hop could cycle while rings disagree; instead the request is refused with FailedPrecondition unless this filer is the key's owner, in which case applying locally is correct anyway. - ringMemberIPs caches resolved member addresses per ring membership so hostname-advertising deployments do not pay a DNS lookup per forwarded request; failed lookups are not cached so a DNS blip self-heals. - DistributedUnlock no longer dereferences the nil response of a failed next-hop RPC. * filer: refuse unverifiable is_moved with PermissionDenied, not FailedPrecondition A routing refusal is different in kind from a write-condition mismatch: remote sync treats FailedPrecondition as a stale stamp and skips it, so reusing that code let a routing failure pass as synced. Owner checks now also run before the peer-IP lookup so the common accept path does no DNS. * filer: expire resolved ring member IPs after 5 minutes A member's hostname can re-resolve to a new IP while its ring address stays unchanged; caching forever would reject its genuine forwards until a membership change or restart. * filer: deduplicate concurrent ring member DNS lookups At cache expiry, parallel forwarded requests would each resolve every member hostname serially; singleflight collapses them into one lookup per ring membership. * filer: detach the shared ring lookup from the caller's context The singleflight winner's ctx is cancelled when its request ends; the shared result would then be an incomplete member list and genuine forwards denied. The lookup now runs on a detached context with its own deadline so a canceled caller cannot poison it. * filer: resolve ring member hostnames in parallel The shared lookup gave every member one serial budget, so a few slow resolutions could leave later members out of the cached list and reject their genuine forwards. Each member now resolves concurrently under its own detached deadline. * filer: gather literal member IPs before spawning lookups A ring mixing IP literals and hostnames raced: the literal appends ran unlocked alongside the resolver goroutines' locked appends. Split into two passes so only hostname results share the mutex. --------- Co-authored-by: jsas <1351492+jsas@users.noreply.github.com> |
||
|
|
80fd3635d2 |
volume: skip TTL last-write scan when it cannot fit its budget (#11472)
* volume: skip TTL last-write scan when it cannot fit its budget * Update weed/storage/volume_checking.go Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> |
||
|
|
7129e1178e |
s3: evaluate bucket policy before ACL public-read for anonymous requests (#11471)
* s3: evaluate bucket policy before ACL public-read for anonymous requests AuthWithPublicRead granted anonymous access on a public-read ACL before consulting the bucket policy, so an explicit Deny (e.g. s3:ListBucket) was skipped for anonymous callers while still enforced for authenticated ones. Run the policy engine first: a matching Deny or Allow is honored, otherwise fall through to the ACL grant as before. * s3: defer object-level anonymous requests to the handler's policy recheck Evaluating the bucket policy with a nil entry at middleware time makes tag conditions like s3:ExistingObjectTag/<key> resolve against missing values, so a conditional Deny could wrongly block anonymous Get/Head on a public bucket whose handler recheck would permit it. Object requests now take the ACL grant and let Get/HeadObjectHandler re-evaluate with the fetched entry; only bucket-level requests (List, HeadBucket), which have no such recheck, are decided by the middleware policy verdict. Reading the bucket config first also refreshes the compiled policy on a cache miss, so a remotely deleted policy cannot leave a stale verdict in the engine for nonresident buckets. * s3: recheck bucket policy before serving directory objects handleDirectoryObjectRequest runs before the object handlers' policy recheck, so directory content on a public-read bucket was served to anonymous callers without any policy evaluation. Evaluate the policy with the directory entry, matching the recheck the file path performs. |
||
|
|
3c1e8ca7a8 |
volume: never finish serving a needle whose data fails its CRC (#11467)
* storage: hold back last chunk until CRC verifies on whole-needle reads Above PagedReadLimit the needle is streamed: headers and body go out before the checksum is computed, so a corrupted needle was served as 200 with bad bytes and readers could not fall back to a replica. The final chunk is now written only after the checksum verifies; on a mismatch the response ends short of Content-Length and the client sees a failed transfer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: verify needle checksum on streamed reads (parity) Mirror the Go fix: carry the needle checksum in NeedleStreamInfo and have StreamingBody accumulate the CRC and verify it before emitting the last frame; a mismatch ends the body with an error so the client sees the transfer fail rather than receiving corrupt bytes that look complete. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: abort the transfer when a committed stream fails A writeFn failure after any byte or WriteHeader call leaves the declared status and Content-Length already sent; http.Error's text then joins the body and can exactly fill the withheld tail of a corrupted needle read — the client sees a complete 200 instead of a failed transfer to retry. Track whether the response is committed (headers sent, or bytes buffered for the deferred flush) and panic with http.ErrAbortHandler instead of appending an error body; pre-commit failures keep the 500 path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * server: drop the response writer wrapper from the committed-response check Counting buffered writes is enough: with no bytes buffered the status and headers cannot have gone out, and the range branches commit via the explicit WriteHeader call before writeFn runs. The extra ResponseWriter wrapper added a new Write sink site that CodeQL flags. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
0f3ba98e11 |
volume: make volume.scrub report a live needle whose stored id is damaged (#11468)
* storage: scrub live needles' stored id against the index key scrubVolumeData only compared the needle's stored id for tombstones, so header damage on a live needle — where the data CRC cannot see it — passed every scrub mode while reads of that needle kept failing or serving the wrong key's data. Compare the id for every indexed needle. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: scrub live needles' stored id against the index key (parity) Mirror the Go scrub fix: compare the stored needle id with the index key for live needles too, not only for deleted ones. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * rust volume: cover damaged live needle id in scrub test The tombstone test proved the index-key check fires for deleted entries; add the live-needle mirror of Go's TestScrubVolumeDataChecksLiveNeedleId so a regression in the live path is caught in Rust too. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
80a26020d7 |
util: serialize all ViperProxy access so startup cannot hit concurrent map read/write (#11470)
* util: serialize every ViperProxy method; stop promoting unlocked viper calls ViperProxy embedded *viper.Viper, so only the five declared methods took the mutex while every promoted call — GetStringMap in backend.LoadConfiguration was the reported crash — touched viper's maps unsynchronized. `weed server` starts the volume server (SetDefault writer) and the master (GetStringMap reader) back to back, and a race build reports the pair on a plain start. The wrapped viper is now a named field: a method must be declared here to exist on the proxy, so unsynchronized access fails at compile time rather than at runtime. Every promoted use in the tree (GetStringMap, GetUint32, GetFloat64, GetDuration, IsSet, AllKeys, Set) gets a locked wrapper; NewViperProxy replaces struct literals for local vipers. GetStringMap deep-copies its result — viper hands back the internal subtree, so iterating it after the lock is released would race the next writer. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * util: take the shared lock while LoadConfiguration merges a config file viper.MergeInConfig rewrites the same maps the proxy serializes; without the lock a merge can race a concurrent SetDefault or reader exactly like the reported startup crash. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * util: deep-copy slice elements in the GetStringMap snapshot A slice of maps inside the returned subtree still shared the inner maps — copy elements recursively so nothing the caller mutates is viper's internal state. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * util: add the missing AutomaticEnv wrapper used by tests sse_reader_test reaches it through GetViper(); without the wrapper the call no longer exists once the viper field stopped being embedded. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * util: return a fresh slice from GetStringSlice A stored []string comes back uncast from viper — the backing array is shared internal state like the GetStringMap subtree, so copy it while holding the lock. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
5389f61cef |
volume server: do not finish a GET when the needle CRC mismatches (#11464)
* volume server: do not finish a GET when the needle CRC mismatches A streamed full-needle read compared the CRC only after every page had been written. Once the response buffer flushed, the client already had a completed 200 and the corrupt bytes. Hold the last page until the checksum matches, and if an earlier page has already been flushed, abort the connection instead of calling http.Error. Fixes #11459 * volume server: abort partial-content bodies on write error too The non-Range path drops the unflushed tail and aborts on a mid-body error; the single-range and multi-range paths still flushed it after WriteHeader(206) was committed, delivering corrupt bytes as a complete body. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: assert the started 200 is aborted in the write-error test The test previously returned on any request error, so it passed without verifying the abort. It now asserts the client got the committed 200 headers and then a failed body read. Also trims comments. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
8ad2f29e3e |
shell: let volume.deleteEmpty drop volumes with no live needles (#11437)
* shell: let volume.deleteEmpty drop volumes with no live needles The candidate check only accepted a .dat at superblock size, so a volume whose every needle was deleted still had to be vacuumed first — minutes of compaction to rewrite bytes that were all garbage anyway. FileCount counts every indexed entry and DeleteCount every entry made garbage by overwrite or delete, so FileCount <= DeleteCount means nothing live remains and the volume can be unlinked directly. The quietFor guard is unchanged. * volume server: add only_garbage VolumeDelete guard VolumeDelete(only_empty) refuses every volume that ever held data, so a volume whose needles are all deleted could only be removed after a vacuum rewrote it. The new only_garbage flag deletes only when the byte counters show nothing live: DeletedSize covering all of ContentSize, the same all-garbage state vacuum measures. Byte counters are used because the file/delete counts drift on index reload. * rust volume: mirror only_garbage VolumeDelete guard Same check as the Go server: a volume deletes under only_garbage when its deleted bytes cover all content bytes. The grpc handler rejects before the store drops the volume from its map, since destroy errors after removal would still unmount it. * volume delete: let either enabled check pass, keep onlyEmpty on the wire An upgraded shell sending only_garbage to a pre-upgrade server would be read as an unconditional delete (field ignored, only_empty false). The request now keeps only_empty set so old servers check emptiness and refuse, while new servers delete when either check passes. * volume.deleteEmpty: skip remote-backed and protected read-only volumes A remote-tiered replica shares its cloud object with the other replicas, so keepRemoteData=false on one delete removes data they still reference. Protected read-only volumes are quarantined or under maintenance, which is exactly when a replica should not be dropped. * volume delete: validate guarded copies across disks before deleting * volume delete: hold copy locks across guarded validate-and-delete CheckVolumeDeletable released each copy's locks before Destroy ran, so a write landing on a later copy between the two passes refused its destroy after earlier copies were already removed. Pin every copy's dataFileAccessLock (and its location's volumesLock) across validation and removal so a refused delete leaves all copies intact. * volume delete: send deleted-volume notices after releasing locks A blocking send on a full DeletedVolumesChan under volumesLock can stall the heartbeat loop that drains it while it waits on the same locks. Collect the notices under the lock span and send after release. * pb: restore generated-file cosmetics to match the repo's protoc version Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
c58bd0dfd3 |
s3: honor assignment fsync in UploadWithRetry (#11449)
* fix: honor assignment fsync in UploadWithRetry * Tests feedback |
||
|
|
f7680cf812 |
volume: one long-lived tokio runtime for blocking tiered S3 reads (#11448)
* volume: one long-lived tokio runtime for blocking tiered S3 reads block_on_tier_future, behind read_range_blocking and delete_file_blocking, spawned an OS thread and built a fresh current-thread tokio runtime on every call, then tore the runtime down. On an S3-tiered volume that is once per needle read, per streamed 64 KiB chunk, per DatScanPlan record and per destroy. The SDK client's pooled HTTPS connections are driven by tasks on the runtime a request ran on, so each teardown dropped the pool and every call re-dialed and re-handshook TLS. A panic inside the SDK was also flattened to the fixed string "tier runtime thread panicked". Now one process-wide runtime (OnceLock, multi_thread, 2 workers named tier-io) drives all tier I/O; block_on_tier_future spawns onto it and parks the caller on an mpsc channel for the JoinHandle result. Blocking the caller is unavoidable (the storage layer is synchronous) and is what the old code did through thread::spawn().join(). Handle::block_on is not used because the wrappers are also reached from inside another runtime's worker, where it panics with "Cannot start a runtime from within a runtime". JoinError panics are downcast to &str/String and the payload is kept in the error. Tests cover runtime reuse (Handle::id equal across calls, thread name tier-io), calls from a std thread, from spawn_blocking, and directly from current-thread and multi-thread runtime contexts, and the panic payload. Against the old body 7 of 9 fail. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: return a tier runtime build failure instead of panicking Review follow-up. tier_runtime() expect'ed the runtime build, so an OS refusing threads panicked inside Volume::destroy (after the volume left the in-memory map, before its files were removed) and inside needle reads, bypassing their error paths. Keep the runtime in a Mutex<Option<Runtime>> behind tier_handle() -> Result<Handle, String>: a failed build is returned to the caller through block_on_tier_future's existing Result and is not cached, so a later call retries once the pressure is gone. The lock is held only while building. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim comments on the shared tier I/O runtime Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4299fdf578 |
volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest (#11446)
* volume server: VolumeEcShardsDelete full teardown unloads every disk and keeps the shard gauge honest
Go's VolumeEcShardsDelete full teardown calls vs.store.UnloadEcVolume in
the blanket path (weed/server/volume_grpc_erasure_coding.go:488) and
location.UnloadEcVolume in the generation-fenced path (:511): each disk
that had the volume registered drops it, closes its shard descriptors
and gives back its ec_shards gauge before the artifacts are unlinked.
The Rust handler used Store::remove_ec_volume / DiskLocation::remove_ec_volume
instead, which only remove the map entry. Store::remove_ec_volume also
stops at the FIRST disk holding the vid, so on a split-disk volume
(shards on several disks) the blanket teardown left the sibling disks'
EcVolume registered with open fds while the unlink loop deleted their
files underneath it: the heartbeat kept advertising shards whose files
were gone, the inodes stayed pinned by the open descriptors, and the
VOLUME_GAUGE{collection,"ec_shards"} never came back down. The fenced
path leaked the gauge and the descriptors the same way on the one disk
it wiped.
Both paths now use the unload_ec_volume helpers from #11413 (every disk
for the blanket teardown, the strictly-older disk for the fenced one),
and the two Status::internal messages name the disk directory like Go's
"... on %s: %w".
Regression tests build a two-disk store with the same vid mounted on
each disk (the SplitDiskEcFixture, which gains a collection knob so the
gauge read is isolated from parallel tests mounting under "") and assert
that a blanket teardown leaves no EcVolume registered on any disk and
returns the gauge to its pre-mount value, and that a fenced teardown
decrements the gauge for the older disk's shard while preserving the
newer disk. Both fail against the previous handler.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume server: trim comments on the EC full-teardown unload path
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
3e9fc9e75b |
volume: reject invalid -minFreeSpace and duration flags at startup like Go (#11447)
* volume: reject invalid -minFreeSpace and duration flags at startup like Go Symptom: the Rust volume server swallowed invalid values for three startup flags and ran with a limit the operator did not ask for. -minFreeSpace=150 became 150 bytes (Go: fatal, percent must be 0..=100), -minFreeSpace=10GiBx and -minFreeSpace=50B fell through to the 1% default (Go: fatal), and -inflightUploadDataTimeout=30sec, =abc or =30 silently became the 60 s default (Go: flag.Duration usage error, exit 2). Go's byte-size units (1,024MB, 42 mib, 100Ki, 1eb) were also not understood and mapped to 1%. Go reference: weed/util/minfreespace.go:73-106 (MustParseMinFreeSpace / ParseMinFreeSpace), weed/util/bytes.go:212-243 (ParseBytes and bytesSizeTable), weed/command/volume.go:226-227 (flag.Duration, i.e. time.ParseDuration grammar). Fix: the two inflight timeout Cli fields are now std::time::Duration with a clap value_parser that ports time.ParseDuration (units ns/us/µs/ μs/ms/s/m/h, fractions, bare 0, int64-nanosecond overflow), so clap prints a usage error naming the value and exits 2 at the same layer as Go's flag package. parse_min_free_spaces returns Result and is built from Go-parity parse_min_free_space / parse_bytes; resolve_config_with_env aborts with Go's "The value specified in -minFreeSpace not a valid value" message, the fail-fast pattern it already uses for --max, --index and --readMode. No new dependency. Negative durations and NaN percent are rejected explicitly (Go quirks with no sane Rust representation). Tests cover the Go fixtures from minfreespace_test.go and bytes_test.go, the duration grammar, and flag-level rejection through Cli::try_parse_from and resolve_config_with_env; the flag-level tests were confirmed to fail against the previous parser. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: say that -minFreeSpace commas separate directories, never thousands Review follow-up. parse_bytes keeps Go's ParseBytes thousands-comma support, but the flag is split on commas into per-directory entries before any entry is parsed (Go's MustParseMinFreeSpace does the same), so -minFreeSpace=1,024MB is the two entries 1 and 024MB. The doc comment and the single-entry tests advertised the comma form as if it were reachable from the flag. Document the split and add a flag-level test: two entries from parse_min_free_spaces, and resolve_config aborting on the directory-count mismatch with one -dir. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim comments on the strict flag parsing Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
afce0a3dd3 |
docs: restore the HTTP REST API reference (#11454)
* docs: restore the HTTP REST API reference The REST API documentation was lost in the README refactor, and the wiki only covers the master server API. Add a complete reference for the three HTTP surfaces — filer paths, master file-id/topology endpoints, and volume-server content endpoints — generated from the actual handlers and query parameters. * docs: correct REST API details flagged in review Tagging uses Seaweed- headers not query params, the filer recursive delete option changes the DELETE default, omitted resize mode does not mean fit, default file mode is 0660, the master redirect is 308, the listing flag is -dirListLimit, TUS is enabled by default at /.tus, and -port.public opens the separate read-only listener. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs: match the tagging delete example to canonical header names Go canonicalizes Seaweed-k1 to Seaweed-K1 on write, and the delete list is compared case-sensitively, so ?tagging=k1,k2 would not match. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
83d44be0f3 |
volume: detect S3 not-found by typed SDK errors, not the "service error" string (#11444)
* volume: detect S3 not-found by typed SDK errors, not the "service error" string
remote_storage/s3.rs decided ObjectNotFound by matching the Display
output of an aws_sdk_s3 SdkError against "NoSuchKey" / "404" /
"NotFound". In the locked SDK (aws-smithy-runtime-api 1.11.6,
src/client/result.rs:487-497) that Display is a fixed string per
variant, "service error" for every S3 error, so ObjectNotFound was
unreachable: every missing remote object surfaced as
Other("s3 get object: service error") with the real cause discarded.
Go (weed/remote_storage/s3/s3_storage_client.go) uses typed checks:
HEAD (373-374): awserr.RequestFailure with StatusCode() == 404;
GET (436-437): awserr.Error with Code() == s3.ErrCodeNoSuchKey.
read_file now matches SdkError::ServiceError whose GetObjectError
is_no_such_key(); a bare 404 on GET stays a generic error, as in Go.
stat_file matches HeadObjectError::is_not_found() or a raw HTTP 404
status, Go's actual condition. Non-service errors fall through to
Other unchanged. Every SdkError message in s3.rs and s3_tier.rs is
formatted with DisplayErrorContext so the S3 error code and message
survive instead of "service error".
Six network-free unit tests drive the client through a canned
HttpClient (404 NoSuchKey, bare 404 on GET and HEAD, 404 with a
foreign body on HEAD, 403 AccessDenied on GET and HEAD). They need
aws-smithy-runtime-api as a dev-dependency; it is already in the lock
at a single version, so no new crates.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume: HEAD not-found is the raw 404 status alone, as in Go
Review follow-up. The HEAD arm also accepted the SDK's NotFound error
code on any status, so a 400 carrying <Code>NotFound</Code> became a
missing object. Go's stat looks only at RequestFailure.StatusCode() ==
404 (weed/remote_storage/s3/s3_storage_client.go:373); do the same. The
raw status still covers the body-less 404 the SDK turns into NotFound
and a 404 whose body names a foreign code. Regression test for the
non-404 NotFound body, which failed against the previous arm.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume: trim comments on the typed S3 not-found checks
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: franchb <ru.balloo@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
975cec9228 |
s3api: exclude marker part in listObjectParts pagination (#11463)
* s3api: exclude marker part in listObjectParts pagination Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com> * s3api: guard listObjectParts marker boundary and enhance pagination test Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com> * s3api: fold in review feedback from the parallel #11462 fix Same core fix; this adds the explanatory comment, tightens the overflow guard to math.MaxInt64, makes the fake filer sort entries like a real listing, and adds the marker-exclusivity assertions alongside the pagination walk. Co-authored-by: yi111 <yi111@users.noreply.github.com> --------- Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> Co-authored-by: yi111 <yi111@users.noreply.github.com> |
||
|
|
f31a026b2a |
master,filer: fix lock ring poisoning after leader change (#11453)
* cluster: never broadcast an empty lock ring An empty member list is never a usable ring state, but a delayed RemoveServer on a former leader can fire after the new leader already broadcast the recovered ring. That late broadcast carries a newer wall-clock version, so clients accept the empty ring and permanently reject the good one. Skip the broadcast entirely when the member list is empty, keeping the last non-empty snapshot for reconnecting clients. * cluster: periodically rebroadcast the lock ring Ring updates are purely event-driven, so one lost or poisoned update is permanent until the next membership change — with a single filer that may never come. Re-arm a per-group timer after every broadcast so the current leader keeps re-sending the ring; clients reject nothing newer than their last accepted version, so a re-sent snapshot always heals a stale view. * filer,s3api: reset the lock ring on master change Ring versions are per-master monotonic — each master stamps wall-clock nanoseconds — so a late high-version update accepted from a former leader makes the new leader's snapshot look stale forever. Detect a leader change across the reconnect gap (currentMaster is cleared between attempts, so remember the last served master) and reset the ring to bootstrap state so the new leader's view always applies. * cluster: fail lock acquisition when no lock server exists retryUntilLocked loops forever, so a filer reporting an empty lock ring wedges every append write indefinitely. Bound only the "no lock server found" case — ordinary contention is still waited out since the holder releases eventually. The constructors now return nil on failure: the filer append path and S3 object writes fail fast, while mounts degrade to their existing lockless mode. * cluster: reset only the ring version on master change Ring versions are per-master monotonic, so a version gate reset is all a leader change needs. Clearing the whole ring made every filer its own write owner until the next update and dropped the prior-owner window for keys the new leader remaps; the last ring now keeps routing until the new leader's snapshot transitions off it. * cluster: skip redundant ring installs and defer rebroadcasts An unchanged member list now only bumps the accepted version instead of installing a snapshot: periodic rebroadcasts no longer fire the topology-change callback or restart the prior-owner window. And a rebroadcast that lands inside a membership stabilization window yields to the pending timer rather than publishing an intermediate ring. * cluster,mount: bound lock unavailability, fail ops that cannot lock Only 'lock already owned' contention retries without bound now; every other failure — no lock server, or a dead ring member refusing connections — shares the same unavailability budget, so a ring naming departed filers can no longer hang a lock forever. Mount open-write, create, and rename fail with EAGAIN when the required lock cannot be acquired instead of proceeding without cross-mount serialization. * cluster: check pending stabilization inside the broadcast critical section rebroadcast released the mutex between the pending-timer check and nextBroadcastUpdate, so a membership change arriving in the gap could arm a stabilization timer while the rebroadcast emitted an intermediate ring. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: acquire path locks before mutating create/rename state Create took the DLM lock only after the filer create, so a lock failure returned EAGAIN with an eagerly persisted file left behind. Rename marked source handles renamed before acquiring locks, so a failed acquisition left them suppressing old-path flushes for a rename that never happened. Both now take the locks first; the create's lock is released again if the entry race loses to another creator and AcquireHandle takes over. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: keep the old-path lock when rename lock migration fails The migration stopped the handle's lock before acquiring the replacement, so a nil result left the handle writing with no lock at all. Acquiring the new-path lock first means failure keeps the existing lock instead of reporting success with serialization dropped. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: skip new-path rename lock when a handle already holds it A target file open for write on this mount already carries a lock on newPath; the lock manager does not grant a second lock to the same owner, so the rename would wait on itself until the handle closed. Also avoid locking twice when old and new paths coincide. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: hand the rename's target lock to the migrating handle The rename holds a lock on newPath for its duration, so the response migration's fresh acquisition waited on that same lock until the handle released — under fhLockTable, blocking the handle's own close. Adopt the rename's lock directly; nested move responses still acquire their own. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: move the replaced target's lock to the renamed handle When the target path was already locked by an open handle on this mount, the migrated source handle kept only its stale old-path lock — the target's close would then release the last lock on the new path while the renamed handle was still open. Adopt the replaced handle's lock instead. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: stop the handle lock inside the fh lock on release ReleaseHandle stopped fh.dlmLock before taking the fhLockTable slot, so a rename migration holding that slot could still observe and adopt a lock that was already stopping. Stopping under the fh lock makes the transfer serialize against the release. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: claim the replaced target's lock for the renamed handle When the target path is already locked by an open handle on this mount, adopting it at migration time keeps the renamed path protected after that handle closes, without waiting on a lock this mount already holds. If the handle was released mid-migration the claimed lock is stopped, and a fresh acquire covers the case where it was already gone. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: claim the target handle's lock before the rename runs Skipping the new-path lock when a handle already holds it let that handle's close release the lock mid-rename, leaving the path unguarded until the response migrated it. Take over the lock at check time and hold it for the rename's duration: the response adopts it for the migrating handle, or it returns to the target handle / is released on failure. The target handle lookup also falls back to the entry's stored inode for a forgotten path mapping. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * mount: read handle locks only under the fh lock during rename The loose dlmLock reads raced ReleaseHandle, which now mutates the lock inside the handle lock; check and claim it under the same hold. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
4914c14982 | docs: regenerate star history chart | ||
|
|
2f6c237238 |
filer: keep lazy remote reads from resurrecting deleted paths (#11452)
* filer: keep lazy remote reads from resurrecting deleted paths Under a remote mount with filer.remote.sync as write-back, a path that was deleted or renamed away could come back as a chunkless remote-only entry: between the local delete and the daemon's remote delete, a store miss made maybeLazyFetchFromRemote trust a bucket that was behind the filer. The ghost then outlived the remote object -- HEAD answered 200, GET failed, and nothing cleaned it up. The filer now tombstones paths it deletes under a remote mount, learned both synchronously from its own delete path and from peer metadata events. The lazy fetch and the lazy listing skip a tombstoned path until the path is written again, until the mount's persisted write-back sync offset has passed the delete event (the remote delete has landed), or until a generous TTL covers a mount without a daemon. Fixes #11440 Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: cover recursive remote deletes with an ancestor tombstone A recursive delete now records the directory tombstone before walking children, so a partial traversal or a store that drops the subtree without listing it still leaves every descendant covered. Directory tombstones also subsume older descendant entries on add, descendant adds covered by a standing ancestor are skipped, and an existing tombstone can be refreshed even at capacity. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: scope remote tombstones to the deleted object's generation A remote object whose own mtime postdates the local delete is a new generation, not the one the tombstone hides, so a recreated directory can surface remote writes made after its delete while old-generation objects stay hidden. Lazy fetch now stats the remote object before deciding, listings pass each child's remote mtime, and a sync offset releases a tombstone once it reaches the delete's own timestamp. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: rebuild remote deletion tombstones after restart In-memory tombstones are lost on restart while remote write-back offsets persist, so a filer boot replays the persisted metadata log from the oldest mount offset and folds deletes back into the tombstone set through the same event handler. Lazy remote reads hold off while the replay runs so a pending delete cannot resurrect in the gap. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: release remote tombstones only after their delete event lands The write-back offset orders against event timestamps, but the synchronous delete path recorded tombstones with the local clock before its event was emitted — a later unrelated event could already have pushed the mount's watermark past that guess, releasing the tombstone before the daemon applied the delete. Tombstones recorded ahead of their event are now marked pending and can only be lifted by the event confirming them or by TTL; event-stamped tombstones release through the offset as before. The remote-mtime generation bypass is dropped: remote and filer clocks are independent, and a pending remote delete removes whatever object sits at the path, so a "newer" remote object would only resurrect as a phantom. Tombstoned lookups now skip the remote stat entirely. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: drop dir tombstone when recursive delete fails before listing The ancestor tombstone is recorded before the child listing; if that listing fails nothing was deleted, and the leftover tombstone would hide still-existing remote children for the whole TTL. Tombstones for children already deleted stay, since their remote deletes are still owed. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: block lazy remote reads on startup tombstone rebuild The rebuild gate is now a done-channel set synchronously before the replay goroutine starts, so no lazy read can slip through in between. Reads wait on it with context cancellation instead of returning an empty miss that makes remote-only objects look deleted. The replay start is floored at now-TTL: mounts without a recorded write-back offset previously replayed the whole persisted history, and events older than the TTL would only build already-expired tombstones. The gate check now runs after the mount lookup so replaying the meta log's own directory listings does not deadlock on the gate, and the replay retries with backoff until it succeeds instead of failing open. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: mark restamped tombstone pending until its delete event lands When a local delete raises an existing tombstone's timestamp, the new value is only a local clock guess ahead of that delete's event. Leaving the tombstone un-pending lets a write-back offset release it before the event is actually consumed, reopening the resurrection window. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: bound tombstone replay to the tombstone TTL Persisted-log replay retried forever, keeping lazy remote reads gated indefinitely when the log cannot be read. Cap retries at the tombstone TTL measured from replay start: past that point every tombstone would have expired anyway, so opening the gate loses no protection. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: re-check deletion tombstone before persisting lazy fetch A delete landing while StatFile is in flight passed the earlier tombstone check but still persisted the fetched entry, resurrecting a path whose remote delete is pending. Re-check right before CreateEntry. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: retract a lazily persisted entry when a delete raced the insert The pre-insert tombstone check still leaves a window between the check and the store insert. Since deletes always record the tombstone before removing the entry, a tombstone visible right after a successful insert means the delete already ran: delete the entry back out so the tombstoned path stays deleted. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: note why the replay deadline can safely open the gate Deletes made after startup are captured by the live delete and event paths, so a stalled replay can only be missing pre-restart deletes, all of which are past the tombstone TTL by the deadline. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: retract only the entry a lazy remote read materialized Deleting by path after a raced delete could remove a legitimate rewrite that replaced the fetched entry. Verify the stored entry still matches the remote object (or the just-created directory shape) before deleting, and apply the same post-insert check to lazy listing children. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * filer: require full-entry equality before retracting a lazy entry Remote-only matching still removed a write that had updated the fetched entry, e.g. appended chunks. Compare the persisted entry against what this read materialized; any change means a real update owns the path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
df4995b894 |
volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock (#11443)
* volume server: VolumeMarkReadonly answers NotFound when the volume vanished under the lock
make_volume_readonly looked the volume up, notified the master (step 1),
then took the store write lock (step 2) and marked the volume only `if
let Some(..)`. When the volume left the store during step 1 -- a master
round trip, during which an unmount or a heartbeat expiry can land --
the missing else meant the RPC reported success for a volume the server
no longer has, and step 3 told the master again that it is read-only.
Go's Store.MarkVolumeReadonly (weed/storage/store.go) returns
"volume %d not found" when findVolume comes back nil, and
makeVolumeReadonly (weed/server/volume_grpc_admin.go) returns that error
before the step-3 notification. The Rust step 2 now does the same:
find_volume_mut(vid) -> Status::not_found("volume {vid} not found"), and
the `?` skips step 3, as it already did for a set_read_only_persist
failure. The scrub caller already matches NotFound to skip such a
volume instead of failing the whole report; it now actually gets it.
volume_mark_writable already returns NotFound under its write lock.
The regression test opens the step-1 window deterministically: step 1
awaits the current_master_url read lock, so the test holds its write
guard, lets make_volume_readonly park there after its own lookup
succeeded, unmounts the volume, then releases the guard. With no master
configured the notification is a no-op, so the write lock in step 2 is
the only place left that can notice the volume is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* volume server: trim comments on the vanished-volume mark-readonly path
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
||
|
|
b750853c42 |
shell: refuse s3.bucket.create on an existing bucket (#11455)
* shell: refuse s3.bucket.create on an existing bucket CreateEntry without o_excl replaces the bucket entry, dropping every extended attribute: lifecycle configuration, owner, versioning and the irreversible Object Lock flag. Send o_excl so a re-run fails with 'bucket already exists' instead of silently resetting the bucket. * filer: fail exclusive creates when the lookup itself fails CreateEntry discards FindEntry errors, so an o_excl create hitting a transient store failure would take the insert path and upsert over the entry it was meant to preserve. Propagate the lookup error when o_excl is set; non-exclusive creates keep their existing semantics. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * shell: test s3.bucket.create requests an exclusive create Exercises the command end to end through a fake filer gRPC server and asserts the OExcl flag reaches the wire along with the already-exists error path. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * shell: synchronize captured requests and assert the exact bucket error The fake filer records CreateEntry requests on the gRPC server goroutine, so reads need the same mutex; the test also now checks for the exact "bucket my-bucket already exists" message rather than any error that mentions existence. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
317e756b9a |
volume server: validate ext and collection in gRPC CopyFile/ReceiveFile (Rust) (#11451)
* volume server: validate ext and collection in gRPC CopyFile Port the Go-side checks (checkVolumeFileExtension, checkVolumeCollection) to the Rust volume server so a client-supplied collection or ext carrying a separator or ".." cannot fold a path outside the volume directory. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * volume server: validate ext and collection in gRPC ReceiveFile Same port on the write path: the file ReceiveFile creates is built from client-supplied fields, so reject traversal there too. Reported through the response error field, matching Go's SendAndClose. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> |
||
|
|
5b79f51e3c |
filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks (#11435)
* filer.remote.sync: stamp entries with IF_CHUNKS_EQUAL so a stale write-back cannot delete live chunks updateLocalEntry records the RemoteEntry stamp after an upload by writing the event's entry back with UpdateEntry. The filer deletes every stored chunk absent from an updated entry, so when the file was rewritten while its upload was in flight (or the event is a replay), the stale snapshot deletes the rewrite's chunks: the entry then points at the new fid with no needle behind it, and the rewrite's own upload fails and is skipped as superseded. The stamp write now carries WriteCondition IF_CHUNKS_EQUAL over the event's chunk fids, evaluated by the filer under the path lock. A refused stamp means the filer moved past this event; the superseding event follows in the log and stamps the current entry, so the refusal is logged and skipped like a superseded upload. Reproduction: weed server -filer plus a weed server -s3 remote, remote.mount, filer.remote.sync; hold the remote (docker pause) so one upload stays in flight, rewrite the file through the filer, unpause. Before: the entry's chunk is 404 on every volume server. After: the stale stamp is refused, the rewrite's chunk stays live and reads back after a vacuum. * filer.remote.sync: stamp entries with IF_ENTRY_EQUAL so stale inline content or metadata cannot be restored The IF_CHUNKS_EQUAL guard compared only the chunk fid multiset, so a rewrite that touched inline content or metadata alone still compared equal and the stale snapshot overwrote the live entry. The new clause compares the whole stored entry against the event's entry under the same path lock. * filer: route conditional UpdateEntry to the entry's owner filer Two filers locking the same path locally could still pass a stale condition on the non-owner while the owner's entry had moved on. When a condition or expected_extended precondition is set, forward the request to the entry's owner the same way conditional CreateEntry does, with is_moved bounding the hop. * filer: compare IF_ENTRY_EQUAL against the normalized expected entry FindEntry grows FileSize to the chunk extent, so a raw event entry with FileSize still zero failed the condition on an unchanged file and the stamp was skipped, letting a replay upload the object again. * filer.remote.sync: classify refused stamps by gRPC status only A FailedPrecondition substring in an unrelated error would have been swallowed as a skipped stamp; status.FromError already unwraps. * remote sync: keep the event entry intact for IF_ENTRY_EQUAL --------- Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
635f69a821 |
test(s3tables): add an OLake Iceberg catalog integration test (#11441)
OLake (github.com/datazip-inc/olake) is a CDC and ingestion engine that writes Apache Iceberg. It covers two paths none of the existing catalog suites reach. It is a strict Java Iceberg client. OLake does not write Iceberg from Go — its Go process spawns a Java sidecar over gRPC and writes through the official Apache Iceberg library, because the Go library has no equality deletes and CDC needs them. That makes it the client class weed/s3api/iceberg's metadata compliance backfill exists to serve: the one that fails with "Cannot parse missing long current-snapshot-id" when spec-required keys are omitted. And it produces equality deletes. Its upsert path commits operation=overwrite with an equality-delete file and a delete manifest. ClickHouse, Doris, Trino, Spark and DuckDB all only append, so nothing else in this directory exercises a delete manifest at all. Six subtests: the destination check reaches SUCCEEDED and actually loads the REST catalog; discover enumerates the source; a full sync commits a snapshot; PyIceberg reads back what the Java writer committed; an update plus a re-sync records an overwrite carrying equality deletes with a delete manifest in the current snapshot; and the catalog does not rewrite manifests the official Java writer produced, which gives the manifest-repair path a negative test to go with the ClickHouse positive one. What it deliberately does not assert is a delete-applied read. PyIceberg refuses to scan a table carrying equality deletes (apache/iceberg#6568) while reading its metadata fine, and an engine that can apply them costs a multi-gigabyte image. Recording the commit correctly is the catalog's contract; applying deletes on read is the engine's. The README says so, says the read half was verified by hand once with StarRocks, and warns against later "upgrading" this to a PyIceberg rows read — which would either fail or, if PyIceberg ever starts skipping deletes instead of raising, pass by not looking. The workflow job asserts the suite actually ran rather than trusting a green exit, for the same reason: at least one top-level PASS and zero SKIP. No product change — the destination config is the generic catalog_type=rest with the standard OAuth2 client-credentials flow, and s3_path_style is not even set, since OLake turns it on itself whenever s3_endpoint is non-empty. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
11791fad6a |
filer: resolve the collection a bucket delete drops (#11439)
* filer: resolve the collection a bucket delete drops A bucket delete dropped the collection named after the bucket, which assumes bucket name is collection name. With a collection rule the write path honors, deleting the bucket either orphaned its collection or, when a bucket was named after a shared collection, removed volumes other buckets still write to. Resolve the collection through the same rule chain the write path uses and drop it only when no other bucket resolves there too. A listing failure keeps the collection, the safe side of an unknown. * filer: prove collection exclusivity across all paths before dropping it The sibling-bucket scan missed every non-bucket writer: a broad rule like '/' or '/buckets/', a rule under a surviving bucket, or a rule on an unrelated path can route into the same collection. Check every storage rule's prefix instead, and mirror the grouped gateway's explicit <group>_<bucket> collection, which otherwise resolves a rule-named collection the bucket never wrote to. * s3: let the filer own the collection decision on bucket delete Both entry points deleted a name-derived collection around the filer's own resolved delete, bypassing its exclusivity check and wiping sibling data. The filer now resolves the collection a bucket actually used, including the grouped form. * filer: keep a collection the default write route also uses Rule-less writes outside buckets land in the filer's default collection, so a bucket resolving there shares it with them. |
||
|
|
56d2f05ccd |
topology: wake the vacuum dispatcher when a worker frees quota (#11436)
* topology: wake the vacuum dispatcher when a worker frees quota The dispatch loop slept a fixed 10s whenever every pending volume was waiting for a per-server quota slot, so a sweep took volumes x 10s regardless of how fast the compactions were. Workers now signal on a buffered channel after crediting quota; the dispatcher waits on it with the 10s sleep kept only as a timeout. * master: add -vacuumIntervalSeconds to tune the automatic sweep interval The 14-minute base interval was a literal inside the refresh loop while every neighbouring vacuum knob was already a flag. Defaults to 840s, unchanged. * topology: keep the 14 minute floor on the vacuum interval A zero-valued MasterOption or a negative -vacuumIntervalSeconds left the sweep sleeping only its jitter, so treat non-positive intervals as the previous default. |
||
|
|
8c1be63c92 |
ecbalancer: let a non-overflow parity shard leave a data-bearing rack (#11438)
The parity pass only queued shards past the per-type cap, so a single parity shard sharing a rack with data was never a move candidate even when an empty data-free rack existed (2+1 over 3 DCs settled 2/1/0). Non-overflow candidates now move too, but only to a rack without data; overflow shards keep the existing data-rack fallback. |
||
|
|
bb9942c646 |
volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes (#11413)
* volume server: sweep stale EC artifacts before VolumeEcShardsGenerate re-encodes The Rust VolumeEcShardsGenerate went straight into write_ec_files: no unload of an already-mounted EC volume and no stale-artifact sweep. Only .ec00..ecNN on the encoding disk were truncated, so a retry could mix two encode runs. A stale N.ec03 left on a sibling disk survived, reconcile later mounted it against the new .ecx, and the new .vif made the encode_ts_ns identity guard pass, so reads served old-run bytes at new-run offsets. Mirror Go's VolumeEcShardsGenerate (#9880 / #9953): UnloadEcVolume on every disk, then removeStaleEcArtifacts on every disk location before encoding. remove_ec_volume_files_full_teardown already has removeStaleEcArtifacts' semantics (.ec00..ec31, .ecx/.ecj/.ecsum[.vN] in both the data and idx dirs, .vif only on a shard-only disk; never the source .dat/.idx), so reuse it. Add Store::unload_ec_volume, which unlike remove_ec_volume does not stop at the first disk and closes the descriptors so the unlink frees the inodes. The store write lock covers only unload + sweep, not the encode. The failure arm now also drops the generation-0 .ecsum, as Go's defer does. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume server: wake the heartbeat after VolumeEcShardsGenerate unloads shards The pre-encode unload drops mounted EC shards from memory, but unlike every other unmount path it did not wake the heartbeat, so the master kept routing reads to shards this server no longer serves until the next pulse. Notify once the store lock is released, and before the sweep error propagates: a failed sweep has unloaded the shards too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume server: clean up encode artifacts when the .vif write fails too Go's shouldCleanup defer covers every error before the .vif commits, not just a failed encode. A serialize or write failure on the .vif left the fresh .ecNN/.ecx/.ecsum behind, which the next generate would have to rely on the new sweep to remove. Extract the cleanup and run it on the .vif error paths as well. * volume server: write the EC .vif atomically Go's SaveVolumeInfo writes a temp file, syncs it, and renames it over the target, so a failed write leaves the previous metadata intact and a read-only .vif fails the save. The direct fs::write truncated the file first, so a write or sync failure could leave an empty .vif even after cleanup_encode removed the generated shards. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
f0afcf904d |
volume: an EC volume needs an .ecx to mount, and a 0-byte stub never outranks a real index (#11415)
* volume: an EC volume needs a non-empty .ecx to mount Two gaps against Go in how the Rust volume server treats the .ecx. EcVolume::new mounted with no index at all. The per-shard VolumeEcShardsMount path picks the disk by shard file alone, so a shard whose .ecx was on no local directory still registered and was advertised to the master; every VolumeEcShardRead then failed with "ecx file not open", and add_shard's 0-byte guard was neutralised because ecx_file_size stayed 0. Go's NewEcVolume returns an error wrapping os.ErrNotExist. EcVolume::new now fails with NotFound, and Store::mount_ec_shard looks up the .ecx owner across all disks first (findEcxIdxDirForVolume) so a shard on a sibling disk of its index still mounts instead of turning into a hard failure. A 0-byte .ecx stub, as left by a failed EC distribute copy, counted as a valid index. Go requires Size() > 0 wherever the file steers a decision: HasEcxFileOnDisk, findEcxIdxDirForVolume, indexEcxOwners (shared by reconcile and mirror), and VolumeEcShardsCopy removes a copied 0-byte .ecx and fails the copy. Mirror each through one is_usable_ecx_file helper. NewEcVolume itself still accepts a lone 0-byte .ecx as a legitimate empty index, but prefers a non-empty copy, local directory first, over a stub in the other directory; the resolution in EcVolume::new now follows the same order. Tests that mounted EC volumes without any .ecx get a real fixture. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: mount_ec_shard tries every disk; reconcile ignores a 0-byte local .ecx mount_ec_shard returned the first disk's error, so an unusable shard copy (a 0-byte .ecNN left by an interrupted move) hid a good copy on the next disk. Like Go's MountEcShards, keep scanning: NotFound means "not this disk", any other failure is collected, and an all-disks-fail error names every disk tried. "No .ecx on any local disk" is now told apart from "shard not on this server". The orphan-shard reconcile took its locally-mirrored fast path whenever a local .ecx existed at all. A 0-byte stub there registered the shards against an empty index while the owner index skipped that same stub. Go gates the fast path on HasEcxFileOnDisk; do the same. ec_local_ecx_path loses its last production caller and becomes test-only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: match Go's mount error text and skip the owner stat on the owning disk MountEcShards in Go skips the HasEcxFileOnDisk stat when the disk's own directories already hold the .ecx, dedups a shared -dir.idx across locations in findEcxIdxDirForVolume, and reports "load failures" with the same wording. Also drop two issue-number references from comments. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
c1ccbcda13 |
volume server: the write queue answers uploads with the needle's real ETag (#11414)
With SEAWEED_WRITE_QUEUE=1 every upload came back with ETag "00000000". The upload handler built the needle with Needle::default(), so its checksum was CRC(0), and handed a clone of it to the queue. The CRC was only computed in the write path, on the worker's clone, and WriteResult carries no checksum back, so n.etag() in the handler formatted the zero checksum. The direct path writes through &mut n and was correct. Compute the checksum in the handler while building the needle, the way Go's CreateNeedleFromRequest does, over the same bytes the write path hashes (the stored data, gzipped or not). The ETag and the has-name flag are read before the write, so the needle is moved into the queue instead of cloned, which also drops a full payload copy per queued upload. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
94a68fa9b9 |
volume: walk_index_file keeps row alignment across short reads (#11412)
* volume: walk_index_file keeps row alignment across short reads walk_index_file issued one Read::read per batch and decoded whatever came back. Read::read may legally return a short count that is not a multiple of the 17-byte entry size (FUSE and network filesystems, a BufReader whose capacity is not a multiple of 17). The split entry at the end of the batch was dropped with no carry and the next read started mid-entry, so every later row was decoded from misaligned bytes and fed to the index as a garbage key/offset/size. This function backs every in-memory index load. Go's WalkIndexFile is immune because it reads through io.ReaderAt, which returns a full buffer or an error. Fill the batch buffer until it is full or the reader reports EOF, retrying ErrorKind::Interrupted, and only then decode whole entries. Reads stay batched at ROWS_TO_READ entries. EOF semantics are unchanged and match Go: on io.EOF Go decodes the whole entries in the final buffer, ignores a trailing partial entry and returns nil. A torn final entry is still skipped without an error here. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * volume: trim walk_index_file comments The batch-fill loop and the ShortReader test helper each carried a paragraph where a sentence suffices. --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com> |
||
|
|
b3a8701989 |
lance: authenticate the catalog with Bearer tokens and x-api-key (#11431)
* lance: accept OAuth2 bearer tokens for catalog auth Lance and LanceDB clients can only send OAuth2 / Bearer / API-Key headers on catalog calls, never SigV4, so behind an auth-enabled S3 gateway every namespace request failed with 403 Access Denied. Mirror the Iceberg catalog's OAuth2 support: POST /oauth/token accepts an S3 access key / secret key as client_id / client_secret, validates them against IAM, and returns a signed JWT. The Auth middleware accepts that token as a Bearer credential before falling through to SigV4. Closes #11430 * lance: accept x-api-key header carrying an S3 credential The Lance namespace spec's third auth scheme maps api_key onto the x-api-key header. Accept "access_key:secret_key" there and validate it against IAM, so clients that only hold static headers can authenticate without minting a token first. * lance: answer invalid_client with the Basic challenge RFC 6749 5.2 requires a 401 from the token endpoint to carry WWW-Authenticate matching the scheme the client used, so it knows how to retry. * lance: cap the token endpoint request body /oauth/token is unauthenticated, so ParseForm needs the same size bound decodeBody applies to every other catalog request. * lance: keep query strings out of request logs /oauth/token rejects a client_secret sent in the query, but the logging middleware and the catch-all wrote RequestURI to the log before that rejection ran. Log the path alone so a mis-sent secret never reaches the log. * lance: log the escaped path, not the decoded one URL.Path decodes percent escapes, so a request like /%0aFORGED could split log lines. EscapedPath keeps the encoding while still dropping the query string. |
||
|
|
b9ad62fc16 |
[Volume] Keep DAT and index state consistent after async batch Sync failure (#11425)
* fix 11400 * persist failed-recovery quarantine and harden rollback - record the unavailable state in a .unavailable marker, fsync it, and re-arm it on load so a restart cannot serve an unverified pair - quarantine the volume so heartbeats stop advertising it - block MarkVolumeWritable while unavailable, rechecked under noWriteLock - fail every request of a failed batch, not only the succeeded ones - restore the needle map and truncate .dat on inline fsync rollback failure - add truncateIndex for the sorted-file needle map - mirror the fail-closed semantics in the Rust volume server * volume: erase rolled-back mappings instead of leaving tombstones A rolled-back batch or failed inline write used Delete() to undo a needle that did not exist beforehand, leaving a tombstoned map entry whose stale offset makes the next write to that needle fail reading a header that no longer exists. Add removeMapping/restoreMapping to the mappers so recovery erases entries that were absent before the batch and reinstates the exact prior offset/size for ones that were, including tombstones. The index row still goes through Delete so a replay forgets the needle. * volume: gate bulk readers on unavailable and fsync the marker's dir - fsync_dir(&self.dir) synced the volume dir's parent, not the dir holding .unavailable; pass the marker path so the create survives a host crash - export UnavailableError and check it in ReadAllNeedles, VolumeTailSender, VolumeIncrementalCopy, and IncrementalBackup so replica-sync paths cannot stream or append data from an unverified .dat/.idx pair; mirror on the Rust side via read_dat_slice, read_all_needles, dat_scan_plan, and the incremental-copy handler * volume: drop issue references from comments near touched code * volume: stop active scans when the volume becomes unavailable The stream entry-point checks ran once per RPC, so a volume quarantined by a failed recovery mid-scan kept serving data. Recheck availability per needle/chunk on the detached read paths: tail scan and heartbeat, read-all, incremental copy, incremental backup writes, and the Rust StreamingBody chunk reads. Rust incremental copy also rejects a quarantined volume before sync_to_disk touches the backend. --------- Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
d848b8ed00 |
rust volume: regression test for makeup_diff replay across a 32 GiB offset boundary (#11410) (#11427)
* rust volume: test makeup_diff replay across a 32 GiB offset boundary Issue #11410 corrupted a replayed write's index offset in Go's makeupDiff by patching only four of the five offset bytes. The Rust makeup_diff already encodes the whole offset through idx_entry_to_bytes and Offset::from_actual_offset; this adds the mirror of TestConcurrentWriteCrossesOffsetBoundary so a regression would fail here the same way it does under -tags=5BytesOffset on the Go side. Sparse-truncate the .dat to 64 GiB, compact, write, commit: the index offset must equal the .cpd size and the needle must stay readable through a second vacuum. Gated on the 5bytes feature since a 64 GiB .dat exceeds the 32 GiB range of 4-byte offsets. * rust volume: skip the offset-boundary replay test on Windows Windows set_len allocates the full 64 GiB extension instead of a sparse range, so the test fails with StorageFull on CI runners. Gate it to unix, where set_len leaves the extension unallocated. |
||
|
|
196c71b613 | docs: regenerate star history chart | ||
|
|
559ec33498 |
build(deps): bump github/codeql-action from 4.38.0 to 4.38.1 (#11423)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.38.0 to 4.38.1. - [Release notes](https://github.com/github/codeql-action/releases) - [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md) - [Commits](https://github.com/github/codeql-action/compare/v4.38.0...v4.38.1) --- updated-dependencies: - dependency-name: github/codeql-action dependency-version: 4.38.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d6397fdc50 |
build(deps): bump dorny/paths-filter from 3 to 4 (#11422)
Bumps [dorny/paths-filter](https://github.com/dorny/paths-filter) from 3 to 4. - [Release notes](https://github.com/dorny/paths-filter/releases) - [Changelog](https://github.com/dorny/paths-filter/blob/master/CHANGELOG.md) - [Commits](https://github.com/dorny/paths-filter/compare/v3...v4) --- updated-dependencies: - dependency-name: dorny/paths-filter dependency-version: '4' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
264030c08c |
build(deps): bump docker/setup-qemu-action from 4.3.0 to 4.4.0 (#11421)
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 4.3.0 to 4.4.0. - [Release notes](https://github.com/docker/setup-qemu-action/releases) - [Commits](https://github.com/docker/setup-qemu-action/compare/v4.3.0...v4.4.0) --- updated-dependencies: - dependency-name: docker/setup-qemu-action dependency-version: 4.4.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
8099e71934 |
build(deps): bump golang.org/x/crypto from 0.56.0 to 0.57.0 (#11419)
Bumps [golang.org/x/crypto](https://github.com/golang/crypto) from 0.56.0 to 0.57.0. - [Commits](https://github.com/golang/crypto/compare/v0.56.0...v0.57.0) --- updated-dependencies: - dependency-name: golang.org/x/crypto dependency-version: 0.57.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
63eb67c9da |
build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts from 1.50.0 to 1.51.0 (#11420)
build(deps): bump github.com/aws/aws-sdk-go-v2/service/sts Bumps [github.com/aws/aws-sdk-go-v2/service/sts](https://github.com/aws/aws-sdk-go-v2) from 1.50.0 to 1.51.0. - [Release notes](https://github.com/aws/aws-sdk-go-v2/releases) - [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.50.0...service/s3/v1.51.0) --- updated-dependencies: - dependency-name: github.com/aws/aws-sdk-go-v2/service/sts dependency-version: 1.51.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |