* filer: keep lazy remote reads from resurrecting deleted paths
Under a remote mount with filer.remote.sync as write-back, a path that
was deleted or renamed away could come back as a chunkless remote-only
entry: between the local delete and the daemon's remote delete, a store
miss made maybeLazyFetchFromRemote trust a bucket that was behind the
filer. The ghost then outlived the remote object -- HEAD answered 200,
GET failed, and nothing cleaned it up.
The filer now tombstones paths it deletes under a remote mount, learned
both synchronously from its own delete path and from peer metadata
events. The lazy fetch and the lazy listing skip a tombstoned path until
the path is written again, until the mount's persisted write-back sync
offset has passed the delete event (the remote delete has landed), or
until a generous TTL covers a mount without a daemon.
Fixes#11440
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: cover recursive remote deletes with an ancestor tombstone
A recursive delete now records the directory tombstone before walking
children, so a partial traversal or a store that drops the subtree
without listing it still leaves every descendant covered. Directory
tombstones also subsume older descendant entries on add, descendant
adds covered by a standing ancestor are skipped, and an existing
tombstone can be refreshed even at capacity.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: scope remote tombstones to the deleted object's generation
A remote object whose own mtime postdates the local delete is a new
generation, not the one the tombstone hides, so a recreated directory
can surface remote writes made after its delete while old-generation
objects stay hidden. Lazy fetch now stats the remote object before
deciding, listings pass each child's remote mtime, and a sync offset
releases a tombstone once it reaches the delete's own timestamp.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: rebuild remote deletion tombstones after restart
In-memory tombstones are lost on restart while remote write-back
offsets persist, so a filer boot replays the persisted metadata log
from the oldest mount offset and folds deletes back into the tombstone
set through the same event handler. Lazy remote reads hold off while
the replay runs so a pending delete cannot resurrect in the gap.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: release remote tombstones only after their delete event lands
The write-back offset orders against event timestamps, but the synchronous
delete path recorded tombstones with the local clock before its event was
emitted — a later unrelated event could already have pushed the mount's
watermark past that guess, releasing the tombstone before the daemon
applied the delete. Tombstones recorded ahead of their event are now
marked pending and can only be lifted by the event confirming them or by
TTL; event-stamped tombstones release through the offset as before.
The remote-mtime generation bypass is dropped: remote and filer clocks
are independent, and a pending remote delete removes whatever object sits
at the path, so a "newer" remote object would only resurrect as a
phantom. Tombstoned lookups now skip the remote stat entirely.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: drop dir tombstone when recursive delete fails before listing
The ancestor tombstone is recorded before the child listing; if that
listing fails nothing was deleted, and the leftover tombstone would hide
still-existing remote children for the whole TTL. Tombstones for children
already deleted stay, since their remote deletes are still owed.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: block lazy remote reads on startup tombstone rebuild
The rebuild gate is now a done-channel set synchronously before the
replay goroutine starts, so no lazy read can slip through in between.
Reads wait on it with context cancellation instead of returning an
empty miss that makes remote-only objects look deleted.
The replay start is floored at now-TTL: mounts without a recorded
write-back offset previously replayed the whole persisted history, and
events older than the TTL would only build already-expired tombstones.
The gate check now runs after the mount lookup so replaying the meta
log's own directory listings does not deadlock on the gate, and the
replay retries with backoff until it succeeds instead of failing open.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: mark restamped tombstone pending until its delete event lands
When a local delete raises an existing tombstone's timestamp, the new
value is only a local clock guess ahead of that delete's event. Leaving
the tombstone un-pending lets a write-back offset release it before the
event is actually consumed, reopening the resurrection window.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: bound tombstone replay to the tombstone TTL
Persisted-log replay retried forever, keeping lazy remote reads gated
indefinitely when the log cannot be read. Cap retries at the tombstone
TTL measured from replay start: past that point every tombstone would
have expired anyway, so opening the gate loses no protection.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: re-check deletion tombstone before persisting lazy fetch
A delete landing while StatFile is in flight passed the earlier
tombstone check but still persisted the fetched entry, resurrecting a
path whose remote delete is pending. Re-check right before CreateEntry.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: retract a lazily persisted entry when a delete raced the insert
The pre-insert tombstone check still leaves a window between the check
and the store insert. Since deletes always record the tombstone before
removing the entry, a tombstone visible right after a successful insert
means the delete already ran: delete the entry back out so the
tombstoned path stays deleted.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: note why the replay deadline can safely open the gate
Deletes made after startup are captured by the live delete and event
paths, so a stalled replay can only be missing pre-restart deletes, all
of which are past the tombstone TTL by the deadline.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: retract only the entry a lazy remote read materialized
Deleting by path after a raced delete could remove a legitimate rewrite
that replaced the fetched entry. Verify the stored entry still matches
the remote object (or the just-created directory shape) before deleting,
and apply the same post-insert check to lazy listing children.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: require full-entry equality before retracting a lazy entry
Remote-only matching still removed a write that had updated the fetched
entry, e.g. appended chunks. Compare the persisted entry against what
this read materialized; any change means a real update owns the path.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* filer: resolve the collection a bucket delete drops
A bucket delete dropped the collection named after the bucket, which
assumes bucket name is collection name. With a collection rule the
write path honors, deleting the bucket either orphaned its collection
or, when a bucket was named after a shared collection, removed volumes
other buckets still write to.
Resolve the collection through the same rule chain the write path uses
and drop it only when no other bucket resolves there too. A listing
failure keeps the collection, the safe side of an unknown.
* filer: prove collection exclusivity across all paths before dropping it
The sibling-bucket scan missed every non-bucket writer: a broad rule like
'/' or '/buckets/', a rule under a surviving bucket, or a rule on an
unrelated path can route into the same collection. Check every storage
rule's prefix instead, and mirror the grouped gateway's explicit
<group>_<bucket> collection, which otherwise resolves a rule-named
collection the bucket never wrote to.
* s3: let the filer own the collection decision on bucket delete
Both entry points deleted a name-derived collection around the filer's
own resolved delete, bypassing its exclusivity check and wiping sibling
data. The filer now resolves the collection a bucket actually used,
including the grouped form.
* filer: keep a collection the default write route also uses
Rule-less writes outside buckets land in the filer's default collection,
so a bucket resolving there shares it with them.
* master: bound each volume server DeleteCollection, and finish the fan-out
A collection delete fanned out to every volume server holding it with
context.Background(), so a server that accepted the connection and then
went quiet held the whole delete open with nothing to end it. Each RPC is
bounded now, on the same budget allocateVolumeTimeout gives the other
master-to-volume-server admin RPC. The volume server runs the delete to
completion regardless of the request context, so giving up costs the
confirmation and not the deletion.
The walk itself is the caller's, not a per-server one:
- It outlives the caller. A cancelled request must not abandon a
destructive fan-out part-done, with volumes left behind and no request
still running to come back for them.
- It no longer stops at the first server that refuses, which left the
collection on every server after it in the list. The first failure is
still what is reported, and the collection stays in the topology so a
later delete comes back for the rest.
- It sends one RPC per server rather than one per replica.
ListVolumeServers reports a node once for every replica it holds, while
DeleteCollection removes the whole collection from the server it
reaches, so a collection with thousands of volumes repeated the same
whole-collection delete thousands of times over.
Both passes run too. Returning after a failed normal pass left the
collection's EC shards in place with nothing left to retry them.
Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
* master: delete the EC shards behind /col/delete too
The HTTP handler carried its own copy of the volume-server walk and only
ever ran the normal pass, so a collection deleted through it kept its EC
shards. It shares the gRPC path now, which also gets it the bounded RPCs
and the one-per-server fan-out.
Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
* filer: bound the collection delete a bucket delete leaves behind
Deleting a bucket entry deletes its collection afterwards, deliberately
detached from the request so a client that hangs up cannot strand the
bucket's volumes. Detached meant unbounded, though: with the master down
or mid-election the wait for a leader has nothing to end it, so the
handler parks, and the client retrying behind it parks another.
It keeps outliving the request and now carries a deadline of its own. The
budget bounds the wait, not the work: the master keeps deleting on its own
fan-out once asked, so giving up costs the confirmation.
Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
* s3api: bound the collection RPCs a bucket creation and deletion issue
Neither carried a deadline, so a transient failure anywhere down the chain
held the S3 request open until the client gave up on it. Both budgets are
taken outside the filer failover walk, so one budget covers the whole walk
rather than granting each filer a fresh one.
The walk itself stops when that budget is spent, and stops without blaming
anyone: the caller's own expiry is not evidence against the filer that was
answering, and the next filer has no time left to answer in either.
Recorded as a filer failure, a slow master upstream would flag every filer
in the walk, and the three failures that open the circuit take unrelated
object reads down with them.
Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
* s3api: a failed collection listing no longer fails a bucket creation
PutBucket lists collections to notice a leftover one it is about to reuse.
The result feeds a warning and nothing else -- s3a.exists is what decides
whether the bucket already exists -- yet a transient failure of that
listing returned 500 and refused the creation. It is advisory now, so a
failure is logged and the creation continues, exactly as it does when the
listing returns false.
Claude-Session: https://claude.ai/code/session_01EnB1fbryyKc2LetRZxQPTP
* util, pb: classify a filer error by the status the server sent
DoSeaweedListWithSnapshot wrapped a failed ListEntries with %v, dropping the
gRPC status, so IsTransientError fell back to matching substrings against a
message that now held the caller's path. Keep the status with %w and let it
decide, reading the server's own text rather than the wrapper's.
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* s3: keep the bucket and prefix out of the list retry decision
A bucket named transport, or a prefix under logs/unavailable/, made a
PermissionDenied listing look transient and got it retried; a key holding the
not-found sentence suppressed a retry that should have run. Both checks now
read the filer's status, and only fall back to the text when there is none.
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* filer, s3: classify a delete failure before the path is wrapped into it
The filer put the non-empty-folder marker behind its own "delete directory %s"
wrapper and the gateway matched it as a substring, so a key named after the
marker turned a real delete failure into the demote-the-marker no-op and the
request answered 204. Keep the marker leading the message that crosses the
wire, turn it back into a sentinel where the response is read, and match that.
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* wdclient: bound the wait for a master leader by the caller's context
WithClient waited on GetMaster with context.Background(), so a caller that
arrived while no master leader was known parked in a 200ms poll loop until one
appeared, whatever deadline it had already set on the RPC. Each retry above it
then left another goroutine in the same wait.
Take the context in WithClient and WithClientCustomGetMaster and hand it to
GetMaster, and stop the retry loop once it is done. The dial keeps
context.Background(): fn brings its own RPC context, so a cancellation seen
here cannot be attributed to the shared connection.
Call sites pass whatever they hold: the request context in the filer's
CollectionList, DeleteCollection and Statistics handlers and in the credential
store's propagation, the operation context in the shell's s3.bucket.delete and
the kafka gateway's broker and filer discovery, and context.Background() where
there is none - the shell commands, the admin dashboard wrapper, and the
exclusive locker's initial lease. The locker's release keeps its own
uncancelled context so a slow unlock cannot turn into a ghost lock.
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* wdclient: test that WithClient gives up with the caller's context
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* wdclient: cut the master retry backoff short when the caller gives up
util.Retry sleeps unconditionally between attempts, so a transient error
arriving just before the caller's deadline still cost it a full backoff step.
Use the context-aware util.RetryWithBackoff, the same helper the volume lookup
in this file already uses.
Two call sites went with it: the shell's lock-holder lookup builds its three
second bound before WithClient so it also covers finding the leader, as its
comment already promised, and the filer's post-delete collection cleanup goes
back to an uncancelled context - the entry is already gone, so a caller that
hung up must not leave the collection behind.
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* wdclient: test that a cancel during backoff ends the retry
Claude-Session: https://claude.ai/code/session_01BjDWtZsCoZY6x4pdDmGWxU
* filer: do not sweep children when deleting a folder non-recursively
doBatchDeleteFolderMetaAndData lists a folder and bails out if it has any
children, then calls Store.DeleteFolderChildren unconditionally. On the
non-recursive path that bulk sweep has nothing legitimate to remove: it only
runs once the listing came back empty, so the sole rows it can delete are
ones inserted after the check.
The S3 empty-folder cleaner deletes through this path, so a PUT landing
between the listing and the sweep loses its entry after the write was already
acknowledged. Neither side sees an error - the client has its 200 and the
cleaner logs an ordinary empty-folder deletion - and the chunks leak, since
the cleaner passes shouldDeleteChunks=false and nothing was enumerated to
collect. Workloads that scatter objects over many shallow prefixes empty and
refill those folders constantly, which is what makes the window reachable.
Sweep only when the delete is recursive, or when the whole-bucket shortcut
skipped the listing and depends on it.
Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r
* filer: pin the folder entry removal left by the racing-child test
The surviving entry is reachable by path but drops out of listings until the
folder comes back, and nothing in the test said so. Assert it, so the exposure
that remains after this change is visible rather than implied.
Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r
Reads of remote-backed entries now record hit or miss in
SeaweedFS_remote_cache_read_total{source,bucket,result} on the filer HTTP
path and the S3 gateway, so cache effectiveness of mounted buckets can be
graphed. Inline-content entries count as hits since they are served
locally without chunks. The filer purges the per-bucket series when the
bucket directory is deleted, so a standalone filer does not accumulate
series across bucket delete/recreate churn.
* filer: propagate lazy metadata deletes to remote mounts
Delete operations now call the remote backend for mounted remote-only entries before removing filer metadata, keeping remote state aligned and preserving retry semantics on remote failures.
Made-with: Cursor
* filer: harden remote delete metadata recovery
Persist remote-delete metadata pendings so local entry removal can be retried after failures, and return explicit errors when remote client resolution fails to prevent silent local-only deletes.
Made-with: Cursor
* filer: streamline remote delete client lookup and logging
Avoid a redundant mount trie traversal by resolving the remote client directly from the matched mount location, and add parity logging for successful remote directory deletions.
Made-with: Cursor
* filer: harden pending remote metadata deletion flow
Retry pending-marker writes before local delete, fail closed when marking cannot be persisted, and start remote pending reconciliation only after the filer store is initialised to avoid nil store access.
Made-with: Cursor
* filer: avoid lazy fetch in pending metadata reconciliation
Use a local-only entry lookup during pending remote metadata reconciliation so cache misses do not trigger remote lazy fetches.
Made-with: Cursor
* filer: serialise concurrent index read-modify-write in pending metadata deletion
Add remoteMetadataDeletionIndexMu to Filer and acquire it for the full
read→mutate→commit sequence in markRemoteMetadataDeletionPending and
clearRemoteMetadataDeletionPending, preventing concurrent goroutines
from overwriting each other's index updates.
Made-with: Cursor
* filer: start remote deletion reconciliation loop in NewFiler
Move the background goroutine for pending remote metadata deletion
reconciliation from SetStore (where it was gated by sync.Once) to
NewFiler alongside the existing loopProcessingDeletion goroutine.
The sync.Once approach was problematic: it buried a goroutine launch
as a side effect of a setter, was unrecoverable if the goroutine
panicked, could race with store initialisation, and coupled its
lifecycle to unrelated shutdown machinery. The existing nil-store
guard in reconcilePendingRemoteMetadataDeletions handles the window
before SetStore is called.
* filer: skip remote delete for replicated deletes from other filers
When isFromOtherCluster is true the delete was already propagated to
the remote backend by the originating filer. Repeating the remote
delete on every replica doubles API calls, and a transient remote
failure on the replica would block local metadata cleanup — leaving
filers inconsistent.
* filer: skip pending marking for directory remote deletes
Directory remote deletes are idempotent and do not need the
pending/reconcile machinery that was designed for file deletes where
the local metadata delete might fail after the remote object is
already removed.
* filer: propagate remote deletes for children in recursive folder deletion
doBatchDeleteFolderMetaAndData iterated child files but only called
NotifyUpdateEvent and collected chunks — it never called
maybeDeleteFromRemote for individual children. This left orphaned
objects in the remote backend when a directory containing remote-only
files was recursively deleted.
Also fix isFromOtherCluster being hardcoded to false in the recursive
call to doBatchDeleteFolderMetaAndData for subdirectories.
* filer: simplify pending remote deletion tracking to single index key
Replace the double-bookkeeping scheme (individual KV entry per path +
newline-delimited index key) with a single index key that stores paths
directly. This removes the per-path KV writes/deletes, the base64
encoding round-trip, and the transaction overhead that was only needed
to keep the two representations in sync.
* filer: address review feedback on remote deletion flow
- Distinguish missing remote config from client initialization failure
in maybeDeleteFromRemote error messages.
- Use a detached context (30s timeout) for pending-mark and
pending-clear KV writes so they survive request cancellation after
the remote object has already been deleted.
- Emit NotifyUpdateEvent in reconcilePendingRemoteMetadataDeletions
after a successful retry deletion so downstream watchers and replicas
learn about the eventual metadata removal.
* filer: remove background reconciliation for pending remote deletions
The pending-mark/reconciliation machinery (KV index, mutex, background
loop, detached contexts) handled the narrow case where the remote
object was deleted but the subsequent local metadata delete failed.
The client already receives the error and can retry — on retry the
remote not-found is treated as success and the local delete proceeds
normally. The added complexity (and new edge cases around
NotifyUpdateEvent, multi-filer consistency during reconciliation, and
context lifetime) is not justified for a transient store failure the
caller already handles.
Remove: loopProcessingRemoteMetadataDeletionPending,
reconcilePendingRemoteMetadataDeletions, markRemoteMetadataDeletionPending,
clearRemoteMetadataDeletionPending, listPendingRemoteMetadataDeletionPaths,
encodePendingRemoteMetadataDeletionIndex, FindEntryLocal, and all
associated constants, fields, and test infrastructure.
* filer: fix test stubs and add early exit on child remote delete error
- Refactor stubFilerStore to release lock before invoking callbacks and
propagate callback errors, preventing potential deadlocks in tests
- Implement ListDirectoryPrefixedEntries with proper prefix filtering
instead of delegating to the unfiltered ListDirectoryEntries
- Add continue after setting err on child remote delete failure in
doBatchDeleteFolderMetaAndData to skip further processing of the
failed entry
* filer: propagate child remote delete error instead of silently continuing
Replace `continue` with early `break` when maybeDeleteFromRemote fails
for a child entry during recursive folder deletion. The previous
`continue` skipped the error check at the end of the loop body, so a
subsequent successful entry would overwrite err and the remote delete
error was silently lost. Now the loop breaks, the existing error check
returns the error, and NotifyUpdateEvent / chunk collection are
correctly skipped for the failed entry.
* filer: delete remote file when entry has Remote pointer, not only when remote-only
Replace IsInRemoteOnly() guard with entry.Remote == nil check in
maybeDeleteFromRemote. IsInRemoteOnly() requires zero local chunks and
RemoteSize > 0, which incorrectly skips remote deletion for cached
files (local chunks exist) and zero-byte remote objects (RemoteSize 0).
The correct condition is whether the entry has a remote backing object
at all.
---------
Co-authored-by: Chris Lu <chris.lu@gmail.com>
* mount: let filer handle chunk deletion decision
Remove chunk deletion decision from FUSE mount's Unlink operation.
Previously, the mount decided whether to delete chunks based on
its locally cached entry's HardLinkCounter, which could be stale.
Now always pass isDeleteData=true and let the filer make the
authoritative decision based on its own data. This prevents
potential inconsistencies when:
- The FUSE mount's cached entry is stale
- Race conditions occur between multiple mounts
- Direct filer operations change hard link counts
* filer: check hard link counter before deleting chunks
When deleting an entry, only delete the underlying chunks if:
1. It is not a hard link
2. OR it is the last hard link (counter <= 1)
This protects against data loss when a client (like FUSE mount)
requests chunk deletion for a file that has multiple hard links.