Files
seaweedfs/weed/filer/filer_lazy_remote.go
T
Chris LuGitHubDevin <158243242+devin-ai-integration[bot]@users.noreply.github.com>Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2f6c237238 filer: keep lazy remote reads from resurrecting deleted paths (#11452)
* filer: keep lazy remote reads from resurrecting deleted paths

Under a remote mount with filer.remote.sync as write-back, a path that
was deleted or renamed away could come back as a chunkless remote-only
entry: between the local delete and the daemon's remote delete, a store
miss made maybeLazyFetchFromRemote trust a bucket that was behind the
filer. The ghost then outlived the remote object -- HEAD answered 200,
GET failed, and nothing cleaned it up.

The filer now tombstones paths it deletes under a remote mount, learned
both synchronously from its own delete path and from peer metadata
events. The lazy fetch and the lazy listing skip a tombstoned path until
the path is written again, until the mount's persisted write-back sync
offset has passed the delete event (the remote delete has landed), or
until a generous TTL covers a mount without a daemon.

Fixes #11440

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: cover recursive remote deletes with an ancestor tombstone

A recursive delete now records the directory tombstone before walking
children, so a partial traversal or a store that drops the subtree
without listing it still leaves every descendant covered. Directory
tombstones also subsume older descendant entries on add, descendant
adds covered by a standing ancestor are skipped, and an existing
tombstone can be refreshed even at capacity.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: scope remote tombstones to the deleted object's generation

A remote object whose own mtime postdates the local delete is a new
generation, not the one the tombstone hides, so a recreated directory
can surface remote writes made after its delete while old-generation
objects stay hidden. Lazy fetch now stats the remote object before
deciding, listings pass each child's remote mtime, and a sync offset
releases a tombstone once it reaches the delete's own timestamp.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: rebuild remote deletion tombstones after restart

In-memory tombstones are lost on restart while remote write-back
offsets persist, so a filer boot replays the persisted metadata log
from the oldest mount offset and folds deletes back into the tombstone
set through the same event handler. Lazy remote reads hold off while
the replay runs so a pending delete cannot resurrect in the gap.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: release remote tombstones only after their delete event lands

The write-back offset orders against event timestamps, but the synchronous
delete path recorded tombstones with the local clock before its event was
emitted — a later unrelated event could already have pushed the mount's
watermark past that guess, releasing the tombstone before the daemon
applied the delete. Tombstones recorded ahead of their event are now
marked pending and can only be lifted by the event confirming them or by
TTL; event-stamped tombstones release through the offset as before.

The remote-mtime generation bypass is dropped: remote and filer clocks
are independent, and a pending remote delete removes whatever object sits
at the path, so a "newer" remote object would only resurrect as a
phantom. Tombstoned lookups now skip the remote stat entirely.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: drop dir tombstone when recursive delete fails before listing

The ancestor tombstone is recorded before the child listing; if that
listing fails nothing was deleted, and the leftover tombstone would hide
still-existing remote children for the whole TTL. Tombstones for children
already deleted stay, since their remote deletes are still owed.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: block lazy remote reads on startup tombstone rebuild

The rebuild gate is now a done-channel set synchronously before the
replay goroutine starts, so no lazy read can slip through in between.
Reads wait on it with context cancellation instead of returning an
empty miss that makes remote-only objects look deleted.

The replay start is floored at now-TTL: mounts without a recorded
write-back offset previously replayed the whole persisted history, and
events older than the TTL would only build already-expired tombstones.
The gate check now runs after the mount lookup so replaying the meta
log's own directory listings does not deadlock on the gate, and the
replay retries with backoff until it succeeds instead of failing open.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: mark restamped tombstone pending until its delete event lands

When a local delete raises an existing tombstone's timestamp, the new
value is only a local clock guess ahead of that delete's event. Leaving
the tombstone un-pending lets a write-back offset release it before the
event is actually consumed, reopening the resurrection window.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: bound tombstone replay to the tombstone TTL

Persisted-log replay retried forever, keeping lazy remote reads gated
indefinitely when the log cannot be read. Cap retries at the tombstone
TTL measured from replay start: past that point every tombstone would
have expired anyway, so opening the gate loses no protection.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: re-check deletion tombstone before persisting lazy fetch

A delete landing while StatFile is in flight passed the earlier
tombstone check but still persisted the fetched entry, resurrecting a
path whose remote delete is pending. Re-check right before CreateEntry.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract a lazily persisted entry when a delete raced the insert

The pre-insert tombstone check still leaves a window between the check
and the store insert. Since deletes always record the tombstone before
removing the entry, a tombstone visible right after a successful insert
means the delete already ran: delete the entry back out so the
tombstoned path stays deleted.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: note why the replay deadline can safely open the gate

Deletes made after startup are captured by the live delete and event
paths, so a stalled replay can only be missing pre-restart deletes, all
of which are past the tombstone TTL by the deadline.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: retract only the entry a lazy remote read materialized

Deleting by path after a raced delete could remove a legitimate rewrite
that replaced the fetched entry. Verify the stored entry still matches
the remote object (or the just-created directory shape) before deleting,
and apply the same post-insert check to lazy listing children.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* filer: require full-entry equality before retracting a lazy entry

Remote-only matching still removed a write that had updated the fetched
entry, e.g. appended chunks. Compare the persisted entry against what
this read materialized; any change means a real update owns the path.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-09-26 08:00:39 +08:00

228 lines
7.6 KiB
Go

package filer
import (
"context"
"errors"
"fmt"
"strings"
"time"
"google.golang.org/protobuf/proto"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/pb/remote_pb"
"github.com/seaweedfs/seaweedfs/weed/remote_storage"
"github.com/seaweedfs/seaweedfs/weed/util"
)
type lazyFetchContextKey struct{}
// maybeLazyFetchFromRemote is called by FindEntry when the store returns no
// entry for p. If p is under a remote-storage mount, it stats the remote
// object, builds a filer Entry from the result, and persists it via
// CreateEntry with SkipCheckParentDirectory so phantom parent directories
// under the mount are not required.
//
// On a CreateEntry failure after a successful StatFile the in-memory entry is
// still returned (availability over consistency); the singleflight key is
// forgotten so the next lookup retries the filer write.
//
// Returns nil without error when: p is not under a remote mount; the remote
// reports the object does not exist; or any other remote error occurs.
func (f *Filer) maybeLazyFetchFromRemote(ctx context.Context, p util.FullPath) (*Entry, error) {
// Prevent recursive invocation: CreateEntry calls FindEntry, which would
// re-enter this function and deadlock on the singleflight key.
if ctx.Value(lazyFetchContextKey{}) != nil {
return nil, nil
}
if f.RemoteStorage == nil {
return nil, nil
}
mountDir, remoteLoc := f.RemoteStorage.FindMountDirectory(p)
if remoteLoc == nil {
return nil, nil
}
// A startup tombstone rebuild may still be replaying the meta log; wait
// for it so a pending delete cannot resurrect here.
if done := f.remoteTombstonesDone.Load(); done != nil {
select {
case <-*done:
case <-ctx.Done():
return nil, ctx.Err()
}
}
if f.isRemoteDeletionPending(ctx, p, mountDir) {
glog.V(2).InfofCtx(ctx, "maybeLazyFetchFromRemote: %s deleted locally, remote delete pending", p)
return nil, nil
}
remoteConf, found := f.RemoteStorage.FindRemoteStorageConf(p)
if !found {
return nil, nil
}
relPath := strings.TrimPrefix(string(p), string(mountDir))
if relPath != "" && !strings.HasPrefix(relPath, "/") {
relPath = "/" + relPath
}
base := strings.TrimSuffix(remoteLoc.Path, "/")
remotePath := "/" + strings.TrimLeft(base+relPath, "/")
objectLoc := &remote_pb.RemoteStorageLocation{
Name: remoteLoc.Name,
Bucket: remoteLoc.Bucket,
Path: remotePath,
}
type lazyFetchResult struct {
entry *Entry
}
key := string(p)
val, err, _ := f.lazyFetchGroup.Do(key, func() (interface{}, error) {
buildCtx := context.WithoutCancel(ctx)
client, clientErr := f.buildRemoteStorageClient(buildCtx, remoteConf)
if clientErr != nil {
glog.V(1).InfofCtx(ctx, "maybeLazyFetchFromRemote: reject %s: %v", p, clientErr)
return lazyFetchResult{nil}, nil
}
remoteEntry, statErr := client.StatFile(objectLoc)
if statErr != nil {
if errors.Is(statErr, remote_storage.ErrRemoteObjectNotFound) {
glog.V(3).InfofCtx(ctx, "maybeLazyFetchFromRemote: %s not found in remote", p)
} else {
glog.Warningf("maybeLazyFetchFromRemote: stat %s failed: %v", p, statErr)
}
return lazyFetchResult{nil}, nil
}
if remoteEntry == nil {
glog.V(3).InfofCtx(ctx, "maybeLazyFetchFromRemote: %s StatFile returned nil entry", p)
return lazyFetchResult{nil}, nil
}
mtime := time.Unix(remoteEntry.RemoteMtime, 0)
entry := &Entry{
FullPath: p,
Attr: Attr{
Mtime: mtime,
Crtime: mtime,
Mode: 0644,
FileSize: uint64(remoteEntry.RemoteSize),
},
Extended: MergeRemoteContentEncoding(remoteEntry, nil),
Remote: remoteEntry,
}
persistBaseCtx, cancelPersist := context.WithTimeout(context.Background(), 30*time.Second)
defer cancelPersist()
persistCtx := context.WithValue(persistBaseCtx, lazyFetchContextKey{}, true)
// A delete may have landed while StatFile was in flight; re-check so
// the fetched object cannot resurrect a path whose delete is pending.
if f.isRemoteDeletionPending(persistCtx, p, mountDir) {
glog.V(2).InfofCtx(ctx, "maybeLazyFetchFromRemote: %s deleted during remote stat", p)
return lazyFetchResult{nil}, nil
}
saveErr := f.CreateEntry(persistCtx, entry, nil, false, false, nil, true, f.MaxFilenameLength)
if saveErr != nil {
glog.Warningf("maybeLazyFetchFromRemote: failed to persist filer entry for %s: %v", p, saveErr)
f.lazyFetchGroup.Forget(key)
return lazyFetchResult{entry}, nil
}
// A delete records its tombstone before removing the entry, so a
// tombstone visible now means the insert raced a delete that already
// ran: retract the persisted entry so the path stays deleted.
if f.isRemoteDeletionPending(persistCtx, p, mountDir) {
glog.V(2).InfofCtx(ctx, "maybeLazyFetchFromRemote: %s deleted while persisting", p)
f.lazyFetchGroup.Forget(key)
f.retractLazyRemoteEntry(persistCtx, entry)
return lazyFetchResult{nil}, nil
}
return lazyFetchResult{entry}, nil
})
if err != nil {
return nil, err
}
result, ok := val.(lazyFetchResult)
if !ok {
return nil, fmt.Errorf("maybeLazyFetchFromRemote: unexpected singleflight result type %T for %s", val, p)
}
return result.entry, nil
}
// retractLazyRemoteEntry deletes the entry at entry.FullPath only when it is
// still the entry a lazy remote read just materialized — a concurrent write
// may have replaced it, and deleting by path alone would take that write down.
func (f *Filer) retractLazyRemoteEntry(ctx context.Context, entry *Entry) {
existing, findErr := f.FindEntry(ctx, entry.FullPath)
if findErr != nil || existing == nil {
return
}
// The stored entry must still be exactly what this read materialized —
// an intervening write (appended chunks, touched attributes) means a
// real update owns the path now.
if !proto.Equal(existing.ToProtoEntry(), entry.ToProtoEntry()) {
return
}
if err := f.doDeleteEntryMetaAndData(ctx, existing, false, false, nil); err != nil && !errors.Is(err, filer_pb.ErrNotFound) {
glog.Warningf("retractLazyRemoteEntry %s: %v", entry.FullPath, err)
}
}
func (f *Filer) maybeDeleteFromRemote(ctx context.Context, entry *Entry) (bool, error) {
if entry == nil || f.RemoteStorage == nil {
return false, nil
}
mountDir, remoteLoc := f.RemoteStorage.FindMountDirectory(entry.FullPath)
if remoteLoc == nil {
return false, nil
}
if !entry.IsDirectory() && entry.Remote == nil {
return false, nil
}
remoteConf, found := f.RemoteStorage.GetRemoteStorageConf(remoteLoc.Name)
if !found {
return false, fmt.Errorf("resolve remote storage client for %s: not found", entry.FullPath)
}
client, clientErr := f.buildRemoteStorageClient(ctx, remoteConf)
if clientErr != nil {
return false, fmt.Errorf("resolve remote storage client for %s: %w", entry.FullPath, clientErr)
}
if client == nil {
return false, fmt.Errorf("resolve remote storage client for %s: initialization failed", entry.FullPath)
}
objectLoc := MapFullPathToRemoteStorageLocation(mountDir, remoteLoc, entry.FullPath)
if entry.IsDirectory() {
if err := client.RemoveDirectory(objectLoc); err != nil {
if errors.Is(err, remote_storage.ErrRemoteObjectNotFound) {
return true, nil
}
return false, fmt.Errorf("remove remote directory %s: %w", entry.FullPath, err)
}
glog.V(3).InfofCtx(ctx, "maybeDeleteFromRemote: deleted directory %s from remote", entry.FullPath)
return true, nil
}
if err := client.DeleteFile(objectLoc); err != nil {
if errors.Is(err, remote_storage.ErrRemoteObjectNotFound) {
return true, nil
}
return false, fmt.Errorf("delete remote file %s: %w", entry.FullPath, err)
}
glog.V(3).InfofCtx(ctx, "maybeDeleteFromRemote: deleted %s from remote", entry.FullPath)
return true, nil
}