Files
seaweedfs/weed/mount/meta_cache/meta_cache_init.go
T
Chris LuandGitHub 65114575eb mount: invalidate hot directory listings by section (#10712)
* mount: invalidate hot directory listings by section

A cached directory used to be dropped whole when it saw 64 changes in
2s: with a continuous writer the listing cycled through wipe, direct
listing and full rebuild for as long as the writer kept going, and
every sibling lookup fell through to the filer in between.

Split each cached listing into name-range sections of 1024 entries. A
burst of foreign changes invalidates just the section it lands in;
entries stay served and events keep applying, and the next readdir
re-lists only that range from the filer, reconciled through the version
gate so it cannot roll back newer applied events. Lookups in an
invalidated section read through until then. The mount's own writes no
longer invalidate anything: they are ground truth for its cache.

* meta_cache: drop the version floor with a deleted or moved directory

The other teardown paths already clear both maps; a floor left behind
here would fence the listing of a directory re-created at the same
path.

* mount: harden section refresh

An unversioned listing (pre-upgrade filer) now only fills gaps instead
of reconciling: without a snapshot to order against, an overwrite or
the deletion sweep could roll back an event applied after the listing.

The section table can be rebuilt or re-split between the listing and
its apply, so the refresh only marks fresh or splits when the section
still covers the range it read. Splicing bounds from a stale range
into a rebuilt table could leave them unsorted.

Bound the wait: a readdir gives a refresh five seconds before serving
the maintained-but-unverified cache. Bound the size: a range grown
past four sections aborts the refresh and drops the directory cache,
re-tiling it with a full rebuild, with that request served direct.

Cover the filer-facing path with a listing server: paging with the
snapshot pinned across pages, the section cutoff, no calls for a
fresh section, and the overgrown-range abort.

* meta_cache: make the section table a self-contained state machine

Churn counting, freshness, stale-range scanning and the refresh
completion with its guard and split now live on dirSections itself,
free of the lock, the store and the apply loop, so they test directly
with synthetic clocks and tables. MetaCache keeps thin wrappers that
hold its mutex and find the directory's table.

* meta_cache: keep section internals out of the apply request

The request now carries the completed build's table and one refresh as
opaque values built by section code, and the boundary-derivation rule
moves out of the build loop into a collector next to the rest of the
section logic.

* mount: fence refreshed sections with a snapshot floor

A refresh versioned the entries it fetched and tombstoned the ones it
swept, but a name absent from both cache and listing kept the old
directory floor, so a delayed event between the two snapshots could
resurrect it into a section already marked fresh. The section now
carries its own floor, consulted next to the directory floor, covering
every name in the range, present or absent — which also retires the
refresh's per-entry version stamps and sweep tombstones.

An unversioned listing sets no floor and vouches for nothing: it may
still fill gaps, but the section stays stale and reads through until a
filer that stamps snapshots re-validates it.

A listing's reach is unknowable up front — a resumed handle can skip
far ahead, and shrunken sections let one batch span many — so a
readdir now re-validates every stale section from its start name to
the end of the directory instead of the next two.

* mount: fence tombstoned names with floors and gate the reconcile

A tombstone answered for its name before the floors were consulted, so
one at an old position let through events the newer listing floor
should have fenced; a build never hit this because it prunes
superseded tombstones, which a section refresh does not. The version
gate now raises a tombstone to the floors like any other record.

With no per-entry versions, only the section floor fences a
reconcile's work, so a range the rebuilt or re-split table no longer
has must not touch the store either: the range check moves ahead of
the mutations, under the same lock the floor install holds.

An unversioned refresh no longer retries: the section is remembered as
unverifiable and skipped by the stale scan, or every batch of every
readdir would re-list the same ranges against a filer that cannot
vouch for them.

* mount: clear beaten unversioned markers and skip refresh mid-build

An unversioned marker outliving the snapshot write that replaced its
content bypassed the section floor the same way an old tombstone did,
letting a delayed pre-snapshot event roll the entry back. The refresh
now clears the marker when its write wins; pinned local-only entries
are not replaced at all, keeping their content and marker.

A rebuild wipes and repopulates the store off the apply loop, so a
refresh reconciling meanwhile could sweep children the build had
already inserted and let it publish the directory incomplete. The
refresh now skips a building directory, as events (buffered) and
purges (skipped) already do; its staleness dies with the build's
fresh table.

* mount: clear the unversioned marker only after its replacement lands

Clearing before the insert meant a failed write left the old local
content claiming the listing floors, fencing the very events that were
still entitled to correct it.

* meta_cache: rename the section state machine to sectionList

dirSections named both the type and the map of them.

* mount: raise the default cacheDirMaxEntries to 100000

The low ceiling guarded against whole-listing rebuild churn: a big
cached directory under writes kept re-streaming everything. Sectioned
invalidation ended that — a burst now costs one range listing — so the
remaining cost of caching a large directory is its one-time build,
comparable to the single direct listing that read-through mode pays on
every enumeration instead.

* meta_cache: cover section border and edge cases

A bound-named entry belongs to the section starting at the bound: the
neighboring refresh's sweep stops before it, its own section's covers
it. Churn past everything the build saw lands in the tail section, a
rename spanning two sections invalidates both, and a listed entry at
the section's end name is cut off with the ones beyond it.
2026-08-11 21:11:18 -07:00

231 lines
7.7 KiB
Go

package meta_cache
import (
"context"
"errors"
"fmt"
"time"
"golang.org/x/sync/errgroup"
"github.com/seaweedfs/seaweedfs/weed/filer"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/util"
)
// DirectoryTooLargeError reports a directory the mount refuses to cache
// locally. Its listings read through to the filer instead.
type DirectoryTooLargeError struct {
Path util.FullPath
}
func (e *DirectoryTooLargeError) Error() string {
return fmt.Sprintf("directory %s is too large to cache locally", e.Path)
}
// maxCacheableEntries is the directory size above which a build gives up, or 0
// to cache everything.
func EnsureVisited(mc *MetaCache, client filer_pb.FilerClient, dirPath util.FullPath, maxCacheableEntries int) error {
// Collect all uncached paths from target directory up to root
var uncachedPaths []util.FullPath
currentPath := dirPath
for {
// If this path is cached, all ancestors are also cached
if mc.isCachedFn(currentPath) {
break
}
if mc.isOversized(currentPath) {
// The directory itself reads through; an ancestor is stepped over,
// or it would wedge every listing beneath it forever.
if currentPath == dirPath {
return &DirectoryTooLargeError{Path: currentPath}
}
} else {
uncachedPaths = append(uncachedPaths, currentPath)
}
// Continue to parent directory
if currentPath != mc.root {
parent, _ := currentPath.DirAndName()
currentPath = util.FullPath(parent)
} else {
break
}
}
if len(uncachedPaths) == 0 {
return nil
}
// Fetch all uncached directories in parallel with context for cancellation
// If one fetch fails, cancel the others to avoid unnecessary work
g, ctx := errgroup.WithContext(context.Background())
for _, p := range uncachedPaths {
path := p // capture for closure
g.Go(func() error {
err := doEnsureVisited(ctx, mc, client, path, maxCacheableEntries)
var tooLarge *DirectoryTooLargeError
if errors.As(err, &tooLarge) && path != dirPath {
// An ancestor found oversized just reads through; failing the
// group here would cancel the builds of its cacheable
// descendants, and the caller would treat the refusal as the
// listed directory's own.
return nil
}
return err
})
}
return g.Wait()
}
// batchInsertSize is the number of entries to accumulate before flushing to LevelDB.
// 100 provides a balance between memory usage (~100 Entry pointers) and write efficiency
// (fewer disk syncs). Larger values reduce I/O overhead but increase memory and latency.
const batchInsertSize = 100
const (
emptyRebuildConfirmations = 2
emptyRebuildConfirmDelay = 50 * time.Millisecond
)
func doEnsureVisited(ctx context.Context, mc *MetaCache, client filer_pb.FilerClient, path util.FullPath, maxCacheableEntries int) error {
// Use singleflight to deduplicate concurrent requests for the same path
_, err, _ := mc.visitGroup.Do(string(path), func() (interface{}, error) {
// Check for cancellation before starting
if ctx.Err() != nil {
return nil, ctx.Err()
}
// Double-check if already cached (another goroutine may have completed)
if mc.isCachedFn(path) {
return nil, nil
}
glog.V(4).Infof("ReadDirAllEntries %s ...", path)
// Use context.Background() for build lifecycle calls so that
// errgroup cancellation of ctx doesn't cause enqueueAndWait to
// return early, which would trigger cleanupBuild while the
// operation is still queued.
if err := mc.BeginDirectoryBuild(context.Background(), path); err != nil {
return nil, fmt.Errorf("begin build %s: %w", path, err)
}
cleanupDone := false
cleanupBuild := func(reason string) {
if cleanupDone {
return
}
cleanupDone = true
if deleteErr := mc.deleteFolderChildrenForRebuild(context.Background(), path); deleteErr != nil {
glog.V(2).Infof("clear %s build %s: %v", reason, path, deleteErr)
}
if abortErr := mc.AbortDirectoryBuild(context.Background(), path); abortErr != nil {
glog.V(2).Infof("abort %s build %s: %v", reason, path, abortErr)
}
}
defer func() {
if !cleanupDone && ctx.Err() != nil {
cleanupBuild("canceled")
}
}()
// reloadFromFiler wipes the cached children and reloads them from the filer.
reloadFromFiler := func() (entryCount int, snapshotTsNs int64, sections sectionBoundsCollector, err error) {
err = util.Retry("ReadDirAllEntries", func() error {
entryCount = 0
sections = sectionBoundsCollector{}
var batch []*filer.Entry // reset on retry, allow GC of previous entries
if err := mc.deleteFolderChildrenForRebuild(ctx, path); err != nil {
return fmt.Errorf("clear existing entries for %s: %w", path, err)
}
var listErr error
snapshotTsNs, listErr = filer_pb.ReadDirAllEntriesWithSnapshot(ctx, client, path, "", func(pbEntry *filer_pb.Entry, isLast bool) error {
entry := filer.FromPbEntry(string(path), pbEntry)
if !mc.includeSystemEntries && IsHiddenSystemEntry(string(path), entry.Name()) {
return nil
}
if maxCacheableEntries > 0 && entryCount >= maxCacheableEntries {
return &DirectoryTooLargeError{Path: path}
}
sections.note(entry.Name())
batch = append(batch, entry)
entryCount++
// flush by size, not isLast: hidden entries can return early
if len(batch) >= batchInsertSize {
if err := mc.doBatchInsertEntries(ctx, batch); err != nil {
return fmt.Errorf("batch insert for %s: %w", path, err)
}
batch = make([]*filer.Entry, 0, batchInsertSize)
}
return nil
})
if listErr != nil {
return listErr
}
if len(batch) > 0 {
if err := mc.doBatchInsertEntries(ctx, batch); err != nil {
return fmt.Errorf("batch insert remaining for %s: %w", path, err)
}
}
return nil
})
return entryCount, snapshotTsNs, sections, err
}
entryCount, snapshotTsNs, sections, fetchErr := reloadFromFiler()
if fetchErr != nil {
var tooLarge *DirectoryTooLargeError
if errors.As(fetchErr, &tooLarge) {
// Remember the refusal so the next visit fails fast instead of
// streaming up to the limit again to rediscover it.
mc.markOversized(path)
glog.V(0).Infof("directory %s exceeds %d entries, reading it through instead of caching", path, maxCacheableEntries)
cleanupBuild("oversized")
return nil, fetchErr
}
cleanupBuild("failed")
return nil, fmt.Errorf("list %s: %w", path, fetchErr)
}
// A transient empty listing would strand a populated directory cached over
// an empty store; re-read to confirm before trusting it. First re-read is
// immediate (a clean-EOF stream glitch clears at once), later ones space out.
// On cancellation the deferred cleanup aborts the build.
for attempt := 0; entryCount == 0 && attempt < emptyRebuildConfirmations; attempt++ {
if ctx.Err() != nil {
return nil, ctx.Err()
}
if attempt > 0 {
select {
case <-time.After(emptyRebuildConfirmDelay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
if entryCount, snapshotTsNs, sections, fetchErr = reloadFromFiler(); fetchErr != nil {
cleanupBuild("failed")
return nil, fmt.Errorf("confirm empty list %s: %w", path, fetchErr)
}
if entryCount > 0 {
glog.Warningf("rebuild of %s saw a transient empty listing, recovered %d entries on confirmation", path, entryCount)
}
}
if err := mc.CompleteDirectoryBuild(context.Background(), path, snapshotTsNs, sections.bounds); err != nil {
cleanupBuild("unreplayed")
return nil, fmt.Errorf("complete build for %s: %w", path, err)
}
cleanupDone = true // Prevent deferred cleanup after successful publish
return nil, nil
})
return err
}
func IsHiddenSystemEntry(dir, name string) bool {
return dir == "/" && (name == "topics" || name == "etc")
}