seaweedfs

mirror of https://github.com/seaweedfs/seaweedfs.git synced 2026-05-20 16:51:31 +00:00

Files

Chris Lu edf7d2a074 fix(filer): eliminate redundant disk reads causing memory/CPU regression (#9039 )

* fix(filer): eliminate redundant disk reads causing memory/CPU regression (#9035)

Since 4.18, LocalMetaLogBuffer's ReadFromDiskFn was set to
readPersistedLogBufferPosition, causing LoopProcessLogData to call
ReadPersistedLogBuffer on every 250ms health-check tick when a
subscriber encounters ResumeFromDiskError.  Each call creates an
OrderedLogVisitor (ListDirectoryEntries on the filer store), spawns a
readahead goroutine with a 1024-element channel, finds no data, and
returns — 4 times per second even on an idle filer.

This is redundant because SubscribeLocalMetadata already manages disk
reads explicitly with its own shouldReadFromDisk / lastCheckedFlushTsNs
tracking in the outer loop.

Set ReadFromDiskFn back to nil for LocalMetaLogBuffer.  When
LoopProcessLogData encounters ResumeFromDiskError with nil
ReadFromDiskFn, the HasData() guard returns ResumeFromDiskError to the
caller (SubscribeLocalMetadata), which blocks efficiently on
listenersCond.Wait() instead of polling.

* fix(filer): add gap detection for slow consumers after disk-read stall

When a slow consumer falls behind and LoopProcessLogData returns
ResumeFromDiskError with no flush or read-position progress, there may
be a gap between persisted data and in-memory data (e.g. writes stopped
while consumer was still catching up). Without this, the consumer would
block on listenersCond.Wait() forever.

Skip forward to the earliest in-memory time to resume progress, matching
the gap-handling pattern already used in the shouldReadFromDisk path.

* fix(filer): clear stale ResumeFromDiskError after gap-skip to avoid stall

The gap-detection block added in the previous commit skips lastReadTime
forward to GetEarliestTime() and continues the outer loop.  On the next
iteration, shouldReadFromDisk becomes true (currentReadTsNs >
lastDiskReadTsNs), the disk read returns processedTsNs == 0, and the
existing gap handler at the top of the loop runs its own gap check.
That check uses readInMemoryLogErr == ResumeFromDiskError as the entry
condition — but readInMemoryLogErr is still the stale error from two
iterations ago.  GetEarliestTime() now equals lastReadTime.Time (we
already advanced to it), so earliestTime.After(lastReadTime.Time) is
false and the handler falls into listenersCond.Wait() — stuck.

Clear readInMemoryLogErr at the gap-skip point, matching the existing
pattern at the earlier gap handler that already clears it for the same
reason.

* fix(log_buffer): GetEarliestTime must include sealed prev buffers

GetEarliestTime previously returned only logBuffer.startTime (the active
buffer's first timestamp).  That is narrower than ReadFromBuffer's
tsMemory, which is the min across active + prev buffers.  Callers using
GetEarliestTime for gap detection after ResumeFromDiskError (the
SubscribeLocalMetadata outer loop's disk-read path, the new gap-skip in
the in-memory ResumeFromDiskError handler, and MQ HasData) saw a time
that was *newer* than the real earliest in-memory data.

Impact in SubscribeLocalMetadata's slow-consumer path:
  - tsMemory = earliest prev buffer time (T_prev)
  - GetEarliestTime() = active startTime (T_active, later than T_prev)
  - Consumer position = T1, with T_prev < T1 < T_active
  - ReadFromBuffer returns ResumeFromDiskError (T1 < tsMemory)
  - Gap detect: GetEarliestTime().After(T1) = T_active.After(T1) = true
  - Skip forward to T_active -- silently drops the prev-buffer data
  - And when T_active happens to equal the stuck position, gap detect
    evaluates false, and the subscriber stalls on listenersCond.Wait()

This reproduces the TestMetadataSubscribeSlowConsumerKeepsProgressing
failure in CI where the consumer stalled at 10220/20000 after writing
stopped -- the buffer still had data in prev[0..3], but gap detection
was comparing against the active buffer's startTime.

Fix: scan all sealed prev buffers under RLock, return the true minimum
startTime.  Matches the min-of-buffers logic in ReadFromBuffer.

* test(log_buffer): make DiskReadRetry test deterministic

The previous test added the message via AddToBuffer + ForceFlush and
relied on a race: the second disk read had to happen before the data
was delivered through the in-memory path.  Under the race detector or
on a slow CI runner, the reader is woken by AddToBuffer's notification,
finds the data in the active buffer or its prev slot, and returns after
exactly one disk read — failing the >= 2 disk reads assertion even
though the loop behaved correctly.

Reproduced on master with race detector (2/5 failures).

Rewrite the test to deliver the data exclusively through the disk-read
path: no AddToBuffer, no ForceFlush.  The test waits until the reader
has issued at least one no-op disk read, then atomically flips a
"dataReady" flag.  The reader's next iteration through readFromDiskFn
returns the entry.  This deterministically exercises the retry-loop
behavior the test was originally written to protect, and removes the
in-memory delivery race entirely.

2026-04-11 23:12:54 -07:00

abstract_sql

Fix chown Input/output error on large file sets (#7996 )

2026-01-09 18:02:59 -08:00

arangodb

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

cassandra

feat: add TLS configuration options for Cassandra2 store (#7998 )

2026-01-14 17:59:59 -08:00

cassandra2

feat: add TLS configuration options for Cassandra2 store (#7998 )

2026-01-14 17:59:59 -08:00

elastic/v7

go fix

2026-02-20 18:42:00 -08:00

empty_folder_cleanup

go fmt

2026-04-10 17:31:14 -07:00

etcd

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

foundationdb

go fix

2026-02-20 18:42:00 -08:00

hbase

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

leveldb

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

leveldb2

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

leveldb3

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

mongodb

fix: comprehensive go vet error fixes and add CI enforcement (#7861 )

2025-12-23 14:48:50 -08:00

mysql

Fix chown Input/output error on large file sets (#7996 )

2026-01-09 18:02:59 -08:00

mysql2

print only adapted url

2024-03-25 12:50:43 -07:00

postgres

fix(filer/postgres): use pgx v5 API for PgBouncer simple protocol (#9010 )

2026-04-09 16:36:15 -07:00

postgres2

fix(filer/postgres): use pgx v5 API for PgBouncer simple protocol (#9010 )

2026-04-09 16:36:15 -07:00

redis

fix(weed/filer/redis): dropped error (#8895 )

2026-04-02 15:39:04 -07:00

redis2

fix(weed/filer/redis2): fix dropped error (#8952 )

2026-04-07 14:59:01 -07:00

redis3

chore: remove ~50k lines of unreachable dead code (#8913 )

2026-04-03 16:04:27 -07:00

rocksdb

Add error list each entry func (#7485 )

2025-11-25 19:35:19 -08:00

sqlite

go fix

2026-02-20 18:42:00 -08:00

store_test

fix(weed/filer/store_test): fix dropped errors (#8782 )

2026-03-26 12:07:48 -07:00

tarantool

go fix

2026-02-20 18:42:00 -08:00

tikv

go fix

2026-02-20 18:42:00 -08:00

ydb

go fix

2026-02-20 18:42:00 -08:00

configuration.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

copy_params.go

Use filer-side copy for mounted whole-file copy_file_range (#8747 )

2026-03-23 18:35:15 -07:00

entry_codec.go

fix(filer,mount): add nanosecond timestamp precision (#9019 )

2026-04-10 11:51:06 -07:00

entry.go

test: add pjdfstest POSIX compliance suite (#9013 )

2026-04-10 09:52:16 -07:00

filechunk_group_test.go

Context cancellation during reading range reading large files (#7093 )

2025-08-06 10:09:26 -07:00

filechunk_group.go

mount: improve read throughput with parallel chunk fetching (#7569 )

2025-11-29 10:06:11 -08:00

filechunk_manifest_test.go

move to https://github.com/seaweedfs/seaweedfs

2022-07-29 00:17:28 -07:00

filechunk_manifest.go

Fix S3 Gateway Read Failover #8076 (#8087 )

2026-01-22 14:07:24 -08:00

filechunk_section_test.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

filechunk_section.go

mount: improve read throughput with parallel chunk fetching (#7569 )

2025-11-29 10:06:11 -08:00

filechunks2_test.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

filechunks_read_test.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

filechunks_read.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

filechunks_test.go

S3 API: Advanced IAM System (#7160 )

2025-08-30 11:15:48 -07:00

filechunks.go

fix: multipart upload ETag calculation (#8238 )

2026-02-06 21:54:43 -08:00

filer_buckets.go

Prevent bucket renaming in filer, fuse mount, and S3 (#8048 )

2026-01-16 19:48:09 -08:00

filer_conf_test.go

Reduce memory allocations in hot paths (#7725 )

2025-12-12 12:51:48 -08:00

filer_conf.go

s3api: fix static IAM policy enforcement after reload (#8532 )

2026-03-06 12:35:08 -08:00

filer_delete_entry.go

filer: propagate lazy metadata deletes to remote mounts (#8522 )

2026-03-12 13:16:28 -07:00

filer_deletion_test.go

Filer: Add retry mechanism for failed file deletions (#7402 )

2025-10-29 18:31:23 -07:00

filer_deletion.go

Filer: Add retry mechanism for failed file deletions (#7402 )

2025-10-29 18:31:23 -07:00

filer_hardlink.go

move to https://github.com/seaweedfs/seaweedfs

2022-07-29 00:17:28 -07:00

filer_lazy_remote_listing.go

feat(filer): add lazy directory listing for remote mounts (#8615 )

2026-03-13 09:36:54 -07:00

filer_lazy_remote_test.go

Add data file compaction to iceberg maintenance (Phase 2) (#8503 )

2026-03-15 11:27:42 -07:00

filer_lazy_remote.go

filer: propagate lazy metadata deletes to remote mounts (#8522 )

2026-03-12 13:16:28 -07:00

filer_notify_append.go

fix: include DiskType in metadata log volume assignment (#7918 )

2025-12-30 17:32:33 -08:00

filer_notify_read.go

chore: remove ~50k lines of unreachable dead code (#8913 )

2026-04-03 16:04:27 -07:00

filer_notify_test.go

refactor filer_pb.Entry and filer.Entry to use GetChunks()

2022-11-15 06:33:36 -08:00

filer_notify.go

fix(filer): eliminate redundant disk reads causing memory/CPU regression (#9039 )

2026-04-11 23:12:54 -07:00

filer_on_meta_event_test.go

Adjust rename events metadata format (#8854 )

2026-03-30 18:25:11 -07:00

filer_on_meta_event.go

Adjust rename events metadata format (#8854 )

2026-03-30 18:25:11 -07:00

filer_rename.go

Prevent bucket renaming in filer, fuse mount, and S3 (#8048 )

2026-01-16 19:48:09 -08:00

filer_search.go

filer: async empty folder cleanup via metadata events (#7614 )

2025-12-03 21:12:19 -08:00

filer.go

fix(filer): eliminate redundant disk reads causing memory/CPU regression (#9039 )

2026-04-11 23:12:54 -07:00

filerstore_hardlink.go

fix(filer): update hard link ctime when nlink changes on unlink (#9018 )

2026-04-10 11:23:52 -07:00

filerstore_translate_path.go

Add error list each entry func (#7485 )

2025-11-25 19:35:19 -08:00

filerstore_wrapper_test.go

fix(filer): remove cancellation guard from RollbackTransaction and clean up #8909 (#8916 )

2026-04-03 17:55:27 -07:00

filerstore_wrapper.go

fix(filer): do not abort entry deletion when hard link cleanup fails (#9022 )

2026-04-10 13:59:58 -07:00

filerstore.go

Add error list each entry func (#7485 )

2025-11-25 19:35:19 -08:00

interval_list_test.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

interval_list.go

Mount concurrent read (#4400 )

2023-04-13 22:32:45 -07:00

meta_aggregator.go

go fmt

2026-04-10 17:31:14 -07:00

meta_replay.go

chore: remove ~50k lines of unreachable dead code (#8913 )

2026-04-03 16:04:27 -07:00

metadata_event_context.go

Adjust rename events metadata format (#8854 )

2026-03-30 18:25:11 -07:00

metadata_event_sink_test.go

mount: make metadata cache rebuilds snapshot-consistent (#8531 )

2026-03-07 09:19:40 -08:00

metadata_event_sink.go

mount: make metadata cache rebuilds snapshot-consistent (#8531 )

2026-03-07 09:19:40 -08:00

read_remote.go

improve: large file sync throughput for remote.cache and filer.sync (#8676 )

2026-03-17 16:49:56 -07:00

read_write.go

s3api: fix static IAM policy enforcement after reload (#8532 )

2026-03-06 12:35:08 -08:00

reader_at_test.go

Context cancellation during reading range reading large files (#7093 )

2025-08-06 10:09:26 -07:00

reader_at.go

mount: improve read throughput with parallel chunk fetching (#7627 )

2025-12-04 23:40:56 -08:00

reader_cache_test.go

chore: execute goimports to format the code (#7983 )

2026-01-07 13:06:08 -08:00

reader_cache.go

mount: improve read throughput with parallel chunk fetching (#7627 )

2025-12-04 23:40:56 -08:00

reader_pattern.go

Fix a few data races when reading files in mount (#3527 )

2022-08-26 16:41:37 -07:00

remote_mapping.go

s3api: fix static IAM policy enforcement after reload (#8532 )

2026-03-06 12:35:08 -08:00

remote_storage_test.go

Add data file compaction to iceberg maintenance (Phase 2) (#8503 )

2026-03-15 11:27:42 -07:00

remote_storage.go

s3api: fix static IAM policy enforcement after reload (#8532 )

2026-03-06 12:35:08 -08:00

s3iam_conf_test.go

s3: fix configuring IAM for the same user

2022-08-30 09:37:52 -07:00

s3iam_conf.go

convert error fromating to %w everywhere (#6995 )

2025-07-16 23:39:27 -07:00

stream_benchmark_test.go

go fmt

2026-04-10 17:31:14 -07:00

stream_prefetch_test.go

feat(s3): add concurrent chunk prefetch for large file downloads (#8917 )

2026-04-03 19:57:30 -07:00

stream_prefetch.go

feat(s3): add concurrent chunk prefetch for large file downloads (#8917 )

2026-04-03 19:57:30 -07:00

stream.go

feat(s3): add concurrent chunk prefetch for large file downloads (#8917 )

2026-04-03 19:57:30 -07:00

topics.go

merge current message queue code changes (#6201 )

2024-11-04 12:08:25 -08:00