mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-09-30 19:55:48 +00:00
The ReaderCache is shared by all streams of a process (every S3 GET, for instance), but a ChunkReadAt released chunks as if it owned them: - moving on to the next chunk called UnCache on the previous one, destroying the buffer even when other streams were still inside it; - since #11384 a buffer is dropped once any reader has consumed it to the end and no read call is in flight. Streams copy out in slices (256 KiB in the S3 gateway), so between two calls a slower stream is not attached and loses the buffer to a faster one. Either way the slower stream refetches the whole chunk from the volume servers. With many clients downloading the same popular object at once, each chunk is fetched over and over; in production we saw the S3 gateway pull ~10 Gbit/s from volume servers while serving ~1 Gbit/s to clients. A ChunkReadAt now pins the chunk it is positioned in. The pin is taken and released only under the ReaderCache lock, since concurrent ReadAt calls on one ChunkReadAt (as in mount) share it. It is released when the stream reads the chunk to its end, moves to another chunk (including one served from the chunk cache), or falls back to random reads. A buffer is dropped once no stream pins it and no read is in progress, if it was consumed or its last stream left it; a read still in flight when the stream leaves drops it on detach, as UnCache did via destroy. Eviction by slot limit and memory budget is unchanged. lastChunkFid is now guarded as well: concurrent ReadAt calls raced on it. Tests: two ChunkReadAt instances streaming one object in interleaved slices fetch each chunk exactly once (2-3 times before); leaving a chunk for a chunk-cache hit or while another read is in flight releases it; concurrent ReadAt calls on one ChunkReadAt leave no pins behind under -race.