mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-08-31 05:07:01 +00:00
da087f77b392e8de065f9863cfec609eee23bf3e
2034
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eb3bbfeb1f |
filer: apply the path's storage rule TTL on every write path (#10963)
* filer: cover the storage rule TTL on the object transaction write path An object written through ObjectTransaction used to land with ttlSec 0 even under an fs.configure TTL rule, while the same object written through CreateEntry got the rule's TTL. Guard the shared stamping so the two paths cannot drift apart again. * filer: apply the path's storage rule to an appended entry AppendToEntry resolved the storage option from the path - so its chunks land on a TTL volume under an fs.configure TTL rule - but never stamped the rule's TTL on the entry it creates, leaving an entry that outlives its data. Route it through applyStorageDefaultsToEntry, which now feeds the entry's own TTL into the option so the placement an existing entry's appended chunks get is unchanged. * filer: apply the path's storage rule to a completed TUS upload The PATCH path resolves the storage option from the target, so a TUS upload into an fs.configure TTL prefix writes its chunks to a TTL volume, but completion built the final entry with ttlSec 0 - the entry outlived the data it pointed at. Stamp it through applyStorageDefaultsToEntry, which also subsumes the hand-rolled read-only check and supplies the rule's name-length limit. * filer: apply the destination's storage option TTL to a copied entry The copy handler re-uploads the source's chunks under the destination's storage option, so a copy into an fs.configure TTL prefix already lands its data on a TTL volume. The entry, though, carried the source's ttlSec - 0 for a source outside the prefix, or the source's own TTL where the two rules differ - so it never expired with the data it pointed at. Take the TTL from the same option the chunks were placed with, after the data-only copy has restored the destination's metadata. |
||
|
|
a02c0024e5 |
master: cap the reported capacity at what the disks hold (#10960)
* master: cap the reported capacity at what the disks hold Statistics reported max volume count times the volume size limit, which is how many volumes the cluster is allowed to place, not how much space it has. A cluster given far more slots than its disks can fill reported a capacity it could never reach -- 65536 slots at 30GB read as 1.9PB on a 460GB disk -- and the number never moved, since writing data changes neither the slot count nor the size limit. The volume servers already report each filesystem's total and free bytes in their heartbeats, so bound the answer by what they say is left. * mount: keep the last known sizes when filer statistics fails A failed Statistics call returned before df's answer was filled in, so a mount whose filer or master was briefly unreachable reported an empty filesystem rather than the sizes it already had. * master: drop the disk ceiling when a volume server does not report A cluster part way through an upgrade has volume servers that predate the disk bytes in the heartbeat. Summing only the ones that answered left the quiet server's free space out of the total, and the server holding the room is exactly the one that could make the cluster read as full. Answer with the disks only when every one of them reported. |
||
|
|
44115c1051 |
filer: stop TUS uploads from turning into garbage (#10945)
* filer: store TUS sub-chunks through the regular chunk writer A TUS sub-chunk was written with one assigned file id, retried up to three times against that same id, and abandoned on failure: an attempt that had landed on some replicas left a needle no session record and no entry ever references, unreclaimable by vacuum. dataToChunkWithSSE, which the regular write path uses per chunk, assigns a fresh file id per attempt and hands back the file ids of failed attempts, which are now freed the way the regular write path frees them. * filer: retry a chunk write on a fresh volume when the server 5xxs The filer's chunk writer assigns a fresh file id per attempt but only retried transient network errors, so a volume filling up and turning read-only mid-write failed the whole request even though the very next assignment would have landed elsewhere. Every other write client already routes this through ShouldReassignUpload; the filer's own write path now does the same, for regular uploads and TUS sub-chunks alike. * filer: export the chunk deletion queue The filer test harness in weed/server builds filer.Filer as a struct literal, so any code path reaching DeleteChunks dereferenced a nil queue. Exported like the neighboring DeletionRetryQueue so the harness can arm it. * filer: complete a TUS upload whose chunk records overlap A PATCH retried while its predecessor was still storing a sub-chunk - a proxy timeout with an immediate retry is enough - records the same range twice. HEAD computes Upload-Offset as the covered watermark and reported the upload fully received, but completion demanded exactly adjacent records and failed every attempt: the client concluded success from offset == length, no entry was created, and the session eventually expired, turning the entire upload into deleted needles for the vacuum to chew through. Completion now validates gapless coverage with the same watermark HEAD uses. A record extending coverage joins the entry - the read path resolves partial overlaps by ModifiedTsNs, and the raced copies carry identical bytes - while a fully covered duplicate is freed once the entry lands. * filer: allow one mutating TUS request per session at a time Nothing stopped two PATCHes from writing the same range concurrently: both loaded the same offset, both passed the conflict check, and both recorded their sub-chunks. A client whose request timed out in a proxy retries immediately while the server side is still storing the buffered sub-chunk, which is exactly that race. A session now accepts one PATCH or DELETE at a time, the way tusd locks uploads; a concurrent one is refused with 423 Locked, which TUS clients retry, and HEAD keeps answering so progress polling is unaffected. The chunk state is loaded under the claim, so a retried PATCH sees every record its predecessor left and conflicts cleanly instead of duplicating data. * test: cover a TUS PATCH raced by its own retry Stalls a PATCH mid-body over a raw connection, retries the same range while it is in flight, and expects the retry refused with 423 Locked; the upload then resumes from the reported offset and the final content must be intact. * filer: never free a TUS duplicate the entry still references Coverage is computed from ranges, so a record fully covered by another is treated as a duplicate no matter which needle it names. A malformed record naming a file id the entry keeps would have had that needle freed right after the entry landed - the corruption this change set exists to stop. The duplicates are now freed in one batch, skipping any file id the entry references; their records go with the session directory. * test: bound the raw TUS connection reads http.ReadResponse on the stalled PATCH's connection blocked until the whole go test timeout if the filer never answered. * filer: free the needles of chunk write attempts a retry replaced A volume server stores the needle locally and only then fans out to the replicas, so a replication failure 5xxs with the data already written. Each attempt assigns its own file id, so once a later attempt lands elsewhere nothing references the earlier ones: the caller only sees the chunk that succeeded, and the failed ids were dropped. They are now freed the way the caller frees them when the whole write fails. Retrying on a 5xx makes this reachable on every read-only or full volume, which is exactly the condition that filled the reporter's volumes. |
||
|
|
c80664ec21 |
s3: propagate storage rule fsync to volume server uploads (#10906)
The storage rule's fsync decision was computed by the filer (detectStorageOption -> rule.Fsync) and applied on the filer's own HTTP write path, but was never carried onto the chunk uploads S3 issues: the AssignVolumeResponse had no fsync field, so the s3api client could not learn the decision, and the chunked upload URL was hardcoded without it. Every S3 write to a path with fsync configured went to the volume server as a non-fsync write. Carry the decision through the assign response: - filer.proto: AssignVolumeResponse gains bool fsync, filled from the storage option the assign resolved. - operation.AssignResult gains Fsync, so uploadChunk can append ?fsync=true to the volume server upload URL (single and replica fan-out paths). - The S3 PUT/UploadPart assignFunc, the S3 copy path, the admin file browser upload, and the Iceberg worker assign functions all forward the response field. Adds TestUploadReaderInChunksAppendsFsyncWhenAssigned. |
||
|
|
74038e1b14 |
master: don't let a dead KeepConnected handler close its successor's channel (#10900)
A client that reconnects before the old handler exits re-registers the same client name, and addClient overwrites the map entry. The old handler's deferred deleteClient then closed whatever channel the map held under that name: the new, live stream's. Receiving from a closed channel returns nil immediately and forever, so the new handler's send loop degenerated into sending empty responses at wire speed, pinning a core on each side until the client killed the connection. deleteClient now closes the channel its own handler registered and leaves the map entry alone unless it still points to that channel. This also closes the previously orphaned old channel, whose drain goroutine used to leak. The send loop treats a closed channel as an exit instead of a message stream. |
||
|
|
0f85d005ad |
server: 416 only when no requested range overlaps, with Content-Range, and the Rust mirror (#10889)
* filer, volume server: return 416 when no requested range overlaps the content * seaweed-volume: return 416 when no requested range overlaps the content * server: check the range test error, use the request context, fix the no-overlap comment boundary |
||
|
|
173adbc291 |
master: never re-seed a raft cluster over committed state under -raftBootstrap (#10883)
* master: never re-seed a raft cluster over committed state -raftBootstrap deleted logs.dat, stable.dat and snapshots on every start and then bootstrapped a fresh cluster. Since hashicorp raft only snapshots after 8192 log entries, the TopologyId lives in the log, not in a snapshot, so the pre-wipe snapshot recovery found nothing and each restart minted a new cluster identity. A master that came up while it could not reach its peers seeded a rival cluster; when the two logs met, SetTopologyId's split-brain guard fatally stopped every master holding the other id, and the master layer crash-looped with no quorum. Bootstrapping is genesis. Drop the wipe and the inline bootstrap. The first master in -peers already mints a cluster once it has confirmed no peer has a leader, so the flag has nothing left to do and is now ignored; keeping that one master the sole bootstrap authority is what stops a partition from minting two clusters, so the flag must not widen it either. A master with state rejoins its peers, and one whose data dir was reset is admitted by the sitting leader instead of forking again. * test: cover -raftBootstrap restarts in the multi-master suite Three masters start with -raftBootstrap, the way the helm chart renders it on every master on every roll, and the cluster has to hold one TopologyId after they all restart. /dir/status is proxied to the leader, so each master's own view of the identity is read out of its log, which is where a fork shows up. Before the fix the hashicorp case minted a new id on each restart. |
||
|
|
5ebc9c9f4b | server: reject a Range start offset equal to the file size (#10898) | ||
|
|
c1a993bc3b |
filer: keep the TUS sub-chunks that already landed when a write fails (#10876)
* filer: keep the TUS sub-chunks that already landed when a write fails A PATCH is split into 4MB sub-chunks, and each one is recorded in the session as soon as it is stored. The session listing is what HEAD reports as Upload-Offset and what the final entry is assembled from, so a record is a promise that the data behind it exists. When a later sub-chunk failed - a read-only volume, or a client that hung up mid-body - the error path deleted the needles of every sub-chunk the same PATCH had written but left their records in place. The resuming client was then told to continue past bytes the filer had just queued for deletion, and the upload completed into a gapless manifest pointing at needles that were gone: HEAD returned the right size, GET died mid-body once a vacuum reclaimed them. Recorded sub-chunks now stay, which is what resumption expects: the client picks up at the offset the session reports, and an upload that is abandoned frees its chunks with the session. * filer: drop a TUS chunk's record before freeing its data filer.CreateEntry can return an error with the entry already inserted - the parent-directory pass runs after the insert and keeps the entry when it fails. A failed saveTusChunk therefore does not mean the record is absent, and deleting the needle outright left the same corruption the resume path used to cause: a session record pointing at data that is gone. Remove the record first and only free the needle once it is gone. A record lost with its data still stored merely leaks, which the vacuum and fsck paths already account for. * test: cover a TUS PATCH that is cut off mid-body Resets the connection after one 4MB sub-chunk has landed, resumes from the offset the session reports, and vacuums before reading the file back, so anything the filer deleted behind a kept record shows up as a short read. |
||
|
|
35d53a20f6 |
master: let the leader admit a master that starts with no raft state (#10865)
* master: answer with the leader raft already knows Topo.Leader() backs off for up to 20 seconds waiting for an election. Callers that a health probe or a client is blocked on cannot afford that: /cluster/status, /cluster/healthz and /readyz all sit past the probe timeout of both the helm chart and the operator, so a master that is still joining looks dead rather than joining, and the kubelet restarts it. informNewLeader and SendHeartbeat hold the client on a master that cannot serve it, exactly when it should move on to find the one that can. Answer these from MaybeLeader instead, which reports what raft knows right now. MaybeLeader takes over the "am I the leader myself" fallback that Leader() used to apply on top of it, so one non-blocking call is still correct; Leader() keeps the backoff for callers that must wait. * master: let the leader admit a master that starts with no raft state Neither raft implementation lets a server outside the configuration campaign: goraft's promotable() requires a non-empty log, and hashicorp rejects vote requests from a candidate that is not in its configuration. A master that comes up with fresh state therefore cannot elect itself in — the leader has to pull it in. Nothing did. The peer list is static, rendered from the replica count, so scaling it up leaves the sitting leader running the old list with no idea the new masters exist. Under goraft they wait forever. Under hashicorp they are worse off: each bootstraps a cluster of its own from the new list, and two of them form a quorum next to the live leader, with their own TopologyId. That is the split brain SetTopologyId kills a master over. Admit the peer where it registers instead. Only the leader gets past the IsLeader check in KeepConnected, and a joining master's client lands there, so that is the moment it joins. The broadcast OnPeerUpdate rides on is not enough on its own: it only reaches masters already connected, which is why a leader that came up first missed both newcomers. RaftAddServer grew a goraft branch on the way, so cluster.raft.add stops silently doing nothing on the default raft, and RaftRemoveServer with it. Bootstrapping is now one call for both implementations, made only after the peers confirm nobody has a leader, and retried until this master is in rather than checked once and dropped. * master: do not evict a peer that is still in -peers The hashicorp leader drops a master from the raft configuration as soon as it stops answering pings. A master that is merely restarting answers nothing, so an ordinary bounce shrinks the quorum behind the operator's back — and then races its own return: the master comes back, registers, gets re-admitted, and the eviction lands after it. A randomized start/stop walk lands on it. Two of three masters running, the leader evicts the one that just went down, the restart re-adds it, the removal commits late and takes the leader's own leadership with it. What is left is a two-server configuration whose other half is down, and a running master that nobody will ask for a vote — no quorum, no way back until the third master returns. -peers is what declares membership. updatePeers already reconciles the configuration against it on every leadership change, and an operator who really means to drop a master can say so with cluster.raft.remove, so keep the eviction for masters that are no longer listed at all. * test: bounce masters at random and hold the election to it Twelve rounds of stopping or starting a random master, on both raft implementations, checking the two things an election must never get wrong: two masters claiming leadership at once, and a quorum that comes back without agreeing on one. The cluster's identity has to survive the whole walk, since a master that re-mints a TopologyId is the split brain SetTopologyId kills its peers over. The seed is random and logged, so a failure names the walk that reproduces it. Below a quorum the walk moves straight on. A master that has lost its quorum cannot commit anything, and goraft only checks whether it still has one on an election-timeout ticker, after its peers have been quiet for a full timeout — measured taking over 30 seconds to step down. That direction belongs to TestTwoMastersDownAndRestart, which was giving it ten seconds and would have started failing on a slower machine; it now waits on that behaviour explicitly rather than sleeping twice and hoping. WaitForTopologyId returns the id it waited for. Reading it separately raced the leader applying the raft entry that carries it, which shows up as an empty id right after an election rather than as a wrong one. |
||
|
|
0c95137528 |
filer: stop aggregated metadata subscribers from spinning on a peer watermark hold (#10863)
* fix(filer): stop logging a held aggregated read as an error An aggregated subscriber may not read past the peers' low-watermark, and it stops at the first entry beyond it by returning a sentinel from the read callback. LoopProcessLogData logs every callback error, so on a cluster that keeps writing - where there is almost always an entry newer than the watermark - every read wrote an ERROR line naming the entry it stopped at, thousands per minute per filer. Mark the stop as control flow: an error wrapping StopReadingError is handed back to the caller unlogged, and the held-read sentinel wraps it. * fix(filer): release an aggregated watermark hold on peer progress A held read waited on the aggregated buffer's data channel, which the next write signalled - but a write cannot release a hold, only a peer reporting further progress can. On a cluster that keeps writing the loop therefore re-ran a whole pass per arriving event, log file listing and all, and held again on the same entry every time. Signal held readers from the meta aggregator instead, whenever a low-watermark rises: a peer reporting, or one dropped past its removal grace. The retry interval stays as the backstop for what no watermark covers. Count the holds so a parked subscriber stays visible. * fix(filer): floor how often an aggregated watermark hold releases Peers advance their delivery watermark on every event they stream, so releasing a hold on every advance is the same pass-per-event storm as releasing on every write, just without the log lines - and each pass lists a day of log files. Floor the release at 20ms. Advances inside the floor collapse into one release, which then delivers everything they covered. * fix(filer): pace a peer's delivery claim by what its subscribers hold at A filer's local metadata stream carries an idle heartbeat to its peer aggregators, and each peer turns it into that filer's delivery low-watermark. Aggregated subscribers hold at the minimum across peers, so a filer quiet enough to fall back on the heartbeat parked every subscriber in the cluster up to a keepalive interval - 5 seconds - behind live writes. With nine filers, most of them quiet at any moment, the minimum sat there permanently. Pace that heartbeat at 200ms once the filer has peers. It stays a keepalive, at the keepalive interval, for a filer with none. * fix(filer): wake each aggregated hold on its own watermark A persisted-log read is held by what the peers have flushed, an in-memory read by what they have delivered, but both parked on one channel closed whenever either minimum rose. Peers advance their delivery watermark on every event they stream, so a flush-held reader woke at the coalescing floor to re-list a day of log files and park again on the same entry - the storm this set out to fix, in the one place asymmetric peer progress still reached. Signal the two separately and park each read on the one that bounds it. |
||
|
|
5d5fcdf07b |
fix(filer): bound aggregated metadata reads by peer watermarks (#10803)
* fix(filer): watermark-bound aggregated metadata subscription against multi-source merge races The aggregated metadata subscription (SubscribeMetadata) merges per-filer sources that become readable at independent paces, but tracks its progress with a single scalar cursor. Once the cursor passes a timestamp T, anything a source materializes below T afterwards is silently skipped: a peer recovering from a stall re-inserts its backlog late (late ring merge), and a source's flush can land a log file, or a later chunk of the same file, after a subscriber's disk pass listed the files (late persisted-log landing). This is the residual documented in #10501. Bound the subscriber's two read paths by what every source has provably made visible, each with its own watermark: - Delivery low-watermark -> in-memory reads. The meta aggregator tracks, per subscribed peer (self included), the newest timestamp received on that peer's stream - real events, or idle heartbeats (peer streams now opt into ClientSupportsIdleHeartbeat). The aggregated ring is complete up to the minimum across peers; in-memory reads hold at it. - Flush low-watermark -> persisted-log reads. Each filer reports its local log-buffer flush watermark on its stream: a new flushed_ts_ns response field, carried on idle heartbeats and on periodic flush reports (gated on ClientSupportsIdleHeartbeat). Disk passes freeze the minimum across peers before listing the log files and hold at it; the day-boundary cursor jump and the metadata-chunks ref listing are bounded the same way, the latter at minute-file granularity. - Held reads keep the cursor at the last entry actually delivered and retry; the retry re-lists the log files, which is what picks up a late-landing file. Both watermarks are relaxed by the settled horizon (2 x LogFlushInterval) as a liveness escape, so a peer stalled beyond it delays subscribers by at most the horizon instead of forever - any loss that escape allows was unconditional before. With reads held at the flush watermark, a disk advance below it is proven complete on every peer's disk, so the unproven-crossing counter now only counts crossings the horizon escape allowed past a stalled peer. Live delivery on the aggregated stream may lag by up to the idle-heartbeat interval when some peers are quiet; SubscribeLocalMetadata consumers are unaffected. * fix(filer): resume evicted aggregated readers from an original-space disk anchor The aggregated ring rewrites out-of-order peer arrivals to its head, so a subscriber tailing it advances its cursor in bumped (arrival) timestamps, while persisted logs keep original timestamps. When a slow reader's unread window is evicted (e.g. a peer backlog flooding in after a stall) and the reader falls back to disk, resuming from the bumped cursor skips every original-space entry below it that memory never delivered - reproduced as a ~66% silent loss on a 3-filer cluster with one peer's stream frozen for ~70s while the subscriber lagged. Track a disk anchor: the newest original-space position the stream is proven complete through. Disk passes advance it directly; contiguous memory reads advance it to the peers' delivery low-watermark observed before the read (per-peer streams are ordered, so everything with an original timestamp at or below that watermark had already arrived and was delivered). A reader kicked off the ring resumes the disk pass from the anchor instead of the bumped cursor - redelivering what memory already sent is within the subscription's at-least-once contract, skipping what it never sent is not. * fix(filer): close review findings on the peer-watermark subscription bounds Four correctness holes found in review, one generated-file cleanup: - The flush-through claim could assert durability for events still on their way into the buffer: an event is timestamped before notification work that can block, and only then appended. Track stamped-but-unappended events on the Filer (the stamp shares a lock with the reader, and appends are bumped monotonically past the buffer head), and cap the reported flush watermark just below the oldest in-flight stamp. - Removing a peer deleted its watermark entries while its stream kept running: its next signal recreated the deleted entry, which then pinned the low-watermark forever once the stream died. Watermarks now advance only for tracked peers, and peer removal cancels the subscription context so the stream stops feeding the aggregated buffer promptly. - The pipelined sender folded flush reports (TsNs 0 reads as far behind) into batch Events tails, where the aggregator's nil-notification guard dropped them - a busy backlog replay could starve the flush watermark until the settled-horizon escape opened a loss window. Control messages are now unbatchable on the sender, and the receiver also reads watermark state off nested batch entries as belt and braces. - A give-up skip's cursor was not anchored, so the next eviction rewind undid the counted decision and re-entered the same park forever when the evicted window carried bumped timestamps. The anchor now follows give-up skips; an anchored cursor makes the rewind a no-op and keeps the gap machinery's re-arm onto the retained window reachable. - Regenerated-file churn from a different protoc-gen-go-vtproto version is dropped: the vtproto file is upstream's, plus only the flushed_ts_ns marshal/size/unmarshal cases in the same generator style. New tests pin the in-flight floor, the no-resurrection rule for removed peers, and that control messages are never nested in batches. * fix(filer): keep a removed peer's watermarks through a grace period Deleting a peer's watermark entries the moment the master removes it reopened the loss the watermarks exist to prevent: a filer frozen or partitioned long enough to miss master heartbeats is removed from the cluster, its unflushed events still exist, and with its entries gone the low-watermarks snap forward to the healthy peers - subscribers advance past the absent peer's window and its late-landing log files are silently skipped. Reproduced on a 3-filer cluster: freezing two filers for ~70s got them removed ~28s in, and a catching-up subscriber lost their entire overlapping window. Removal now only marks the peer; its watermarks keep participating in the low-watermarks for a grace period (2 x LogFlushInterval, matching the subscribe loops' settled horizon, which already bounds a stale watermark's influence meanwhile). A re-added peer clears the mark and continues its values monotonically - the flap case costs nothing. A peer that stays gone is dropped when the grace expires, so a decommission cannot pin the low-watermarks, and a dropped peer's straggling signals cannot resurrect its entry. * fix(filer): cap delivery heartbeats by the in-flight floor; harden stamps Second review pass on the watermark bounds: - Idle heartbeats on the local stream claimed delivery-completeness through "now" while an event could still sit stamped-but-unappended behind blocking notification work. A peer aggregator turns that claim into its delivery low-watermark, so it could advance (and anchor credits with it) past an event that had not been streamed yet. The heartbeat timestamp is now capped just below the oldest in-flight stamp, like the flush claim already was. - In-flight stamps are forced monotonic against the registry's own history, so a wall-clock step backwards cannot slip a new stamp under an already-sampled floor. The cross-goroutine ordering still shares the meta log's global forward-clock assumption; the comments now say so instead of overclaiming. - Duplicate removal notifications no longer refresh a removed peer's grace deadline: the first removal time wins, so a decommissioned peer cannot sit in the watermark sets forever on repeated updates. - A failed buffer append clears the event's in-flight stamp on purpose: the event is dropped from the change stream entirely (a pre-existing defect of the append path, loudly logged), and a watermark waiting for it would pin this filer's claims forever. The comments now state the decision instead of implying the failure cannot happen. * docs(filer): tighten the watermark comments Comment-only: compress the narrative comments added on this branch down to their load-bearing invariants, and fix one stale sentence (peer removal no longer deletes the watermark entries immediately). No code changes. * fix(filer): subscribe to the local filer before remote peers Self's events reach the aggregated buffer only through the aggregator's own subscription to it, but bootstrap only seeded the peers the master already listed - and self's master registration races that listing, so the watermark set could hold remote peers without self. Once the remotes signalled, the low-watermarks would claim completeness for a stream that was still missing a merge source, letting aggregated subscribers advance past the local filer's events before its subscription started. Seed self first, unconditionally: before that the watermark set is empty (a documented safe state - reads hold at the settled horizon), and after it the set can never be remotes-only. The later master update for self, or a duplicate in the listed peers, is a no-op via the already-followed check in OnPeerUpdate. * fix(filer): fence watermark claims against wall-clock regression Record issued heartbeat/flush claims in the in-flight registry and stamp later events above them, so a backward clock step cannot land an event under a watermark a peer has already advanced to. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(filer): re-check the buffer head after fencing heartbeat claims An event appended between the caught-up check and the delivery claim was covered by the claim but not yet sent on the stream. The claims fence later stamps, so re-checking the head after them proves every covered event was already sent before the heartbeat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(filer): cross the aggregated ring's pre-subscription range only on proof The eviction gate and the gap proofs read "nothing evicted yet" as "memory holds everything after the cursor". That is false for the merge-fed aggregated ring, which is born empty while every peer's history sits on disk: before the ring's first real eviction, a subscriber whose cursor was still below the bounded chunk pass's listing stop was served the ring's earliest entry inclusively, silently skipping the withheld pre-restart files - and the idle-wait callback credited the delivery low-watermark to the disk anchor in the same disconnected state. Mark everything at or below the subscriptions' start as evicted when the aggregator is built, credit the anchor only once the run is connected to the ring, and give the aggregated gap pass a real proof to cross the marked boundary with: each disk pass's proven coverage (the peer flush low-watermark capped by the pass's listing bound). An empty pass whose proof reaches the eviction watermark crosses to it silently - no park, no loss counter - so the mark costs a bounded catch-up delay instead of the 15-minute give-up. * fix(filer): keep shipped chunk tails at or below the hold point A log file spans past its named minute (window start plus up to a flush interval), and chunk-mode clients apply a shipped file whole - so a file tail past the hold point can become a persisted client checkpoint beyond what every peer has proven, and a crash inside that window resumes past another peer's late-but-in-contract flush. Stop the ref listing a minute plus a flush interval below the hold; the withheld band is served by the memory pass (ring retention far exceeds it) or by later passes as the hold advances, so freshness is unchanged. A frozen peer flushing one window that spans its whole freeze can still overshoot; that residual is bounded by the freeze and needs a crash inside it. * docs(filer): trim the review-fix comments --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
3cf7d306a5 |
Give the WebDav chunk reader a bounded, invalidatable location cache (#10801)
* mount: re-resolve volume locations after a failed chunk read NewChunkGroup passed nil as the ReaderCache's CacheInvalidator, so retryFetchAfterCacheInvalidation was dead code on the FUSE read path. A mount that cached a volume's locations while one server was down kept retrying that server after it died, then returned EIO, even though the master and filer both resolved the live replica. The S3 gateway already passes its filerClient; do the same for the mount. * test: FUSE integration tests for volume server failover One mount appends while a second tails, and a volume server is killed, started or restarted mid-stream against a 001-replicated cluster of three volume servers. Automates the scenario matrix reported for Docker Swarm mounts, including the large-file variant and a no-chaos control. * test: report the filer's own view when append content mismatches A mismatch between what the writer wrote and what the reader sees can come from either side's cache. Read the file back through the filer's HTTP handler as well, and let the mount verbosity be raised from the environment, so a failing run says which layer lost the data. * test: wait for the reader mount to converge before comparing A mount caches metadata for about a second, so reading the file the instant the writer's last close returned can legitimately come back short. Poll the reader until it matches or the timeout expires; content that is wrong rather than merely late never converges and still fails, now with the writer's mount and the filer's own view alongside it. * test: detect a failover cluster child that exited at startup Signal(0) succeeds for a zombie and nothing reaped these children until shutdown, so a process that died on startup looked alive until the readiness timeout expired. Reap each child as it is started and consult the result. * test: read a file the killed volume server actually holds Placement decides which two of three servers back each volume, so killing volume N and reading readfile-N could pass without the victim ever holding a replica of it. Resolve each file's volumes through the filer and the master, and pick one the victim backs, preferring a file the reader has not cached. * ci: stop persisting checkout credentials in the failover workflow The job does not use the token after cloning. Also tag the README's command block as bash and match the timeout the workflow actually uses. * test: discard the ignored errors errcheck flags in the failover harness * test: resolve manifests when mapping a file to its volumes A manifest chunk's own fid names the volume holding the manifest, not the volumes holding the data, so a large enough file would point the failover victim at the wrong server. * test: pin the stale-location recovery path with a primed reader Reading a file for the first time after a server dies proves nothing: the lookup is fresh and returns the survivor. Kill one holder and wait for the master to drop it, read a file on that volume so the reader caches the lone survivor, restart the first server, then kill the survivor. The reader's only cached location is now dead while the data is live elsewhere, which is the case the invalidator exists for: EIO without it, recovery with it. * filer: re-look-up a chunk's locations as soon as they all fail A read that fails against every location it was given is far more likely to be holding a stale list than to be hitting a cluster that is briefly slow, but the retry loops spent the whole backoff ladder, about 13 s, before the caller got a chance to invalidate and look the chunk up again. Give the loops a refresh hook and let the reader cache invalidate on the first fully failed pass, so recovery starts in milliseconds. Clients without an invalidator keep the old behavior. The filer's streaming read path has its own fetch loop and is not covered. * webdav: give the chunk reader a bounded, invalidatable location cache WebDav resolved chunk locations through filer.LookupFn, whose own doc asks long-running processes to prefer wdclient.FilerClient: its cache is unbounded, and it has no way to invalidate an entry, so the reader cache was constructed with a nil invalidator and a WebDav server that had cached a location kept reading from it after the volume moved or died. Use FilerClient, as the mount and the S3 gateway already do. * filer: refresh locations on the random-read path too readChunkSliceAt bypasses the chunk cacher in random-access mode and fetches the range directly, which left it without the invalidation the cacher does: a random reader parked on a stale location had no way back at all. Hoist the refresh hook onto the reader cache so both paths share it. * filer: compare chunk locations as a set, not in order Lookups shuffle the locations they return, so comparing positionally reads a reshuffle of the very same replicas as a fresh set and spends an immediate retry on locations that just failed. weed/filer already had an order-independent comparison for this; move it next to the retry loops so both callers share one helper. |
||
|
|
6fda8c67f3 |
Guard the gcs credential path in FetchAndWriteNeedle like the other backends (#10796)
* volume: accept only static-key gcs credentials on the fetch request An inline credentials document of a federated type points the SDK at a url, file or executable of the caller's choosing for the token exchange, so the request-supplied value is no longer just a key. * volume: guard the gcs token endpoint like the other remote endpoints Inline credentials pick where the token request goes, so route the gcs client through the same deny-list and rebinding-safe dialer used for S3 and azure. * rust volume: pin that gcs has no credential-driven dial path * volume: only check gcs credentials on a gcs remote conf Only the gcs backend reads that field, so another backend carrying a stale value should not fail the request. * gcs: load credentials with the type the caller expects The untyped loader is deprecated because it reads whatever the document claims to be; callers handling credentials they do not control now name the types they accept. |
||
|
|
fbd85d31b0 |
ec.decode: check the rebuilt .dat is complete before the shards can be deleted (#10768)
A decode ends by deleting the shards it read, and the only thing standing between that and a bad reconstruction is verifyDecodedVolumeBeforeDelete, which asks whether .dat and .idx are non-empty. A .dat truncated to a single byte passes, and the shards -- the only other copy of everything past the cut -- are deleted on the strength of it. The server already knows the answer it never checks: FindDatFileSize returns the extent the EC index references, and WriteDatFile rebuilds to it. Compare the two once the file is written and fail the decode instead of reporting a short volume as a good one. Longer than the extent still verifies -- padding is not missing data -- so only a genuinely short rebuild is rejected. Needle counts cannot answer this: .idx is written from .ecx, so the count matches by construction and a truncated .dat still reports every needle. |
||
|
|
602746f51d |
test: EC lifecycle chaos harness, with four fixes it found (#10763)
* ec: let the encode's balance see a migrating volume's shards across disk-type buckets Shard generation writes beside the source .dat, so a cross-tier encode (source on hdd, -diskType=ssd) leaves the fresh shards in the source disk-type bucket. The encode's internal balance ingested only the target bucket, saw no shards, and planned no moves; the spread guard then correctly aborted the encode (and before that guard existed, the shards silently stayed clumped on the generation host in the wrong tier). EcBalance now takes the encode batch as migratingVolumeIds and ingests those volumes' shards from every bucket, while everything else keeps the bucket filter so a plain ec.balance never drags deliberately tiered shards onto another disk type. The in-memory model delete also becomes bucket-agnostic: a node holds a given shard in exactly one bucket, and a bucket-scoped delete missed cross-bucket moves in the dry-run model. * volume: decode reads shard 0 from its resolved path, not the EC volume's base dir On a multi-disk server a volume's shards can sit on several disks; the store registers each shard with its own path and CollectEcShards resolves them, but FindDatFileSize derived the .ec00 path from the EcVolume's base directory. When shard 0 lived on a sibling disk, VolumeEcShardsToVolume failed with 'open ...ec00: no such file or directory' and ec.decode aborted. * ec: decode re-copies shards the topology claims but the target does not hold An interrupted earlier decode or balance can leave the master believing the decode target holds a shard whose file never landed: the mount registered but the partial copy was cleaned, or the file was swept. The collect step took the topology's word for it, excluded the shard from the copy set, and the decode failed with 'missing shard'. Probe the target's live inventory (VolumeEcShardsInfo) and treat anything it cannot serve as still-to-copy. * ec: decode discovers shards across disk-type buckets Shards sit wherever encode generation and balance left them: a cross-tier encode leaves them in the source disk-type bucket, a partial migration straddles buckets. ec.decode scoped its shard discovery to the -diskType bucket and reported a decodable volume as having no shards at all. Union across buckets, the way the encode's shard verification already does. * test: EC chaos lifecycle harness Randomized, seeded sequences of the EC lifecycle against a live cluster in the production-shaped layout: multiple data disks per server, a separate -dir.idx directory so .ecx/.ecj sidecars are shared across disks, and a tagged ssd tier. Operations cover encode (hdd and ssd targets), balance, shard damage plus rebuild, decode, re-encode, deletes, scrub, tier moves, crash-restarts, sidecar fault injections (a data-dir .vif pushed into the shared idx dir; a stale-generation shard planted beside a newer encode), and interruptions: a real weed shell subprocess killed mid-encode, mid-decode, and mid-balance, with the recovery re-run required to converge. One invariant holds after every step: every stored byte reads back identical and every deleted needle stays deleted. EC_CHAOS_SEED and EC_CHAOS_STEPS make runs reproducible and scalable. A known gap is tolerated and logged rather than fixed here: a shard mounted on two disks of one node (orphan adoption after an interrupted copy) is invisible to ec.balance's dedup and unaddressable by ec.shard.unmount's shard@address form, so no cleanup path exists yet. * test: fail payload-corruption checks on the test goroutine t.Fatalf inside require.Eventually's condition runs on the poller's goroutine, where Goexit kills only that goroutine and the corruption message can be lost behind a generic timeout. Record the mismatch, end the polling, and fail on the test goroutine. Also assert the full shard count in the cross-bucket decode-discovery test. |
||
|
|
d713ab49f9 |
volume: validate replica targets and restrict gcs credentials in FetchAndWriteNeedle (#10755)
* volume: validate replica upload targets in FetchAndWriteNeedle The replica leg forwarded the fetched needle to a caller-supplied address without checking it, so a malformed target could redirect the upload to an unintended host or path. Require each replica target to be a bare host:port whose host is not loopback / link-local / unspecified, reusing the address deny-list; cluster peers legitimately sit on private networks, so RFC 1918 / CGNAT stay allowed and -volume.allowUntrustedRemoteEndpoints still opts out. Validate every target up front so a bad one fails the request before the local write, and upload through a client that re-checks the resolved address at connect time so a replica hostname cannot rebind to a blocked address after validation. Mirrored in Rust (validation moved ahead of the local write; the Rust S3 path's connect-time re-check is still a follow-up there). * volume: only accept inline gcs credentials in FetchAndWriteNeedle The gcs credentials value on this request could name a local filesystem path, which the SDK reads from disk. Accept only inline JSON here; the server-side GOOGLE_APPLICATION_CREDENTIALS env var still supplies a path. The Rust volume server has no gcs backend, so there is nothing to mirror. |
||
|
|
9125b9c835 |
volume: extend the remote-endpoint guard to the azure backend (#10754)
* remote_storage/azure: allow a per-request HTTP client Thread an optional *http.Client through NewAzBlobClient and add azure.MakeWithHTTPClient, mirroring the S3 backend. When set, the client overrides the azblob transport so a caller can pin the dial path. The existing makers pass nil, so behavior is unchanged. * volume: extend the remote-endpoint guard to the azure backend The endpoint validation and rebinding-safe dialer in FetchAndWriteNeedle covered the S3-SDK backends. The azure backend also dials a caller-supplied AzureEndpoint, so route both families through a single guardedRemoteClient helper that returns the endpoint each backend dials and a constructor bound to the guarded HTTP client. azure is guarded only when AzureEndpoint is set; an empty endpoint derives the public host from the account. -volume.allowUntrustedRemoteEndpoints still opts out. * rust volume: assert the azure endpoint has no remote-client path The Rust volume server has no azure backend, so make_remote_storage_client rejects the type before any client is built. Add a regression test pinning that invariant. |
||
|
|
7f27c572c4 |
log_buffer: end bounded reads that find the buffer empty (#10750)
A bounded LoopProcessLogData (stopTsNs set) on a buffer that never took a write since process start fell into the ResumeFromDiskError branch, which never checks stopTsNs when ReadFromDiskFn is nil and HasData() is false. The read parked on the notification loop forever while the subscription's idle heartbeats kept the stream looking alive, so a bounded SubscribeMetadata pass on a freshly restarted idle filer never completed. Terminate like the caught-up path does, returning a nil error: leaking the pending ResumeFromDiskError would latch the filer's outer loop into its gap machinery, which parks the bounded subscriber all over again. |
||
|
|
4f50c5b0d4 |
feat: throughput limits for replicate, EC shard, and worker-driven moves (#10749)
* feat: throughput limits for replicate, EC shard, and worker-driven moves VolumeCopy was the only rate-limitable transfer; EC shard copies, replica creation, and worker-driven moves all ran at whatever the receiving server's maintenance rate allowed, with no per-operation control. - proto: VolumeEcShardsCopyRequest and the balance / ec_balance task params and configs gain io_byte_per_second; 0 keeps today's behavior (the volume server's own maintenance rate governs). - volume server: VolumeEcShardsCopy throttles with one WriteThrottler per request, shared across the shard, .ecx, .ecj, .vif, and .ecsum copies so the limit caps the transfer as a whole - the same shape as VolumeCopy. - volume_move: ReplicateVolume accepts the limit; EcMoveOptions carries it through MoveEcShards/CopyAndMountEcShards into the copy request, with fake-client tests asserting propagation. - shell: ec.balance gains -ioBytePerSecond; volume.tier.move's replication top-up honors the command's existing -ioBytePerSecond instead of running unthrottled. - worker: balance and ec_balance configs gain io_byte_per_second (surfaced in the admin config schema), carried through detection and plugin job parameters into task params and handed to the shared mover; batch balance jobs inherit the limit from their detection results. The limit is per copy stream, so maxParallelization multiplies the aggregate ceiling. * worker plugins: expose io_byte_per_second in the plugin config and derive it The plugin-driven detection path derives its task Config from the plugin configuration values, and both balance and ec_balance left IoBytePerSecond at zero there - a configured limit silently reverted to the server maintenance rate. Both derive functions now read the field (clamped at zero), and the plugin descriptors expose it with defaults so the configuration form carries it. |
||
|
|
db5a086d04 |
read cold remote objects straight from the origin while caching (#10731)
* refactor: extract remote mount resolution into shared helpers * refactor: share the adaptive remote cache wait policy * filer: stream cold remote reads from the origin while caching * s3: stream cold remote reads from the origin instead of 503 retries * test: cover the S3 origin stream-through path * remote mounts: match on path components and prefer the longest mount * fail short origin streams instead of silently truncating * s3: try the origin before failing a cold read on a local cache error * s3: gate origin streaming on the entry's resolved version * return the cache RPC's NotFound as a canonical status and classify it everywhere * filer: keep multipart-range cold reads on the retry path |
||
|
|
5b519489c1 |
remote_storage: build all S3-compatible clients through one constructor (#10720)
* remote_storage: build S3-compatible clients through one constructor The eight non-s3 S3-SDK providers each duplicated the AWS session setup and only the s3 maker could take a custom *http.Client. Route every S3-compatible type (s3, wasabi, b2, aliyun, tencent, baidu, filebase, storj, contabo) through MakeWithHTTPClient with a single options table, and add S3CompatibleEndpoint so callers can resolve the endpoint a given type dials. No behavior change. * volume: apply the remote-endpoint check to all S3-compatible providers FetchAndWriteNeedle validated the endpoint and used the pinned dialer only for type "s3". Every S3-SDK backend (wasabi, b2, aliyun, tencent, baidu, filebase, storj, contabo) dials a caller-supplied endpoint through the same client, so gate on S3CompatibleEndpoint to apply the same check uniformly. -volume.allowUntrustedRemoteEndpoints still opts out. * volume: don't route the guarded remote-endpoint client through a proxy The guarded client exists to dial the validated endpoint directly and re-check the resolved IP at connect time. With http.ProxyFromEnvironment set, the dialer only validates the proxy's address while the proxy re-resolves the endpoint host, which reopens the rebinding window. Drop the proxy on this path; operators that need one can opt out with -volume.allowUntrustedRemoteEndpoints. |
||
|
|
c6e1387f59 |
shell: multi-target fs.mergeVolumes and volume.mark -readonlyCanDelete (#10706)
* shell: fs.mergeVolumes distributes one volume across multiple -toVolumeId targets * volume: volume.mark -readonlyCanDelete rejects writes but keeps accepting deletes * seaweed-volume: mirror readonlyCanDelete volume state |
||
|
|
365d3e9e87 |
filer: TUS concatenation extension (#10702)
* filer: TUS creation accepts Upload-Concat partial uploads * filer: TUS final uploads concatenate completed partials * filer: TUS concatenation tests * filer: consumed marker pins TUS chunk ownership on completion * filer: TUS session delete decides chunk ownership after removing the session info * filer: TUS completion persists the consumed marker before creating the entry * filer: TUS completion re-verifies the session after persisting the consumed marker * filer: serialize TUS session ownership transitions per filer * filer: surface failed TUS consumed-marker rollbacks |
||
|
|
753cb8cda8 |
master: stop copying the cluster to name it (#10700)
* topology: name a node's volumes without copying them ToVolumeLocations reads a volume id off every volume in the cluster, and got there through GetVolumes, which copies a whole storage.VolumeInfo per volume to be read for four bytes of it. Every client that connects asks for this. At 800k volumes the walk goes from 94.6MB to 16.0MB, which is the ids themselves. * master: log why a client send failed, not what was sent The message names every volume on a newly connected node, so a client going away had the master format a protobuf that size into text -- through the one log level that is always on. The error is the part worth having. |
||
|
|
46ce8cbe84 |
master: stream volume listings (#10676)
* master: stream volume listings A listing of 800k volumes is 36MB on the wire but 305MB as messages, and the master built all of it, then held it while grpc encoded it. Two of those at once is most of a small master's heap, and the maintenance scanner asks every 30 minutes. The topology goes out first, listing nothing, then its volumes in batches, so the master holds a batch rather than a cluster: 341MB of live heap for one listing becomes 4.4MB. It allocates much the same either way -- what changes is how much of it has to be live at once, which is what sets the heap ceiling. Batches are built under their disk's lock and sent outside it, so a slow reader stalls the stream rather than the topology. They therefore do not share one instant, which a single listing did not either: it takes each disk's lock in turn, so a volume moving during either can be seen twice or not at all. The client helper hides which kind of master answered: one too old for the stream is asked the old way and its reply cut into the same batches. Either way the topology handed over lists no volumes, so a caller cannot come to depend on finding them there. * admin: stream the listing the maintenance scan reads It asks for every volume in the cluster every 30 minutes. Reassembling it client-side keeps the scan identical -- ActiveTopology splits disks by the disk ids on the volumes, so it needs them in the topology -- while the master no longer builds the whole reply to send it. |
||
|
|
00c5572e8c |
volume: decode IPv6 transition addresses in the remote-endpoint guard (#10683)
* volume: decode IPv6 transition addresses in the remote-endpoint guard checkBlockedIP normalized only ::ffff: mapped IPv4, so NAT64 (64:ff9b::/96), 6to4 (2002::/16), Teredo (2001:0000::/32), and IPv4-compatible (::/96) addresses that embed an internal IPv4 (loopback, 169.254.169.254, RFC 1918) passed the endpoint guard even though the plain IPv4 forms are refused. Extract the embedded IPv4 from those forms and re-check it against the deny list, which covers both the up-front validation and the dial-time guard. Mirrored in the Rust volume server. * volume: require the full NAT64 well-known prefix before decoding Only 64:ff9b::/96 carries the embedded IPv4 in the low 32 bits, so also require bytes 4-11 to be zero before treating an address as NAT64; other 64:ff9b: prefixes place the IPv4 elsewhere and are left untouched. Add public-target coverage for 6to4, Teredo, and IPv4-compatible so every decoder is exercised on both a blocked and an allowed destination. Mirrored in the Rust volume server. |
||
|
|
3911e4c548 | master: keep a racing registration out of a dying collection (#10677) | ||
|
|
a2ff9cca27 |
master: let VolumeList ask for the volumes it wants (#10674)
* master: let VolumeList ask for the volumes it wants The request carried nothing, so every caller was answered with the whole cluster. A dashboard opening one volume's page, or a capacity probe adding up one bucket, was served all 800k of them and threw away the rest -- and the master built every one of those messages first. The topology, its disks and their counters are still reported in full: a caller reading free space or replica placement needs the cluster whichever volumes it asked about. Only what is listed under a disk is selected, ec shards included. An empty collection and a zero volume id take everything, the way volume.list already reads its own -collectionPattern and -volumeId, so a caller that forgets to narrow is answered too much rather than answered wrongly. That leaves the default collection unnameable, since it is the one the empty string names, so it gets a field of its own. An older client sends none of it and is answered exactly as before. * admin: ask the master for the volume the page is showing A volume's detail page was pulling every volume in the cluster to find one and its replicas, and discarding the rest. * admin: ask the master for the ec volume the page is showing Same as the volume detail page: one volume's shards were found by pulling every ec shard in the cluster. * s3: ask the master for the bucket's own collection The SOSAPI capacity probe summed one collection's volumes out of a listing of every volume in the cluster. Cluster capacity still comes out the same: it is read from the disk counters, which a filtered listing reports in full. * topology: read the disk usage counters atomically They are written with atomic.AddInt64 from heartbeats but were read plainly by the two listings and by FreeSpace, and the map they sit in was iterated without the lock its neighbour takes. Under -race a listing concurrent with a heartbeat trips on both. |
||
|
|
f09e8345c6 |
storage: stop keeping the remote storage key on the master (#10672)
A master decides nothing from it. Every caller that read it was asking whether a volume is remote, which the backend name answers, and the value itself is reported on demand by the server holding the volume, through the volume info in ReadVolumeFileStatus. It is also the one string here that cannot be shared: unique per volume, so unlike the collection and backend names it carries its own characters for every volume a master tracks. VolumeInfo goes from 136 bytes to 120. 800k volumes registered from a heartbeat that has been over the wire go from 214 to 163 B/volume when tiered. The volume server's own status page keeps showing the key, now read from the volume it holds rather than relayed through a master, which is also where the other volume server implementation reads it. The heartbeat digest drops it on the same grounds: a change to something the master does not hold cannot make its copy stale. Both implementations and their shared vectors move together, and the field-coverage test now names what is deliberately not retained rather than being loosened. |
||
|
|
567052bfb6 |
s3: take bucket sizes from the master's summary (#10664)
* pb: ask the master what each collection holds Callers tracking usage were sent every volume in the cluster to add up themselves, which is the master's largest single allocation. * topology: summarise what each collection holds One pass over the topology, allocating per collection rather than per volume. Regular volumes count once each for logical totals and once per replica for physical, taken from the lookup index, which is already keyed by volume and so needs no set of seen ids. Ec shards are node-local so their sizes sum, while the file and delete counts describe the volume and resolve once every holder has been seen. Replicas of one volume disagree while a write is landing or a heartbeat is late. Walking a full listing took whichever replica the map iteration reached first, so the answer moved between runs; this takes the largest, which is stable and never reports usage below what some replica already holds. * s3: take bucket sizes from the master's summary The bucket size metrics pulled the whole volume list once a minute and added it up, which cost the master 184.6MB of allocation and 17.8MB on the wire for six numbers per collection. VolumeList over 550k volumes 184.6 MB allocated, 17.8 MB on the wire CollectionStatistics 176 bytes allocated, 47 bytes on the wire The aggregation moves to the master with it, so the cases the removed tests covered are now asserted against it directly. * topology: count the replica holding the most live data Quotas are enforced on size less deletions, and the replica with the biggest raw size can be the one that has deleted the most. Counting it reported a bucket smaller than it is and would leave one writable over its quota, which is the opposite of what picking the largest was meant to guarantee. * topology: cap a volume's deletions at what it holds Live usage is read as a collection's size less its deletions, so a volume reporting more deleted bytes than it has cancels live bytes belonging to other volumes in the same bucket and reports it smaller than it is. Replica selection already floored that volume's own live size at zero; the totals have to agree with it. |
||
|
|
a2ffc7aadf |
heartbeat: keep the master current through collection churn (#10657)
* heartbeat: name departed volumes in delta heartbeats * master: release the lookup index with a deleted collection * master: keep a fresh grow safe from the report that raced it * volume: name the volumes a deleted collection took with it Deleting a collection left the master to work out what went by omission from the next full volume list, which it no longer gets: heartbeats carry the whole list only when the master asks for it. The volumes a bucket's churn creates and destroys between two of those requests are never named in either direction, so the master keeps counting their slots as occupied and a cluster that creates and drops collections quickly runs its free-slot accounting dry -- assigns fail with no free volumes left while the disk holds a handful of volumes. The destroy path already knows exactly which volumes it removed, so send them down the same channel every other deletion uses. * rust: name the volumes a deleted collection took with it Mirrors the Go volume server. The notify path derives its deltas by diffing snapshots, so a collection delete that does not wake it is invisible until the master next asks for the whole list. |
||
|
|
3a61debaa5 |
filer: rebuild peer metadata subscriptions after a master reconnect (#10648)
* filer: keep the existing peer subscription on a repeated add A cluster node add for a peer that is already followed restarted the subscription, dropping the metadata events between the two runs. * master: tell a connecting client the current cluster membership Cluster node updates are only broadcast to the clients connected at that moment. A filer that lost its master stream while a peer came back never learned about the peer, and stopped replicating its metadata for good. * test: a filer joining the master learns about the filers already there * test: a filer resubscribes to a peer that registered while it was disconnected Runs the reported sequence against real processes: filer2 leaves, filer1 is paused and its master stream is broken, filer2 registers again, and filer1 has to replicate from it after reconnecting. |
||
|
|
37f3dff677 |
volume: validate the file extension in CopyFile and ReceiveFile (#10644)
* volume: validate the file extension in CopyFile and ReceiveFile
CopyFile and ReceiveFile build an on-disk path from the client-supplied
Ext. Both are intentionally ungated for cluster-internal peers, so a
value like "/../../x" is joined onto the volume directory and, once
path-cleaned, resolves outside it -- an EC-shard receive can then write,
and CopyFile read, anywhere the process can reach.
Constrain Ext to a real suffix (a leading dot followed by alphanumerics)
before it is used to build any path, so it can no longer carry a
separator or a parent reference.
* test: use an alphanumeric missing-file extension in the copy variants
The not-found and stop-offset-zero cases used ".definitely-missing" as a
deliberately absent source. The extension is now validated, and the hyphen
makes it invalid, so switch to ".missing" -- still a nonexistent file, but a
real extension shape.
* volume: validate the collection in CopyFile and ReceiveFile
The client-supplied Collection is folded into the on-disk path as
"<collection>_<vid>" by VolumeFileName and EcShardBaseFileName, both joined
with path.Join / util.Join. A Collection carrying a separator, e.g.
"../../x", therefore path-cleans to a target outside the volume directory,
the same escape the extension check just closed. Reject a collection that is
a bare parent reference or holds a separator; ordinary names ('.', '-' and
all) still pass.
|
||
|
|
e9cde3e4b1 |
master: gate raft membership RPCs behind the admin whitelist (#10649)
* master: evict a dead peer via the local raft handle OnPeerUpdate only runs on the leader, and the AddVoter branch right above mutates the local raft directly. The remove branch instead dialed our own RaftRemoveServer back over gRPC. Drop the self-dial and remove the peer through the local handle, matching the add path. This also leaves operator tooling as the only caller of the RaftRemoveServer RPC. * master: require whitelist auth for raft membership RPCs RaftAddServer, RaftRemoveServer and RaftLeadershipTransfer rewrite raft quorum but had no caller check beyond "am I the leader". Any client that could reach the master gRPC port could add an unreachable phantom voter and stall the write path. Gate the three on the admin whitelist, mirroring the volume server's checkGrpcAdminAuth. With no whitelist configured the guard allows every caller, so default and single-master deployments are unaffected; operators who set -whiteList get these RPCs locked down to it. The leader's own dead-peer eviction no longer dials these RPCs, so the only remaining callers are operator tooling. |
||
|
|
ce7d388639 |
heartbeat: send only the volumes that changed (#10640)
* pb: let a heartbeat carry only the volumes that changed A partial list cannot travel in volumes: a master that did not understand it would read the absences as deletions. So changes get their own field, used only once the master has said it compares digests and can tell when it has fallen behind. * master: apply the volumes a heartbeat reports as changed Only the named volumes are touched. A full report says the server holds exactly these; a changed report says nothing about the ones it leaves out, so absence must not read as removal. Also advertises that the master compares digests, which is what lets a server stop sending its whole list. Advertising it once per connection means a server reconnecting to a master that does not is back to full lists straight away. * volume: send only the volumes that changed once the master accepts them The whole list goes on every heartbeat until the master says it compares digests, and again whenever it asks, so a master that cannot tell when it has fallen behind never has to. has_no_volumes stays derived from a full list alone. Deriving it from what a heartbeat happens to carry would make a quiet one read as a server that had lost every volume, and the master would drop them all. The digest still covers every volume held rather than the ones sent, which is what lets the master confirm that applying the changes left it current. Reporting state is per-connection: a server that reconnects, or reaches a different master, starts again from the full list. * volume: let the zero reporting state stand for having told no master anything A Store built as a literal, which tests do, left the reporting state nil and panicked on the first heartbeat. As a value its zero form already means nothing has been reported to anyone, which is exactly the state that sends the whole list. * rust: send only the volumes that changed once the master accepts them Mirrors the Go volume server, with one hazard the Go side does not have: mount and unmount deltas here are derived by diffing successive heartbeats, so a heartbeat that carries a partial list would report every volume it left out as unmounted. Collecting now returns the full set alongside the message, and every site that diffs uses that rather than what went on the wire. * volume: do not let a full-list request be lost to the heartbeat it raced The request arrived while a heartbeat was already being built as a delta, and committing that heartbeat cleared it, so the master waited for another digest mismatch before asking again. Count the requests and clear only the one the heartbeat answered. * rust: stop marking volumes reported by a heartbeat that is thrown away The state-notify path collected a heartbeat only to diff its volume list, then sent a delta message of its own and dropped the one it had collected. Once collecting recorded what the master had been told, every mount or unmount silently marked the changed volumes as sent, and the master learned of them only after a digest mismatch. Snapshotting no longer records anything, and no longer expires ec volumes whose deletion that path was already discarding. * master: announce only the volumes a change actually brought Every changed volume was broadcast as a new location. Volumes grow constantly and growth moves no location, so on a busy cluster that told every connected client about volumes it could already reach, filling bounded broadcast queues and pushing out the topology updates that matter. * master: ask for the full list when only one can repair the master Delta heartbeats stop the full report, and with it the only thing that re-registers a volume the lookup index lost. The volume server cannot see that divergence and its digest cannot show it, so the master now checks its own two indexes agree and asks for the list when they do not. A node reporting one volume id twice is kept on full lists for the same reason rather than merely skipped: its digest can never be verified, so nothing else would tell the master what it had stopped holding. * master: keep the volume options on every heartbeat response A volume server takes them from whatever response arrives, and preallocate is a bare bool with no way to tell off from unmentioned. A response sent to ask for the volume list therefore turned preallocation off until the server reconnected. Responses sent mid-stream now start from the configured options rather than being built field by field. * master: announce a volume the lookup index had lost Repairing the index makes the volume servable again, but clients were told it went when the node dropped out and nothing told them otherwise: the disk map still held it, so it did not count as an arrival. Reaching the lookup index is what makes a volume servable, so recovering an entry there is an arrival as far as clients are concerned, on both the full report and the changed-volume path. |
||
|
|
cab666fca1 |
filer: configurable TUS max upload size and session expiry (#10638)
* make TUS max upload size and session expiry configurable * default TUS session expiry to 24h |
||
|
|
5ec813b4f1 |
topology: follow a volume that moved between a server's disks (#10628)
* topology: follow a volume that moved between a server's disks The heartbeat diff asked only whether a volume id was reported anywhere on the node, so a volume that moved to a disk of another type stayed on the disk it left as well. The master then held two copies of it forever: the volume count was overstated, and GetVolumesById returned whichever disk the map iterated first, so lookups could hand back the disk the volume had already left. Track which disk types the heartbeat named each volume on, and treat a volume named on another disk as absent from this one. Disk types are interned to an index because a server reports a handful of them across hundreds of thousands of volumes. A volume named on two disks at once is a stale twin rather than a move, and is still kept on both -- dropping one would tell the master a replica vanished. Only a volume named twice on one disk type is unrepresentable, so that is now what marks the node, rather than any repeat of an id. * master: do not tell clients a moved volume left the node A volume moved between a node's disks is removed from one and added to the other, so it lands in both lists of the same heartbeat. Clients apply additions before deletions, so the removal wins and they end up with no location for a volume that never went anywhere. Skip removals for volumes the node still holds, as the ec shard paths already do, and update the topology before judging the delta removals so an unmount that really did happen is still reported. * trim the comments on this change to the parts that are not evident * master: judge a volume removal on normal replicas alone HasVolumesById answers for ec shards as well, so a replica encoded into ec shards looked like it was still on the node and clients were never told the normal location had gone. They hold normal and ec locations separately and prefer the normal one from the same generation, so that location would have gone on shadowing the shards. |
||
|
|
6d08b08f37 |
heartbeat: carry a volume digest and verify it (#10627)
* pb: carry a volume digest on the heartbeat The full volume list is the only way a master notices a volume that vanished without a delta, so it cannot simply be dropped. A digest gives the same guarantee without the list, and a way back to the list when they disagree. The digest has explicit presence: a server holding no volumes reports 0, which has to stay distinguishable from a server that does not compute one at all. * volume: report a digest of the volumes each heartbeat carries Digests exactly what goes on the wire: volumes skipped as quarantined, phantom or expired are absent from both the list and the digest, so the master compares against the same set the server meant to report. Runs the master's own hash over the master's own conversion of the message, so the two ends cannot drift into disagreeing about a field. * master: check the reported volume digest and ask for the list on a mismatch Compared after everything the heartbeat carried has been applied, so agreement means the master is current rather than that nothing changed. Servers reporting no digest are untouched, and a mismatch on a heartbeat that already carried the full list is reported rather than answered: there is nothing further to ask for, so asking again would loop. Nodes reporting one volume id twice are skipped for the same reason. * rust: report the heartbeat volume digest Mirrors the Go volume server. The master compares this against a digest it computes itself, so the hash has to agree byte for byte across the two implementations, not merely be a hash of the same fields: report_hash_vectors pins it against values generated by the Go side, and the ttl and replica placement narrowing the master applies when it decodes a message is applied here too rather than assumed away. A drift there would not corrupt anything, but every volume server on this implementation would report a digest the master can never match and fall back to sending its whole volume list forever, which is the cost the digest exists to avoid. * master: pin what the digest check does to each kind of report The upgrade story rests on these: a server that reports no digest is never asked for anything, so the two sides can be upgraded in either order, and a disagreement that resending cannot fix is reported rather than re-asked, so it cannot loop. * topology: enumerate the digest coverage test from the message The list of fields was written out by hand, so a field added to VolumeInformationMessage later would fall outside the digest while the test went on passing, and a change to it would never reach the master. Walk the message descriptor instead. Some fields are narrowed or normalised on the way into VolumeInfo, so the smallest change to the wire value can land back on the stored one; the test offers several values per field and asks only that some change is visible. |
||
|
|
b46946ece5 |
filer: list directories without decoding chunk lists (#10616)
* filer: decode a listed entry without building its chunk list
A readdir reads attributes and never looks at chunks, but decoding an
entry builds the whole chunk list first: four allocations per chunk, all
of it thrown away. On a directory of ordinary 4MB-chunked files that is
most of what listing costs.
DecodeAttributesOnly walks the wire format and hands everything except
the chunks to the generated unmarshaller, so new fields in filer.proto
need no attention here. The chunks are still measured, because the S3
copy and multipart paths deliberately store a zero FileSize and let the
chunk extents define the size, but nothing is allocated to do it.
The blob is only re-encoded once a chunk is actually seen, so an entry
without any -- every directory, for one -- is unmarshalled where it lies
and pays nothing for the walk.
Listings opt in through the context, the way the lazy remote paths
already do; a store that ignores it stays correct.
chunks full attrs-only allocs
0 312.8n 310.1n ~ 1 -> 1
1 686.1n 411.1n -40.07% 7 -> 1
4 1.742u 667.4n -61.69% 24 -> 1
16 5.770u 1.544u -73.25% 86 -> 1
64 25.23u 6.004u -76.20% 328 -> 1
* mount: list directories with chunk lists omitted
The two meta cache listings behind a readdir are the only callers, and
neither reads a chunk. On 200k single-chunk files one enumeration goes
from 364ms to 277ms and drops a million allocations.
The read-through listing still fetches whole entries from the filer,
which would need the request to say it wants attributes only.
* mount: give the readdir benchmark's entries a chunk
Chunkless entries made the decode look far cheaper than it is, which is
the part of a listing worth measuring.
* filer: let a listing ask for entries without their chunk lists
The read-through readdir fetches whole entries over gRPC, and for a wide
directory the chunk lists are most of what crosses the wire and most of
what the client then unmarshals. A 4MB-chunked file is 113 bytes of
entry against 46 without its chunk.
ListEntriesRequest gains omit_chunks. The size a client needs is already
in the attributes, where the store decode folded the chunk extents in,
so dropping the list costs the client nothing.
The filer still reads the entries whole. A listing is where a TTL-expired
entry gets collected and deleted, and deleting one needs its chunks to
find the data, so omitting them there would leak. Only the response is
trimmed.
The hint moves to filer_pb so one context flag serves both transports:
the gRPC request sets omit_chunks, and a listing served from the local
store skips building the chunks. Cache population is unaffected either
way, since EnsureVisited starts from its own context.
* filer: reject a chunk the full decoder would reject
The walk skipped a chunk's bytes without looking inside them, so a
FileChunk carrying a corrupt nested fid, or a string that is not valid
UTF-8, sailed past the listing decoder while every other read of the same
entry still failed. The file listed with a plausible size and then gave
EIO on open, and corruption that used to fail the listing loudly was
hidden instead.
The chunk bytes are the one part of the blob the generated unmarshaller
never sees, so the two checks it would have made are made here: a
submessage has to parse, and a proto3 string has to be valid UTF-8.
FileChunk's only submessages are FileIds of scalars, so walking them is a
complete check. A descriptor-driven test fails if FileChunk ever gains a
field of either kind that the walk does not know to check, which is the
part that keeps this honest as filer.proto grows.
Taking the scratch buffer lazily, only once a chunk is actually dropped,
also takes the pool out of the path for entries that have none. Those
were measurably slower than the full decoder before; they are now level
with it. Each chunk's length prefix is parsed once rather than twice.
chunks full attrs-only vs base
0 171.4n 176.9n ~ (p=0.670)
1 366.6n 259.4n -29.24%
4 1.034u 500.2n -51.60%
16 3.905u 1.464u -62.52%
64 13.48u 5.195u -61.46%
* filer: carry the size before dropping chunks over the wire
Dropping the chunk list assumed every store folds the chunk extents into
FileSize when it decodes. A store that keeps entries as JSON rather than
as an encoded Entry never re-derives it, so an object written with a zero
FileSize kept its real size only in the chunks, and stripping them left
the client reading the file as empty. Stamp the size into the attributes
first, which costs nothing and does not depend on how the store loaded
the entry.
* mount: test that the readdir context reaches the store decode
Everything else exercises the decoder directly, so a refactor that
stopped threading the context would have reverted the whole thing with
every test still passing.
The benchmark's chunks also carried a constant legacy FileId, which
BeforeEntrySerialization reparses over Fid on the way in, so all 200k
entries stored one byte-identical chunk rather than the varying fixture
it looked like.
|
||
|
|
d1f503181b |
[s3] force filer apply s3 expiry metadata (#10469)
* fix: apply S3 Expiry Metadata * add test Header X-Seaweedfs-Expires-S3 * resolve comments * test entry lookup by mtime * filer: skip s3 expiry stamp on versioned entries The s3 expiry path skips entries carrying a version id, so stamping one takes away its expiry rather than moving it onto mtime. Files under .versions/ are written once, so crtime already tracks their needles. --------- Co-authored-by: Konstantin Lebedev <whitefox@mayflower.work> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
f46b2a1925 |
Stop the filer test helpers from pinning gigabytes of log buffers (#10560)
* log buffer: wake the interval loop on shutdown instead of sleeping through it loopInterval parked in time.Sleep(flushInterval) and only re-checked IsStopping when it woke, so a buffer shut down early kept both loop goroutines - and the PreviousBufferCount+1 slabs of BufferSize they reach - alive for up to a full interval afterwards. Select on shutdownCh against a ticker instead, and give the loops a WaitGroup so a test can observe that they exit. * test: release the filers the server tests build Every helper here left its filer's meta log buffer running, so each test pinned PreviousBufferCount+1 buffers of BufferSize for the rest of the run: ~3.5GB of live heap across the package, which overruns the address space on linux/386 and kills the 32-bit job with an out-of-memory throw. Thread the test through the helpers so the buffer is shut down on cleanup, and shut the subscribe harness's filer down outright - its deletion loop keeps the whole filer reachable otherwise. That harness quiesces its flush path first, since Filer.Shutdown closes the store a flush still in flight would write through. |
||
|
|
a9de90ae29 |
test: wait for every queued flush before deleting the log files (#10546)
TestSubscribeLoop_FlushProvenGapSkipsToRetained deleted the log files once the eviction watermark's own flush had landed, while the windows sealed after it were still queued. Those flushes then wrote their files back, and the subscriber served them from disk instead of taking the gap-skip path, so the windows whose files really were gone came out missing. Wait through the last sealed window instead. |
||
|
|
813a4b0711 |
fix(filer): stop skipping recent unflushed events on metadata subscription gaps (#10501)
* fix(filer): don't skip unflushed events on metadata subscription gaps A subscriber that falls behind the in-memory log ring during a write burst could have its read position jumped past events that were evicted but not yet flushed to disk — silently lost for filer.backup/filer.sync/ mount subscribers. Route all three gap-skip sites through resolveDiskGapResume: only skip past windows older than a settled horizon (2*LogFlushInterval); recent gaps wait for the flush and re-read disk. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(filer): harden metadata gap-skip guard Address review findings on the settled-horizon guard: - local subscriptions: gate the skip on the buffer's flush watermark observed before the disk read (resolveLocalGapResume) — a disk miss is then proof the gap is empty, with no wall-clock assumptions - aggregated subscriptions: cap skips strictly below the horizon boundary (persisted reads exclude ts <= cursor) and pace capped advances, so the sliding horizon cannot cause disk-probe spinning - replace unbounded sync.Cond waits with a bounded select on the buffer's subscriber channel + retry timer + ctx cancellation, eliminating the lost-wakeup stall Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(filer): close remaining gap-skip loss paths from re-review - log_buffer: a sentinel-offset (time-based) read below the earliest in-memory entry silently started at earliest, skipping a window that may hold evicted-but-unflushed events. Track the ring's eviction watermark (lastEvictedTsNs) and keep the inclusive fast path only when nothing at/after the position was ever evicted; otherwise return ResumeFromDiskError so the subscription gap guard decides. - capped horizon advances stay on the disk-probe path (never expose a mid-gap position to the memory read) and keep pacing - gap jumps land just below earliest: positions are exclusive, so the earliest entry itself is still delivered - subscriber notification keys include clientId/epoch so a replacement stream never inherits a channel the old stream's cleanup closes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(log_buffer): generalize the eviction-watermark read gate Third-review findings: - epoch/zero-time reads bypassed the eviction watermark: gate them the same way, so a SinceNs=0 subscriber cannot silently start at the earliest retained entry after an unflushed window was evicted - apply the watermark gate regardless of the cursor's batch offset (batch offsets carry no meaning for time-based reads); this also serves adjacent cursors (earliest == current+1) from memory instead of stalling them in the gap loop - LoopProcessLogData reader names include clientId/epoch, since they are registered as subscriber keys internally (same collision as the outer notification keys) - test: pin the gate (below/at watermark, epoch-after-eviction) and update the slow-consumer test to the sharper contract — complete in-memory history is served from memory; disk only once evicted Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(log_buffer): watermark equality is unsafe for inclusive sentinel cursors Sentinel (Offset <= 0) time-based cursors search from ts-1ns, i.e. they read inclusively of their own timestamp — and the evicted window may end exactly at that timestamp. Allow watermark equality only for exclusive (positive-offset) cursors; sentinel cursors must be strictly above it. Also shut down the test buffer and pin the equality cases. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(filer): wait on an empty aggregated buffer instead of falling through When a disk read finds nothing for a ResumeFromDiskError gap and the aggregated buffer has no readable entries (zero earliest time), resolveDiskGapResume declines to advance and control fell through to the in-memory read — which returns ResumeFromDiskError again immediately, spinning through the probe cycle without any wait. Fold the case into the existing recent-gap branch so it waits (notification, cancellation, or the retry interval) before re-probing, matching the local subscription path, which already waits unconditionally. * fix(filer): serve the exact eviction-boundary entry without another flush wait When the flush watermark has passed the earliest in-memory entry but that entry sits exactly one nanosecond above the cursor, the exclusive resume target collapses onto the cursor and resolveLocalGapResume declines to advance — while a sentinel cursor at the eviction watermark keeps deferring the in-memory read to disk. Progress then depends on the next flush cycle. Re-arm the cursor with a positive (exclusive) offset instead: ReadFromBuffer explicitly allows positive-offset cursors at the watermark, so the boundary entry is served from memory immediately. Also promote the aggregated path's horizon-capped gap skip to a warning: that skip may pass events a stalled flush lands later (the aggregated ring has no flush watermark to gate on), so operators should see when a flush stall outlasts the settled horizon. * fix(filer): resume a gap skip with an exclusive cursor The resume position landed one nanosecond below the earliest in-memory entry but kept the inclusive sentinel offset. When that timestamp is also the eviction watermark, the read gate answers ResumeFromDiskError for an inclusive cursor, the disk read finds nothing, and the resume target collapses onto the cursor, so neither helper can advance: the subscriber parks on timed waits forever. A skip is only taken once the gap is proven empty, so nothing remains to deliver at the resume timestamp. Resume exclusively instead, and let an inclusive cursor already sitting on the target count as progress. * fix(filer): gate aggregated gap skips on eviction, not wall clock Wall-clock age never proved persistence. The aggregated ring has no flush watermark because peers persist their own local logs, so a horizon of 2 * LogFlushInterval was standing in for one. While the disk stayed empty the horizon kept sliding forward and walked the cursor past evicted-but-unflushed events one window at a time - the volume-outage stall this change set exists to survive. The ring does carry a real proof: its eviction watermark. If nothing at or after the cursor was ever dropped, memory still holds every entry after it and the gap is provably empty. Below the watermark entries were dropped and only the producing peer's flush can supply them, so wait for that flush instead of advancing. The skip that remains is the one the ring can prove, which keeps the infinite-loop guard for a genuinely empty gap. * perf(filer): stop re-probing disk on every metadata append A subscriber parked on a gap woke on the log buffer's subscriber channel, which fires on every append. On the aggregated path each wake re-ran a full ReadPersistedLogBuffer - a ListDirectoryEntries plus a readahead goroutine - so one parked subscriber turned every cluster-wide metadata event into a store query, exactly while the cluster is already struggling with the flush stall that parked it. An append also cannot settle the gap: what this waits on is a peer persisting its own log, which nothing local signals. Wait on the timer alone there. The local buffer does signal its flush on that channel, so keep it, but drain the stale token first: otherwise the appends riding the same channel spin the wait at write rate. The retry interval covers the notification the drain discards. * fix(filer): surface a metadata subscriber parked on a gap Refusing to skip an unresolved gap trades silent loss for a silent stall, and a stall is no easier to diagnose: filer.sync and mount followers just stop advancing, with no error on either end. The only trace was a V(3) line nobody runs with. Warn on entry and once a minute after, and carry a subscribe_gap_stalled gauge for the duration, so a flush that never lands shows up as a stalled subscriber rather than a consumer that mysteriously went quiet. * refactor(filer): drop the metadata listener cond with no waiters left Both listenersCond.Wait() sites are gone, so listenersWaits never leaves zero, the guard in the filer's notify callback never fires, and the two Broadcast calls around client registration wake nobody. Remove the cond, its counter and its lock, and pass a nil notify func: the log buffer already skips a nil one, and the subscriber channels carry these wakeups now. * test(log_buffer): drive the eviction watermark through a real seal Both eviction tests wrote lastEvictedTsNs directly, leaving the one line that sets it uncovered: copyToFlushInternal has to read slot 0 before SealBuffer shifts it out, and moving that read one statement later still passed every test. Append past the ring instead, and check the watermark appears only on the seal that drops a window, matches that window's stop, and then advances. The buffer under test never flushes, matching the aggregated meta ring where the watermark is the only emptiness proof. * refactor(filer): build the subscriber reader name once per stream It was rebuilt on every loop iteration from values that cannot change for the life of the stream. Hoist it next to the notification key it mirrors. * docs(filer): tighten the gap-handling comments Several ran to eight lines restating the same reasoning at each site. Keep the non-obvious why, drop the retelling. * fix(filer): mark a confirmed disk position as an exclusive cursor A disk read returns the timestamp of the last entry it handed to the subscriber, but the cursor built from it stayed inclusive. Land that cursor exactly on the eviction watermark and the read gate sends it back to disk for an entry the disk just delivered; the re-read finds nothing, neither emptiness proof holds for an inclusive cursor there, and the subscriber parks - until some later flush, or forever while one stays stalled. Everything after the watermark was sitting in memory the whole time. Carry the offset instead, so it also covers the ring evicting onto the cursor after the read rather than before it. The constant is no longer gap-specific, so it is now named for what it asserts. * perf(filer): re-park a gap wait woken by an ordinary append Draining one stale token did not bound anything: the subscriber channel carries an append per metadata write, so under continuous writes the next one satisfies the wait immediately. During the write burst these stalls come from, each parked subscriber ran the recovery loop at write rate rather than the intended two-second cadence. Only a flush can settle a gap, so check the flush watermark on wake-up and go back to waiting if an append is all that arrived. The retry timer is created once, so re-parking does not extend the interval. * fix(filer): keep a disk-derived cursor inclusive Marking every disk position exclusive assumed the entry at that timestamp was the only one there. On the aggregated stream it is not: disk can deliver one filer's persisted event at T while another filer's event at the same T is still unflushed in the aggregate ring. The exclusive cursor skipped it, and since the persisted reader also excludes ts <= its start, nothing would ever bring it back. Take the weaker guarantee instead. The reason the exclusive cursor was introduced - an inclusive one landing on the eviction watermark parks forever - is better answered in the resolvers: at the watermark memory holds nothing (retained windows start strictly after it) and the persisted reader cannot return that entry at any later time either, so refusing the gap buys nothing and never ends. Skip it whatever the cursor's inclusivity. That proof does not depend on a flush, so the local resolver takes it as a second, independent disjunct alongside its flush watermark. * perf(filer): wake a parked gap wait on flushes only Re-parking on an append bounded the work but not the wake-ups: the subscriber channel carries one per metadata write, so a parked subscriber still took a scheduling round-trip per write. Worse, draining it kept the channel empty, so every writer's non-blocking notify succeeded instead of falling through - the burst paid for the wake-ups too. Give the log buffer a flush-only subscriber list, notified from loopFlush, and park on that. The append channel now fills once and stays full, which is exactly the state the non-blocking send is designed for. * fix(filer): stop the gap-stall gauge from leaking a series per connection clientName is req.ClientName + "@" + peer address, so it carries the client's ephemeral source port and changes on every reconnect. Labelling the gauge with it minted a new series per connection, and clearing it only ever Set(0), so nothing was ever released - a client in a reconnect loop grows the filer's metric map and /metrics payload without bound. Key it on the stable client-supplied name, as the neighbouring subscribe gauge already does, and delete the series on teardown. Two logging fixes ride along, both in the same reporter: the resume warning was unpaced while the park warning throttles to one a minute, so a burst that parks and resumes every couple of seconds warned on every cycle; and clear() doubled as the teardown path, announcing "resumed after 14m0s parked" for a client that actually gave up and disconnected still behind. * fix(filer): give a parked gap wait the exits the read loop has A park never re-enters the read loop, so every exit that loop relies on stopped working while a subscriber was parked. It kept scanning the filer store every 2s for a client that a higher-epoch reconnect had already superseded, until the TCP connection finally died - hours, on a half-open one. It never reached the only code that honors UntilNs, so bounded callers like `weed shell fs.verify` and `filer.meta.tail -until=` hung instead of exiting. And a notification channel closed out from under it turned the bounded wait into a spin, since a receive on a closed channel returns instantly. Check all three where the subscriber actually waits. Bound the wait too: waiting is productive while the window is still queued for flush, but a peer that never returns, or whose filer store this filer cannot read at all, makes it permanent - and a subscriber that silently stops delivering is no better than one that silently skips. Fail the stream after that instead of hanging, loudly enough to say which gap and for how long. * fix(log_buffer): keep the eviction gate out of the shared read path The gate belonged to the filer's subscribe loops but was installed in ReadFromBuffer, which the message queue shares and which has no gap handling of its own. Three MQ paths broke on it: GetUnflushedMessages asks for everything in memory past the flush watermark and got ResumeFromDiskError instead, so SQL results silently dropped unflushed rows; a RESET_TO_EARLIEST consumer's epoch cursor was sent to disk, and the MQ disk reader resets an empty read back to epoch, so the whole partition replayed on every pass; and a disk cursor landing on the watermark carried offset -2, which the gate refused and the disk reader's ts <= start filter also skips, so neither side could ever serve it. The rewrite also dropped the old Offset <= 0 requirement, letting a stale positive-offset cursor jump to memory over on-disk history, and refused any negative-timestamp cursor even with nothing evicted. Restore the read path exactly as it was and keep only lastEvictedTsNs, which is the useful primitive. The filer loops now consult the watermark themselves before reading memory, which is where the gap handling that makes the refusal actionable already lives - and they check it every pass, not just when the disk came up empty, since a disk read can leave the cursor short of the watermark too. * fix(filer): read the log file whose window spans past its own name A log file is named for the start of the window it holds, but a window runs up to a flush interval longer, so "12-30" can hold entries through 12:31:58. File selection compared the cursor's minute against that name and skipped anything sorting earlier, so a subscriber resuming at 12:31:10 never saw the rest of that file: the read reported nothing on disk while the entries sat in it. That was survivable when a miss only meant "wait and retry", but the gap resolvers now read a miss as proof the range is empty and move the cursor past it, which turns those entries into silent loss. Start the file scan a flush interval early; entries are still filtered against the exact cursor, so this only opens one more file and never re-delivers. * fix(filer): resume a gap with a cursor the memory read will serve Reverting the eviction gate put ReadFromBuffer back to refusing every positive-offset cursor below the in-memory window, but the resolvers still handed back one - earliest-1 marked exclusive. So the resume bounced straight to ResumeFromDiskError, the disk had nothing, and the resolver saw its own cursor as no progress and parked: a subscriber stalled, and after the new bound failed outright, with the whole gap sitting in the ring the entire time. The exclusive marking only existed to dodge a park the gate itself caused, and the gate is gone. Resume at earliest with the sentinel offset, which is the position master used and which case 2.1 reads inclusively. That makes the resume unconditionally ahead of the cursor, so the progress check it needed goes away with it. * fix(filer): stop dropping a log file whose window outruns its name Listing the earlier file was not enough: the iterator then decided whether to read it by comparing the *following* file's name against the cursor, which treats that name as an upper bound on this file's contents. It is not one. A file is named for the start of the window it holds, minute-truncated, and the window runs up to a flush interval longer, so "12-30" can hold an event at 12:31:20 while "12-31" sits right after it - and a cursor at 12:31:10 skipped straight past the event. Bound the decision on the file itself: skip it only when its name plus the minute truncation plus a flush interval still lands at or before the cursor. That also covers the last file in the queue, which the old check never skipped because it had no successor to compare against. * fix(filer): stop a chunk-ref read from rewinding the subscriber CollectLogFileRefs reports the minute-level name of the last file it shipped, and the caller assigns that straight to the read position. Since the scan now reaches back a flush interval to catch a spanning file, a request at 12:31:10 that picks up the 12-30 ref moved the cursor to 12:30:00 - so the memory read that followed replayed events older than the client's own SinceNs and re-sent what the chunk reader had already been handed. The same rewind was reachable before, within a minute, whenever the last file's name sorted behind the request. Clamp the reported position to the one that was asked for. It still under-advances by design, since the server never reads the entries it ships refs for and cannot know where they end. * fix(filer): resume below earliest so a one-entry window is not skipped Moving the resume onto earliest itself was wrong for the smallest window there is. A sealed window holding a single entry has startTime == stopTime, and the sealed-buffer lookup only enters a window whose stopTime is strictly after the cursor, so a cursor sitting exactly on earliest walks past it and its sole event is never delivered. Low-volume metadata windows are routinely one entry, which is precisely when losing it is hardest to notice. Go back to one nanosecond below, which takes the startTime.After branch and returns the whole window, and keep the sentinel offset the memory read requires. The earlier test used an active multi-entry buffer, where both cursors happen to work; the new one seals single-entry windows and compares which entry comes back. * fix(filer): count a persisted read as progress only when it moves A chunk-ref read reports the minute-level name of the last file it shipped, now clamped so it never rewinds, so it comes back non-zero even when it names the position the subscriber already held. The loops read non-zero as progress: they cleared the stall timer, then found the cursor still short of the eviction watermark and parked again. Every retry re-shipped the same refs and reset the timer, so the bound that is supposed to end an unrecoverable stall was never reached - and a chunk-capable client buffers those refs waiting for an event that never comes, so it just accumulates duplicates. Require the reported position to be strictly ahead of the cursor. * fix(filer): count the evicted ranges the aggregated stream cannot prove The eviction watermark belongs to the merged ring, but the disk it gets checked against is the union of every peer's own log and each peer flushes on its own schedule. A read that lifts the cursor from below the watermark to above it may have done so entirely on a peer that is already ahead, while a lagging peer still holds unflushed events inside the range just crossed; when it flushes them they sit behind the cursor and are never delivered. One aggregate maximum is not proof that every peer persisted the range. Nothing available locally separates that from the ordinary case where every peer had in fact persisted it: the aggregator tracks peers by address while log files carry a random per-filer id, so "has this peer flushed through T" cannot be answered here at all. Deciding it needs the source filer's own flush watermark carried on the subscribe stream, which is a wire change this does not make. Count and log the crossing so the window is at least measurable instead of invisible. * fix(filer): send log file refs through the pipelined sender sendLogFileRefs wrote on the raw gRPC stream while pipelinedSender's goroutine concurrently calls Send for queued memory events - two senders on one stream, which gRPC forbids. The window used to open once per fall-behind; the gap retry loop now reopens it every pass. Routing refs through the sender also restores ordering: refs used to overtake up to 1024 queued events, and the client treats any non-ref message as the signal to process buffered refs, so overtaken refs were applied against the wrong position. The reason refs bypassed the sender was the batcher: their TsNs of 0 reads as far behind, and the client recognizes refs by the top-level field alone - a refs envelope would drop its Events tail, and refs inside Events would be applied as an empty event. Teach the batcher instead: refs messages always go solo, and one drained mid-batch is sent solo right after that batch. * fix(filer): count parked subscribers instead of flagging them by name The per-client gauge series could not work. Its label was rebuilt from the peer address at first, which leaks a series per reconnect; keyed on the client-supplied name instead, it collides - every mount registers as "mount" - so one stream's teardown deleted a parked sibling's live series, and the sibling never re-created it because its own park state said the gauge was already set. Either way the alert this gauge exists to drive goes dark. A count needs no identity: Inc on park, Dec on resume or teardown, scope as the only label. Client details stay in the logs. Also start warning only once a stall has outlived the warn interval. park() warned immediately on every first park, so catch-up churn that parks and resumes every couple of seconds logged a warning pair per cycle - exactly the flood the pacing was supposed to prevent, burying the long-stall warnings that matter. * fix(filer): one gap resolver, and re-arm an unservable adjacent cursor The two resolvers were the same function - the aggregated one is the local one with a flush watermark of zero, since its ring never flushes - duplicated down to the comment justifying the resume target. A fix applied to one and not the other is how the two streams drift; merge them. The merge carries the one behavioral fix both copies needed. Timestamp collision bumps make adjacent entries exactly 1ns apart, so an entry ending an evicted window leaves the cursor exactly one below earliest with a positive batch offset. The resume target then equals the cursor and both copies refused it as no progress - but that cursor cannot be served (ReadFromBuffer refuses positive offsets below the window) while the sentinel resume at the same timestamp is, and both deliver exactly the entries after it. Refusing parked a subscriber whose data was entirely in memory: until the next flush locally, and through a 15-minute stall failure on the aggregated path. Advance on equal target when the held cursor is exclusive; a sentinel there is already served, so it still refuses. * fix(filer): a bounded subscription parked exactly on UntilNs is finished The park's UntilNs check was strict while the bound is inclusive and cursors are exclusive: a disk read whose last entry sits exactly on UntilNs leaves the cursor there with everything up to the bound already delivered. If the next range was an unprovable gap, the completed subscription parked anyway and eventually failed - fs.verify hanging and then erroring on a healthy cluster. * fix(filer): re-derive the subscribe loops as one state machine The loops had grown three generations of gap handling - a post-disk-read branch, a post-memory-read branch, and an every-pass guard bolted on in front of the memory read - each consulting state the others mutated. The worst interaction wedged the aggregated stream permanently: diskExhausted compared the disk result against the cursor that result had just updated, and paired it with a ResumeFromDiskError latch that only a resolver advance cleared, so one fall-behind sent every later pass into the gap block and LoopProcessLogData never ran again. Mounts kept a healthy-looking stream and applied nothing. Both loops now run the same derived sequence. One disk pass; progress is the pre-update cursor against the result, freshly each pass. One gap decision before the memory read: a cursor the ring evicted past either keeps draining the disk (it just advanced), resolves forward (the gap is proven empty), or parks - and a cursor memory refused with nothing evicted after it re-arms onto the retained window. The stale-latch branch is gone: the error is only consulted on a pass whose own disk read came up empty. All four parks go through one parkOnGap helper, so the exits live in one place. The eviction watermark is read after the disk read and the same value feeds both the guard and the unproven-crossing report, which previously compared against a snapshot taken before a potentially minutes-long backlog read and missed evictions landing during it. The next-day jump now clears the stall reporter - it used to leave a stale park epoch that could kill the next brief park instantly at the 15-minute bound - and counts its own watermark crossing. Chunk-ref reads get two rules the old loops lacked. Refs are sent once per position: retries re-sent identical batches every two seconds into a client that only drains them on a non-ref message, growing an unbounded pending list. And when refs cannot advance the cursor - they report file start minutes, which sit below the content the client was actually given, pinning the cursor under the watermark forever during bursts - the pass falls through to entry reads, which move the cursor by real timestamps and double as the client's drain signal. * fix(filer): give up on an unprovable gap instead of failing the stream Failing after maxGapStall assumed the client could do something better, but every consumer just reconnects at the same SinceNs and hits the same wall, so an unprovable gap - a dead peer, or a peer whose filer store this filer cannot read at all - turned into a permanent 15-minute fail/reconnect loop delivering nothing. Master handled the same state by skipping instantly and silently. Take the middle: wait the full bound, then abandon the gap and resume at the eviction watermark, where everything retained starts strictly after, so the loss is exactly the range that could not be proven. The skip shares the unproven-crossing counter and logs at error level - loss is bounded, recorded, and the stream keeps working. A stall with nothing evicted past the cursor loses nothing by waiting, so it restarts the clock and keeps parking rather than skipping. * test(filer): pin the file-skip bound through the production predicate The spanning-file test asserted against its own copy of the arithmetic, so a regression in the iterator - restoring the next-file-name comparison or dropping the flush-interval term - would keep CI green while re-introducing the silent loss the fix closed. Extract the bound into logFileMayContainAfter, call it from the iterator, and point the test at it; breaking the production expression now fails the test. * test(log_buffer): pin the flush-subscriber contract The registry the filer's gap parks wait on had no test at all. Cover the observable contract: an append never wakes a flush subscriber, a flush does with the watermark already stored, unregistering closes the channel so an abandoned waiter unblocks, and double or unknown unregisters are harmless. The store-before-notify ordering in loopFlush is what makes the parks' wake-up re-check sound, and it is not black-box testable - reordering leaves a same-goroutine window of nanoseconds that hundreds of tight round-trips never catch. Mark it load-bearing at the site instead; a reorder now at least has to argue with the comment it deletes. * fix(filer): a -1 SinceNs is a position, not the refs-gate sentinel The once-per-position refs gate used -1 as "never sent", but a client may legally subscribe with SinceNs=-1, whose cursor timestamp is exactly -1: the very first pass then believed refs were already sent there and fell back to streaming the whole persisted history entry by entry - the bootstrap load chunks mode exists to avoid. Use MinInt64, which no cursor can carry. * refactor(filer): drop the aggregator's listener cond with no waiters left Same shape as the FilerServer cond already removed: nothing increments ListenersWaits and nothing ever calls Wait, so the three Broadcasts wake nobody and the notify callback's guard is always false. Aggregated subscribers wake through the buffer's subscriber channels now. * fix(filer): make the gap metrics say what they count The crossing counter's help text described only the aggregated peer case, but give-ups increment it for local stalls too - a wedged local flush - which sends an operator chasing peer replication when the problem is the local store. Label it by scope and say both. The stalled gauge counted every park, including waits with nothing evicted and nothing at risk, while its help text promised evicted-but-unpersisted events; describe it as what it is, a count of subscribers parked on a gap. * test(filer): make the flush and stall tests assert what they claim The flush-subscriber rounds were vacuous: the probe entry sat above the round timestamps, so every round entry was collision-bumped and the stored watermark exceeded the local value each assertion compared against - the same bump mistake this test suite already made once. Put the probe below the rounds and guard each round against bumping, so a vacuous setup fails instead of passing. The stall-outcome test wrote the reporter's park epoch directly, bypassing the gauge Inc that gaveUp() later Decs - leaving the shared process gauge at -1 for every test that runs after it. Park through the real path, age the park by hand, release what the test holds, and assert the gauge lands back where it started. * fix(filer): finish the stream checks before marking it parked parkOnGap stamped the reporter before waitOnGap ran its instant done exits, so a bounded subscription completing inside a gap state was marked parked for the microsecond before done fired - a phantom gauge blip and a false "disconnected still behind" warning on every healthy completion. Fold waitOnGap into parkOnGap so the done exits run first and the park mark only ever covers a stream that actually waits. While the park owns its timer, back the retry off as the stall ages - 2s probes growing toward one a minute - since every retry re-reads the persisted log, and probing the store each 2s for 15 minutes per parked subscriber during the very outage that parked them makes the bad time worse. The subtest still named for the old fail-the-stream stall behavior goes with the merge. * fix(log_buffer): gate filer cursors against eviction under the read lock The subscribe loops checked the eviction watermark and then read memory, but a seal can land between the two: the read then served a sentinel cursor from the earliest retained window, silently skipping the window just evicted - the loss class this PR exists to make loud, surviving as a race. The only place the check is atomic with the serve decision is inside ReadFromBuffer, under the lock seals take to evict. Rather than put the policy back into the shared read path - which broke four message-queue readers last time - add a new sentinel offset that opts into it: EvictionGatedOffset reads inclusively exactly like -2, except below the watermark it is refused to disk. The filer loops stamp it on every cursor they hand the memory read; a refusal lands in the same gap machinery the loop-side check feeds, so the race collapses into the handled path. MQ cursors never carry it and keep master behavior byte-for-byte. * fix(filer): gate the aggregated gap on received, not bumped, timestamps The aggregated ring rewrites an out-of-order arrival to its head plus a nanosecond, so after any bump-heavy interval - a peer history replay following a restart is enough - its eviction watermark lives above every timestamp that exists on any peer's disk. Comparing a disk cursor against it parked subscribers that had in fact drained every peer's log: a 15-minute delivery freeze ending in a give-up skip and a false loss alarm, on a healthy cluster where master resumed instantly. Track a second watermark in the received timestamp space - the highest pre-bump timestamp among evicted entries - and gate the aggregated loop on that. Disk cursors and received timestamps are the same space, so the comparison means what it says: at or past it, every evicted entry's original was at or below the cursor, and everything flushed of them was already delivered. The bumped watermark keeps guarding the in-ring read gate, whose cursors live in ring space. The local buffer is untouched: it flushes its own bumped timestamps, so there the two spaces are one. * fix(filer): ship each log chunk once and stop echoing ref'd files inline Chunk mode duplicated data through two doors. Consecutive ref collections overlap by design - the scan backs off a flush interval to catch a spanning file, and a filer appends chunks to its newest file - so the same file was shipped again on every pass that re-listed it: the client re-downloaded its chunks, and a duplicated file mid-batch rewinds timestamps inside the client's per-filer merge, which reads each stream as sorted - transiently resurrecting deleted entries during catch-up. Track per subscription how many chunks of each file were shipped and send only the unsent suffix; state prunes with the scan window, so it holds a few files per filer. The second door was the entry fallback: when refs cannot advance the minute-named cursor, the pass streamed the ref'd file's tail inline, and the client applies inline events unfiltered - the same tail it already applied from chunks. Entry passes for chunk clients now advance the cursor without delivering; everything they skip is covered by the refs already sent or the deltas the next collection ships. * refactor(filer): one gap decision shared by both subscribe loops The post-disk gap tree - guard, drain, resolve, two parks, the re-arm - existed twice, differing only in buffer, watermark space, flush getter, park channels, and reason strings. Four rounds of review fixes have shown the copies drift the moment one is edited alone. gapPass now carries the five differences and the tree lives once; the loops shrink to a three-way switch between reading memory, restarting the pass, and ending the stream. * docs(filer): trim the gap-machinery comments to the why Several blocks had grown to ten-plus lines restating what the tests already pin or retelling one rationale at multiple sites. Keep the non-obvious why - the load-bearing flush ordering, the two timestamp spaces, the refusal-at-equality argument - in a few lines each. * fix(log_buffer): credit an entry's received timestamp to its own window The received-ts capture ran before the rollover check, so an append that sealed the previous window stamped its timestamp onto that window and then lost it in the reset of the new one. The eviction watermark this feeds broke both ways: the sealed window's value was inflated by an entry it does not contain - parking aggregated subscribers on gaps that were drained - and the entry's real window was deflated, proving gaps empty that still held its event on some peer's unflushed path. Credit the timestamp only after the entry lands, when its window is known. * fix(filer): rebase a shipped chunk suffix to logical offset zero A grown file's delta kept the chunks' original file offsets, but the client's chunk reader starts at logical zero and a list opening higher reads as instant EOF - a successfully empty replay, and since chunk clients no longer receive disk entries inline, the appended events were silently dropped. Clone the suffix chunks with offsets rebased to zero; the cut is record-aligned because each append is one uploaded chunk of whole entries, so the suffix decodes as a file of its own. * fix(filer): finish every chunk refs batch with a transition the client acts on Both chunk consumers buffer refs until a non-ref message arrives, so a source with historical logs and a quiet ring - a mount reconnecting after a filer restart is the common case - shipped its backlog and then went silent: the client sat on the refs until the next metadata mutation anywhere in the cluster. The disk step now ends every batch with the empty-notification marker the client already treats as a resume-cursor advance. The same step closes the inline replay: the cursor used to stay at the last file's minute name, so the memory read re-delivered the retained tail of a file the client had just read via chunks - T1..Tn applied twice. The advance-only entry read now runs on every chunk pass, moving the cursor to the true disk content end before memory is consulted. Ordering inside the pass is load-bearing: the entry read can outrun the shipped refs by a chunk appended between collection and read, and the transition timestamp becomes the client's refs filter - stamping it past unshipped content would silently drop that chunk's events on the next delta. The pass therefore re-ships the delta after the entry read, so the transition never exceeds shipped content. Bump-displaced aggregated entries can still arrive inline above the cursor with originals below it; that duplication is bounded and stays within the documented at-least-once residual. * fix(filer): prune ref state at the minute the scan actually stops at The collector compares file names at minute granularity while the prune used the exact-nanosecond scan bound, so for a cursor at 12:31:20 the 12-30 file was still collected but its sent state was already deleted - the next pass reshipped the whole file, re-creating the duplicate-refs class the state exists to prevent. Truncate the bound to the minute the file names live in. * fix(filer): derive the chunk cursor from the shipped refs themselves The advance-only entry read left the three positions that must agree in each other's blind spots. Its snapshot could trail the second delta's, so a chunk appended between them shipped events newer than the cursor and the memory pass sent them again. And it made the filer decode the tail range on every pass, serialized ahead of the client's own reads by the transition marker - re-introducing a slice of the replay work chunk mode exists to offload. Compute the cursor from the shipped set instead: the final entry timestamp of each filer's last shipped chunk, decoded once through the shared chunk cache. Refs coverage, transition marker, and memory start are then the same number by construction - nothing is decoded twice, nothing is dropped, and the per-pass server cost falls to one cached chunk decode per filer. The second delta and the once-per-position refs gate existed to patch the entry read's snapshot races, so both go with it; the range read survives only as a fallback for legacy chunks that do not decode standalone. * fix(filer): keep the chunk-cursor probe inside the shipped snapshot Three holes in the tail probe, all variations of stepping outside what was shipped. A permanently missing chunk failed the stream before the transition marker, so the client discarded its pending refs and reconnected to the same failure forever - blocking all later metadata behind one dead volume, where every other replay path (including the client's own reader) skips such chunks; the probe now walks back to the last readable chunk, and a filer with nothing readable simply contributes no cursor. The legacy fallback re-listed the logs after the refs were collected, so a concurrent append could push the range end over an unshipped chunk and the marker past events the client never received; it now streams the shipped chunk list itself, so no snapshot other than the shipped one is ever consulted. And a file selected before UntilNs can hold entries past it, which the client filters while still adopting the marker as its checkpoint - a later bounded request then skipped them; the marker is clamped to the bound. * fix(filer): make the cursor probe an exact mirror of the client's reader The probe answered from the server's view of the chunks; the marker's correctness depends on the client's. Its backward walk found the last readable chunk, but the client reads forward and stops at the first unreadable one, never resuming within a file - for readable, missing, readable the marker claimed the suffix the client never applied, losing those entries permanently. Keeping only each filer's final file ref discarded the progress of earlier readable files when that file was wholly missing, rewinding the marker to the start cursor. And a torn trailing size prefix - what a crashed writer leaves - failed the probe where the client reads a clean end, blocking the marker forever on data the client accepts. The probe is now shaped like the reader it answers for: per file the readable prefix, per filer the newest file with content, and no condition escapes as an error - understating the marker only re-ships, overstating loses events, and a probe failure must never block the transition the client is waiting on. Each rule is pinned by a test that fails against the previous shape. * fix(filer): judge chunk readability at the volumes, not the decode cache Two ways the probe's answer could drift from what the client experiences. A chunk this server decoded earlier stays warm in the shared cache after its volume dies, so the probe sailed past a chunk the direct-reading client stops at - marker beyond the unread suffix, entries lost. Every chunk now passes a volume lookup before the cache is consulted; the lookup rides the master client's in-memory map, so the probe stays cheap. And a probe stop was treated as harmless understatement, but the delta had already marked the whole ref sent: a transient server-side failure left the cursor stranded behind shipped content for the life of the connection, parking aggregated streams below the watermark for data the client already holds. The pass now rolls back the sent state of every ref above the file that answered the probe, so unreached refs re-ship and re-probe until the cursor gets there. Re-shipped entries at or below the client's checkpoint are filtered client-side, and batches are marker-separated, so a re-shipped file cannot rewind a merge mid-batch. * test(filer): end-to-end subscribe-loop harness and wire-contract tests Every escaped bug across this change's review rounds lived in an interaction the unit tests could not see: the loop state machine, the disk/memory handoff, or the server/client contract. The harness runs the real SubscribeLocalMetadata loop against a real leveldb-backed filer, faking only the volume layer behind the existing test hooks, and asserts the delivered stream itself. Eight scenarios, each pinning a class this change was reviewed for: the headline evicted-unflushed gap parks and then delivers in full; a ring that evicted nothing serves memory promptly; a backlog-to-live handoff with 1ms-adjacent timestamps across every boundary delivers exactly once; a flush-proven gap over vacuumed log files skips to the retained ring including a single-entry window; a bounded subscription terminates at its bound; a permanently wedged flush ends in the give-up skip with the stream still alive; and chunk mode is checked against the real client code - pb.ReadLogFileRefs applied to the shipped refs must cover everything the transition marker claims, with and without a dead volume in the middle. Validated by re-introducing three fixed bugs: the missing eviction guard delivers during the unproven gap, a 2ms cursor error at the handoff drops exactly one event, and resuming at rather than below the earliest retained window loses a single-entry window's sole event - each caught by the scenario built for it. The gap timing knobs become vars so parks run at test speed, a small filer hook swaps the volume-touching read functions, and a sender test pins the refs wire rules the client depends on: never batched, never an envelope, everything in order. * fix(filer): re-ship a partially read answering file, pin the probe's limits The sent-state rollback stopped at files newer than the one that answered the probe. When the answering file itself was only prefix-readable - a dead or transient chunk mid-file - its unread suffix stayed marked sent, and the next append advanced the cursor past it for good. The probe now reports whether the answering file was read through to its end, and a prefix-limited answer re-ships that file too; a torn tail counts as complete, since the client's read ends there as well. The rollback rules live in one predicate with a table test - files below a complete answer stay sent, because the client has moved past them and re-shipping cannot rewind its filter. Two test honesty fixes ride along. The loop harness derived its timestamp base from time.Now() per call, so expectations recomputed across a second boundary drifted by exactly one second; the base is now fixed per harness. And the probe's liveness boundary is pinned as a test instead of a comment: a volume lookup cannot see a dead needle or a stale location inside a resolvable volume, so a warm cache can answer past a chunk the client fails on - accepted because metadata log chunks die volume-at-a-time and the alternative is a real read per probe, which is what the probe exists to avoid. The test states the boundary so changing it is a decision, not an accident. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Chris Lu <chris.lu@gmail.com> |
||
|
|
3514925581 |
filer: let a nested path rule turn worm off (#10503)
* filer: let a nested path rule turn worm off mergePathConf ORs the booleans, so worm set on a bucket could never be lifted on a directory under it, while every string field is overridden by the more specific rule. Make worm tri-state instead: unset inherits, set wins. readOnly, fsync and disableChunkDeletion keep the OR, so a nested rule still cannot escape a lock the bucket set. Configurations written before this carry an explicit "worm": false on every rule, because they are marshalled with EmitUnpopulated. Reading those back as an override would quietly drop worm from nested paths, so filer.conf is now stamped with a version and the flag is dropped to unset when the version predates it. * filer: copy the worm value out of the matched rule mergePathConf aliased the pointer into the merged result, so a caller that wrote through it would reach into the stored rule. |
||
|
|
ccf5dc34e9 |
test: stop comparing two JWTs minted a second apart (#10495)
TestProxyReadDropsCallerJwtQueryParam mints a read token up front and requires the token the volume server would evaluate to equal it byte for byte. The expiry claim has one-second resolution -- GenJwtForVolumeServer sets it from jwt.NewNumericDate(time.Now().Add(...)) -- so two mints on either side of a tick produce different strings for the same authority and the same file, and the assertion fails for a reason the test is not about. It surfaces on the 32-bit job, where the runner is slow enough that the HEAD subtest (the second one, after a full proxy round trip) lands in a later second than the mint at the top of the test. Confirmed directly: minting the same file id with the same key either side of a boundary yields different tokens. Assert what the test is actually about instead -- that the credential decodes against the read key and authorizes this file id -- which holds whatever second it is minted in, and is a closer statement of the property than string equality. |
||
|
|
78ed665557 |
webdav: answer PROPFIND child stats from the listing (#10492)
* webdav: answer PROPFIND child stats from the listing golang.org/x/net/webdav discards the FileInfo that Readdir returned and stats every child again, five times over, so a PROPFIND on a directory costs five sequential filer lookups per entry: 15006 lookups and 1.4s for 3000 subdirectories, 24s for 60000. Windows Explorer times out well before that. Hold the entries a listing already fetched for the lifetime of the request and serve those stats from them - 6 lookups and 0.03s for the same 3000 entries. WebDavFile.Stat has to stop dropping its request context for the held entries to be reachable. * webdav: keep the request context on the lookups a listing drives stat, Readdir and Seek reached the filer on context.Background(), so a client that walked away left the listing streaming and the lookups running. Seek also missed the entries the listing had already fetched. Write and cleanup paths keep their own context - a cancelled request must not abandon a flush half done. |
||
|
|
9351202ca9 |
volume: scan for on-disk EC shards when staging a decoded volume (#10465)
The staged-new-volume placement skipped a disk holding the vid's EC shards using only the in-memory ecVolumes map, missing a shard present on disk but not mounted. Also scan the candidate disk for <vid>.ecNN files, so the promise holds regardless of mount state. Claude-Session: https://claude.ai/code/session_01Ks16jnt4S7gdDk8cheQ3xu |
||
|
|
3b3e8af430 |
volume: skip a shard-holding disk when staging a decoded volume (Go+Rust) (#10464)
volume: skip a shard-holding disk when staging a decoded volume ReceiveFile staged-new-volume mode picked any free disk of the target medium. Skip a disk that already holds the vid's EC shards (Go DiskLocation.FindEcVolume / Rust ec_volumes), so a decoded .dat never lands in the same directory as a shard. This lets a caller safely stage onto a shard host that has a spare disk, instead of requiring a host with no shard of the vid at all. Claude-Session: https://claude.ai/code/session_01Ks16jnt4S7gdDk8cheQ3xu |
||
|
|
9c37e52c9b |
volume: EC decode onto a clean peer via staged-new-volume adopt (Go+Rust) (#10463)
Decoding EC shards back to a normal volume in place reconstructs <vid>.dat
in the shards' own directory, so the vid is momentarily registered as both
an EC and a normal volume in one location — the load/scan path then sees it
as both, risking mount ambiguity and needle loss. VolumeEcShardsToVolume
still supports that in-place path; this adds the primitives to decode onto
a *clean* peer instead:
- ReceiveFile gains a staged-new-volume mode: when the volume does not
exist here and ReceiveFileInfo.disk_type is set, pick a free-slot disk
of that medium and write <base><ext>.copying (not a valid volume name,
so the scanner never half-loads a partial push).
- VolumeEcShardsToVolume gains from_staged: adopt the pushed .dat/.idx/
.vif — rename .copying into place under a .note in-progress marker,
then mount — so <vid> lands on the peer only as a normal volume.
The caller decodes the shards off-box and streams the finished volume to a
peer holding no shard of the vid on the target medium. Go and Rust volume
servers get identical handlers. Proto: ReceiveFileInfo.disk_type (12; 8-11
reserved for versioned-EC), VolumeEcShardsToVolumeRequest.from_staged (3) +
disk_type (4).
Claude-Session: https://claude.ai/code/session_01Ks16jnt4S7gdDk8cheQ3xu
|