mirror of
https://github.com/versity/scoutfs.git
synced 2026-07-19 22:42:40 +00:00
c042bf68a7
dirty_alloc_blocks saves the freed list state before modifying it so it can roll back on error. But it only saved the block_ref (blkno + seq), not the full alloc_list_head which also includes first_nr, total_nr, and flags. When the freed list's head block is nearly full, dirty_alloc_blocks forces allocation of a new empty head block. It zeros alloc->freed.ref and sets alloc->freed.first_nr = 0. dirty_list_block then sees the empty ref, allocates a fresh block, and updates alloc->freed.ref to point at it. first_nr = 0 correctly describes that new empty block. All good so far. If the subsequent avail dirty_list_block fails, the error path restores alloc->freed.ref to orig_freed but leaves alloc->freed.first_nr at 0. The original head block still holds its N entries on disk, but the in-memory head now claims first_nr = 0 -- the recovery should have rewound first_nr back to N along with the ref, but had no saved copy to rewind from. On the next dirty_alloc_blocks call the threshold check sees first_nr = 0! (full empty space) and skips the new-block path. dirty_list_block CoWs the existing head block and list_block_add writes new entries past the N already-present blknos while incrementing first_nr from 0. The head's first_nr drifts permanently below the block's actual nr. This leads to the alloc list head/block mismatch BUG_ON at alloc.c:375 and downstream extent overlap errors that permanently stall the server. Fix by saving and restoring the full scoutfs_alloc_list_head struct instead of just the scoutfs_block_ref. Exposed by stress testing. Signed-off-by: Auke Kok <auke.kok@versity.com>