The server commit deadlocks when its freed allocator list head block
fills to near capacity. Both gates that decide whether a transaction has
room measure it as free slots in the clean freed head block:
- hold_commit() admits a holder only if scoutfs_alloc_meta_remaining()
reports enough freed room (2 * COMMIT_HOLD_ALLOC_BUDGET slots).
- empty_list()/fill_list()'s list_has_blocks() proceed only if the head
has extent_mod_blocks() slots.
The clean head is not what a transaction gets: the first dirtying
allocation runs dirty_alloc_blocks(), which rotates in a fresh head block
when the current head is under EMPTY_FREED_THRESH. A clean, nearly-full
head has a full block's worth of room as soon as it's touched. With the
gates refusing on the clean full head, no holder is admitted and the
drains never start, so the rotation in dirty_alloc_blocks() is never
reached. The server spins applying empty commits and the filesystem can't
mount or recover (observed at ~11k empty commits/sec; freed head first_nr
8148 of an 8184 capacity).
Fix the accounting in scoutfs_alloc_meta_remaining(): when the freed list
isn't dirtied yet and the clean head is under EMPTY_FREED_THRESH, report
the room the pending rotation will give, SCOUTFS_ALLOC_LIST_MAX_BLOCKS - 2
(a fresh block, less the old avail and freed head blocks the rotation
frees into it). list_has_blocks() routes through the same function so
fill_list()/empty_list() use identical accounting; otherwise an
avail-low, freed-full commit could still wedge because fill_list()
couldn't refill avail past the clean full freed head.
hold_commit() then admits a holder (or a drain starts) and the first
allocation rotates the full head. The avail gate and the meta_low()
loop-stop are left conservative, so genuine ENOSPC still fails and freeing
loops still commit before overflowing a head.
Signed-off-by: Auke Kok <auke.kok@versity.com>
A failed mount tears down with scoutfs_put_super(), which calls
scoutfs_srch_destroy() first. srch_destroy() does cancel_work_sync() on
the srch compact worker, but that worker can be parked in an
uninterruptible scoutfs_net_sync_request() to a server that will never
respond (e.g. the server is stuck in recovery). Nothing completes the
request: the forced-unmount drain in the net shutdown path only runs for
umount -f, and a failed mount never calls ->umount_begin, so the request
sits on the resend queue and cancel_work_sync() waits forever.
The result is an unkillable D-state mount that survives the SIGKILL a
mount timeout sends, and only clears on reboot:
systemd[1]: data-archive.mount: Killing process 15740 (mount) with signal SIGKILL.
systemd[1]: data-archive.mount: Mount process still around after SIGKILL. Ignoring.
cat /proc/26717/stack
[<0>] __flush_work+0x16f/0x240
[<0>] __cancel_work_sync+0x135/0x1a0
[<0>] scoutfs_srch_destroy+0x33/0x70 [scoutfs]
[<0>] scoutfs_put_super+0x4f/0x1a0 [scoutfs]
[<0>] scoutfs_fill_super+0x260/0x520 [scoutfs]
[<0>] mount_bdev+0xf9/0x150
[<0>] do_new_mount+0x17a/0x310
[<0>] __x64_sys_mount+0x107/0x140
The worker it waits on, blocked in the sync request that never returns:
task:kworker/u269:1 state:D
Workqueue: scoutfs_srch_compact scoutfs_srch_compact_worker [scoutfs]
Call Trace:
__wait_for_common+0x90/0x1d0
scoutfs_net_sync_request+0xdb/0xf0 [scoutfs]
scoutfs_client_srch_get_compact+0x2e/0x40 [scoutfs]
scoutfs_srch_compact_worker+0x64/0x3d0 [scoutfs]
On the fill_super error path, mark forced_unmount and shut the client
connection down before teardown. That drains pending requests with
-ECONNABORTED so the worker returns and srch_destroy()'s cancel_work_sync
completes. It is done for any failure, before the direct put_super call
and before returning to generic_shutdown_super (which calls put_super
when s_root was set), so both teardown paths are covered. sbi is
allocated before any goto out, and scoutfs_client_net_shutdown() is a
no-op when the client or connection was never set up, so early failures
are safe.
Signed-off-by: Auke Kok <auke.kok@versity.com>
A node only needs to be fenced once, but scoutfs_fence_start() can be
called for the same rid more than once. When a new leader starts it
fences the previous leader as it removes it from the quorum
(quorum_block_leader), and that same rid can also be a mounted client
that then fails to recover within the timeout (client_recovery).
The second fence call collides on that name, sysfs returns -EEXIST,
and the error is propagated to fence_pending_recov_worker() which
treats any error as fatal and shuts the server down. On the next
mount a new leader hits the same stale set and the same collision,
so the filesystem can never finish recovery.
Jun 15 09:22:35 kernel: scoutfs f.000000.r.222222: fencing previous leader f.000000.r.111111 at term 183942 in slot 3 with address x.x.x.x:6000
Jun 15 09:22:36 scoutfs-fenced[9194]: [2026-06-15 09:22:36.588037842] server f.000000.r.222222 fencing rid 1111111111111111 at IP x.x.x.x for quorum_block_leader
Jun 15 09:23:09 kernel: scoutfs f.000000.r.222222 error: 30000 ms recovery timeout expired for client rid 1111111111111111, fencing
Jun 15 09:23:09 kernel: sysfs: cannot create duplicate filename '/fs/scoutfs/f.000000.r.222222/fence/1111111111111111'
Jun 15 09:23:09 kernel: scoutfs f.000000.r.222222 error: fence returned err -17, shutting down server
The reclaim path keys off the rid, not the reason, so one pending fence
covers both needs. Reserve the rid by checking for an existing entry
and adding the new pending_fence to the list under the same lock; a
concurrent caller for the same rid then sees it and returns success
without touching sysfs. The sysfs create runs after the list insert,
so any remaining -EEXIST reflects stale on-disk state and is left to
shut the server down rather than masked.
Restructure the function to a single exit: track on-list with a bool,
NULL the local on success, and let the out path remove from the list
and free if a failure left the entry behind.
Signed-off-by: Auke Kok <auke.kok@versity.com>
Scoutfs print segfaults walking a log_merge btree that has more than one
level. print_btree_block() prints parent (level > 0) items via
print_block_ref(), which invokes the item callback with a NULL value to
print the key portion before printing the child ref:
func(key, 0, 0, NULL, 0, arg);
print_log_merge_item immediately casts val and reads a field,
dereferencing NULL. A log_merge of a single leaf block (height 1) never
hits the parent path; one with height > 1 crashes on the first parent
item:
scoutfs[22043]: segfault at 8 ip 0000000000408ef0 sp 00007fffc5edd8b0 error 4
#0 0x0000000000408ef0 in print_log_merge_item ()
#1 0x000000000040958d in print_btree_block.constprop.0.isra ()
#2 0x000000000040a471 in print_cmd ()
#3 0x0000000000404264 in cmd_execute ()
#4 0x00000000004025c9 in main ()
print_mounted_client_entry has the same bug: it casts val and reads
mcv->addr / mcv->flags with no NULL guard, so a mounted_clients btree of
height > 1 segfaults the same way. print_srch_root_item guards NULL but
casts to scoutfs_srch_compact or scoutfs_srch_file without bounds
checking, so a short or malformed item reads past its end.
Fix all three: return early when val is NULL (printing just the key for
the parent ref where applicable), and bounds-check val_len before each
cast so a short item is reported instead of read past its end.
Signed-off-by: Auke Kok <auke.kok@versity.com>
RHEL7 was the only conditional user of this define, but since
support for that is removed, these can be dropped.
Signed-off-by: Auke Kok <auke.kok@versity.com>
scoutfs_rename2 was a thin shim that validated flags and forwarded
to scoutfs_rename_common. It existed because the old el7
RHEL_IOPS_WRAPPER path used a non-flag-taking .rename op alongside
.rename2; with that path gone there is only one rename method,
and the wrapper has no purpose.
Move the RENAME_NOREPLACE flag validation into scoutfs_rename_common
and point the directory inode_operations .rename slot at it directly.
The symlink inode_operations already used scoutfs_rename_common, so
this also makes symlink rename consistently reject unknown flags
instead of silently accepting them.
Signed-off-by: Auke Kok <auke.kok@versity.com>
el8 already provides stack_trace_save() and stack_trace_print()
in linux/stacktrace.h, so the legacy save_stack_trace/print_stack_trace
fallback inlines are dead. Drop the detection stanza and the inlines.
Signed-off-by: Auke Kok <auke.kok@versity.com>
el8 already provides vm_fault_t and vmf_error(), so the fallback
typedef and inline are dead. Drop the detection stanza and the
two function-signature ifdefs in data.c that switched between the
pre-v4.11 and modern fault handler prototypes.
Signed-off-by: Auke Kok <auke.kok@versity.com>
Only el7 was capable of testing this formatversion. And there
no longer is el7 support. Remove the test.
Signed-off-by: Auke Kok <auke.kok@versity.com>
We still need to keep KC_GENERIC_PERFORM_WRITE_KIOCB_IOV_ITER for
el8. So kc_generic_perform_write remains in place for now.
Signed-off-by: Auke Kok <auke.kok@versity.com>
This is a pre-el7 remnant, possibly from a really old rhel9
version, and was already effectively a stub.
Signed-off-by: Auke Kok <auke.kok@versity.com>
This depends on nfs-utils being installed on the host. Without it
it will skip, and count as a failure. It starts nfs-server and
does a bare exportfs.
- Tests basic read/write/stage/release/data wait.
- Tests setfacl/getfacl.
Signed-off-by: Auke Kok <auke.kok@versity.com>
Scoutfs has supported posix ACLs through the xattr handler table,
which allowed NFS to fetch them through this sideband, which worked
for older kernels.
With recent changes we've pulled in .get_acl because the mainline
kernel is changing how ACL ops are called. But we still left .set_acl
unreachable. This meant that on el9.7 nfs clients could now reach
.get_acl, but still not set them.
With this change, we're finally exposing .set_acl consistently
across all el releases and allowing nfs clients to both get and set
posix ACLs.
Signed-off-by: Auke Kok <auke.kok@versity.com>
El9.8 backported the upstream v6.15.* rename of from_timer to
timer_container_of. Switch the two callers in fence.c and recov.c
to the new style and add a simple kcompat define for older kernels.
Signed-off-by: Auke Kok <auke.kok@versity.com>
Ramps up data preallocation based on the number of online
blocks. This results in a simple 2<<n block allocation pattern
until n=11 (2048) - the default value of data_prealloc_blocks.
Signed-off-by: Auke Kok <auke.kok@versity.com>