Auke Kok 69b09eafe3 Use raft pre-vote to avoid election term inflation.
A quorum member that loses contact with the majority used to increment
and persist its term on every election timeout while it could never win.
A network-partitioned former leader could climb its term far above the
cluster's, then on reboot read that inflated term back from its quorum
block and usurp the leader that legitimately replaced it.

The previous term-persist approach (deferring the durable term write
until a vote arrived) did not cover a partition that splits off an
intercommunicating sub-majority: in a five node cluster a two node
minority that can still reach each other exchange votes, so both the
voter (adopting the higher term) and the candidate (seeing a vote)
persist the climbing term anyway.

Add a raft pre-vote round.  On election timeout a member becomes a
pre-candidate and probes for support at its prospective next term
without incrementing or persisting it.  Peers grant a pre-vote only when
they do not currently see a live leader, and pre-vote messages are
excluded from the adopt-greater-term path so granting changes no state.
Only once a majority of pre-votes is collected does the member increment
its term, persist it, and stand for election for real.  A partitioned
member or sub-majority can never reach a majority of pre-votes, so it
loops pre-vote rounds at a frozen term, never inflating it, and follows
the majority's leader once the partition heals.

Signed-off-by: Auke Kok <auke.kok@versity.com>
2026-07-10 22:28:25 -07:00
2026-06-05 09:49:45 -07:00
2020-12-07 09:47:12 -08:00
2020-12-07 10:39:20 -08:00
2021-11-05 11:16:57 -07:00
2026-06-03 11:34:21 -07:00

Introduction

scoutfs is a clustered in-kernel Linux filesystem designed to support large archival systems. It features additional interfaces and metadata so that archive agents can perform their maintenance workflows without walking all the files in the namespace. Its cluster support lets deployments add nodes to satisfy archival tier bandwidth targets.

The design goal is to reach file populations in the trillions, with the archival bandwidth to match, while remaining operational and responsive.

Highlights of the design and implementation include:

  • Fully consistent POSIX semantics between nodes
  • Atomic transactions to maintain consistent persistent structures
  • Integrated archival metadata replaces syncing to external databases
  • Dynamic seperation of resources lets nodes write in parallel
  • 64bit throughout; no limits on file or directory sizes or counts
  • Open GPLv2 implementation

Community Mailing List

Please join us on the open scoutfs-devel@scoutfs.org mailing list hosted on Google Groups

S
Description
No description provided
Readme 6.6 MiB
Languages
C 86.3%
Shell 10%
Roff 2.5%
TeX 0.8%
Makefile 0.4%