From 1bc3f178e751d8b90efc111bb1e2d056bdaefbc8 Mon Sep 17 00:00:00 2001 From: Gleb Chesnokov Date: Thu, 20 Aug 2026 16:11:47 +0300 Subject: [PATCH] scst: Replace obsolete cluster setup guidance README.dlm embeds a 2015 Pacemaker recipe that disables quorum and fencing. README.drbd also says persistent reservation state is never shared between nodes. The first is unsafe for shared storage and the second contradicts the current cluster_mode implementation. Commit 6420070cb992 ("scst: Synchronize persistent reservation state between cluster nodes") added DLM-backed reservation synchronization. The current contract lives in scst_dlm.c and scst_pres.c, while fencing and resource-manager policy remain outside SCST. Replace the fixed Pacemaker commands with DLM dependencies and startup and shutdown ordering. Require fencing, describe cluster_mode, retain the DRBD protocol C requirement and leave multipath policy to the installed initiator stack. --- scst/README.dlm | 151 ++++++++++------------------------------------- scst/README.drbd | 26 +++++--- 2 files changed, 48 insertions(+), 129 deletions(-) diff --git a/scst/README.dlm b/scst/README.dlm index 7ca663c0f..07293ca17 100644 --- a/scst/README.dlm +++ b/scst/README.dlm @@ -14,134 +14,43 @@ scst_dlm.c uses the DLM to keep PR data synchronized across all nodes in a cluster. -Software Components -------------------- +Dependencies and deployment +--------------------------- -The following software components are needed by the code in scst_dlm.c: -* The DLM kernel driver (dlm.ko). This driver is only built if CONFIG_DLM - has been set. -* The DLM control daemon (dlm_controld.pcmk). This daemon passes cluster - node IDs and IP addresses to the DLM kernel driver via the configfs - interface of the DLM kernel driver. -* Corosync to manage cluster membership of the cluster nodes and to assign - a node ID to each cluster node. -* A facility to start the DLM control daemon, e.g. Pacemaker. +The SCST DLM integration requires a kernel built with DLM support and a +working DLM userspace and cluster-membership stack. SCST consumes the DLM +interfaces; it does not configure membership, quorum or fencing. Follow the +documentation shipped with the selected distribution and cluster stack for +those parts of the deployment. -On most Linux distributions the software packages that contain this software -have the names kernel, dlm, corosync and pacemaker. +Do not disable quorum handling or fencing merely to bring up an example. +Incorrect membership or failed-node exclusion can allow multiple nodes to +write shared storage without the required coordination and can corrupt data. +Validate fencing and recovery on a disposable cluster before exporting real +storage. -NOTE! You might need to apply a DLM bugfix patch, see scst-devel mailing list -thread https://sourceforge.net/p/scst/mailman/scst-devel/thread/CADHfD59FK6seaammL8b9LM3U3tw5HvYp3kPTk_r1OYkPR7bPhg@mail.gmail.com/#msg34761854 -for more details. +Every node that accesses the same logical device must use the same stable +t10_dev_id. Enabling cluster_mode selects DLM-backed persistent-reservation +handling and may create or join the lockspace derived from that identifier. -DLM Configuration ------------------ +Startup and shutdown ordering +----------------------------- -The DLM kernel module supports the TCP and SCTP communication protocols. An -advantage of SCTP for H.A. purposes is that it supports multihoming. One of -these protocols can be selected via the -r option of dlm_controld. -That option can be set via the "args" argument of the Pacemaker dlm_controld -resource. For more information, see also: -* The dlm_controld(8) man page. -* In the "Pacemaker 1.1, Clusters from Scratch" guide, the section "Configure - the Cluster for the DLM". -* The dlm_controld resource agent: /usr/lib/ocf/resource.d/pacemaker/controld +Use the following ordering as an integration constraint, not as a deployment +script: +* Configure and validate cluster membership, quorum, fencing and DLM according + to the current cluster-stack documentation. +* Configure SCST while its target ports remain disabled. +* Enable cluster_mode only for shared devices and verify that every node uses + the same t10_dev_id before enabling target ports. +* During shutdown, disable target ports and wait for initiator sessions and + commands to drain before disabling cluster_mode. +* Stop the DLM and cluster stack only after SCST has released its lockspaces. -Here is an example of how to set up a cluster with two nodes and how to -configure and start the DLM control daemon: - 1. If a network switch is present between the two nodes, enable IPv4 multicast - on that switch. - 2. Copy /etc/corosync/corosync.conf.example into /etc/corosync/corosync.conf - and edit that file. - 3. If a file /etc/default/corosync exists, enable Corosync in that file. - 4. Start Corosync: - systemctl start corosync || /etc/init.d/corosync start - 5. Check that all configured Corosync rings have two members: - corosync-cfgtool -s && { corosync-cmapctl | grep members; } - 6. Start pcsd: - systemctl start pcsd || /etc/init.d/pcsd start - 7. Set up cluster authentication: - pcs cluster auth centos7-vm centos7b-vm - 8. Start Pacemaker: - systemctl start pacemaker || /etc/init.d/pacemaker start - 9. If the cluster has only two nodes, disable the Pacemaker quorum policy and - disable STONITH: - crm_attribute -t crm_config -n no-quorum-policy -v ignore - crm_attribute -t crm_config -n stonith-enabled -v false -10. Check the cluster status: - pcs status -11. Create a Pacemaker resource for dlm_controld: - pcs resource delete dlm - pcs resource create dlm ocf:pacemaker:controld \ - args="-q0 -f0" allow_stonith_disabled=true \ - op monitor timeout=60 \ - --clone interleave=true -12. Check the Pacemaker status: - pcs status - - -Startup and Shutdown --------------------- - -The startup sequence is as follows: -* Load the DLM kernel module. If not loaded explicitly, "modprobe scst" will - load the DLM kernel module implicitly. -* Load and configure SCST with all target ports disabled. -* Enable cluster mode for all SCST devices that can be accessed through more - than one cluster node: - for x in /sys/kernel/scst_tgt/handlers/*/*/; do - echo 1 >$x/cluster_mode & - done - wait -* Start Corosync and Pacemaker. -* Wait until Pacemaker has reached the idle state: - pacemaker_dc_status() { - local dc - - dc="$(crmadmin -D 2>/dev/null | sed 's/Designated Controller is: //')" - [ -n "$dc" ] && - crmadmin -S "$dc" 2>/dev/null | - sed 's/^Status of crmd@[^[:blank:]]*:[[:blank:]]\([^[:blank:]]*\).*/\1/' - } - timeout=300 - for ((i=0;i $x & - done - wait - while ls -Ad /sys/kernel/scst_tgt/targets/*/*/sessions/* 2>&1 | - grep -vE '/sys/kernel/scst_tgt/targets/(copy_manager|scst_local)/'; do - sleep 1 - done -* Tell SCST to release the DLM lockspaces: - while grep -q '^1$' /sys/kernel/scst_tgt/devices/*/cluster_mode 2>/dev/null - do - for x in /sys/kernel/scst_tgt/devices/*/cluster_mode; do - { [ -e "$x" ] && echo 0 > "$x"; } & - done - wait - sleep 1 - done -* Stop Pacemaker and Corosync -* Unload the SCST kernel modules -* Unload the DLM kernel driver +Loading modules, writing cluster_mode, enabling targets and exercising failover +all change live storage or cluster state. Perform them only with an explicit +host, device, fencing and recovery plan. Lockspace names diff --git a/scst/README.drbd b/scst/README.drbd index 25fea3a56..eb1faa2a1 100644 --- a/scst/README.drbd +++ b/scst/README.drbd @@ -1,6 +1,13 @@ SCST, DRBD and Dual Primary Mode ******************************** +Scope and safety +---------------- +This document describes SCST requirements for a dual-primary DRBD setup. It is +not a complete deployment recipe. Validate the DRBD, cluster-manager, fencing, +and initiator configuration against the documentation for the versions in use +before exporting real storage. + Introduction ------------ One possible approach to build a highly available storage solution is to @@ -26,20 +33,23 @@ settings for each DRBD block device exported via SCST: DRBD Configuration ------------------ -As documented in the DRBD manual, dual primary mode only works with the DRBD -replication protocol C and not with protocol A nor with protocol B. +As documented in the DRBD 9 User's Guide, dual-primary mode requires +synchronous replication protocol C and suitable fencing. Protocols A and B do +not provide the required completion semantics. Initiator Configuration ----------------------- -When using VMware, do NOT use multipath I/O in round-robin mode. This mode -namely depends on persistent reservations to guarantee data -consistency. When running an SCST cluster, the SCST persistent reservation -state is not shared between the two nodes. +Do not select a multipath policy that permits concurrent access until the +complete configuration has been validated. In particular, enable SCST +cluster_mode for every shared device, use the same stable t10_dev_id on both +nodes, and verify persistent-reservation synchronization and fencing on every +path. Also follow the current initiator vendor's requirements; these are not +defined by SCST. References ---------- -* DRBD authors, [Dual Primary Mode], DRBD Manual, 2011. -* VMware, [iSCSI SAN Configuration Guide 4.1], VMware website, September 2010. +* LINBIT, DRBD 9 User's Guide: + https://linbit.com/drbd-user-guide/drbd-guide-9_0-en/ * Phil White, [VMware ESX, MPIO, SCST and DRBD in Dual Primary Mode], SCST Developers Mailing List, June 2010. * Stefan Berger, [SCST and DRBD in Dual Primary Mode], DRBD users mailing