* Add readWriteOncePod backupPVC config to enable mount-level SELinux labeling On SELinux-enabled clusters the kubelet recursively relabels every file of the backupPVC at mount time, which can take hours on volumes with a high file count. Kubernetes avoids this when the volume is ReadWriteOncePod and the CSI driver advertises SELinux mount support, by mounting with -o context= instead. Add an opt-in per-storage-class 'readWriteOncePod' backupPVC config option that creates the backupPVC with the ReadWriteOncePod access mode and sets the backup pod's SecurityContext.SELinuxChangePolicy to MountOption. It is mutually exclusive with 'readOnly', which takes precedence. Fixes #9873 Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com> * Add changelog file Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com> * Do not set SELinuxChangePolicy for the backup pod Live testing on OCP 4.22 (k8s 1.35) showed that setting SecurityContext.SELinuxChangePolicy to MountOption makes backup pod creation fail outright when the SELinuxMount feature gate is disabled, which is the default on current clusters: Pod is invalid: spec.securityContext.seLinuxChangePolicy: Unsupported value: "MountOption": supported values: "Recursive" The field is also unnecessary. For ReadWriteOncePod volumes the kubelet already performs mount-level SELinux labeling via the SELinuxMountReadWriteOncePod feature gate, which has been on by default since k8s 1.28. Setting the backupPVC access mode to ReadWriteOncePod is sufficient on its own, and is portable to clusters where the broader SELinuxMount gate is still off. Verified on-cluster that the backupPVC is mounted with context="system_u:object_r:container_file_t:s0:c22,c28" instead of the recursive seclabel mount used without the flag. Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com> --------- Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
8.2 KiB
title, layout
| title | layout |
|---|---|
| BackupPVC Configuration for Data Movement Backup | docs |
BackupPVC is an intermediate PVC to access data from during the data movement backup operation.
In some scenarios users may need to configure some advanced options of the backupPVC so that the data movement backup operation could perform better. Specifically:
- For some storage providers, when creating a read-only volume from a snapshot, it is very fast; whereas, if a writable volume
is created from the snapshot, they need to clone the entire disk data, which is time consuming. If the
backupPVC'saccessModesis set asReadOnlyMany, the volume driver is able to tell the storage to create a read-only volume, which may dramatically shorten the snapshot expose time. On the other hand,ReadOnlyManyis not supported by all volumes. Therefore, users should be allowed to configure theaccessModesfor thebackupPVC. - Some storage providers create one or more replicas when creating a volume, the number of replicas is defined in the storage class.
However, it doesn't make any sense to keep replicas when an intermediate volume used by the backup. Therefore, users should be allowed
to configure another storage class specifically used by the
backupPVC. - In SELinux-enabled clusters, such as OpenShift, when using the above-mentioned readOnly access mode setting, SELinux relabeling of the
volume is not possible. Therefore for these clusters, when setting
readOnlyfor a storage class, users must also disable relabeling. Note that this option is not consistent with the Restricted pod security policy, so if Velero pods must run with a restricted policy, disabling relabeling (and therefore readOnly volume mounting) is not possible.
Velero introduces a new section in the node agent configuration ConfigMap (the name of this ConfigMap is passed using --node-agent-configmap velero server argument)
called backupPVC, through which you can specify the following
configurations:
-
storageClass: This specifies the storage class to be used for the backupPVC. If this value does not exist or is empty then by default the source PVC's storage class will be used. -
readOnly: This is a boolean value. If set totruethenReadOnlyManywill be the only value set to the backupPVC's access modes. OtherwiseReadWriteOncevalue will be used. -
spcNoRelabeling: This is a boolean value. If set totrue, thenpod.Spec.SecurityContext.SELinuxOptions.Typewill be set tospc_t. From the SELinux point of view, this will be considered a "Super Privileged Container" which means that selinux enforcement will be disabled and volume relabeling will not occur. This field is ignored ifreadOnlyisfalse. -
readWriteOncePod: This is a boolean value. If set totrue, thenReadWriteOncePodwill be the only value set to the backupPVC's access modes. On SELinux-enabled clusters the kubelet applies the SELinux label to aReadWriteOncePodvolume at mount time (-o context=) instead of recursively relabeling every file on the volume, which can take hours on volumes with a high file count. It requires a CSI driver that advertises SELinux mount support (CSIDriver.spec.seLinuxMount: true) and a storage class that supports creatingReadWriteOncePodPVCs from a snapshot. This field is ignored ifreadOnlyistrue.
The users can specify the ConfigMap name during velero installation by CLI:
velero install --node-agent-configmap=<ConfigMap-Name>
-
annotations: permits to set annotations on the backupPVC itself. typically useful for some CSI provider which cannot mount a VolumeSnapshot without a custom annotation. -
secretNames: a list of secret names to copy from the source PVC's namespace to the Velero namespace before the backupPVC is created, and delete after the DataUpload completes. This is needed for CSI drivers that require namespace-scoped secrets to provision the volume, for example ODF/ceph-csi encrypted volumes that fetch a KMS token secret (ceph-csi-kms-token) from the PVC's namespace. Without this, the backupPVC created in the Velero namespace fails to provision because the secret only exists in the source namespace. -
configMapNames: a list of configmap names to copy from the source PVC's namespace to the Velero namespace before the backupPVC is created, and delete after the DataUpload completes. This is needed for CSI drivers that require namespace-scoped configmaps to provision the volume, for example a tenant-specific ceph-csi KMS connection override configmap (ceph-csi-kms-config).
A sample of backupPVC config as part of the ConfigMap would look like:
{
"backupPVC": {
"storage-class-1": {
"storageClass": "backupPVC-storage-class",
"readOnly": true
},
"storage-class-2": {
"storageClass": "backupPVC-storage-class"
},
"storage-class-3": {
"readOnly": true,
"annotations": {
"some-csi.provider.io/readOnlyClone": true
}
},
"storage-class-4": {
"readOnly": true,
"spcNoRelabeling": true
},
"ocs-storagecluster-ceph-rbd-encrypted": {
"secretNames": ["ceph-csi-kms-token"],
"configMapNames": ["ceph-csi-kms-config"]
},
"storage-class-5": {
"readWriteOncePod": true
}
}
}
Note on encrypted volumes: the copied secrets/configmaps are labeled velero.io/backup-pvc-secret=<DataUpload UID> and
deleted when the DataUpload completes (or on failure). If concurrent DataUploads from different namespaces need a secret with the same
name but different content in the Velero namespace, they conflict. For ceph-csi,
this can be avoided by configuring a unique
tenantTokenName per tenant.
Note:
- Users should make sure that the storage class specified in
backupPVCconfig should exist in the cluster and can be used by thebackupPVC, otherwise the corresponding DataUpload CR will stay inAcceptedphase until timeout (data movement prepare timeout value is 30m by default). - If the users are setting
readOnlyvalue astruein thebackupPVCconfig then they must also make sure that the storage class that is being used forbackupPVCshould support creation ofReadOnlyManyPVC from a snapshot, otherwise the corresponding DataUpload CR will stay inAcceptedphase until timeout (data movement prepare timeout value is 30m by default). - In an SELinux-enabled cluster, any time users set
readOnly=truethey must also setspcNoRelabeling=true. There is no need to setspcNoRelabeling=trueif the volume is not readOnly. readWriteOncePodandreadOnlyare mutually exclusive. If both are set totrue,readOnlywins,readWriteOncePodis ignored and a warning is logged.readWriteOncePodis an alternative toreadOnly+spcNoRelabelingfor SELinux-enabled clusters whose storage does not supportReadOnlyMany(for example Ceph RBD in Filesystem mode or LVM). Users must make sure the storage class used forbackupPVCsupports creating aReadWriteOncePodPVC from a snapshot, otherwise the corresponding DataUpload CR will stay inAcceptedphase until timeout.- If any of the above problems occur, then the DataUpload CR is
canceledafter timeout, and the backupPod and backupPVC will be deleted, and the backup will be marked asPartiallyFailed.
Related Documentation
- Node-agent Configuration - Complete reference for all configuration options
- Node-agent Concurrency - Configure concurrent operations per node
- Node Selection for Data Movement - Configure which nodes run data movement
- Data Movement Pod Resource Configuration - Configure pod resources
- BackupPVC Configuration - Configure backup storage
- RestorePVC Configuration - Configure restore storage
- Cache PVC Configuration - Configure restore data mover storage