Get the volume ID before creating the restore PVC, otherwise the existing PV may be deleted during the creation of restore PVC
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Update CRDs and CLI to support in-place restore (#10038)
Update CRDs(Restore, DataDownload, PodVolumeRestore) and restore create CLI to support in-place restore
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Update Kopia(filesystem) uploader to support incremental and deleteExtraFile during restore (#10066)
Update Kopia(filesystem) uploader to support incremental and deleteExtraFile during restore
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Update Restore Exposer and PVC CSI to support in-place restore (#10104)
1. Update Restore Exposer to support exposing with existing PV for in-place restore
2. Update PVC CSI RIA to continue the restore process for in-place restore
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Update Block uploader to support increase restore (#10244)
Update Block uploader to support increase restore
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Update Exposer to recreate the target PV if the volume mode is different with the restore PVC (#10257)
Update Exposer to recreate the target PV if the volume mode is different with t
he restore PVC
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
* Preserve PVC selected-node annotation via carrier annotation for in-place restore
For in-place volume data restore, the existing PVC is deleted and
recreated. For StorageClasses with the WaitForFirstConsumer volume
binding mode, losing the volume.kubernetes.io/selected-node annotation
could let the scheduler place the recreated workload Pod in a different
zone than the original PV, leaving it stuck in ContainerCreating.
Instead of relying on RestoreItemAction execution order (the generic
PVC RIA unconditionally strips the selected-node annotation), the PVC
CSI RIA now captures the annotation from the existing PVC right before
deleting it and carries it on the target PVC via the Velero-internal
restore.velero.io/inplace-restore-selected-node annotation. The restore
engine translates the carrier back to the Kubernetes annotation after
all RestoreItemActions have run and always strips the carrier so it
never lands on the cluster.
This makes the behavior independent of RIA ordering: the Kubernetes
annotation is stripped by default on every path (including when the
target PVC does not exist and Velero falls back to provisioning a new
PVC), and preservation only happens when the CSI RIA explicitly
captured a value from the existing PVC.
Signed-off-by: chlins <chlins.zhang@gmail.com>
* Update the control path to make the in-place incremental restore with block data mover work E2E (#10410)
Update the control path to make the in-place incremental restore with block data mover work E2E
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
---------
Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
Signed-off-by: chlins <chlins.zhang@gmail.com>
Co-authored-by: chlins <chlins.zhang@gmail.com>
* Add readWriteOncePod backupPVC config to enable mount-level SELinux labeling
On SELinux-enabled clusters the kubelet recursively relabels every file of
the backupPVC at mount time, which can take hours on volumes with a high
file count. Kubernetes avoids this when the volume is ReadWriteOncePod and
the CSI driver advertises SELinux mount support, by mounting with
-o context= instead.
Add an opt-in per-storage-class 'readWriteOncePod' backupPVC config option
that creates the backupPVC with the ReadWriteOncePod access mode and sets
the backup pod's SecurityContext.SELinuxChangePolicy to MountOption. It is
mutually exclusive with 'readOnly', which takes precedence.
Fixes#9873
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
* Add changelog file
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
* Do not set SELinuxChangePolicy for the backup pod
Live testing on OCP 4.22 (k8s 1.35) showed that setting
SecurityContext.SELinuxChangePolicy to MountOption makes backup pod
creation fail outright when the SELinuxMount feature gate is disabled,
which is the default on current clusters:
Pod is invalid: spec.securityContext.seLinuxChangePolicy:
Unsupported value: "MountOption": supported values: "Recursive"
The field is also unnecessary. For ReadWriteOncePod volumes the kubelet
already performs mount-level SELinux labeling via the
SELinuxMountReadWriteOncePod feature gate, which has been on by default
since k8s 1.28. Setting the backupPVC access mode to ReadWriteOncePod is
sufficient on its own, and is portable to clusters where the broader
SELinuxMount gate is still off.
Verified on-cluster that the backupPVC is mounted with
context="system_u:object_r:container_file_t:s0:c22,c28" instead of the
recursive seclabel mount used without the flag.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
---------
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
The copied secret/configmap label value was the owner (DataUpload/
DataDownload) name, which is derived from the Backup/Restore name and
can exceed the 63-char Kubernetes label-value limit or contain invalid
characters. That would make the copy label and the cleanup selector
diverge and orphan the copied resources.
Use string(ownerObject.UID) consistently for the label value in both
copy and cleanup (backup and restore exposers). The UID is a stable,
always-valid label value.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
Mirror the backup-side fix on the restore path. The generic restore
exposer creates the intermediate restore PVC in the Velero namespace
using the target PVC's StorageClass. For encrypted volumes this fails
because ceph-csi looks up the KMS token secret in the PVC's namespace
(the Velero namespace), where it does not exist.
Add SecretNames/ConfigMapNames to the RestorePVC config. When set, the
generic restore exposer copies the named secrets/configmaps from the
target namespace to the Velero namespace before creating the restore
PVC, and cleans them up in CleanUp(). Reuses the same copy/delete
helpers and label as the backup path.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
- Make CopySecret/CopyConfigMap accept a generic labels map instead of
hardcoding the backup-pvc-secret label, aligning with the generic
DeleteSecretsWithLabel helper. Move the BackupPVCSecretLabel constant
from util/kube to the exposer package where it is used.
- Move the secret/configmap copy in Expose() to after
WaitVolumeSnapshotReady and before createBackupVS. That is the most
likely failure point, and nothing needs cleanup before it.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
- Fix premature deletion of shared secrets/configmaps: check owner label
in addition to data equality. Same data + different owner is now a
collision, preventing one DataUpload's CleanUp from removing resources
another DataUpload is still using.
- Copy BinaryData in CopyConfigMap and include it in the equality check,
so configmaps with binary payloads (e.g., CA bundles) are not silently
truncated.
- Add UID preconditions to DeleteSecretsWithLabel and
DeleteConfigMapsWithLabel to avoid TOCTOU races where a recreated
object with the same name could be deleted.
- Move secret/configmap copy to the beginning of Expose(), before any
intermediate objects are created, so failure doesn't require cleanup.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
Move the secret and configmap copy logic from the DataUpload controller
into the CSI snapshot exposer's Expose() method. This keeps all
CSI-specific logic in the exposer and maintains symmetry with CleanUp()
which already handles the cleanup of copied resources.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
Add label-based secret cleanup in CleanUp() to delete any secrets
that were copied to the Velero namespace for backup PVC provisioning.
Uses the velero.io/backup-pvc-secret label to find secrets associated
with the DataUpload being cleaned up.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
* refactor: use k8s.io/api well-known label constants
Several well-known Kubernetes label strings were hardcoded across the
codebase instead of using the constants already exported by
k8s.io/api/core/v1, which is an existing dependency:
"kubernetes.io/hostname" -> corev1api.LabelHostname
"kubernetes.io/os" -> corev1api.LabelOSStable
"topology.kubernetes.io/zone" -> corev1api.LabelTopologyZone
The local kube.NodeOSLabel and zoneLabel consts, which duplicated the
upstream values verbatim, are now defined in terms of the upstream
constants rather than repeating the literal. Both are kept: NodeOSLabel
is exported and referenced from four packages alongside NodeOSLinux and
NodeOSWindows, which have no upstream equivalent, and zoneLabel sits
beside the deprecated-label fallback it is compared against.
No functional change - every replacement is a constant with an identical
value.
Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
* Add changelog for #10279
Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
* Cover the selected-node path in createRestorePod
TestCreateRestorePod only exercised selectedNode == "", so the branch
that pins the restore pod to a node was never executed. Add a case with
a selected node and assert the resulting pod carries the hostname label
in its node selector.
Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
* Also use constants for the arch and deprecated zone labels
Extends the same replacement to the two remaining well-known labels
raised on the issue:
"kubernetes.io/arch" -> corev1api.LabelArchStable
"failure-domain.beta.kubernetes.io/zone" -> corev1api.LabelFailureDomainBetaZone
zoneLabelDeprecated in item_backupper.go was the last local const still
repeating a literal that upstream already exports, so the zone pair now
reads consistently against k8s.io/api. The deprecation note upstream
applies to the label itself, not the constant; Velero reads that label
deliberately as the fallback for PVs created before the topology labels
existed.
Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
---------
Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
Detect the host path volumes by their source instead of their name, so the customized ones are excluded as well. Drop the container capability changes since the data mover needs them to access the data, and fix the copyright headers.
Signed-off-by: chlins <chlins.zhang@gmail.com>
The CSI snapshot and generic restore exposers access data through PVCs, so they no longer inherit the node-agent host path volumes. Also drop all capabilities on the data mover container.
Signed-off-by: chlins <chlins.zhang@gmail.com>
* Add change-id and volume-id retrieve logic for both vks and vanilla k8s environment.
* Add change-id and volume-id support code in exposer.
Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
Change errors.Cause to errors.Is, because github.com/cockroachdb/errors
New() function create a error with error stack with depth 1, but
github.com/pkg/errors's New() function create error with no depth.
Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>