Commit Graph
992 Commits
Author SHA1 Message Date
57bddf7f23 Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds (#9646)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 9s
Run the E2E test on kind / get-go-version (push) Failing after 10s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds

Previously, the backup finalizer controller set backup.Status.Phase to
Completed/PartiallyFailed in-memory BEFORE calling PutBackupMetadata and
PutBackupContents. When these uploads failed (e.g., due to object lock
or immutability), the deferred patch function still wrote the terminal
phase to the Kubernetes API server, preventing the controller from
retrying the upload on the next reconcile.

This fix moves the phase assignment to AFTER both uploads succeed. A
DeepCopy of the backup is used to encode the JSON with the final phase
for object storage, while the in-memory backup object retains the
Finalizing phase until uploads complete.

Caveats:
- CompletionTimestamp is now captured before upload but only committed to
  the API server after upload succeeds. On retry after a transient
  failure, a new timestamp is generated, so the completion time reflects
  when the upload finally succeeded rather than when finalization
  processing completed.
- Metrics (RegisterBackupSuccess/RegisterBackupPartialFailure) are now
  recorded after uploads succeed, so they accurately reflect only fully
  persisted backups.
- The metadata uploaded to object storage contains the final phase and
  completion timestamp via DeepCopy, so storage state is correct even
  before the API server is patched.

Fixes #9645

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Add changelog for #9646

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix testifylint: use require.Error instead of assert.Error

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Address review feedback on backup-finalizer fix

- Add default guard for unhandled phase values in finalPhase switch
- Add retry with DefaultBackoff for PutBackupMetadata per reviewer request
- Replace brittle framework.BackupItemActionResolverV2{} mock with mock.Anything
- Add FinalizingPartiallyFailed test case for PutBackupContents failure

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Use bounded, object-storage-tuned backoff for backup-finalizer uploads

retry.DefaultBackoff is tuned for API server optimistic-concurrency
conflicts (4 steps, ~1.25s total) and gives up far too quickly for
object storage calls, which can see longer transient outages or
throttling (review feedback from blackpiglet). Replace it with a
dedicated, bounded backoff (1s base, 2x factor, 5 steps, ~31s total)
applied to both PutBackupMetadata and PutBackupContents.

Being bounded (rather than retrying forever) means a persistent
failure, e.g. an object-lock/immutability policy denying every write,
surfaces as an error within a bounded time instead of hanging the
reconcile indefinitely; controller-runtime requeues on error, so
retries continue across reconciles (review feedback from priyansh17).

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix PutBackupMetadata retry to re-read backupJSON each attempt

backupJSON is a bytes.Buffer, so passing it directly to
PutBackupMetadata drains it on the first read attempt. A retry after
a transient failure would then upload empty content instead of the
backup metadata. Wrap it in bytes.NewReader(backupJSON.Bytes()) inside
the retry closure so every attempt gets a fresh reader.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-09-22 20:06:46 -04:00
lyndon-liandGitHub 0a2f5278b2 Move the progress messages to activities (#10552)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 10s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* move the progress messages to activities

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>

* update doc for activities in data mover CR

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>

---------

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-22 14:24:05 +08:00
lyndon-liandGitHub 5f14d304c1 Merge pull request #10275 from kaovilai/namespace-selection-by-label
Implement namespace selection by label in resource policy
2026-09-22 09:22:38 +08:00
Tiger KaovilaiandClaude Sonnet 5 40af5efdd0 Implement namespace selection by label in resource policy
Add includedNamespacesByLabel, excludedNamespacesByLabel, and
labelSelectorLogic to IncludeExcludePolicy in the ResourcePolicy
ConfigMap (realizes design in velero-io/velero#9772), letting a backup
select or exclude namespaces by label instead of (or in addition to)
name/wildcard.

The backup controller resolves label selectors against the live
namespace list once per backup, merges the results into
spec.includedNamespaces/excludedNamespaces, then proceeds through the
existing name-based filtering unchanged. A defaulted "*" include list
is replaced by the resolved set; an explicitly-configured include list
(including an explicit "*") is unioned with it instead, and stays
canonical rather than widening. Namespaces matching an exclude
selector are always subtracted from the merged includes, regardless of
how the includes were populated.

Because Velero's namespace-includes/excludes model requires at least
one name (an empty list means "match everything"), a selector that
resolves to zero namespaces is represented with a sentinel glob
pattern ("[-]*") guaranteed to match no real namespace, rather than an
empty list that would silently fall back to including/excluding
everything.

labelSelectorLogic ("AND"/"OR", case-insensitive) controls whether
multiple included/excluded label selectors are combined by
intersection or union; it is validated up front, including inside
ResolveNamespacesByLabel itself, so an invalid value fails fast instead
of silently falling through to OR semantics.

Namespace-selection-by-label and resource-selection-by-label act as
independent axes and do not affect each other, matching the design
discussion in #9772.

Known limitations:
- Selectors are evaluated once per backup against the namespace list
  at that point in time; namespaces created or relabeled mid-backup
  are not picked up.
- Backup-only for now; restore-side namespace mapping is unaffected.

Testing:
- Unit coverage in internal/resourcepolicies for validation, selector
  resolution (including AND/OR logic, case-insensitivity, and
  malformed-selector/invalid-logic error paths), and the no-match
  sentinel.
- Unit coverage in pkg/controller for the merge logic between resolved
  label selections and explicit/defaulted includes and excludes.
- End-to-end coverage in pkg/backup exercising the full backup
  pipeline with label-selected namespaces, including the
  velero.io/exclude-from-backup hard-exclusion interaction and the
  zero-match/fully-excluded sentinel path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
2026-09-18 23:04:55 -04:00
Xun Jiang 5afc20d03e Support skipped PVC in VolumeInfos.
* Modify according to comments.
* Rename the pvSkipTracker and related fields to indicate both PVC and PV are supported.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-18 16:51:00 +08:00
f40112aedf Fix schedule reconciler to compare SkipImmediately and LastSkipped by value (#10490)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Fix schedule reconciler to compare SkipImmediately and LastSkipped by value

The reconciler compared spec.skipImmediately (*bool) and
status.lastSkipped (*metav1.Time) between the live object and its
DeepCopy by pointer, so both checks always reported a change once the
fields were set, and every reconcile issued a redundant Patch. Compare
by value instead (ptr.Equal / equality.Semantic.DeepEqual) and add a
regression test pinning that the skip-once flip persists with exactly
one patch, later reconciles of persisted state patch zero times, and
an explicit false is a no-op.

Signed-off-by: HeonJe LEE <lhjnano@gmail.com>

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>

---------

Signed-off-by: HeonJe LEE <lhjnano@gmail.com>
Signed-off-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>
Co-authored-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-16 17:00:29 -04:00
473f7529e1 Fix backup queue permanently stuck when a dequeued backup completes during the patch (#10521)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 5s
Run the E2E test on kind / get-go-version (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
* Fix backup queue permanently stuck when a dequeued backup completes during the patch

Motivation: backupQueueReconciler patched a dequeued backup to ReadyToStart and
only called backupTracker.AddReadyToStart on the next line. If backupReconciler
picked up that patch and completed the backup (e.g. immediate FailedValidation
while a BackupStorageLocation is briefly unavailable) before the queue
controller reached that line, backupReconciler's Add + deferred Delete ran
first, and the later AddReadyToStart re-inserted a tracker key nothing would
ever delete again. backupTracker is in-memory and never reconciled against
actual Backup phases, so RunningCount() stayed stuck at concurrentBackups and
every later reconcile, including the periodic recheck, was refused at that
gate -- the queue stopped dequeuing permanently until the deployment restarted.

Approach: record the backup as ReadyToStart in the tracker before patching it,
and roll that back if the patch itself fails, so the tracker entry always
exists before the backup can become visible to any other reconciler. Also
folds in two related fixes: the concurrency-refusal log line is now Info
instead of Debug so a stuck queue is visible at the default log level, and the
queue-position renumbering loop's error log (which built a logrus.Entry via
log.WithError(errors.Wrapf(...)) but never called a terminal method on it, so
it never actually logged anything) now emits properly.

Validation: go build ./pkg/controller/..., go vet ./pkg/controller/..., and
gofmt -l on both changed files are all clean. golangci-lint run
./pkg/controller/... reports no findings. go mod tidy produces a zero diff to
go.mod/go.sum, matching this repo's verify-modules check. Mirrored this repo's
own hack/test.sh invocation for this package (-short -vet=... -skip TestAPIs)
and it passes; TestAPIs is a separate envtest suite that needs a local
kubebuilder etcd binary not installed on this machine and fails identically on
an unmodified checkout, so it is a pre-existing environment gap, not a
regression. Added TestBackupQueueReconcilerTrackerNotLeakedWhenBackupCompletesDuringPatch,
which uses a controller-runtime fake client with a Patch interceptor to
simulate a racing reconciler completing the backup right after the
ReadyToStart patch lands; it fails (RunningCount leaks to 1) against the
pre-fix ordering and passes (RunningCount returns to 0) against the fix.

Report: https://github.com/velero-io/velero/issues/10519
Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Assisted-by: claude-sonnet-5 (via Claude Code)

* Add changelog file for PR #10521

Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>

* Add test covering tracker rollback when ReadyToStart patch fails

Addresses review comment: verify backupTracker.RunningCount() returns
to 0 when the ReadyToStart patch itself errors, covering the Delete
rollback path alongside the existing race-condition regression test.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Co-authored-by: Tiger Kaovilai <tkaovila@redhat.com>
2026-09-16 20:34:20 +08:00
Wenkai Yin(尹文开)andGitHub 1010112c34 Update code to support namespace mapping when perform the in-place restore with block data mover (#10461)
* Update code to support namespace mapping when perform the in-place restore with block data mover

Update code to support namespace mapping when perform the in-p
lace restore with block data mover

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
2026-09-16 15:16:29 +08:00
872f903091 Add configurable tolerations for PodVolumeBackup and data mover pods (#9575)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Remove toleration whitelist for PodVolumeBackup and data mover pods

Instead of filtering tolerations through a hardcoded allowlist
(ThirdPartyTolerations), inherit all tolerations from the node-agent
daemonset for PodVolumeBackup/Restore and DataUpload/Download pods,
and from the Velero deployment for maintenance jobs.

This enables backups and restores on nodes with custom NoExecute taints,
which was previously impossible since only two specific toleration keys
were whitelisted.

Fixes #9476

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>

* Fix codespell: replace 'whitelist' with 'allowlist' in changelog

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>

* Implement deduplication of tolerations and add unit tests for the new function

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Merge node-agent-configmap tolerations with third-party allowlist

Add a `tolerations` field to the node-agent-configmap so operators can
declare hosting-pod tolerations explicitly, per blackpiglet's review
feedback that tolerations shouldn't be read from the DaemonSet alone.
These are merged with (and deduplicated against) DaemonSet tolerations
matching the existing third-party allowlist
(kubernetes.azure.com/scalesetpriority, CriticalAddonsOnly), restoring
that allowlist per the follow-up suggestion to keep inheriting it
alongside the new config option.

The toleration dedup helper is moved from pkg/exposer to
pkg/util/kube (exported as DeduplicateTolerations) so it can be
shared with pkg/nodeagent without an import cycle.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix testifylint finding in TestGetTolerations

golangci-lint v2.12.0 (pinned in pr-linter-check.yml) flagged the
shared assert.Equal after the if/else as require-error: use require
for the error assertion so each branch is self-contained, matching
the pattern used elsewhere in this file.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Document toleration merge priority in GetTolerations

Per blackpiglet's review feedback: clarify that configured tolerations
take priority over allowlisted daemonset tolerations because they're
appended first and DeduplicateTolerations keeps only the first
occurrence of each exact (Key, Operator, Value, Effect) combination.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-09-14 18:05:01 -04:00
Lyndon-Li a827f607b5 update fallbackFull in restore finalizer
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-11 16:59:51 +08:00
Lyndon-Li 13a630a15a Merge branch 'main' into add-fallback-full-to-volume-info 2026-09-11 16:19:27 +08:00
lyndon-liandGitHub 7e67f03796 Merge pull request #10513 from ywk253100/cli
Run the E2E test on kind / setup-test-matrix (push) Failing after 4s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
Update volume info in restore finalizing stage
2026-09-11 16:02:50 +08:00
Lyndon-Li 55c756ed0f data path support fallbackFull
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-10 16:12:22 +08:00
Wenkai Yin (尹文开) a3d744db6a Update volume info in restore finalizing stage
Update volume info in restore finalizing stage to record info from DataDownload result

Signed-off-by: Wenkai Yin (尹文开) <wenkai.yin@broadcom.com>
2026-09-10 15:04:54 +08:00
lyndon-liandGitHub 87b45ed7fd Merge pull request #10506 from Lyndon-Li/save-source-size-to-backup
Save source size to volume info
2026-09-10 14:29:48 +08:00
Lyndon-Li e0e600d715 save source size for PVB
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-09 16:09:08 +08:00
Lyndon-Li 6b549f35b3 Merge branch 'main' into save-source-size-to-backup 2026-09-09 15:50:06 +08:00
Lyndon-Li 44f09189c2 Merge branch 'main' into report-incremental-fallback 2026-09-09 14:19:59 +08:00
Xun Jiang/Bruce JiangandGitHub c7a93be95a Add MustIncludeAdditionalItemPVCs to help track BIA added PVC's PVB creation. (#10501)
* Add MustIncludeAdditionalItemPVCs structure in backup. It's used to track PVCs returned by BIA with mustIncluded annotaion and PVC is excluded from backup by global filter.
* Modfiy the volumeHelper interface to add a parameter function for ShouldPerformFSBackup.
* Modify to support fine-grained backup filters.
* Modify according to comments. Use a read-only interface to replace the parameter function.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-09 14:04:57 +08:00
Lyndon-Li 6c86921876 du support source size
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-04 19:49:12 +08:00
Lyndon-Li df709d39d6 report incremental fallback message
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-04 17:21:57 +08:00
Lyndon-Li 5146992b5b add UT for progress message
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-04 16:04:38 +08:00
Lyndon-Li 1048f26c20 controllers support message in progress
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-04 14:28:31 +08:00
Chlins ZhangandGitHub 5bcac16213 Merge pull request #10464 from chlins/fix/error-message-context
Add operation context to user-facing error messages
2026-09-03 10:34:21 +08:00
chlins e03ff894ff Add operation context to user-facing error messages
Prefix raw err.Error() strings surfaced in CR statuses and CLI stderr with the failed operation.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-09-02 15:38:36 +08:00
Xun Jiang/Bruce JiangandGitHub 51e3075e78 Merge pull request #10453 from opbot-xd/fix-ginkgo-v2-cleanup-10440
test: resolve remaining Ginkgo V2 and Gomega anti-patterns
2026-09-02 14:19:58 +08:00
Wenkai Yin(尹文开)andGitHub 07768e7b33 Add "IncrementalBytes" field to status of DataDownload and PVR to indicate data transferred by the incremental restore (#10421)
Add "IncrementalBytes" field to status of DataDownload and PVR to indicate data transferred by the incremental restore

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
2026-09-01 15:43:53 +08:00
opbot_xd a607892eb0 test: resolve remaining Ginkgo V2 and Gomega anti-patterns (#10440)
- Replaced Expect().Should() and Expect().ShouldNot() with .To() and .ToNot() across 12 files (Task 1).
- Replaced synchronously evaluated Eventually() with Expect() in server_status_request_controller_test.go (Task 2B).
- Extracted Skip() calls inside lazy callbacks into conditional checks using slices.Contains() in enable_api_group_extentions.go (Task 3).

Signed-off-by: opbot_xd <awasthikrishna23052005@gmail.com>
2026-09-01 06:43:04 +05:30
b7d83a6f2b Cherry pick the in-place restore implementation PRs from feature branch to main (#10415)
* Update CRDs and CLI to support in-place restore (#10038)

Update CRDs(Restore, DataDownload, PodVolumeRestore) and restore create CLI to support in-place restore

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

* Update Kopia(filesystem) uploader to support incremental and deleteExtraFile during restore (#10066)

Update Kopia(filesystem) uploader to support incremental and deleteExtraFile during restore

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

* Update Restore Exposer and PVC CSI to support in-place restore (#10104)

1. Update Restore Exposer to support exposing with existing PV for in-place restore
2. Update PVC CSI RIA to continue the restore process for in-place restore

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

* Update Block uploader to support increase restore (#10244)

Update Block uploader to support increase restore

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

* Update Exposer to recreate the target PV if the volume mode is different with the restore PVC (#10257)

Update Exposer to recreate the target PV if the volume mode is different with t
he restore PVC

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

* Preserve PVC selected-node annotation via carrier annotation for in-place restore

For in-place volume data restore, the existing PVC is deleted and
recreated. For StorageClasses with the WaitForFirstConsumer volume
binding mode, losing the volume.kubernetes.io/selected-node annotation
could let the scheduler place the recreated workload Pod in a different
zone than the original PV, leaving it stuck in ContainerCreating.

Instead of relying on RestoreItemAction execution order (the generic
PVC RIA unconditionally strips the selected-node annotation), the PVC
CSI RIA now captures the annotation from the existing PVC right before
deleting it and carries it on the target PVC via the Velero-internal
restore.velero.io/inplace-restore-selected-node annotation. The restore
engine translates the carrier back to the Kubernetes annotation after
all RestoreItemActions have run and always strips the carrier so it
never lands on the cluster.

This makes the behavior independent of RIA ordering: the Kubernetes
annotation is stripped by default on every path (including when the
target PVC does not exist and Velero falls back to provisioning a new
PVC), and preservation only happens when the CSI RIA explicitly
captured a value from the existing PVC.

Signed-off-by: chlins <chlins.zhang@gmail.com>

* Update the control path to make the in-place incremental restore with block data mover work E2E (#10410)

Update the control path to make the in-place incremental restore with block data mover work E2E

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>

---------

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
Signed-off-by: chlins <chlins.zhang@gmail.com>
Co-authored-by: chlins <chlins.zhang@gmail.com>
2026-08-26 10:21:32 -04:00
cc7b1dbaef Embed CRD manifests via go:embed instead of codegen (#10329)
config/crd/{v1,v2alpha1}/crds/crds.go were generated files that
gzip-compressed the CRD YAML bases into committed []byte literals via
hack/crd-gen, requiring `go generate` and a dedicated CI drift check
(hack/verify-generated-crd-code.sh). This made the files large,
unreviewable in diffs, and a frequent source of merge conflicts.

Replace the generated files with config/crd/{v1,v2alpha1}/crds.go
using `//go:embed bases/*.yaml` to embed the already-committed YAML
manifests directly, decoding them the same way at init. Since Go's
go:embed can't reach outside a file's own directory tree, the crds
package now lives alongside bases/ instead of in a bases-sibling
subdirectory; import paths in pkg/install and pkg/controller were
updated accordingly.

Drop hack/crd-gen and hack/verify-generated-crd-code.sh entirely, and
trim their references from update-3generated-crd-code.sh and the
codespell skip-list. No codegen step remains, so no drift is possible.

Fixes #10328

AI-Tool-Used: Claude Code
AI-Tool-Use-Level: Category 1 (High)
AI-Code-Category: Category 1 (Production)

Signed-off-by: lubronzhan <lubron.zhan@broadcom.com>
Co-authored-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-24 14:17:25 -04:00
09c1656df9 Skip DeleteSnapshot when ProviderSnapshotID is empty (#9795)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
When CreateSnapshot fails (e.g. quota limit), the snapshot is recorded
with an empty ProviderSnapshotID. During backup deletion, velero was
calling DeleteSnapshot("") which produces unnecessary 404 API calls.

Skip the DeleteSnapshot call when ProviderSnapshotID is empty and log
a warning instead.

Fixes #9429

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-08-24 10:27:34 -04:00
1e26cf7ca0 fix(restore_finalizer): bound WaitRestoreExecHook poll with resourceT… (#10280)
* fix: reuse DefaultResourceTimeout from server config for hook wait

Signed-off-by: Nitish Malang <71919457+nitishmalang@users.noreply.github.com>

* Add changelog for PR 10280

Signed-off-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>

---------

Signed-off-by: Nitish Malang <71919457+nitishmalang@users.noreply.github.com>
Signed-off-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>
Co-authored-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>
2026-08-21 15:15:30 -04:00
3cd6c2e533 Skip signing a download URL when no artifacts can exist yet (#10252)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 9s
Run the E2E test on kind / get-go-version (push) Failing after 11s
push.yml / extract (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Skip signing a download URL when no artifacts can exist yet

Reported in #10232: a DownloadRequest for a backup that never ran still
reaches Processed with a signed URL, and fetching it returns 404.

The controller already has the backup, and the restore for restore
targets, in hand before it signs, so checking the phase costs no extra
call to the object store.

The check is deliberately narrow. It refuses only the pre-execution
phases, where nothing has been written for any target kind: New, Queued,
ReadyToStart and FailedValidation for backups, New and FailedValidation
for restores. InProgress onwards may hold a partial log or other
artifacts, and Deleting may still hold all of them, so those keep the
behaviour callers have today.

That matters because velero backup download has no client side phase
check of its own, unlike backup logs and restore logs. Reusing the
allowlist from pkg/cmd/cli/backup/logs.go would have changed what
backup download can fetch; this does not.

A backup with an empty phase is left alone as well, since that state is
transient and the caller can retry.

Refs #10232

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Derive the phase coverage test from the generated CRDs

The previous test built a slice of phases by hand and asserted its own
length, so it passed no matter what the API did. Adding a fourteenth
backup phase would not have failed it.

This reads the status.phase enum out of the generated CRDs, via the
exported v1crds.CRDs that pkg/install already uses. The enum comes from
the same kubebuilder markers as the Go constants, so a phase added to
the API fails here until it is classified.

Verified by removing Deleting from the expectations, which now fails with
'BackupPhase "Deleting" is served by the CRD but not classified'.

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Use US spelling in comments to satisfy the misspell linter

golangci-lint runs misspell, which flags behaviour as a misspelling of
behavior. Comments only, no functional change.

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Set a Failed phase with a reason when the guard refuses to sign

The guard added in the previous commit left the request at New with no URL, so
the CLI polled until its own timeout and then reported that the backup storage
location may be unavailable. The BSL is fine; the backup never ran.

DownloadRequestPhase gains Failed and DownloadRequestStatus gains Message. The
controller sets both where it refuses, and the CLI stops as soon as it sees the
phase and surfaces the message instead of its generic timeout error.

Adding an enum value is additive, per the direction on the PR discussion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: saral <ilovegojo2580@gmail.com>

---------

Signed-off-by: saral <ilovegojo2580@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 17:26:21 +08:00
chlins 9afa3964ba Only sync finished backups from object storage
Backup metadata with an empty or New phase was synced into the cluster as a pending backup, which the queue controller then ran as if it were newly requested. Hooks are dropped as well, since a synced backup never executes them.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-08-20 10:37:41 +08:00
Shubham Pampattiwar 2a920ab946 Add RBAC for secrets/configmaps to datamover controllers
The DataUpload and DataDownload controllers now copy and delete
namespace-scoped secrets/configmaps for backup/restore PVC provisioning.
Add the corresponding kubebuilder RBAC markers (get;list;create;delete
on secrets and configmaps) and regenerate the ClusterRole.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 986350a6e5 Move secret/configmap copy from controller to CSI snapshot exposer
Move the secret and configmap copy logic from the DataUpload controller
into the CSI snapshot exposer's Expose() method. This keeps all
CSI-specific logic in the exposer and maintains symmetry with CleanUp()
which already handles the cleanup of copied resources.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar c15cf084e3 Add configmap copy support and move secret copy after accept
- Add ConfigMapNames field to BackupPVC config for copying tenant
  configmaps (e.g., ceph-csi-kms-config with Vault connection overrides)
- Add CopyConfigMap, DeleteConfigMapIfAny, DeleteConfigMapsWithLabel
  utilities mirroring the secret copy functions
- Move secret/configmap copy after acceptDataUpload() so only the
  accepting node handles it, avoiding multi-node contest
- Clean up copied configmaps in CleanUp() alongside secrets

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 2a44048024 Copy secrets before Expose in DataUpload controller
Copy configured secrets from the source namespace to the Velero
namespace in the New phase of the DataUpload reconcile loop, before
calling Expose(). This is done in the controller rather than the
exposer because Expose() errors are non-retryable (marked as permanent
failure), while the controller can requeue on collision.

On secret collision (same name, different data from another
DataUpload), the controller requeues with a 5s delay, matching the
existing pattern used for VGDP constraint checking.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
lyndon-liandGitHub da5bee7097 Use thread safe map for cancel recorder (#10255)
* use thread safe map for cancel recorder

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>

* use atomic load and store

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>

---------

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-17 16:55:01 +08:00
harshit sainiandGitHub 11a071637b Use k8s.io/api well-known label constants instead of hardcoded strings (#10279)
* refactor: use k8s.io/api well-known label constants

Several well-known Kubernetes label strings were hardcoded across the
codebase instead of using the constants already exported by
k8s.io/api/core/v1, which is an existing dependency:

  "kubernetes.io/hostname"        -> corev1api.LabelHostname
  "kubernetes.io/os"              -> corev1api.LabelOSStable
  "topology.kubernetes.io/zone"   -> corev1api.LabelTopologyZone

The local kube.NodeOSLabel and zoneLabel consts, which duplicated the
upstream values verbatim, are now defined in terms of the upstream
constants rather than repeating the literal. Both are kept: NodeOSLabel
is exported and referenced from four packages alongside NodeOSLinux and
NodeOSWindows, which have no upstream equivalent, and zoneLabel sits
beside the deprecated-label fallback it is compared against.

No functional change - every replacement is a constant with an identical
value.

Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>

* Add changelog for #10279

Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>

* Cover the selected-node path in createRestorePod

TestCreateRestorePod only exercised selectedNode == "", so the branch
that pins the restore pod to a node was never executed. Add a case with
a selected node and assert the resulting pod carries the hostname label
in its node selector.

Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>

* Also use constants for the arch and deprecated zone labels

Extends the same replacement to the two remaining well-known labels
raised on the issue:

  "kubernetes.io/arch"                      -> corev1api.LabelArchStable
  "failure-domain.beta.kubernetes.io/zone"  -> corev1api.LabelFailureDomainBetaZone

zoneLabelDeprecated in item_backupper.go was the last local const still
repeating a literal that upstream already exports, so the zone pair now
reads consistently against k8s.io/api. The deprecation note upstream
applies to the label itself, not the constant; Velero reads that label
deliberately as the fallback for PVs created before the topology labels
existed.

Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>

---------

Signed-off-by: Harshit saini <harshitsaini1188@gmail.com>
2026-08-17 13:37:44 +08:00
lyndon-liandGitHub ae6f40c286 Merge pull request #10273 from Lyndon-Li/fix-pvr-regression
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
e2e-test-kind.yaml / extract (push) Failing after 5s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 8s
Main CI / get-go-version (push) Failing after 9s
Main CI / Build (push) Skipped
Fix PVR regression
2026-08-15 10:36:46 +08:00
Daniel JiangandGitHub 293f6f6a63 Harden "patchDynamicPVWithVolumeInfo" (#10271)
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
This commit hardens the func "patchDynamicPVWithVolumeInfo":
1. Add nil checks for storageClass and the attributes.
2. Remove the double reported errors.

Signed-off-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-14 19:26:15 +08:00
Lyndon-Li 913ec9f325 fix PVR regression
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-14 18:11:32 +08:00
V PrajwalandGitHub 194404971a Fix schedule reconciler aliasing server-wide skipImmediately default (#10242)
e2e-test-kind.yaml / extract (push) Failing after 5s
Run the E2E test on kind / get-go-version (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
When a Schedule has no explicit spec.skipImmediately, the reconciler
assigned &c.skipImmediately directly into the Schedule's spec pointer.
The subsequent write-through-pointer (*ptr = false) mutated the
reconciler's own shared field, silently disabling
--schedule-skip-immediately for every schedule reconciled afterward
for the life of the process.

Fix: copy the value into a fresh bool before taking its address.

Adds TestReconcileDoesNotCorruptReconcilerSkipImmediately, which
reconciles two schedules against one reconciler instance and asserts
the shared default is preserved.

Signed-off-by: Prajwal <percy38621@gmail.com>
2026-08-14 16:18:00 +08:00
Krishna AwasthiandGitHub 41b95b5919 Refactor: Replace context.TODO() with properly plumbed contexts in CSI actions (#10247)
Signed-off-by: opbot_xd <awasthikrishna23052005@gmail.com>
2026-08-14 16:17:29 +08:00
AftAb-25andGitHub 9a1d2e6eb0 Fix switch case ordering bug in filterBackupOwnerReferences (#10161)
* Fix switch case ordering in filterBackupOwnerReferences (Issue #10160)

When client.Get returns a transient (non-NotFound) error, the previous
case ordering caused the UID mismatch case to fire against a zero-value
struct, silently dropping the owner reference and logging a misleading
'mismatched UIDs' warning instead of the intended error log.

Fix: move the general error handler before the UID mismatch check so
it is evaluated while err is still relevant. The UID check now only
runs when err == nil (i.e. the Schedule was successfully fetched).

Also add a test case that injects a transient Get error via the fake
client interceptor to verify the owner reference is preserved.

Signed-off-by: aftab <aftab123215@gmail.com>

* Add changelog for #10160

Signed-off-by: aftab <aftab123215@gmail.com>

---------

Signed-off-by: aftab <aftab123215@gmail.com>
2026-08-12 15:12:58 +08:00
513e93ff4b fix: correct typos in log messages and status strings (#10192)
- Fix 'dataudownload' typo in DataDownload warning log message
  (data_download_controller.go:696)
- Fix 'datadownlad' misspelled structured log field key to 'datadownload'
  (data_download_controller.go:700) - this caused the log field to be
  unqueryable by the correct key name
- Fix 'retrieveable' -> 'retrievable' in BackupRepository maintenance
  status messages (maintenance.go:354, 417)
- Update corresponding test assertion to match corrected string
  (maintenance_test.go:792)

Signed-off-by: shellyco-code <shellyco-code@users.noreply.github.com>
Co-authored-by: shellyco-code <shellyco-code@users.noreply.github.com>
2026-08-10 19:40:02 +00:00
AftAb-25andGitHub ca72c2e7e2 Fix missing gcFailureBSLUnavailable label during garbage collection (#10154)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
* Fix gcFailureBSLUnavailable label not applied (Issue #10153)

Signed-off-by: aftab <aftab123215@gmail.com>

* Fix linter error: use require.NoError before checking label

Signed-off-by: aftab <aftab123215@gmail.com>

---------

Signed-off-by: aftab <aftab123215@gmail.com>
2026-08-04 13:06:45 -04:00
Xun Jiang 11545ee63c Fix logs, CRD, and GetDataMover for CBT features.
Modify the logs.
Modify the CRD's data mover's comment.
Modify the resource policy's GetDataMover for default data mover case.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-08-04 06:25:05 +00:00
Shubham Pampattiwar bd95633967 Add skip log and improve test coverage
- Log when SkipDefaultResourceModifier skips the default modifier
- Add test for unsupported ResourceModifier Kind (warns, does not
  apply default)
- Add test for default ConfigMap with invalid rules (validation
  failure is non-fatal)
- loadResourceModifierConfigMap now at 100% coverage

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-03 13:19:48 -07:00