Commit Graph
2107 Commits
Author SHA1 Message Date
57bddf7f23 Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds (#9646)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 9s
Run the E2E test on kind / get-go-version (push) Failing after 10s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds

Previously, the backup finalizer controller set backup.Status.Phase to
Completed/PartiallyFailed in-memory BEFORE calling PutBackupMetadata and
PutBackupContents. When these uploads failed (e.g., due to object lock
or immutability), the deferred patch function still wrote the terminal
phase to the Kubernetes API server, preventing the controller from
retrying the upload on the next reconcile.

This fix moves the phase assignment to AFTER both uploads succeed. A
DeepCopy of the backup is used to encode the JSON with the final phase
for object storage, while the in-memory backup object retains the
Finalizing phase until uploads complete.

Caveats:
- CompletionTimestamp is now captured before upload but only committed to
  the API server after upload succeeds. On retry after a transient
  failure, a new timestamp is generated, so the completion time reflects
  when the upload finally succeeded rather than when finalization
  processing completed.
- Metrics (RegisterBackupSuccess/RegisterBackupPartialFailure) are now
  recorded after uploads succeed, so they accurately reflect only fully
  persisted backups.
- The metadata uploaded to object storage contains the final phase and
  completion timestamp via DeepCopy, so storage state is correct even
  before the API server is patched.

Fixes #9645

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Add changelog for #9646

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix testifylint: use require.Error instead of assert.Error

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Address review feedback on backup-finalizer fix

- Add default guard for unhandled phase values in finalPhase switch
- Add retry with DefaultBackoff for PutBackupMetadata per reviewer request
- Replace brittle framework.BackupItemActionResolverV2{} mock with mock.Anything
- Add FinalizingPartiallyFailed test case for PutBackupContents failure

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Use bounded, object-storage-tuned backoff for backup-finalizer uploads

retry.DefaultBackoff is tuned for API server optimistic-concurrency
conflicts (4 steps, ~1.25s total) and gives up far too quickly for
object storage calls, which can see longer transient outages or
throttling (review feedback from blackpiglet). Replace it with a
dedicated, bounded backoff (1s base, 2x factor, 5 steps, ~31s total)
applied to both PutBackupMetadata and PutBackupContents.

Being bounded (rather than retrying forever) means a persistent
failure, e.g. an object-lock/immutability policy denying every write,
surfaces as an error within a bounded time instead of hanging the
reconcile indefinitely; controller-runtime requeues on error, so
retries continue across reconciles (review feedback from priyansh17).

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix PutBackupMetadata retry to re-read backupJSON each attempt

backupJSON is a bytes.Buffer, so passing it directly to
PutBackupMetadata drains it on the first read attempt. A retry after
a transient failure would then upload empty content instead of the
backup metadata. Wrap it in bytes.NewReader(backupJSON.Bytes()) inside
the retry closure so every attempt gets a fresh reader.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-09-22 20:06:46 -04:00
Xun Jiang/Bruce JiangandGitHub 2aea706117 Merge pull request #10441 from Ralthos/bsl-printcolumns
Run the E2E test on kind / setup-test-matrix (push) Failing after 5s
e2e-test-kind.yaml / extract (push) Failing after 7s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Add printer columns for BackupStorageLocation provider and access mode
2026-09-22 17:23:07 +08:00
Adam ZhangandGitHub c3fe97745a Merge pull request #10549 from shoemoney/fix/schedule-create-drops-annotations
Fix schedule create dropping annotations
2026-09-22 16:40:28 +08:00
Chlins ZhangandGitHub 3793d9ab4a Fail the in-place restore pre-flight check when the backed-up pod already exists on the file system restore path (#10550)
PodVolumeRestores are only created for pods that Velero creates, so when
the pod already exists in the cluster the PVC-not-in-use pre-flight check
never runs and the volume data restore is skipped silently, while the
existing pod keeps consuming the PVC. Report an explicit pre-flight error
for such pods, aligned with the PVC CSI RIA behavior.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-09-22 11:53:35 +08:00
Adam ZhangandGitHub 36935902b9 Merge pull request #10536 from abhayrajjais01/fix/kube-node-context-propagation
fix(kube): propagate context to node client in GetNodeOS
2026-09-22 11:25:25 +08:00
lyndon-liandGitHub b66f12024c Merge pull request #10554 from pujitha24/auto/issue-10551
Propagate caller context in repository manager Forget/BatchForget
2026-09-22 11:25:05 +08:00
lyndon-liandGitHub 5f14d304c1 Merge pull request #10275 from kaovilai/namespace-selection-by-label
Implement namespace selection by label in resource policy
2026-09-22 09:22:38 +08:00
Pujitha Paladugu 8a2d9db687 Add changelog for PR 10554
CI's changelog check requires a changelogs/unreleased/<PR#>-<login>
file; this PR didn't have one since the PR number wasn't known until
after it was opened.

Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
2026-09-21 07:37:08 -07:00
Abhayraj Jaiswal b9a662d672 fix(kube): propagate context to node client in GetNodeOS
Signed-off-by: Abhayraj Jaiswal <abhayraj916146@gmail.com>
2026-09-21 05:29:26 +00:00
Xun Jiang a71bc4befc Reset the PVC binding information in pvc_action.go when it has referenced PVB.
Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-21 10:54:59 +08:00
Jeremy Schoemaker e41ba0cafa Add changelog for schedule create annotations fix
Signed-off-by: Jeremy Schoemaker <jeremy@shoemoney.com>
2026-09-19 17:34:06 -05:00
Tiger KaovilaiandClaude Sonnet 5 40af5efdd0 Implement namespace selection by label in resource policy
Add includedNamespacesByLabel, excludedNamespacesByLabel, and
labelSelectorLogic to IncludeExcludePolicy in the ResourcePolicy
ConfigMap (realizes design in velero-io/velero#9772), letting a backup
select or exclude namespaces by label instead of (or in addition to)
name/wildcard.

The backup controller resolves label selectors against the live
namespace list once per backup, merges the results into
spec.includedNamespaces/excludedNamespaces, then proceeds through the
existing name-based filtering unchanged. A defaulted "*" include list
is replaced by the resolved set; an explicitly-configured include list
(including an explicit "*") is unioned with it instead, and stays
canonical rather than widening. Namespaces matching an exclude
selector are always subtracted from the merged includes, regardless of
how the includes were populated.

Because Velero's namespace-includes/excludes model requires at least
one name (an empty list means "match everything"), a selector that
resolves to zero namespaces is represented with a sentinel glob
pattern ("[-]*") guaranteed to match no real namespace, rather than an
empty list that would silently fall back to including/excluding
everything.

labelSelectorLogic ("AND"/"OR", case-insensitive) controls whether
multiple included/excluded label selectors are combined by
intersection or union; it is validated up front, including inside
ResolveNamespacesByLabel itself, so an invalid value fails fast instead
of silently falling through to OR semantics.

Namespace-selection-by-label and resource-selection-by-label act as
independent axes and do not affect each other, matching the design
discussion in #9772.

Known limitations:
- Selectors are evaluated once per backup against the namespace list
  at that point in time; namespaces created or relabeled mid-backup
  are not picked up.
- Backup-only for now; restore-side namespace mapping is unaffected.

Testing:
- Unit coverage in internal/resourcepolicies for validation, selector
  resolution (including AND/OR logic, case-insensitivity, and
  malformed-selector/invalid-logic error paths), and the no-match
  sentinel.
- Unit coverage in pkg/controller for the merge logic between resolved
  label selections and explicit/defaulted includes and excludes.
- End-to-end coverage in pkg/backup exercising the full backup
  pipeline with label-selected namespaces, including the
  velero.io/exclude-from-backup hard-exclusion interaction and the
  zero-match/fully-excluded sentinel path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
2026-09-18 23:04:55 -04:00
Xun Jiang 5afc20d03e Support skipped PVC in VolumeInfos.
* Modify according to comments.
* Rename the pvSkipTracker and related fields to indicate both PVC and PV are supported.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-18 16:51:00 +08:00
lyndon-liandGitHub 60163e0827 Merge pull request #10537 from Lyndon-Li/doc-for-block-data-mover
Run the E2E test on kind / setup-test-matrix (push) Failing after 2s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Add doc for block data mover
2026-09-18 11:08:49 +08:00
d074047b06 Fix: Strip ServiceAccount token volume mounts from EphemeralContainers in PodAction (#10349)
Run the E2E test on kind / setup-test-matrix (push) Failing after 2s
e2e-test-kind.yaml / extract (push) Failing after 7s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 5s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Signed-off-by: opbot_xd <awasthikrishna23052005@gmail.com>
Co-authored-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-09-17 14:16:39 -04:00
fa01151214 docs: Add documentation about --write-sparse-files flag for disk space issues during restore (#9221)
- Added cross-references in troubleshooting.md and file-system-backup.md to the Write Sparse files documentation
- Documented how to use --write-sparse-files flag when restores fail due to disk space constraints
- Clarified important limitation: only works if PV had sparse files during backup that would free up space
- Removed enhanced error messages and code changes (documentation-only approach per feedback)

This addresses issue #2812 by providing clear guidance to users when restores fail due to disk space constraints.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-17 15:31:29 +08:00
Lyndon-Li e677428a7a add doc for block data mover
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-17 11:13:18 +08:00
f40112aedf Fix schedule reconciler to compare SkipImmediately and LastSkipped by value (#10490)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Fix schedule reconciler to compare SkipImmediately and LastSkipped by value

The reconciler compared spec.skipImmediately (*bool) and
status.lastSkipped (*metav1.Time) between the live object and its
DeepCopy by pointer, so both checks always reported a change once the
fields were set, and every reconcile issued a redundant Patch. Compare
by value instead (ptr.Equal / equality.Semantic.DeepEqual) and add a
regression test pinning that the skip-once flip persists with exactly
one patch, later reconciles of persisted state patch zero times, and
an explicit false is a no-op.

Signed-off-by: HeonJe LEE <lhjnano@gmail.com>

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>

---------

Signed-off-by: HeonJe LEE <lhjnano@gmail.com>
Signed-off-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>
Co-authored-by: lyndon-li <98304688+Lyndon-Li@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-16 17:00:29 -04:00
Chlins ZhangandGitHub e164dc5984 Identify the backed-up volume by CSI volume handle in the in-place restore pre-flight check (#10530)
The pre-flight check that verifies the existing PVC is still bound to the
backed-up volume compared PV names. A block data mover restore of a file
system volume recreates the PV under a new name (the volumeMode field is
immutable), so a second in-place restore of the same workload failed the
check even though the PVC was bound to the very same volume.

Record the CSI volume handle in the backup volume info (PVInfo) and
compare handles when both the backup and the bound PV record one; the
PV name remains the fallback for non-CSI volumes and for backups taken
before the handle was recorded. The handle reaches the PVC CSI RIA
through the same carrier annotation mechanism as the source size.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-09-16 16:55:27 -04:00
473f7529e1 Fix backup queue permanently stuck when a dequeued backup completes during the patch (#10521)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 5s
Run the E2E test on kind / get-go-version (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
* Fix backup queue permanently stuck when a dequeued backup completes during the patch

Motivation: backupQueueReconciler patched a dequeued backup to ReadyToStart and
only called backupTracker.AddReadyToStart on the next line. If backupReconciler
picked up that patch and completed the backup (e.g. immediate FailedValidation
while a BackupStorageLocation is briefly unavailable) before the queue
controller reached that line, backupReconciler's Add + deferred Delete ran
first, and the later AddReadyToStart re-inserted a tracker key nothing would
ever delete again. backupTracker is in-memory and never reconciled against
actual Backup phases, so RunningCount() stayed stuck at concurrentBackups and
every later reconcile, including the periodic recheck, was refused at that
gate -- the queue stopped dequeuing permanently until the deployment restarted.

Approach: record the backup as ReadyToStart in the tracker before patching it,
and roll that back if the patch itself fails, so the tracker entry always
exists before the backup can become visible to any other reconciler. Also
folds in two related fixes: the concurrency-refusal log line is now Info
instead of Debug so a stuck queue is visible at the default log level, and the
queue-position renumbering loop's error log (which built a logrus.Entry via
log.WithError(errors.Wrapf(...)) but never called a terminal method on it, so
it never actually logged anything) now emits properly.

Validation: go build ./pkg/controller/..., go vet ./pkg/controller/..., and
gofmt -l on both changed files are all clean. golangci-lint run
./pkg/controller/... reports no findings. go mod tidy produces a zero diff to
go.mod/go.sum, matching this repo's verify-modules check. Mirrored this repo's
own hack/test.sh invocation for this package (-short -vet=... -skip TestAPIs)
and it passes; TestAPIs is a separate envtest suite that needs a local
kubebuilder etcd binary not installed on this machine and fails identically on
an unmodified checkout, so it is a pre-existing environment gap, not a
regression. Added TestBackupQueueReconcilerTrackerNotLeakedWhenBackupCompletesDuringPatch,
which uses a controller-runtime fake client with a Patch interceptor to
simulate a racing reconciler completing the backup right after the
ReadyToStart patch lands; it fails (RunningCount leaks to 1) against the
pre-fix ordering and passes (RunningCount returns to 0) against the fix.

Report: https://github.com/velero-io/velero/issues/10519
Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Assisted-by: claude-sonnet-5 (via Claude Code)

* Add changelog file for PR #10521

Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>

* Add test covering tracker rollback when ReadyToStart patch fails

Addresses review comment: verify backupTracker.RunningCount() returns
to 0 when the ReadyToStart patch itself errors, covering the Delete
rollback path alongside the existing race-condition regression test.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Co-authored-by: Tiger Kaovilai <tkaovila@redhat.com>
2026-09-16 20:34:20 +08:00
Wenkai Yin(尹文开)andGitHub 1010112c34 Update code to support namespace mapping when perform the in-place restore with block data mover (#10461)
* Update code to support namespace mapping when perform the in-place restore with block data mover

Update code to support namespace mapping when perform the in-p
lace restore with block data mover

Signed-off-by: Wenkai Yin(尹文开) <yinw@vmware.com>
2026-09-16 15:16:29 +08:00
Xun Jiang/Bruce JiangandGitHub 7c860515f3 Merge pull request #10507 from mmustafasenoglu/fix/install-wait-default
Run the E2E test on kind / setup-test-matrix (push) Failing after 2s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
install: clarify --wait flag behavior, default false
2026-09-16 11:04:48 +08:00
214b63f6f7 Do not assume a port name is a string when clearing node ports (#10132)
Run the E2E test on kind / setup-test-matrix (push) Failing after 2s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Do not assume a port name is a string when clearing node ports

deleteNodePorts reads the last-applied-configuration annotation, which
is free-form JSON controlled by whoever produced the backup, and cast
p["name"] to string without checking. A port whose name is a number
crashed the restore of that Service with an interface conversion panic.
Every sibling in the same loop already uses the comma-ok form.

Signed-off-by: Arpit Jain <arpitjain099@gmail.com>

* Add changelog file

Signed-off-by: Arpit Jain <arpitjain099@gmail.com>

* Convert name to string by Sprint.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>

---------

Signed-off-by: Arpit Jain <arpitjain099@gmail.com>
Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
Co-authored-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-15 09:50:10 -04:00
Lyndon-Li 4acad69133 Merge branch 'main' into fix-issue-10429 2026-09-15 17:39:47 +08:00
Lyndon-Li a93e4769bc issue 10429: sync the calls to IsConstrained
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-15 17:38:30 +08:00
Mustafa Senoglu c1d6ff9c1d install: clarify --wait flag behavior, default false
Keep the existing opt-in behavior (consistent with velero backup create and velero restore create, where waiting is disabled by default and enabled with --wait), and clarify it in the command help text. Add tests pinning the --wait default (false) and flag parsing.

Signed-off-by: Mustafa Senoglu <mmustafasenoglu0@gmail.com>
2026-09-15 11:14:29 +03:00
Chlins ZhangandGitHub 53a6c37d2f Merge pull request #10449 from PratikMane0112/fix/snapshot-location-label-selector
Fix snapshot-location get --selector flag to actually filter VolumeSnapshotLocations by label
2026-09-15 16:04:34 +08:00
lyndon-liandGitHub c4f73126e8 Merge pull request #10526 from Daniel-1600/fix/schedule-create-backup-type
Fix schedule create dropping backup type
2026-09-15 14:33:50 +08:00
872f903091 Add configurable tolerations for PodVolumeBackup and data mover pods (#9575)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Remove toleration whitelist for PodVolumeBackup and data mover pods

Instead of filtering tolerations through a hardcoded allowlist
(ThirdPartyTolerations), inherit all tolerations from the node-agent
daemonset for PodVolumeBackup/Restore and DataUpload/Download pods,
and from the Velero deployment for maintenance jobs.

This enables backups and restores on nodes with custom NoExecute taints,
which was previously impossible since only two specific toleration keys
were whitelisted.

Fixes #9476

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>

* Fix codespell: replace 'whitelist' with 'allowlist' in changelog

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>

* Implement deduplication of tolerations and add unit tests for the new function

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Merge node-agent-configmap tolerations with third-party allowlist

Add a `tolerations` field to the node-agent-configmap so operators can
declare hosting-pod tolerations explicitly, per blackpiglet's review
feedback that tolerations shouldn't be read from the DaemonSet alone.
These are merged with (and deduplicated against) DaemonSet tolerations
matching the existing third-party allowlist
(kubernetes.azure.com/scalesetpriority, CriticalAddonsOnly), restoring
that allowlist per the follow-up suggestion to keep inheriting it
alongside the new config option.

The toleration dedup helper is moved from pkg/exposer to
pkg/util/kube (exported as DeduplicateTolerations) so it can be
shared with pkg/nodeagent without an import cycle.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Fix testifylint finding in TestGetTolerations

golangci-lint v2.12.0 (pinned in pr-linter-check.yml) flagged the
shared assert.Equal after the if/else as require-error: use require
for the error assertion so each branch is self-contained, matching
the pattern used elsewhere in this file.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Document toleration merge priority in GetTolerations

Per blackpiglet's review feedback: clarify that configured tolerations
take priority over allowlisted daemonset tolerations because they're
appended first and DeduplicateTolerations keeps only the first
occurrence of each exact (Key, Operator, Value, Effect) combination.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-09-14 18:05:01 -04:00
Chlins ZhangandGitHub 2e38432244 Merge pull request #10512 from chlins/feat/inplace-preflight-check3
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Failing after 2s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
Add in-place restore pre-flight check: PVC must be large enough for the backed-up data
2026-09-14 13:57:04 +08:00
lyndon-liandGitHub 16e79d87a3 Merge pull request #10520 from krishhna24/docs-vgs-class-example
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 5s
Run the E2E test on kind / get-go-version (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
docs: fix VolumeGroupSnapshotClass example
2026-09-14 11:23:00 +08:00
lyndon-liandGitHub 8c9236cb5b Merge pull request #10517 from Lyndon-Li/add-fallback-full-to-volume-info
Add fallback full to volume info
2026-09-14 11:07:41 +08:00
Daniel Mungai b1c1c145c2 Add changelog for #10526
Signed-off-by: Daniel Mungai <chegedan699@gmail.com>
2026-09-12 11:41:04 +03:00
KrishhnaTandGitHub 4e481fb7c2 test: use the Kind constant instead of the string literal (#10524)
e2e-test-kind.yaml / extract (push) Failing after 7s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
* test: use the Kind constant instead of the string literal

test/types.go defines `const Kind = "kind"` and most of the suite compares
against it, but three sites still use the bare string. deletion.go is
inconsistent with itself: the BeforeEach skip uses Kind while the one in
runBackupDeletionTests uses "kind", and its skip message hardcodes the
provider name where the other formats it.

namespace-mapping.go dot-imports test/e2e/test rather than test, so the
constant was not in scope there. Import the test package by name, as
test/e2e/migration/migration.go already does alongside its framework
import, and reference test.Kind.

No behavioural change: the constant's value is the string being replaced.

Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>

* Add changelog for #10524

Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>

---------

Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>
2026-09-11 14:21:41 -07:00
chlins bb01691e6b Add in-place restore pre-flight check: PVC must be large enough for the source volume
Compare the existing PVC's capacity against the source volume size
recorded in the backup volume info (#10506) before any side effect and
skip the volume when it is too small, so the restore fails early instead
of running out of space midway. For the block data mover the source size
is the device size; for the file system data movers it is the logical
size of the backed-up files, a lower bound since file system metadata is
not accounted for.

The file system path reads the size from the volume info already carried
in RestoreData. The PVC CSI RIA has no access to the volume info, so the
restore engine carries the size on the PVC item through a Velero-internal
annotation, the same mechanism as the selected-node carrier; both carrier
annotations are stripped before the item is created in the cluster.

The check is skipped when the source size is unknown (backups taken
before it was recorded) or the PVC's capacity is not reported.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-09-11 17:10:14 +08:00
Lyndon-Li 13a630a15a Merge branch 'main' into add-fallback-full-to-volume-info 2026-09-11 16:19:27 +08:00
lyndon-liandGitHub 7e67f03796 Merge pull request #10513 from ywk253100/cli
Run the E2E test on kind / setup-test-matrix (push) Failing after 4s
Scorecard supply-chain security / Scorecard analysis (push) Skipped
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
Update volume info in restore finalizing stage
2026-09-11 16:02:50 +08:00
lyndon-liandGitHub 34a9f500cc let uploader to control fallback centrally (#10523)
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-11 00:34:13 -04:00
krishhna24 d4f6eee899 Add changelog for #10520
Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>
2026-09-10 18:39:04 +05:30
Lyndon-Li 6b3e7ef5dd add fallback to volume info and backup/restore describe
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-10 18:22:35 +08:00
Wenkai Yin (尹文开) a3d744db6a Update volume info in restore finalizing stage
Update volume info in restore finalizing stage to record info from DataDownload result

Signed-off-by: Wenkai Yin (尹文开) <wenkai.yin@broadcom.com>
2026-09-10 15:04:54 +08:00
lyndon-liandGitHub 87b45ed7fd Merge pull request #10506 from Lyndon-Li/save-source-size-to-backup
Save source size to volume info
2026-09-10 14:29:48 +08:00
lyndon-liandGitHub e8af012ac5 Merge pull request #10479 from Lyndon-Li/report-incremental-fallback
Report incremental fallback message
2026-09-10 14:29:19 +08:00
Max Freedom PollardandGitHub 4c007c0af4 Scope schedule and repo CLI list calls to the Velero namespace (#10482)
Run the E2E test on kind / setup-test-matrix (push) Failing after 4s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
Scorecard supply-chain security / Scorecard analysis (push) Skipped
* Scope schedule and repo CLI list calls to the Velero namespace

velero schedule get, velero schedule describe, velero schedule
pause/unpause and velero repo get built a ctrlclient.ListOptions with a
LabelSelector but no Namespace, so the list ran across every namespace in
the cluster. Their single-name paths in the same functions already scope
to f.Namespace(), and every sibling command (backup get, restore get,
backup describe, restore describe, snapshot-location get, schedule
delete) passes Namespace too, so the omission was an oversight rather
than intent.

The read commands print another installation's Schedules and
BackupRepositories. runPause is worse: velero schedule pause --all and
velero schedule unpause --all fetch Schedules from every namespace and
then write Spec.Paused on each, so pausing one installation's schedules
pauses every other installation's schedules as well.

Add Namespace: f.Namespace() to the four List calls:

  pkg/cmd/cli/schedule/get.go:61
  pkg/cmd/cli/schedule/describe.go:59
  pkg/cmd/cli/schedule/pause.go:114
  pkg/cmd/cli/repo/get.go:61

Neither pkg/cmd/cli/schedule nor pkg/cmd/cli/repo had any tests, so the
regression tests are new files. Each seeds a fake client with one object
in the Velero namespace and one in another-velero, and asserts the second
is neither listed, described, nor paused.

Signed-off-by: Max Freedom Pollard <272618364+MaxFreedomPollard@users.noreply.github.com>

* Add changelog for PR 10482

Signed-off-by: Max Freedom Pollard <272618364+MaxFreedomPollard@users.noreply.github.com>

---------

Signed-off-by: Max Freedom Pollard <272618364+MaxFreedomPollard@users.noreply.github.com>
2026-09-09 16:52:11 -04:00
HeonJe LeeandGitHub cbd9059f80 Fix user version priorities parsing to handle CRLF line endings (#10496)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 8s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 5s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
formatUserPriorities stripped only spaces, so an enableapigroupversions
ConfigMap written with CRLF line endings (common for files edited on
Windows) kept a trailing carriage return in each version string, e.g.
"v2beta1\r". versionsContain compares versions with equality, so the
user priority never matched and was silently ignored. Trim the trailing
carriage return so stored versions match again.

Signed-off-by: HeonJe LEE <lhjnano@gmail.com>
2026-09-09 10:24:42 -04:00
Lyndon-Li daed42b1d8 save source size to volume info for DU and PVB
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-09-09 17:34:17 +08:00
Tiger KaovilaiandGitHub 193cfdc58f Fix datamover backup arg mismatch for CSI CBT service account name (#10318)
Run the E2E test on kind / setup-test-matrix (push) Failing after 3s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 8s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
* Fix datamover backup arg mismatch for CSI CBT service account name

The exposer built the pod command with --csi-snapshot-metadata-service-sa,
but the datamover backup command only registered --cbt-sa-name. cobra
rejects unknown flags, so the data mover pod exited immediately whenever
a dedicated CBT service account was configured -- and the reverse also
held: since the flags never matched, the SA name never actually reached
the pod, so any code path depending on it stayed unreachable.

Not limited to the block data mover: this line sits outside the
DataMoverTypeVeleroBlock gate and the cbtInfo != nil gate, so it fires
for any CSI snapshot data-movement backup.

Fix: emit --cbt-sa-name (already consumed by the backup command), naming
it consistently with the other CBT flags on the same line (--change-id,
--volume-id, --snapshot-id).

Add a regression test asserting the emitted flag string parses cleanly
against NewBackupCommand's own flag set, so the two sides can't drift
apart again without a test failure.

* Add changelog for #10318

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
2026-09-09 06:34:06 +00:00
Lyndon-Li 44f09189c2 Merge branch 'main' into report-incremental-fallback 2026-09-09 14:19:59 +08:00
Xun Jiang/Bruce JiangandGitHub c7a93be95a Add MustIncludeAdditionalItemPVCs to help track BIA added PVC's PVB creation. (#10501)
* Add MustIncludeAdditionalItemPVCs structure in backup. It's used to track PVCs returned by BIA with mustIncluded annotaion and PVC is excluded from backup by global filter.
* Modfiy the volumeHelper interface to add a parameter function for ShouldPerformFSBackup.
* Modify to support fine-grained backup filters.
* Modify according to comments. Use a read-only interface to replace the parameter function.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-09-09 14:04:57 +08:00
lyndon-liandGitHub 88da86fb67 Merge pull request #10500 from Lyndon-Li/add-id-to-repo-snapshot
Add ID to repo snapshot
2026-09-09 11:02:49 +08:00