Commit Graph
6712 Commits
Author SHA1 Message Date
Xun Jiang/Bruce JiangandGitHub 7bc03b4632 Merge pull request #10338 from blackpiglet/jxun/resolve_duplicate_initCotainer_name
Run the E2E test on kind / setup-test-matrix (push) Successful in 6s
e2e-test-kind.yaml / extract (push) Failing after 9s
Run the E2E test on kind / get-go-version (push) Failing after 10s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Avoid duplicated InitContainer names generated in velero install CLI.
2026-08-25 13:27:57 +08:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f902e09050 Bump github/codeql-action in the github-actions group (#10366)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 11s
Run the E2E test on kind / get-go-version (push) Failing after 11s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 9s
Main CI / Build (push) Skipped
Bumps the github-actions group with 1 update: [github/codeql-action](https://github.com/github/codeql-action).


Updates `github/codeql-action` from 4.37.6 to 4.37.7
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.6...v4.37.7)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-actions
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-24 15:43:13 -04:00
R4mboandGitHub d374854b0e fix log format string mismatches that produce wrong or mangled output (#10370)
* fix log format string mismatches that produce wrong or mangled output

Signed-off-by: samay43 <samayrbhat43@gmail.com>

* add changelog entry

Signed-off-by: samay43 <samayrbhat43@gmail.com>

---------

Signed-off-by: samay43 <samayrbhat43@gmail.com>
2026-08-24 15:32:56 -04:00
bc49963f1e prevent panic when the restore hook init container command annotation is empty (#10371)
* prevent panic when the restore hook init container command annotation is empty

Signed-off-by: samay43 <samayrbhat43@gmail.com>

* add changelog entry

Signed-off-by: samay43 <samayrbhat43@gmail.com>

---------

Signed-off-by: samay43 <samayrbhat43@gmail.com>
Co-authored-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-24 11:24:04 -07:00
cc7b1dbaef Embed CRD manifests via go:embed instead of codegen (#10329)
config/crd/{v1,v2alpha1}/crds/crds.go were generated files that
gzip-compressed the CRD YAML bases into committed []byte literals via
hack/crd-gen, requiring `go generate` and a dedicated CI drift check
(hack/verify-generated-crd-code.sh). This made the files large,
unreviewable in diffs, and a frequent source of merge conflicts.

Replace the generated files with config/crd/{v1,v2alpha1}/crds.go
using `//go:embed bases/*.yaml` to embed the already-committed YAML
manifests directly, decoding them the same way at init. Since Go's
go:embed can't reach outside a file's own directory tree, the crds
package now lives alongside bases/ instead of in a bases-sibling
subdirectory; import paths in pkg/install and pkg/controller were
updated accordingly.

Drop hack/crd-gen and hack/verify-generated-crd-code.sh entirely, and
trim their references from update-3generated-crd-code.sh and the
codespell skip-list. No codegen step remains, so no drift is possible.

Fixes #10328

AI-Tool-Used: Claude Code
AI-Tool-Use-Level: Category 1 (High)
AI-Code-Category: Category 1 (Production)

Signed-off-by: lubronzhan <lubron.zhan@broadcom.com>
Co-authored-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-24 14:17:25 -04:00
CopilotGitHubcopilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
5c7270a5ad chore: pin helm/kind-action to commit with curl retry fix (#10034)
* Initial plan

* chore: pin helm/kind-action to commit with curl retry fix (PR#165)

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-08-24 11:15:42 -07:00
09c1656df9 Skip DeleteSnapshot when ProviderSnapshotID is empty (#9795)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
When CreateSnapshot fails (e.g. quota limit), the snapshot is recorded
with an empty ProviderSnapshotID. During backup deletion, velero was
calling DeleteSnapshot("") which produces unnecessary 404 API calls.

Skip the DeleteSnapshot call when ProviderSnapshotID is empty and log
a warning instead.

Fixes #9429

Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
2026-08-24 10:27:34 -04:00
lyndon-liandGitHub ef9f3ed883 fix repo connection contest of the two repositories with the same storage type (#10344)
- fix repo connection contest between two BSL
- add UT for repo connection contest

Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-24 09:21:06 +00:00
R4mboandGitHub 20e24a5d33 translate parent snapshot "auto" to an empty parent snapshot in both data mover micro services (#10357)
* translate parent snapshot "auto" to an empty parent snapshot in both data mover micro services

Signed-off-by: samay43 <samayrbhat43@gmail.com>

* add changelog entry

Signed-off-by: samay43 <samayrbhat43@gmail.com>

---------

Signed-off-by: samay43 <samayrbhat43@gmail.com>
2026-08-24 16:13:53 +08:00
LubronandGitHub d9c25173f7 Fix e2e kind matrix misparsing pre-release node tags (#10359)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 12s
Run the E2E test on kind / get-go-version (push) Failing after 13s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 8s
Main CI / get-go-version (push) Failing after 9s
Main CI / Build (push) Skipped
* Fix e2e kind matrix misparsing pre-release node tags

The setup-test-matrix step excluded "alpha|beta" pre-release tags but
not "rc" ones. A tag like v1.37.0-rc.1 slipped through to the awk
field-splitter, which treats "." as the only separator: splitting
"v1.37.0-rc.1" yields ["v1","37","0-rc","1"], and printing
$1"."$2"."$NF produced the bogus version "v1.37.1" - an image that
was never published, since the real tag is v1.37.0-rc.1.

Replace the two greps with a single anchored pattern that only
matches well-formed vX.Y.Z tags, so any hyphenated pre-release
suffix (rc, alpha, beta, or otherwise) is excluded before reaching
the awk step.

Fixes #10358

AI-Tool-Used: Claude Code
AI-Tool-Use-Level: Category 2 (Medium)
AI-Code-Category: Category 2 (Non-Production)

Signed-off-by: lubronzhan <lubron.zhan@broadcom.com>

* Add changelog entry for e2e matrix fix

AI-Tool-Used: Claude Code
AI-Tool-Use-Level: Category 3 (Low)
AI-Code-Category: Category 2 (Non-Production)

Signed-off-by: lubronzhan <lubron.zhan@broadcom.com>

---------

Signed-off-by: lubronzhan <lubron.zhan@broadcom.com>
2026-08-24 13:45:41 +08:00
d2241adba0 Log the discovered parent snapshot ID, not the empty lookup parameter (#10305)
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
getParentBackupInfo's parent-selection log lines interpolated the
parentSnapshot parameter. On the discovery branch (no explicit parent
passed by the caller) that parameter is empty by definition, so every
message about which parent was chosen, or why a run fell back to full,
printed no identifier at all -- e.g. "Using parent snapshot , start
time ...". This is the normal path for scheduled/incremental backups,
so the omission hit the common case, not an edge one.

Bind a parentID local that starts as the parameter but is overwritten
once a parent is actually resolved (explicit or discovered), and log
that instead. No behavior change.

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 16:47:40 -04:00
1e26cf7ca0 fix(restore_finalizer): bound WaitRestoreExecHook poll with resourceT… (#10280)
* fix: reuse DefaultResourceTimeout from server config for hook wait

Signed-off-by: Nitish Malang <71919457+nitishmalang@users.noreply.github.com>

* Add changelog for PR 10280

Signed-off-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>

---------

Signed-off-by: Nitish Malang <71919457+nitishmalang@users.noreply.github.com>
Signed-off-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>
Co-authored-by: Tiger Kaovilai <passawit.kaovilai@gmail.com>
2026-08-21 15:15:30 -04:00
PranjalandGitHub a4b8261ffe fix(cli): show n/a for expiration on stalled New backups (#10326)
Backups stuck in New state have no Status.Expiration yet. The CLI
was estimating expiration from CreationTimestamp + TTL, which made
long-queued backups appear already expired.

Fixes #3555

Signed-off-by: PranjalManhgaye <manhgayepranjal@gmail.com>
2026-08-21 15:12:58 -04:00
R4mboandGitHub c20b09e281 fix nil pointer dereference in WaitUntilVSCHandleIsReady when a VSC error has no message (#10352)
* fix nil pointer dereference in WaitUntilVSCHandleIsReady when a VSC error has no message

Signed-off-by: samay43 <samayrbhat43@gmail.com>

* add changelog entry

Signed-off-by: samay43 <samayrbhat43@gmail.com>

---------

Signed-off-by: samay43 <samayrbhat43@gmail.com>
2026-08-21 15:11:09 -04:00
Xun Jiang 00d7e7d022 Avoid duplicated InitContainer names generated in velero install CLI.
Add random string at the end when there is name collision detected.

Signed-off-by: Xun Jiang <xun.jiang@broadcom.com>
2026-08-21 17:41:30 +08:00
3cd6c2e533 Skip signing a download URL when no artifacts can exist yet (#10252)
Run the E2E test on kind / setup-test-matrix (push) Successful in 4s
e2e-test-kind.yaml / extract (push) Failing after 9s
Run the E2E test on kind / get-go-version (push) Failing after 11s
push.yml / extract (push) Failing after 6s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
Main CI / get-go-version (push) Failing after 7s
Main CI / Build (push) Skipped
* Skip signing a download URL when no artifacts can exist yet

Reported in #10232: a DownloadRequest for a backup that never ran still
reaches Processed with a signed URL, and fetching it returns 404.

The controller already has the backup, and the restore for restore
targets, in hand before it signs, so checking the phase costs no extra
call to the object store.

The check is deliberately narrow. It refuses only the pre-execution
phases, where nothing has been written for any target kind: New, Queued,
ReadyToStart and FailedValidation for backups, New and FailedValidation
for restores. InProgress onwards may hold a partial log or other
artifacts, and Deleting may still hold all of them, so those keep the
behaviour callers have today.

That matters because velero backup download has no client side phase
check of its own, unlike backup logs and restore logs. Reusing the
allowlist from pkg/cmd/cli/backup/logs.go would have changed what
backup download can fetch; this does not.

A backup with an empty phase is left alone as well, since that state is
transient and the caller can retry.

Refs #10232

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Derive the phase coverage test from the generated CRDs

The previous test built a slice of phases by hand and asserted its own
length, so it passed no matter what the API did. Adding a fourteenth
backup phase would not have failed it.

This reads the status.phase enum out of the generated CRDs, via the
exported v1crds.CRDs that pkg/install already uses. The enum comes from
the same kubebuilder markers as the Go constants, so a phase added to
the API fails here until it is classified.

Verified by removing Deleting from the expectations, which now fails with
'BackupPhase "Deleting" is served by the CRD but not classified'.

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Use US spelling in comments to satisfy the misspell linter

golangci-lint runs misspell, which flags behaviour as a misspelling of
behavior. Comments only, no functional change.

Signed-off-by: saral <ilovegojo2580@gmail.com>

* Set a Failed phase with a reason when the guard refuses to sign

The guard added in the previous commit left the request at New with no URL, so
the CLI polled until its own timeout and then reported that the backup storage
location may be unavailable. The BSL is fine; the backup never ran.

DownloadRequestPhase gains Failed and DownloadRequestStatus gains Message. The
controller sets both where it refuses, and the CLI stops as soon as it sees the
phase and surfaces the message instead of its generic timeout error.

Adding an enum value is additive, per the direction on the PR discussion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: saral <ilovegojo2580@gmail.com>

---------

Signed-off-by: saral <ilovegojo2580@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 17:26:21 +08:00
Xun Jiang/Bruce JiangandGitHub 42c2c0f1c7 Merge pull request #10325 from PranjalManhgaye/docs/e2e-readme-additional-bsl-mappings
docs(e2e): fix additional BSL make variable mappings
2026-08-21 15:32:23 +08:00
Xun Jiang/Bruce JiangandGitHub 7db7ce7abb Merge pull request #10313 from PranjalManhgaye/fix/pvc-builder-typo
test(e2e): fix pvcBuilder variable typo in pvc helpers
2026-08-21 15:31:45 +08:00
Xun Jiang/Bruce JiangandGitHub cfe733102e Merge pull request #10324 from PranjalManhgaye/test/e2e-apiextensions-typo
test(e2e): fix apiextensions typo in ginkgo description
2026-08-21 15:31:08 +08:00
Xun Jiang/Bruce JiangandGitHub da09de0e34 Merge pull request #10323 from PranjalManhgaye/docs/e2e-readme-typos
docs(e2e): fix typos in e2e test README
2026-08-21 15:30:42 +08:00
Daniel JiangandGitHub 8bbd546167 Double check the label for backup when deleting VSC (#10346)
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
This commit double checks the label of the VSC on the cluster before
deleting it to avoid mis-deletion.

Signed-off-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-21 11:19:57 +08:00
Shubham PampattiwarandGitHub f27a4ad8c0 Fix LoadAffinity mutation accumulating OS node selector terms (#10342)
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 9s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Successful in 2s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
* Fix LoadAffinity mutation accumulating OS node selector terms

The node-agent parses the loadAffinity configuration once at startup and
keeps it in memory. GetLoadAffinityByStorageClass returned a pointer to one
of the elements of that cached list rather than a copy, so the exposers,
which append a kubernetes.io/os match expression to the returned affinity,
were mutating the shared configuration. Every DataUpload or DataDownload
appended another OS term, growing the data mover pod spec until it could
eventually exceed the object size limit.

Return a deep copy from GetLoadAffinityByStorageClass so that callers can
safely modify the result. A shallow copy is not enough because the
MatchExpressions slice header would still be shared with the source.

Fixes #10341

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>

* Add changelog

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>

---------

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-20 14:15:56 -04:00
Shubham PampattiwarandGitHub 8359870e2a Merge pull request #10337 from shubham-pampattiwar/docs/backup-restore-pvc-secret-copy
Document secretNames/configMapNames for backup/restore PVC config
2026-08-20 10:49:08 -07:00
Krishna AwasthiandGitHub 1c6d758281 Testing: Implement missing unit tests for pkg/backup/snapshots.go (#10315)
Signed-off-by: opbot_xd <awasthikrishna23052005@gmail.com>
2026-08-20 23:40:59 +08:00
Chlins ZhangandGitHub 9a6346abba Merge pull request #10343 from chlins/fix/no-rerun-synced-backups
Run the E2E test on kind / setup-test-matrix (push) Successful in 7s
e2e-test-kind.yaml / extract (push) Failing after 8s
Run the E2E test on kind / get-go-version (push) Failing after 8s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 8s
Main CI / Build (push) Skipped
Only sync finished backups from object storage
2026-08-20 17:21:37 +08:00
lyndon-liandGitHub e7b15d07d1 Merge pull request #10322 from Lyndon-Li/fill-error-to-cr-when-data-mover-pod-evicted
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
e2e-test-kind.yaml / extract (push) Failing after 7s
Run the E2E test on kind / get-go-version (push) Failing after 8s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 7s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Issue 10321: fill the error to the corresponding CR when data mover pod is evicted
2026-08-20 13:21:18 +08:00
chlins 9afa3964ba Only sync finished backups from object storage
Backup metadata with an empty or New phase was synced into the cluster as a pending backup, which the queue controller then ran as if it were newly requested. Hooks are dropped as well, since a synced backup never executes them.

Signed-off-by: chlins <chlins.zhang@gmail.com>
2026-08-20 10:37:41 +08:00
Shubham Pampattiwar 32fdc9591a Clarify tenant token guidance for encrypted volumes
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-19 14:22:31 -07:00
Shubham Pampattiwar 1e1c9e2648 Document secretNames/configMapNames for backup/restore PVC config
Document the new backupPVC/restorePVC secretNames and configMapNames
options (velero#9920) that copy namespace-scoped secrets/configmaps to
the Velero namespace so datamover can back up and restore encrypted CSI
volumes (e.g. ODF/ceph-csi with Vault KMS). Updates the main and v1.18
docs.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-19 09:33:08 -07:00
Shubham PampattiwarandGitHub fd870efe67 Merge pull request #9920 from shubham-pampattiwar/backup-pvc-secret-copy
Run the E2E test on kind / setup-test-matrix (push) Successful in 5s
e2e-test-kind.yaml / extract (push) Failing after 6s
Run the E2E test on kind / get-go-version (push) Failing after 7s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 6s
Main CI / Build (push) Skipped
Support copying namespace-scoped secrets/configmaps for backup and restore PVC provisioning
2026-08-19 08:01:33 -07:00
Lyndon-Li 223aa0d282 add UT for evicted pod
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-19 18:10:53 +08:00
Daniel JiangandGitHub e9e3054276 Enforce namespace of the "musthave" resources in restore (#10333)
e2e-test-kind.yaml / extract (push) Failing after 11s
Run the E2E test on kind / get-go-version (push) Failing after 11s
Run the E2E test on kind / build (push) Skipped
Run the E2E test on kind / setup-test-matrix (push) Successful in 3s
Run the E2E test on kind / run-e2e-test (push) Skipped
push.yml / extract (push) Failing after 6s
Main CI / get-go-version (push) Failing after 9s
Main CI / Build (push) Skipped
This commit ensures the resources in the set "resourceMustHave" can only
be created in the namespace of velero deployment if it's namespace
scoped.

Signed-off-by: Daniel Jiang <daniel.jiang@broadcom.com>
2026-08-19 17:32:50 +08:00
Lyndon-Li ee6f7b5a97 fix UT error
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-19 16:04:47 +08:00
Lyndon-Li 234a2cb288 handle evicted event
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-19 15:37:56 +08:00
Lyndon-Li 534b2720c2 use atomic.Bool for thread safety
Signed-off-by: Lyndon-Li <lyonghui@vmware.com>
2026-08-19 15:20:47 +08:00
339c8edda9 Detect block uploader cancellation through wrapped errors (#10308)
* Detect block uploader cancellation through wrapped errors

Cancelling a block data mover backup was reported as a failure: the
DataUpload ended Failed with an error message and the Backup went
PartiallyFailed, for a user-requested cancel.

The cause is a sentinel equality check. block.ErrCanceled is raised in
the write loop and then wrapped twice before it reaches the provider --
once in block/uploader.go ("error backing up bdev %s") and again in
block/snapshot.go ("Failed to run uploader backup for si %v") -- so
`err == block.ErrCanceled` can never be true and the ErrorCanceled
returns are unreachable. The filesystem provider avoids this by asking
the uploader for its state (kpUploader.IsCanceled()) rather than
inspecting the error.

Use errors.Is at both the backup and restore sites.

Adds TestBlockProviderCancelThroughWrappedError, which injects the
doubly-wrapped sentinel exactly as production builds it. Note the
assertion is require.ErrorIs, not ErrorContains: provider.ErrorCanceled
and block.ErrCanceled carry identical message text, so a substring
assertion passes whether or not the sentinel was recognised -- which is
why the existing test, injecting the bare sentinel, did not catch this.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
(cherry picked from commit 9d6c5da7a893068d424b0c7896638787c636e213)
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* Add changelog for #10308

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

* lint: fix misspelling (recognised -> recognized)

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>

---------

Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 15:03:33 +08:00
Shubham Pampattiwar 2a920ab946 Add RBAC for secrets/configmaps to datamover controllers
The DataUpload and DataDownload controllers now copy and delete
namespace-scoped secrets/configmaps for backup/restore PVC provisioning.
Add the corresponding kubebuilder RBAC markers (get;list;create;delete
on secrets and configmaps) and regenerate the ClusterRole.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar cf04db2705 Use owner UID as backup-pvc-secret label value
The copied secret/configmap label value was the owner (DataUpload/
DataDownload) name, which is derived from the Backup/Restore name and
can exceed the 63-char Kubernetes label-value limit or contain invalid
characters. That would make the copy label and the cleanup selector
diverge and orphan the copied resources.

Use string(ownerObject.UID) consistently for the label value in both
copy and cleanup (backup and restore exposers). The UID is a stable,
always-valid label value.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar b8944bda53 Update changelog to cover restore path
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 9fd5365b15 Copy namespace-scoped secrets/configmaps for restore PVC provisioning
Mirror the backup-side fix on the restore path. The generic restore
exposer creates the intermediate restore PVC in the Velero namespace
using the target PVC's StorageClass. For encrypted volumes this fails
because ceph-csi looks up the KMS token secret in the PVC's namespace
(the Velero namespace), where it does not exist.

Add SecretNames/ConfigMapNames to the RestorePVC config. When set, the
generic restore exposer copies the named secrets/configmaps from the
target namespace to the Velero namespace before creating the restore
PVC, and cleans them up in CleanUp(). Reuses the same copy/delete
helpers and label as the backup path.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar f457a95802 Remove unused DeleteSecretIfAny/DeleteConfigMapIfAny helpers
These single-object delete helpers were introduced earlier but are no
longer called in production code: DeleteSecretsWithLabel and
DeleteConfigMapsWithLabel now delete inline with UID preconditions.
Remove the dead functions and their tests.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar e498c5f79b Address review: generic labels param and copy placement
- Make CopySecret/CopyConfigMap accept a generic labels map instead of
  hardcoding the backup-pvc-secret label, aligning with the generic
  DeleteSecretsWithLabel helper. Move the BackupPVCSecretLabel constant
  from util/kube to the exposer package where it is used.
- Move the secret/configmap copy in Expose() to after
  WaitVolumeSnapshotReady and before createBackupVS. That is the most
  likely failure point, and nothing needs cleanup before it.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar eda695ae3a Fix linter issues: gofmt alignment and require.Error
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 2df1386d08 Add tests for secret/configmap copy in Expose and CleanUp
Add test cases for the CSI snapshot exposer:
- TestExpose_SecretCopy: verifies secret copy, configmap copy, and
  error on missing source secret during Expose()
- TestCleanUp_SecretsAndConfigMaps: verifies label-based cleanup
  deletes owned resources and preserves unrelated ones

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 7db0b391ff Address review feedback: ownership, BinaryData, preconditions, placement
- Fix premature deletion of shared secrets/configmaps: check owner label
  in addition to data equality. Same data + different owner is now a
  collision, preventing one DataUpload's CleanUp from removing resources
  another DataUpload is still using.
- Copy BinaryData in CopyConfigMap and include it in the equality check,
  so configmaps with binary payloads (e.g., CA bundles) are not silently
  truncated.
- Add UID preconditions to DeleteSecretsWithLabel and
  DeleteConfigMapsWithLabel to avoid TOCTOU races where a recreated
  object with the same name could be deleted.
- Move secret/configmap copy to the beginning of Expose(), before any
  intermediate objects are created, so failure doesn't require cleanup.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 6e536a451a Add tests for configmap copy/delete utilities
Add unit tests for CopyConfigMap, DeleteConfigMapIfAny, and
DeleteConfigMapsWithLabel mirroring the existing secret tests.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 986350a6e5 Move secret/configmap copy from controller to CSI snapshot exposer
Move the secret and configmap copy logic from the DataUpload controller
into the CSI snapshot exposer's Expose() method. This keeps all
CSI-specific logic in the exposer and maintains symmetry with CleanUp()
which already handles the cleanup of copied resources.

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar c15cf084e3 Add configmap copy support and move secret copy after accept
- Add ConfigMapNames field to BackupPVC config for copying tenant
  configmaps (e.g., ceph-csi-kms-config with Vault connection overrides)
- Add CopyConfigMap, DeleteConfigMapIfAny, DeleteConfigMapsWithLabel
  utilities mirroring the secret copy functions
- Move secret/configmap copy after acceptDataUpload() so only the
  accepting node handles it, avoiding multi-node contest
- Clean up copied configmaps in CleanUp() alongside secrets

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar f91f669e77 Fix linter issues in secret utilities
- Fix import ordering in test file (gofmt)
- Add nolint:gosec for BackupPVCSecretLabel constant (not a credential)
- Use assert.Error instead of assert.True(err != nil) (testifylint)

Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00
Shubham Pampattiwar 6a7b5872b0 Add changelog for PR #9920
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
2026-08-18 10:28:27 -07:00