locations.md linked to customize-locations.md, which does not exist; the page is customize-installation.md (same heading, and the same line already links there). minio.md linked debugging-install.md without the ../ the rest of that file uses, resolving to a nonexistent contributions/debugging-install.md. restore-reference.md had two dead same-page anchors: #resource-restore-order (heading is Restore order) and #durable-snapshot-pv-restore (heading is Snapshot PV Restore).
Signed-off-by: avneetbansal-aws <284363899+avneetbansal-aws@users.noreply.github.com>
Co-authored-by: avneetbansal-aws <284363899+avneetbansal-aws@users.noreply.github.com>
* Fix unchecked type assertion panic in ChangeImageNameAction
replaceImageName reads a restored container's image field out of the
unstructured object and asserts it to string without checking ok. The
comma-ok map lookup on the line above only confirms the "image" key is
present -- it says nothing about the value's type. A restored resource
whose image field is present but not a JSON string (e.g. a number,
bool, null, array, or object) causes an unrecovered
"interface conversion: interface {} is not string" panic in this
RestoreItemAction plugin whenever the optional image-remapping
ConfigMap feature is configured.
Switch to the comma-ok form of the assertion and skip (with a log
message) any container whose image field isn't a string, instead of
panicking.
Signed-off-by: Kaizhe Huang <derek0405@gmail.com>
* Add regression test for non-string image field panic
Covers the comma-ok assertion fix: replaceImageName operates on
generic unstructured content decoded from a backup tarball, which
isn't validated against the Pod schema before this code runs, so
"image" isn't guaranteed to be a string.
Signed-off-by: Kaizhe Huang <derek0405@gmail.com>
* Guard container-entry type assertion, use unstructured.NestedString
Addresses reviewer feedback: container.(map[string]any) was also an
unchecked assertion, and unstructured.NestedString gives safer,
more idiomatic type-checking than a manual comma-ok assertion. Also
switches the skip-path logging from Info to Warn per review.
Signed-off-by: Kaizhe Huang <derek0405@gmail.com>
---------
Signed-off-by: Kaizhe Huang <derek0405@gmail.com>
PodVolumeBackupStatus has no Node field, so the Node column was always
empty. The node is recorded in the spec.
Fixes#10444
Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
* e2e: add kind VolumeGroupSnapshotClass test data
Velero selects a VolumeGroupSnapshotClass by the
velero.io/csi-volumegroupsnapshot-class label, and csi-driver-host-path
ships no VolumeGroupSnapshotClass at all, so there is nothing for the
selector to find on a kind cluster. Any VolumeGroupSnapshot e2e coverage
needs a class to exist first.
Add the kind entry alongside the existing volume-snapshot-class test
data, following the same layout and naming. Verified against a kind
cluster running external-snapshotter v8.6.0 and csi-driver-host-path:
applying this file and creating a VolumeGroupSnapshot that selects two
labelled PVCs reaches readyToUse with both member snapshots ready.
Nothing applies this file yet. It is a prerequisite for the
VolumeGroupSnapshot specs tracked in #7507, kept separate so the class
can be reviewed on its own.
Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>
* Add changelog for #10582
Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>
---------
Signed-off-by: krishhna24 <krishhnatupedev@gmail.com>
* Add e2e test for namespace selection by label in resource policy
Covers design step 8 of #9772: a backup with no explicit
--include-namespaces (the same shape a Schedule with no
includedNamespaces produces), relying entirely on a ResourcePolicy
ConfigMap's includedNamespacesByLabel to select which namespaces to
back up.
Creates labeled and unlabeled namespaces, backs up with a
ResourcePolicy ConfigMap setting includedNamespacesByLabel, and
verifies only the labeled namespaces are restored - closing the e2e
coverage gap #10275 deferred to velero-io/velero#10564.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Add changelog entry for e2e test PR
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Fix e2e test: create Backup directly, not via CLI's implicit --include-namespaces=*
The kind e2e run showed every unlabeled namespace getting backed up
and restored anyway. velero backup create's --include-namespaces flag
defaults to ["*"] when omitted (pkg/cmd/cli/backup/create.go), so
skipping the flag still sent an *explicit* wildcard - and
mergeNamespacesByLabel deliberately leaves an explicit "*" untouched
rather than narrowing it, so includedNamespacesByLabel never got a
chance to replace anything.
Create the Backup object directly via the controller-runtime client
instead, leaving BackupSpec.IncludedNamespaces genuinely unset - the
only way to exercise the "defaulted empty" narrowing path the CLI's
own default makes unreachable.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
---------
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit fixes the issue that when a namespace with label velero.io/exclude-from-backup: "true"
is added to the exludeNamespaces of a backup CR. There will be
duplicated entries of the namespace in the spec of the backup.
Signed-off-by: Daniel Jiang <daniel.jiang@broadcom.com>
Links to plugin-versioning.md and general-progress-monitoring.md broke when
those docs moved to design/Implemented/; two other relative paths had one
directory level wrong.
Signed-off-by: Zain <43629888+ZainnQureshii@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Add the maintainer information required by the CNCF Incubation criteria:
- State that all maintainers share collective responsibility for the entire
project, consistent with CODEOWNERS assigning ownership to the maintainer
group as a whole (no siloed per-area owners).
- Add a Contacting the maintainers section listing GitHub, Slack, the mailing
list, and the security disclosure process.
This makes MAINTAINERS.md cover names, contact information, domain of
responsibility, and affiliation as required for the Incubation application.
Signed-off-by: Shubham Pampattiwar <spampatt@redhat.com>
* Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds
Previously, the backup finalizer controller set backup.Status.Phase to
Completed/PartiallyFailed in-memory BEFORE calling PutBackupMetadata and
PutBackupContents. When these uploads failed (e.g., due to object lock
or immutability), the deferred patch function still wrote the terminal
phase to the Kubernetes API server, preventing the controller from
retrying the upload on the next reconcile.
This fix moves the phase assignment to AFTER both uploads succeed. A
DeepCopy of the backup is used to encode the JSON with the final phase
for object storage, while the in-memory backup object retains the
Finalizing phase until uploads complete.
Caveats:
- CompletionTimestamp is now captured before upload but only committed to
the API server after upload succeeds. On retry after a transient
failure, a new timestamp is generated, so the completion time reflects
when the upload finally succeeded rather than when finalization
processing completed.
- Metrics (RegisterBackupSuccess/RegisterBackupPartialFailure) are now
recorded after uploads succeed, so they accurately reflect only fully
persisted backups.
- The metadata uploaded to object storage contains the final phase and
completion timestamp via DeepCopy, so storage state is correct even
before the API server is patched.
Fixes#9645
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Add changelog for #9646
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Fix testifylint: use require.Error instead of assert.Error
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Address review feedback on backup-finalizer fix
- Add default guard for unhandled phase values in finalPhase switch
- Add retry with DefaultBackoff for PutBackupMetadata per reviewer request
- Replace brittle framework.BackupItemActionResolverV2{} mock with mock.Anything
- Add FinalizingPartiallyFailed test case for PutBackupContents failure
Generated with [Claude Code](https://claude.ai/code)
via [Happy](https://happy.engineering)
Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Happy <yesreply@happy.engineering>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Use bounded, object-storage-tuned backoff for backup-finalizer uploads
retry.DefaultBackoff is tuned for API server optimistic-concurrency
conflicts (4 steps, ~1.25s total) and gives up far too quickly for
object storage calls, which can see longer transient outages or
throttling (review feedback from blackpiglet). Replace it with a
dedicated, bounded backoff (1s base, 2x factor, 5 steps, ~31s total)
applied to both PutBackupMetadata and PutBackupContents.
Being bounded (rather than retrying forever) means a persistent
failure, e.g. an object-lock/immutability policy denying every write,
surfaces as an error within a bounded time instead of hanging the
reconcile indefinitely; controller-runtime requeues on error, so
retries continue across reconciles (review feedback from priyansh17).
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
* Fix PutBackupMetadata retry to re-read backupJSON each attempt
backupJSON is a bytes.Buffer, so passing it directly to
PutBackupMetadata drains it on the first read attempt. A retry after
a transient failure would then upload empty content instead of the
backup metadata. Wrap it in bytes.NewReader(backupJSON.Bytes()) inside
the retry closure so every attempt gets a fresh reader.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
---------
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Happy <yesreply@happy.engineering>
Update Request struct initialization in backup_test.go to use
SkippedVolumeTracker and NewSkipVolumeTracker following the rename
from SkippedPVTracker.
Signed-off-by: Adam Zhang <adam.zhang@broadcom.com>
PodVolumeRestores are only created for pods that Velero creates, so when
the pod already exists in the cluster the PVC-not-in-use pre-flight check
never runs and the volume data restore is skipped silently, while the
existing pod keeps consuming the PVC. Report an explicit pre-flight error
for such pods, aligned with the PVC CSI RIA behavior.
Signed-off-by: chlins <chlins.zhang@gmail.com>
CI's changelog check requires a changelogs/unreleased/<PR#>-<login>
file; this PR didn't have one since the PR number wasn't known until
after it was opened.
Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Motivation:
Forget and BatchForget in pkg/repository/manager/manager.go accept a
caller-provided context.Context but ignore it, hardcoding
context.Background() when calling into the repository provider. This
means cancellation and timeouts set by callers (e.g. the backup
deletion controller) are silently dropped during repository connection
and snapshot deletion. Additionally, BatchForget returned a wrapped nil
instead of the real connection error when prd.BoostRepoConnect failed,
because it referenced an unrelated, already-nil err variable instead of
connectErr.
Approach:
Pass the caller's ctx through to prd.BoostRepoConnect, prd.Forget, and
prd.BatchForget in both Forget and BatchForget, instead of substituting
context.Background(). Fix BatchForget's connection-failure branch to
wrap and return connectErr instead of the stale err. Other methods on
manager (InitRepo, ConnectToRepo, PrepareRepo, PruneRepo, UnlockRepo)
don't accept a ctx parameter at all, so they are unaffected and out of
scope for this change.
Validation:
- go build ./pkg/repository/... and go build ./... pass.
- go vet ./pkg/repository/... is clean.
- go test ./pkg/repository/... passes, including three new tests added
to pkg/repository/manager/manager_test.go.
- golangci-lint run ./pkg/repository/... is clean.
- Confirmed the new tests reproduce both bugs: temporarily reverting
only manager.go and re-running go test ./pkg/repository/manager/...
made all three new tests fail (missing propagated context value and
cancellation, and a nil error returned where the real connect error
was expected); re-applying the fix makes them pass. This is a silent
behavior bug (broken context propagation and a swallowed error), not
a crash.
Report: https://github.com/velero-io/velero/issues/10551
Signed-off-by: Pujitha Paladugu <10557236+pujitha24@users.noreply.github.com>
Assisted-by: claude-sonnet-5 (via Claude Code)
velero schedule create registers --annotations through BackupOptions.BindFlags,
but the Schedule ObjectMeta it builds only sets Labels, so the flag was accepted
and silently discarded.
This matters beyond the Schedule object itself: BackupBuilder.FromSchedule
falls back to schedule.Annotations when the template carries none, so every
backup generated by the schedule lost the annotations too. velero schedule
describe already prints these fields and the Schedule CRD already carries them,
so the create path was the only gap.
Same shape as #10526, which fixed the backup type being dropped on the same
struct literal.
Signed-off-by: Jeremy Schoemaker <jeremy@shoemoney.com>
Add includedNamespacesByLabel, excludedNamespacesByLabel, and
labelSelectorLogic to IncludeExcludePolicy in the ResourcePolicy
ConfigMap (realizes design in velero-io/velero#9772), letting a backup
select or exclude namespaces by label instead of (or in addition to)
name/wildcard.
The backup controller resolves label selectors against the live
namespace list once per backup, merges the results into
spec.includedNamespaces/excludedNamespaces, then proceeds through the
existing name-based filtering unchanged. A defaulted "*" include list
is replaced by the resolved set; an explicitly-configured include list
(including an explicit "*") is unioned with it instead, and stays
canonical rather than widening. Namespaces matching an exclude
selector are always subtracted from the merged includes, regardless of
how the includes were populated.
Because Velero's namespace-includes/excludes model requires at least
one name (an empty list means "match everything"), a selector that
resolves to zero namespaces is represented with a sentinel glob
pattern ("[-]*") guaranteed to match no real namespace, rather than an
empty list that would silently fall back to including/excluding
everything.
labelSelectorLogic ("AND"/"OR", case-insensitive) controls whether
multiple included/excluded label selectors are combined by
intersection or union; it is validated up front, including inside
ResolveNamespacesByLabel itself, so an invalid value fails fast instead
of silently falling through to OR semantics.
Namespace-selection-by-label and resource-selection-by-label act as
independent axes and do not affect each other, matching the design
discussion in #9772.
Known limitations:
- Selectors are evaluated once per backup against the namespace list
at that point in time; namespaces created or relabeled mid-backup
are not picked up.
- Backup-only for now; restore-side namespace mapping is unaffected.
Testing:
- Unit coverage in internal/resourcepolicies for validation, selector
resolution (including AND/OR logic, case-insensitivity, and
malformed-selector/invalid-logic error paths), and the no-match
sentinel.
- Unit coverage in pkg/controller for the merge logic between resolved
label selections and explicit/defaulted includes and excludes.
- End-to-end coverage in pkg/backup exercising the full backup
pipeline with label-selected namespaces, including the
velero.io/exclude-from-backup hard-exclusion interaction and the
zero-match/fully-excluded sentinel path.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com>