* Fix backup-finalizer: do not set backup phase to Completed before PutBackupMetadata succeeds Previously, the backup finalizer controller set backup.Status.Phase to Completed/PartiallyFailed in-memory BEFORE calling PutBackupMetadata and PutBackupContents. When these uploads failed (e.g., due to object lock or immutability), the deferred patch function still wrote the terminal phase to the Kubernetes API server, preventing the controller from retrying the upload on the next reconcile. This fix moves the phase assignment to AFTER both uploads succeed. A DeepCopy of the backup is used to encode the JSON with the final phase for object storage, while the in-memory backup object retains the Finalizing phase until uploads complete. Caveats: - CompletionTimestamp is now captured before upload but only committed to the API server after upload succeeds. On retry after a transient failure, a new timestamp is generated, so the completion time reflects when the upload finally succeeded rather than when finalization processing completed. - Metrics (RegisterBackupSuccess/RegisterBackupPartialFailure) are now recorded after uploads succeed, so they accurately reflect only fully persisted backups. - The metadata uploaded to object storage contains the final phase and completion timestamp via DeepCopy, so storage state is correct even before the API server is patched. Fixes #9645 Generated with [Claude Code](https://claude.ai/code) via [Happy](https://happy.engineering) Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Happy <yesreply@happy.engineering> Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> * Add changelog for #9646 Generated with [Claude Code](https://claude.ai/code) via [Happy](https://happy.engineering) Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Happy <yesreply@happy.engineering> Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> * Fix testifylint: use require.Error instead of assert.Error Generated with [Claude Code](https://claude.ai/code) via [Happy](https://happy.engineering) Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Happy <yesreply@happy.engineering> Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> * Address review feedback on backup-finalizer fix - Add default guard for unhandled phase values in finalPhase switch - Add retry with DefaultBackoff for PutBackupMetadata per reviewer request - Replace brittle framework.BackupItemActionResolverV2{} mock with mock.Anything - Add FinalizingPartiallyFailed test case for PutBackupContents failure Generated with [Claude Code](https://claude.ai/code) via [Happy](https://happy.engineering) Co-Authored-By: Claude <noreply@anthropic.com> Co-Authored-By: Happy <yesreply@happy.engineering> Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> * Use bounded, object-storage-tuned backoff for backup-finalizer uploads retry.DefaultBackoff is tuned for API server optimistic-concurrency conflicts (4 steps, ~1.25s total) and gives up far too quickly for object storage calls, which can see longer transient outages or throttling (review feedback from blackpiglet). Replace it with a dedicated, bounded backoff (1s base, 2x factor, 5 steps, ~31s total) applied to both PutBackupMetadata and PutBackupContents. Being bounded (rather than retrying forever) means a persistent failure, e.g. an object-lock/immutability policy denying every write, surfaces as an error within a bounded time instead of hanging the reconcile indefinitely; controller-runtime requeues on error, so retries continue across reconciles (review feedback from priyansh17). Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> * Fix PutBackupMetadata retry to re-read backupJSON each attempt backupJSON is a bytes.Buffer, so passing it directly to PutBackupMetadata drains it on the first read attempt. A retry after a transient failure would then upload empty content instead of the backup metadata. Wrap it in bytes.NewReader(backupJSON.Bytes()) inside the retry closure so every attempt gets a fresh reader. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> --------- Signed-off-by: Tiger Kaovilai <tkaovila@redhat.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Happy <yesreply@happy.engineering>
Overview
Velero (formerly Heptio Ark) gives you tools to back up and restore your Kubernetes cluster resources and persistent volumes. You can run Velero with a public cloud platform or on-premises.
Velero lets you:
- Take backups of your cluster and restore in case of loss.
- Migrate cluster resources to other clusters.
- Replicate your production cluster to development and testing clusters.
Velero consists of:
- A server that runs on your cluster
- A command-line client that runs locally
Documentation
The documentation provides a getting started guide and information about building from source, architecture, extending Velero and more.
Please use the version selector at the top of the site to ensure you are using the appropriate documentation for your version of Velero.
Troubleshooting
If you encounter issues, review the troubleshooting docs, file an issue, or talk to us on the #velero-users and #velero-dev channel on the Kubernetes Slack server.
Community
Velero is an open community and we welcome your participation. The best way to get involved is to join our bi-weekly community meetings:
- Join the Velero community meetings, held bi-weekly, alternating between Beijing-friendly and US/Europe-friendly time zones.
- Subscribe to the project meeting calendar.
- Chat with us on the Kubernetes Slack
#velero-userschannel and join the mailing list.
See the community page for the full schedule and details.
Contributing
If you are ready to jump in and test, add code, or help with documentation, follow the instructions on our Start contributing documentation for guidance on how to setup Velero for development.
Governance
Velero's governance describes how the project is run, including the decision-making process, the roles and responsibilities of maintainers, and how to become a maintainer. Governance applies across the Velero org and is maintained at velero-io/.github.
Changelog
See the list of releases to find out about feature changes.
Velero compatibility matrix
The following is a list of the supported Kubernetes versions for each Velero version.
| Velero version | Expected Kubernetes version compatibility | Tested on Kubernetes version |
|---|---|---|
| 1.18 | 1.18-latest | 1.33.7, 1.34.1, and 1.35.0 |
| 1.17 | 1.18-latest | 1.31.7, 1.32.3, 1.33.1, and 1.34.0 |
| 1.16 | 1.18-latest | 1.31.4, 1.32.3, and 1.33.0 |
| 1.15 | 1.18-latest | 1.28.8, 1.29.8, 1.30.4 and 1.31.1 |
| 1.14 | 1.18-latest | 1.27.9, 1.28.9, and 1.29.4 |
Velero supports IPv4, IPv6, and dual stack environments. Support for this was tested against Velero v1.8.
The Velero maintainers are continuously working to expand testing coverage, but are not able to test every combination of Velero and supported Kubernetes versions for each Velero release. The table above is meant to track the current testing coverage and the expected supported Kubernetes versions for each Velero version.
If you are interested in using a different version of Kubernetes with a given Velero version, we'd recommend that you perform testing before installing or upgrading your environment. For full information around capabilities within a release, also see the Velero release notes or Kubernetes release notes. See the Velero support page for information about supported versions of Velero.
For each release, Velero maintainers run the test to ensure the upgrade path from n-2 minor release. For example, before the release of v1.10.x, the test will verify that the backup created by v1.9.x and v1.8.x can be restored using the build to be tagged as v1.10.x.
Cloud Native Computing Foundation
Velero is a Cloud Native Computing Foundation sandbox project.
Copyright Contributors to Velero, established as Velero a Series of LF Projects, LLC. For website terms of use, trademark policy and other project policies please see https://lfprojects.org/policies/.
