Merge branch 'main' into copilot/update-filter-docs-pattern

This commit is contained in:
Tiger Kaovilai
2026-08-26 10:53:48 -04:00
committed by GitHub
429 changed files with 18515 additions and 5443 deletions
@@ -41,7 +41,7 @@ This creates three critical gaps for common backup scenarios:
- Maintain full backward compatibility — existing backups with no `namespacedFilterPolicies` behave exactly as they do today
- Define clear precedence rules for how per-namespace filters interact with global filters
- Add corresponding validation within the Resource Policies validation pipeline using existing Velero wildcard validation functions
- Update `velero backup describe` output to display per-namespace filter information when present
- Update `velero backup describe` output to display the referenced ResourcePolicy ConfigMap name when configured
- Ensure the restore process works correctly with backups produced by namespace-scoped filters, without requiring restore-side code changes in the initial phase
## Non-Goals
@@ -77,7 +77,8 @@ clusterScopedFilterPolicy:
names: ["my-app-*"]
- kinds: [CustomResourceDefinition]
labelSelector:
app: my-app
matchLabels:
app: my-app
namespacedFilterPolicies:
# NEW: per-namespace filter overrides
- namespaces:
@@ -85,7 +86,8 @@ namespacedFilterPolicies:
resourceFilters:
- kinds: [ConfigMap, Secret, Deployment]
labelSelector:
app: my-app
matchLabels:
app: my-app
- namespaces:
- ns-b
resourceFilters:
@@ -93,7 +95,8 @@ namespacedFilterPolicies:
names: [app-1, app-2]
- kinds: [ConfigMap]
labelSelector:
app: my-service
matchLabels:
app: my-service
```
All four sections coexist in the same ConfigMap. They are independent — `volumePolicies` handles volume backup strategy, `includeExcludePolicy` handles global resource type filtering, `clusterScopedFilterPolicy` handles cluster-scoped resource filtering by kind/name/label, and `namespacedFilterPolicies` handles per-namespace, per-kind overrides.
@@ -107,7 +110,9 @@ namespacedFilterPolicies:
- namespaces: [ns-a]
resourceFilters:
- kinds: [ConfigMap, Secret] # these kinds share a selector
labelSelector: {app: my-app}
labelSelector:
matchLabels:
app: my-app
names: ["app-*"]
- kinds: [Deployment] # this kind has its own selector
names: [workload-1, workload-2]
@@ -116,6 +121,24 @@ namespacedFilterPolicies:
This model has one way to express filters — there is no ambiguity about how to structure the configuration. Only resource kinds listed in `resourceFilters` entries are included in the backup for the matched namespaces; unlisted kinds are implicitly excluded.
#### Label selectors (`matchLabels` / `matchExpressions`)
`labelSelector` and each entry of `orLabelSelectors` use the standard Kubernetes selector shape (same as `BackupSpec.labelSelector`):
```yaml
labelSelector:
matchLabels:
app: my-app
matchExpressions:
- key: environment
operator: In
values: [prod, staging]
- key: do-not-backup
operator: DoesNotExist
```
Supported `matchExpressions` operators: `In`, `NotIn`, `Exists`, `DoesNotExist`. Prefer `In` for value-OR on one key; use `orLabelSelectors` for OR across independent multi-key groups. `labelSelector` and `orLabelSelectors` cannot co-exist in the same `resourceFilters` entry.
#### Catch-All Resource Filter (Empty `kinds` or `["*"]`)
A `ResourceFilter` entry with an empty (or omitted) `kinds` field, or a field explicitly set to `["*"]`, acts as a **catch-all**. Its `labelSelector` or `orLabelSelectors` (if provided) is applied to **all resource types in the namespace that are not already matched by a kind-specific filter entry**. If no selectors are provided, all unlisted resources are included. Using `["*"]` is highly recommended as it makes the catch-all intention explicit and self-documenting.
@@ -319,9 +342,10 @@ resourceFilters:
resourceFilters:
- kinds: ["Pod"]
labelSelector:
"invalid label key!": "value" # invalid key syntax
matchLabels:
"invalid label key!": "value" # invalid key syntax
```
**Behavior:** Validation error during backup creation when `labels.SelectorFromSet()` fails:
**Behavior:** Validation error during backup creation when `metav1.LabelSelectorAsSelector()` fails:
```
namespacedFilterPolicies[0].resourceFilters[0]: invalid label selector: "invalid label key!" is not a valid label key
```
@@ -340,7 +364,33 @@ This is consistent with how other discovery-dependent features handle this error
## ResourceFilter Field Notes
**`labelSelector`** supports equality-based selectors only (`key=value`). Set-based requirements (e.g., `environment in (prod, staging)`) are not supported. To match resources with any of several label combinations, use `orLabelSelectors` with multiple maps — each map is AND-evaluated internally, and the maps are OR-evaluated across the list. `labelSelector` and `orLabelSelectors` cannot co-exist in the same entry.
**`labelSelector`** uses the standard Kubernetes shape: `matchLabels` (equality) and `matchExpressions` (set-based: `In`, `NotIn`, `Exists`, `DoesNotExist`). All requirements within one selector are AND-ed. Example:
```yaml
labelSelector:
matchLabels:
app: my-app
matchExpressions:
- key: environment
operator: In
values: [prod, staging]
- key: do-not-backup
operator: DoesNotExist
```
**`orLabelSelectors`** is a list of the same selector shape. Match if **any** entry matches (AND within each entry, OR across the list). Prefer `In` for value-OR on one key; use `orLabelSelectors` for OR of independent multi-key groups. `labelSelector` and `orLabelSelectors` cannot co-exist in the same entry.
```yaml
orLabelSelectors:
- matchLabels:
tier: frontend
matchExpressions:
- key: track
operator: In
values: [canary]
- matchLabels:
tier: backend
```
**`names` / `excludedNames`** accept exact resource names or glob patterns. If `names` is empty, all resource names are included (subject to label filters). `excludedNames` takes precedence over `names` when a name matches both.
@@ -420,7 +470,8 @@ data:
resourceFilters:
- kinds: [ConfigMap, Secret, Deployment]
labelSelector:
app: my-app
matchLabels:
app: my-app
# ns-b has no filter policy entry, so global filters apply (include everything)
```
@@ -462,10 +513,12 @@ data:
resourceFilters:
- kinds: [Deployment]
labelSelector:
app: production-workload-1
matchLabels:
app: production-workload-1
- kinds: [StatefulSet]
labelSelector:
app: production-workload-2
matchLabels:
app: production-workload-2
```
### Per-Kind Exact Names
@@ -561,7 +614,8 @@ data:
resourceFilters:
- kinds: ["*"] # catch-all: applies to every kind not listed below
labelSelector:
backup: "true" # back up any resource carrying this label
matchLabels:
backup: "true" # back up any resource carrying this label
```
**Result:** Every resource type in `production` that has the label `backup=true` is backed up. Resources without that label are excluded. No kind enumeration is required.
@@ -589,7 +643,8 @@ data:
names: [db-credentials, tls-cert] # these exact Secrets by name
- kinds: ["*"] # catch-all for all other kinds
labelSelector:
backup: "true" # back up by label
matchLabels:
backup: "true" # back up by label
```
**Result:**
@@ -666,7 +721,8 @@ data:
names: [workload-1, workload-2]
- kinds: [StatefulSet]
labelSelector:
app: my-app
matchLabels:
app: my-app
- kinds: [ConfigMap, Secret]
names: ["app-*"]
excludedNames: ["*-tmp-*", "*-debug-*", "*-tmp", "*-debug"]
@@ -697,7 +753,7 @@ spec:
### `velero backup describe`
The output is extended to display namespace-scoped filter policies when present in the ResourcePolicy ConfigMap:
The output displays the referenced ResourcePolicy ConfigMap name when configured on the backup. It intentionally avoids resolving and displaying the live ConfigMap contents, because the ConfigMap content in the cluster may be modified or deleted after the backup execution, which could lead to displaying out-of-sync or inaccurate information:
```
Name: selective-backup
@@ -721,46 +777,9 @@ Resources:
Label selector: <none>
Resource Policy: backup-filter-policy
Namespace-Scoped Filter Policies:
ns-a:
Resource Filters:
ConfigMap, Secret, Deployment:
Label selector: app=my-app
Included names: <none>
Excluded names: <none>
target-namespace:
Resource Filters:
Deployment:
Label selector: app=production-workload-1
Included names: <none>
Excluded names: <none>
StatefulSet:
Label selector: app=production-workload-2
Included names: <none>
Excluded names: <none>
production:
Resource Filters:
Deployment:
Label selector: <none>
Included names: [api-server, worker]
Excluded names: <none>
<catch-all> (all other kinds):
Label selector: backup=true
Included names: <none>
Excluded names: <none>
Fine-Grained Global Filter Policy:
Resource Filters:
ClusterRole, ClusterRoleBinding:
Label selector: <none>
Included names: [my-app-*]
Excluded names: <none>
CustomResourceDefinition:
Label selector: app=my-app
Included names: <none>
Excluded names: <none>
Resource policies:
Type: configmap
Name: backup-filter-policy
Storage Location: default
@@ -795,7 +814,7 @@ Notes:
- Global filters (--include-resources, --selector, etc.) apply to all included namespaces
- Namespace-scoped filters defined in --resource-policies-configmap override global filters for matching namespaces
- Fine-grained global filter policies defined in --resource-policies-configmap override global filters for cluster-scoped resources
- Use 'velero backup describe' to view resolved filter policies after backup creation
- Use 'velero backup describe' to view the referenced ResourcePolicy ConfigMap name after backup creation
```
### CLI Integration Points
@@ -808,12 +827,12 @@ Notes:
**Help and Discovery:**
- `velero backup create --help` includes updated filtering documentation
- `velero backup describe` shows resolved filter policies for troubleshooting
- `velero backup describe` shows the referenced ResourcePolicy ConfigMap name
- Validation errors include ConfigMap field references for easy debugging
**Configuration Discovery:**
- `velero backup create --help` includes namespace-scoped filtering documentation
- `velero backup describe` shows resolved filter policies for verification
- `velero backup describe` shows the referenced ResourcePolicy ConfigMap name for verification
## User Perspective
@@ -823,7 +842,7 @@ This design provides fine-grained, per-namespace, per-kind control over backup f
- **For users adopting namespace-scoped filter policies**: Create a ConfigMap with the `namespacedFilterPolicies` section and reference it via `BackupSpec.ResourcePolicy` (or the existing `--resource-policies-configmap` flag). The backup will selectively include/exclude resources per namespace based on the filter rules.
- **For users already using ResourcePolicy for volume policies**: Add the `namespacedFilterPolicies` section to the same ConfigMap. Both volume policies and namespace-scoped filters coexist.
- **For restore from a namespace-filtered backup**: No changes to restore workflow. Restore processes whatever is in the archive. Users can use existing `RestoreSpec.IncludedNamespaces` for additional filtering at restore time.
- **`velero backup describe` output**: Extended to show per-namespace, per-kind filter details when the ResourcePolicy ConfigMap contains `namespacedFilterPolicies`.
- **`velero backup describe` output**: Displays the referenced ResourcePolicy ConfigMap name when configured on the backup.
- **Validation errors**: Reported at backup start when the ResourcePolicy ConfigMap contains invalid `namespacedFilterPolicies` configurations. Consistent with how volume policy validation errors are reported today.
## Alternatives Considered
@@ -108,14 +108,16 @@ clusterScopedFilterPolicy:
names: ["my-app-*"]
- kinds: [CustomResourceDefinition]
labelSelector:
app: my-app
matchLabels:
app: my-app
namespacedFilterPolicies:
- namespaces:
- ns-a
resourceFilters:
- kinds: [ConfigMap, Secret, Deployment]
labelSelector:
app: my-app
matchLabels:
app: my-app
- namespaces:
- ns-b
resourceFilters:
@@ -123,7 +125,8 @@ namespacedFilterPolicies:
names: [app-1, app-2]
- kinds: [ConfigMap]
labelSelector:
app: my-service
matchLabels:
app: my-service
```
The restore-side ConfigMap does **not** require `volumePolicies` or `includeExcludePolicy` sections. Those are backup-specific. The YAML parser will ignore unknown fields gracefully, so a user can technically point to the same ConfigMap used for backup — the restore pipeline will only read `namespacedFilterPolicies` and `clusterScopedFilterPolicy`.
@@ -137,7 +140,9 @@ namespacedFilterPolicies:
- namespaces: [ns-a]
resourceFilters:
- kinds: [ConfigMap, Secret] # these kinds share a selector
labelSelector: {app: my-app}
labelSelector:
matchLabels:
app: my-app
names: ["app-*"]
- kinds: [Deployment] # this kind has its own selector
names: [workload-1, workload-2]
@@ -146,6 +151,36 @@ namespacedFilterPolicies:
Only resource kinds listed in `resourceFilters` entries are restored for the matched namespaces; unlisted kinds are implicitly excluded (globally excluded kinds cannot be re-included — see precedence model).
#### Label selectors (`matchLabels` / `matchExpressions`)
`labelSelector` and each entry of `orLabelSelectors` use the standard Kubernetes selector shape (same as `RestoreSpec.labelSelector`):
```yaml
labelSelector:
matchLabels:
app: my-app
matchExpressions:
- key: environment
operator: In
values: [prod, staging]
- key: do-not-restore
operator: DoesNotExist
```
Supported `matchExpressions` operators: `In`, `NotIn`, `Exists`, `DoesNotExist`. Prefer `In` for value-OR on one key; use `orLabelSelectors` for OR across independent multi-key groups. `labelSelector` and `orLabelSelectors` cannot co-exist in the same `resourceFilters` entry.
```yaml
orLabelSelectors:
- matchLabels:
tier: frontend
matchExpressions:
- key: track
operator: In
values: [canary]
- matchLabels:
tier: backend
```
#### Peek-and-Map Fallback for Unresolved Kinds
The `kinds` field accepts both plural resource names (e.g., `configmaps`, `mycustomkinds.mygroup.io`) and singular `Kind` names (e.g., `ConfigMap`, `MyCustomKind`).
@@ -382,9 +417,10 @@ resourceFilters:
resourceFilters:
- kinds: ["Deployment"]
labelSelector:
"invalid label key!": "value" # invalid key syntax
matchLabels:
"invalid label key!": "value" # invalid key syntax
```
**Behavior:** Validation error during restore creation when `labels.ValidatedSelectorFromSet()` fails:
**Behavior:** Validation error during restore creation when `metav1.LabelSelectorAsSelector()` fails:
```
namespacedFilterPolicies[0].resourceFilters[0]: invalid label selector: "invalid label key!" is not a valid label key
```
@@ -420,7 +456,8 @@ namespacedFilterPolicies:
resourceFilters:
- kinds: [ConfigMap, Secret] # Secret listed here is ineffective — globally excluded
labelSelector:
app: my-app
matchLabels:
app: my-app
- kinds: [Deployment]
```
@@ -461,8 +498,8 @@ After existing filter setup, the filter policies are resolved into the runtime m
The `resolveRestoreNamespacedFilterPolicies` function:
- For each `NamespacedFilterPolicy`, iterates its `ResourceFilters` entries
- Resolves kind names to fully-qualified group-resource strings using the discovery helper
- Converts `labelSelector` maps into `labels.Selector` objects using `labels.ValidatedSelectorFromSet()`
- Converts `orLabelSelectors` maps into `[]labels.Selector`
- Converts `labelSelector` into a `labels.Selector` via `ToMetaV1LabelSelector` + `metav1.LabelSelectorAsSelector()`
- Converts `orLabelSelectors` into `[]labels.Selector` the same way
- Creates `IncludesExcludes` instances for `names`/`excludedNames` patterns
- Identifies catch-all entries (empty or `["*"]` kinds) and stores them in `catchAllFilter`
- Builds a `resourceFilterMap` keyed by the resolved group-resource string
@@ -537,7 +574,8 @@ data:
resourceFilters:
- kinds: [Deployment, ConfigMap]
labelSelector:
app: my-app
matchLabels:
app: my-app
# ns-b has no filter policy entry, so global filters apply (restore everything)
```
@@ -631,7 +669,8 @@ data:
names: [db-credentials, tls-cert] # these exact Secrets by name
- kinds: ["*"] # catch-all for all other kinds
labelSelector:
backup: "true" # restore by label
matchLabels:
backup: "true" # restore by label
```
**Result:**
@@ -658,7 +697,8 @@ data:
names: ["my-app-*"]
- kinds: [CustomResourceDefinition]
labelSelector:
app: my-app
matchLabels:
app: my-app
namespacedFilterPolicies:
- namespaces:
- production
@@ -0,0 +1,357 @@
# RestoreItemAction Must-Include Additional Items
## Abstract
Backup Item Actions (BIAs) can already mark additional items as must-include via `backup.velero.io/must-include-additional-items`, so Velero bypasses resource and namespace exclusion filters when backing those dependencies up.
This proposal adds the same plugin-controlled escape hatch on restore: `restore.velero.io/must-include-additional-items`, so Restore Item Actions (RIAs) can force-restore declared `AdditionalItems` even when they would otherwise be dropped by global restore filters.
## Glossary & Abbreviation
**Additional Item**: A resource identifier returned by a Backup/Restore Item Action's `Execute()` result that Velero should process as a dependency of the current item.
**BIA**: Backup Item Action plugin.
**RIA**: Restore Item Action plugin.
**Must-Include**: A plugin-set annotation on the action's `UpdatedItem` that tells Velero to bypass global include/exclude filters for that action's `AdditionalItems`.
**Global Restore Filter**: `RestoreSpec` filters applied uniformly — `IncludedNamespaces`/`ExcludedNamespaces`, `IncludedResources`/`ExcludedResources`, `IncludeClusterResources`, and label selectors.
**Fine-Grained Restore Filter**: Per-namespace / cluster-scoped policies from `RestoreSpec.ResourcePolicy` (`namespacedFilterPolicies`, `clusterScopedFilterPolicy`), as described in [Fine Grained Restore Filters via Resource Policies](https://github.com/velero-io/velero/blob/main/design/restore-filter-enhancement/fine-grained-restore-filters-design.md).
**`resourceMustHave`**: A small hardcoded server-side set of resource types that bypass resource and namespace I/E checks inside `restoreItem()` today (but not `IncludeClusterResources=false`).
## Background
### Backup-side precedent
On backup, a BIA may set `backup.velero.io/must-include-additional-items: "true"` on the returned `UpdatedItem`.
Velero strips that annotation (it is an internal signal, not intended to land on the live object) and passes `mustInclude=true` into recursive `backupItem` calls for that action's `AdditionalItems`.
When `mustInclude` is true, `itemInclusionChecks` skips namespace/resource exclusion checks (and related exclusion labels / fine-grained name filters) so plugin-declared dependencies are not dropped by the user's backup filters.
In-tree CSI BIAs already rely on this for VolumeSnapshot / VolumeSnapshotContent / VolumeSnapshotClass style dependency chains.
### Restore-side gap
On restore, RIAs can return `AdditionalItems`, and Velero recursively calls `restoreItem()` for each of them.
That path already bypasses fine-grained restore filters and global label selectors, because those are evaluated earlier in `getOrderedResourceCollection` / `getSelectedRestoreableItems`.
However, `restoreItem()` still enforces global resource includes/excludes, namespace includes/excludes, and `IncludeClusterResources=false`.
The fine-grained restore filters design explicitly documents this remaining floor:
> Note that these additional items must still pass global resource/namespace exclusions.
There is no restore-side equivalent of the BIA must-include annotation.
Plugins that need a hard dependency restored despite a selective restore configuration have no opt-in way to express that, short of relying on the server-side `resourceMustHave` list (which is global, not plugin-scoped, and does not bypass `IncludeClusterResources=false`).
### Motivating scenario
Consider a selective restore that includes only application namespaces and excludes storage/snapshot resource types, while a plugin knows that restoring a PVC correctly requires a related cluster-scoped or cross-namespace dependency that exists in the backup archive.
Today the RIA can request that dependency as an `AdditionalItem`, but Velero will skip it at the global exclusion checks inside `restoreItem()`.
With a restore must-include annotation, the plugin can declare the dependency as required and Velero will restore it (provided the object is present in the backup tarball).
## Goals
- Add `restore.velero.io/must-include-additional-items` with the same parent-annotation contract as the backup-side must-include annotation.
- When an RIA sets the annotation on `UpdatedItem`, bypass global resource I/E, namespace I/E, and `IncludeClusterResources=false` for that RIA's `AdditionalItems`.
- Keep the change opt-in and backward compatible: restores and plugins that do not set the annotation behave exactly as today.
- Document the trust model, precedence rules, and interaction with existing restore gates for plugin authors and operators.
## Non-Goals
- Changing the plugin protobuf / `RestoreItemAction` interface shape (no new RPC fields).
- Changing CRDs or adding CLI flags.
- Changing the `resourceMustHave` list (including any narrowing related to VolumeSnapshotContent).
- Updating in-tree RIAs (CSI or otherwise) to set the new annotation as part of this change.
- Per-additional-item granularity (the annotation applies blanket to all `AdditionalItems` from that RIA invocation, matching BIA).
- Materializing items that were never backed up.
## High-Level Design
Mirror the backup workflow:
1. Introduce annotation constant `restore.velero.io/must-include-additional-items`.
2. After each RIA `Execute()`, if `UpdatedItem` carries the annotation with value `"true"`, strip it and set `mustIncludeAdditionalItems=true`.
3. Pass that boolean into recursive `restoreItem(..., mustInclude)` calls for the action's `AdditionalItems`.
4. When `mustInclude` is true, skip the global resource/namespace/`IncludeClusterResources` exclusion checks inside `restoreItem()`.
5. Keep all non-filter gates unchanged (tarball presence, already-restored, completed Jobs, API errors, wait-for-additional-items, etc.).
Top-level items from the archive continue to be restored with `mustInclude=false`, so user filters still apply to the primary restore set.
```mermaid
flowchart TD
startRestore[Start Restore] --> readTarball[Read Item from Backup Tarball]
readTarball --> topLevelRestoreItem["restoreItem(..., mustInclude=false)"]
topLevelRestoreItem --> checkMustInclude{"mustInclude == true?"}
checkMustInclude -- No --> checkFilters{"Pass Global Resource/Namespace Filters?"}
checkFilters -- No --> skipItem[Skip Restore]
checkFilters -- Yes --> nonFilterGates["Other gates: isCompleted, already-restored, ..."]
checkMustInclude -- Yes --> nonFilterGates
nonFilterGates --> executeRIA[Execute RestoreItemAction]
executeRIA --> checkSkip{"SkipRestore?"}
checkSkip -- Yes --> skipItem
checkSkip -- No --> checkAnnotation{"Has must-include annotation?"}
checkAnnotation -- Yes --> stripAnnotation[Strip Annotation]
stripAnnotation --> setFlagTrue["mustIncludeAdditionalItems = true"]
checkAnnotation -- No --> setFlagFalse["mustIncludeAdditionalItems = false"]
setFlagTrue --> loopAdditionalItems[Loop over AdditionalItems]
setFlagFalse --> loopAdditionalItems
loopAdditionalItems --> existsInBackup{"Item file in tarball?"}
existsInBackup -- No --> warnSkip[Warn and skip]
existsInBackup -- Yes --> recursiveRestoreItem["restoreItem(..., mustInclude=mustIncludeAdditionalItems)"]
recursiveRestoreItem --> checkMustInclude
```
> The edge `recursiveRestoreItem --> checkMustInclude` is a recursive call (new `restoreItem` stack frame), not a same-frame loop.
## Detailed Design
### Annotation constant
In `pkg/apis/velero/v1/labels_annotations.go`, next to the existing backup constant:
```go
// Velero checks this annotation to determine whether to skip resource excluding check.
MustIncludeAdditionalItemAnnotation = "backup.velero.io/must-include-additional-items"
// MustIncludeAdditionalItemRestoreAnnotation is set by RestoreItemActions on the UpdatedItem
// to tell Velero to bypass global resource/namespace exclusion checks (and IncludeClusterResources=false)
// for that action's AdditionalItems. Value must be "true". The annotation is stripped before
// the item is applied to the cluster.
//
// Notice: SkipRestore on the Execute output takes precedence. If SkipRestore is true, the
// annotation is never inspected and AdditionalItems are not processed.
MustIncludeAdditionalItemRestoreAnnotation = "restore.velero.io/must-include-additional-items"
```
Only the string value `"true"` enables the bypass (same as backup).
### `restoreItem` signature
```go
func (ctx *restoreContext) restoreItem(
obj *unstructured.Unstructured,
groupResource schema.GroupResource,
namespace string,
mustInclude bool,
) (results.Result, results.Result, bool)
```
Call sites:
| Site | `mustInclude` value |
|---|---|
| Top-level restore loop | `false` |
| Recursive additional-item restore after an RIA | derived from that RIA's `UpdatedItem` annotation |
### Bypass exclusion checks; keep namespace creation
Today, namespace exclusion and `EnsureNamespaceExistsAndIsReady` share one `if namespace != ""` block in `restoreItem()`.
If must-include only skipped the exclusion check without refactoring, an additional item targeting an excluded namespace would fail because its target namespace was never ensured.
Required structure:
```go
if mustInclude {
restoreLogger.Info("Skipping the resource/namespace exclusion checks because the item is marked as must-include")
} else {
if !ctx.resourceIncludesExcludes.ShouldInclude(groupResource.String()) && !ctx.resourceMustHave.Has(groupResource.String()) {
restoreLogger.Info("Not restoring item because resource is excluded")
return warnings, errs, itemExists
}
if namespace != "" {
if !ctx.namespaceIncludesExcludes.ShouldInclude(obj.GetNamespace()) && !ctx.resourceMustHave.Has(groupResource.String()) {
restoreLogger.Info("Not restoring item because namespace is excluded")
return warnings, errs, itemExists
}
} else {
if boolptr.IsSetToFalse(ctx.restore.Spec.IncludeClusterResources) {
restoreLogger.Info("Not restoring item because it's cluster-scoped")
return warnings, errs, itemExists
}
}
}
// Namespace creation runs regardless of mustInclude.
if namespace != "" {
nsToEnsure := getNamespace(restoreLogger, archive.GetItemFilePath(ctx.restoreDir, "namespaces", "", obj.GetNamespace()), namespace)
_, nsCreated, err := kube.EnsureNamespaceExistsAndIsReady(nsToEnsure, ctx.namespaceClient, ctx.resourceTerminatingTimeout, ctx.resourceDeletionStatusTracker)
// ... existing error handling and restoredItems bookkeeping ...
}
```
Namespace remapping is unchanged: exclusion checks use the original namespace (`obj.GetNamespace()`); namespace creation uses the remapped target `namespace` parameter.
### Process the annotation after each RIA
Inside the applicable-actions loop in `restoreItem()`, after `SkipRestore` handling and type-asserting `UpdatedItem`:
```go
obj = unstructuredObj
mustIncludeAdditionalItems := false
if annotations := obj.GetAnnotations(); annotations != nil &&
annotations[velerov1api.MustIncludeAdditionalItemRestoreAnnotation] == "true" {
mustIncludeAdditionalItems = true
restoreLogger.Info("RestoreItemAction marked additional items as must-include; bypassing resource/namespace exclusion checks for them")
delete(annotations, velerov1api.MustIncludeAdditionalItemRestoreAnnotation)
obj.SetAnnotations(annotations)
}
for _, additionalItem := range executeOutput.AdditionalItems {
// existing tarball stat / unmarshal / namespace mapping ...
w, e, additionalItemExists := ctx.restoreItem(
additionalObj,
additionalItem.GroupResource,
additionalItemNamespace,
mustIncludeAdditionalItems,
)
// existing merge / filteredAdditionalItems bookkeeping ...
}
```
### Filter bypass matrix
| Gate | Plain AdditionalItem | `resourceMustHave` | RIA `mustInclude=true` | BIA `mustInclude=true` (parity target) |
|---|---|---|---|---|
| Fine-grained policies (kind/name/label) | Bypass (never enter selection Phase B filters) | N/A in `restoreItem` | Bypass (same) | Bypass |
| Global label selectors | Bypass (never re-enter selection) | N/A in `restoreItem` | Bypass (same) | Bypass |
| Global resource I/E | Honored | Bypass | Bypass | Bypass |
| Global namespace I/E | Honored | Bypass | Bypass | Bypass |
| `IncludeClusterResources=false` | Honored | Honored (not bypassed) | Bypass | Bypass |
| Item must exist in backup tarball | Required | Required | Required | N/A (fetched from cluster) |
| `isCompleted` / already-restored / API errors | Still apply | Still apply | Still apply | `DeletionTimestamp` still applies on backup |
RIA must-include is intentionally a **stronger** override than `resourceMustHave` because it also bypasses `IncludeClusterResources=false`.
That matches BIA must-include semantics (plugin-trusted hard dependencies), rather than widening the hardcoded server list.
### Interaction with fine-grained restore filters
Per [Fine Grained Restore Filters via Resource Policies](../restore-filter-enhancement/fine-grained-restore-filters-design.md), plugin additional items already bypass `namespacedFilterPolicies` / `clusterScopedFilterPolicy` kind, name, and label checks.
Those filters live in the selection phases; additional items enter `restoreItem()` directly.
This proposal only changes the remaining global gates inside `restoreItem()`.
With must-include set, an additional item effectively bypasses **all** restore filters (fine-grained and global).
Without the annotation, behavior is unchanged: fine-grained filters are still bypassed, global exclusions still apply.
### Interaction with existing restore gates
#### `SkipRestore` precedence
If `Execute()` returns `SkipRestore: true`, `restoreItem()` returns before inspecting the annotation, and no `AdditionalItems` are processed.
This mirrors backup-side precedence where `velero.io/skip-from-backup` outranks must-include.
#### Multi-RIA semantics
Annotation handling is per RIA invocation inside the actions loop:
1. RIA N executes → inspect/strip annotation on that `UpdatedItem` → restore that RIA's `AdditionalItems` with the derived flag.
2. RIA N+1 sees the already-stripped object unless it sets the annotation again.
A later RIA does not inherit an earlier RIA's must-include decision.
#### Transitive propagation
The parent's `mustInclude` flag admits the child additional item through filters.
It does **not** automatically force-include grandchildren.
Each RIA level that needs the escape hatch must set the annotation on its own `UpdatedItem`, matching BIA behavior.
#### Non-filter gates that still apply
Even when `mustInclude=true`:
- Missing archive file → warn and skip (existing behavior).
- `isCompleted` resources (e.g. completed Jobs) → skip.
- Already present in `ctx.restoredItems` → skip.
- Create/update API failures → errors as today.
- `WaitForAdditionalItems` / `AreAdditionalItemsReady` polling after the additional-item loop → unchanged.
### Relationship to `resourceMustHave`
| Mechanism | Who decides | Bypasses resource/ns I/E | Bypasses `IncludeClusterResources=false` |
|---|---|---|---|
| `resourceMustHave` | Velero server (hardcoded) | Yes | No |
| RIA must-include | Plugin author (annotation) | Yes | Yes |
The two mechanisms coexist.
This proposal does not migrate in-tree CSI (or other) RIAs onto the annotation.
Doing so would be a separate behavior change: it could force-restore types users explicitly excluded, and would newly restore cluster-scoped dependencies even when `IncludeClusterResources=false`.
### Plugin usage sketch
```go
func (p *myRestoreAction) Execute(input *velero.RestoreItemActionExecuteInput) (*velero.RestoreItemActionExecuteOutput, error) {
item := input.Item.(*unstructured.Unstructured)
annotations := item.GetAnnotations()
if annotations == nil {
annotations = map[string]string{}
}
annotations[velerov1api.MustIncludeAdditionalItemRestoreAnnotation] = "true"
item.SetAnnotations(annotations)
return &velero.RestoreItemActionExecuteOutput{
UpdatedItem: item,
AdditionalItems: []velero.ResourceIdentifier{
{GroupResource: schema.GroupResource{Group: "example.io", Resource: "dependencies"}, Namespace: "dep-ns", Name: "dep-1"},
},
}, nil
}
```
Plugin authors must ensure the additional item was actually captured in the backup (typically via the corresponding BIA also using `backup.velero.io/must-include-additional-items`).
### Tests
Extend restore coverage (existing `TestRestoreActionAdditionalItems` patterns / focused cases) for:
1. Resource exclusion bypass with annotation; still skipped without annotation.
2. Namespace exclusion bypass **and** target namespace creation.
3. `IncludeClusterResources=false` bypass for cluster-scoped additional items.
4. Annotation stripped from the object applied to the cluster.
5. `SkipRestore: true` prevents additional-item processing even if the annotation is set.
6. Missing tarball entry still warns and skips.
7. Transitive case: child RIA must re-set the annotation for grandchildren.
8. Top-level restore path still passes `mustInclude=false` and honors filters.
### Documentation
- Constant doc comment (including `SkipRestore` precedence).
- Plugin-author docs for Restore Item Actions: annotation key/value, blanket scope, filter-bypass matrix, namespace-creation side effect, tarball requirement.
## Security Considerations
Installing an RIA that sets this annotation grants that plugin authority to restore dependencies outside the operator's restore filters, including:
- resources in namespaces the restore excluded (and creation of those target namespaces if needed);
- resource types the restore excluded;
- cluster-scoped resources even when `IncludeClusterResources=false`.
This matches the existing BIA trust model: item-action plugins are already privileged components of the Velero deployment.
Operators should treat RIA installation as a trust decision.
The annotation is stripped before apply so it does not persist as attacker-controlled cluster state from the backup archive alone; a matching RIA must run and return `AdditionalItems` for the bypass to take effect.
## Compatibility
- No CRD or plugin interface changes.
- Existing restores unchanged when no RIA sets the annotation.
- Existing tests that assert additional items are dropped under namespace filters / `IncludeClusterResources=false` remain valid for the no-annotation path.
- Compatible with fine-grained restore filters: additional items already bypass those filters; this proposal only addresses the documented global-exclusion floor.
## Alternatives Considered
### Per-item must-include on each `ResourceIdentifier`
Pros: selective control within one `AdditionalItems` list.
Cons: requires API changes to `ResourceIdentifier` or a parallel structure; diverges from BIA; plugins that need selectivity can already split across actions or omit non-required items.
Rejected for this proposal; may be revisited later if plugin authors demonstrate a concrete need.
### Widen `resourceMustHave` instead of a plugin annotation
Pros: no plugin contract change.
Cons: server-forced, global, not scoped to a plugin call; does not give third-party plugins a general tool; does not match BIA; conflicts with efforts to keep hardcoded force-include lists narrow.
Rejected — wrong trust model for a general plugin escape hatch.
@@ -0,0 +1,368 @@
# Volume Data In-place Full/Incremental Restore
## Table of Contents
- [Background](#background)
- [Goals](#goals)
- [Non-Goals](#non-goals)
- [Overview](#overview)
- [Detailed Design](#detailed-design)
- [CRD Changes](#crd-changes)
- [CLI](#cli)
- [Workload Management](#workload-management)
- [Handling Cross-Zone Scheduling (WaitForFirstConsumer)](#handling-cross-zone-scheduling-waitforfirstconsumer)
- [Namespace Mapping](#namespace-mapping)
- [Pre-flight Checks](#pre-flight-checks)
- [1. PVC is Not Actively Used by a Running Pod](#1-pvc-is-not-actively-used-by-a-running-pod)
- [2. PVC is Bound to the Original PV](#2-pvc-is-bound-to-the-original-pv)
- [3. Volume Size Validation](#3-volume-size-validation)
- [Error Handling](#error-handling)
- [Restore Workflow Update](#restore-workflow-update)
- [In-place Incremental Restore for CSI Snapshot with Block Data Move for Block Volumes](#in-place-incremental-restore-for-csi-snapshot-with-block-data-move-for-block-volumes)
- [In-place Full Restore for CSI Snapshot with Block Data Move for Block Volumes](#in-place-full-restore-for-csi-snapshot-with-block-data-move-for-block-volumes)
- [In-place Incremental Restore for CSI Snapshot with File System Data Move for File System Volumes](#in-place-incremental-restore-for-csi-snapshot-with-file-system-data-move-for-file-system-volumes)
- [In-place Full Restore for CSI Snapshot with File System Data Move for File System Volumes](#in-place-full-restore-for-csi-snapshot-with-file-system-data-move-for-file-system-volumes)
- [In-place Incremental Restore for CSI Snapshot with Block Data Move for File System Volumes](#in-place-incremental-restore-for-csi-snapshot-with-block-data-move-for-file-system-volumes)
- [In-place Full Restore for CSI Snapshot with Block Data Move for File System Volumes](#in-place-full-restore-for-csi-snapshot-with-block-data-move-for-file-system-volumes)
- [In-place Incremental Restore for File System Backup for File System Volumes](#in-place-incremental-restore-for-file-system-backup-for-file-system-volumes)
- [In-place Full Restore for File System Backup for File System Volumes](#in-place-full-restore-for-file-system-backup-for-file-system-volumes)
- [Installation](#installation)
- [Upgrade](#upgrade)
## Background
Currently, Velero only supports restoring volume data to a newly provisioned PVC. If the target PVC already exists in the cluster, Velero skips the data restoration entirely and leaves the existing volume untouched.
This design introduces the "in-place restore" capability, allowing Velero to restore volume data directly into an existing, bound PVC. When performing an in-place restore, users can choose to either overwrite the volume entirely (in-place full restore) or only restore the modified data to optimize performance (in-place incremental restore).
To ensure data consistency and allow Velero to safely recreate the PVC during the process, users must manually delete any pods consuming the target volume before initiating an in-place restore.
## Goals
- Enable Velero to restore volume data directly into an existing, bound PVC without requiring the user to manually delete the PVC and PV.
- Support both Full (overwrite all) and Incremental (overwrite only changed data) in-place restores.
- Support in-place restores for Windows workloads.
- Ensure data consistency and correct Kubernetes scheduling constraints (e.g., handling `WaitForFirstConsumer` and zonal topologies) are respected during and after the restore.
## Non-Goals
- Automating the deletion of workloads before the restore. It remains the user's responsibility to ensure the volume is not actively consumed and the Pods are completely removed before triggering the restore to prevent data corruption and allow PVC recreation.
- In-place restore for CSI snapshot without data move.
- In-place restore for Native Snapshots (cloud provider snapshots without CSI).
- Fine-grained, per-volume control over in-place restores. The newly introduced in-place restore policies apply globally to all volumes within a single restore operation. Allowing users to specify different restore strategies for individual volumes is deferred to a future enhancement.
## Overview
This design focuses exclusively on volume data restoration. To support this, we are introducing a new field, `ExistingVolumeDataPolicy`, to the `Restore` spec. This feature operates independently of Kubernetes resource restoration, which remains controlled by the existing `ExistingResourcePolicy` field.
Depending on how unchanged data is handled during the restoration process, in-place volume data restores are categorized into two types:
- **In-place full restore**: Overwrites the volume with the backup data, regardless of whether the existing data has changed.
- **In-place incremental restore**: Optimizes the process by restoring only the data that has changed since the backup, leaving unmodified data intact. This is achieved by leveraging Changed Block Tracking (CBT) for block data and file metadata comparisons for file system data.
Support for in-place full and incremental restores varies depending on the underlying backup method, as detailed in the following table:
| Backup Method | In-place Full Restore | In-place Incremental Restore |
| --------------------------------------- | --------------------- | ---------------------------- |
| CSI Snapshot with Block Data Move | Yes | Yes |
| CSI Snapshot with File System Data Move | Yes | Yes |
| CSI Snapshot without Data Move | No | No |
| File System Backup | Yes | Yes |
| Native Snapshot | No | No |
Additionally, a new boolean field `DeleteExtraFiles` is added to the `UploaderConfig` within the `Restore` spec. When performing a file system restore (either via PodVolumeBackup or CSI File System Data Move), this flag controls whether files present in the target volume but absent in the backup should be deleted. Setting this to `true` ensures the target volume's file system exactly mirrors the backup state. Note that this setting is ignored for block data mover restores, as block-level operations inherently overwrite the entire file system structure.
Because Velero must create a temporary restore Pod in the Velero namespace to mount the volume and restore the data, it cannot directly use the existing PVC, which resides in the workload namespace. Velero must delete the existing PVC, recreate a temporary restore PVC in the Velero namespace, and bind it to the existing PV. The core strategy for implementing an in-place restore involves the following sequence:
```mermaid
flowchart TD
subgraph PVC CSI RIA
A[Patch existing PV's reclaim policy to Retain] --> B[Delete existing PVC]
end
subgraph Exposer
B --> C[Create temporary restore PVC in Velero namespace<br>and bind it to existing PV]
C --> D[Create temporary restore Pod<br>that mounts temporary restore PVC]
end
subgraph Block/File System Uploader
D --> E[Restore data directly into the volume]
end
subgraph Exposer Post-Restore
E --> F[Delete temporary restore Pod and PVC]
end
F --> G[Target workload Pod mounts target PVC<br>once it is recreated]
```
When restoring a file system volume using the block data mover, the PV must temporarily have its `volumeMode` set to `Block` so the restore Pod can mount it as a raw block device. Because the `volumeMode` field in a PV spec is immutable, reusing the existing PV directly is not possible. Instead, Velero must delete the existing PV and create a temporary one. The sequence for this scenario is as follows:
```mermaid
flowchart TD
subgraph PVC CSI RIA
A[Patch existing PV's reclaim policy to Retain] --> B[Delete existing PVC]
end
subgraph Exposer
B --> C[Delete existing PV]
C --> D[Create temporary restore PV with volumeMode: Block<br>using same volume handle]
D --> E[Create temporary restore PVC in Velero namespace<br>with volumeMode: Block and bind to temporary PV]
E --> F[Create temporary restore Pod<br>that mounts temporary restore PVC]
end
subgraph Block Uploader
F --> G[Restore data directly into the volume]
end
subgraph Exposer Post-Restore
G --> H[Delete temporary restore Pod, PVC, and PV]
H --> I[Recreate original PV with volumeMode: Filesystem]
end
I --> J[Recreate original PVC in workload namespace<br>and allow it to bind to recreated PV]
```
## Detailed Design
### CRD Changes
To support the new in-place restore policies and incremental data transfer, several Custom Resource Definitions (CRDs) will be updated.
**Restore CRD**
A new field `existingVolumeDataPolicy` is added to the `Restore` spec to allow users to define how existing volume data should be handled. Additionally, a new field `deleteExtraFiles` is added to the `uploaderConfig` to control file deletion during file system restores.
```yaml
spec:
existingVolumeDataPolicy: "" # Valid values: "", none, full, incremental
uploaderConfig:
deleteExtraFiles: false
```
- `existingVolumeDataPolicy`:
- `""` (default) or `none`: Do not restore volume data if the target PVC already exists.
- `full`: Perform an in-place full restore, overwriting all existing data on the volume.
- `incremental`: Perform an in-place incremental restore, only overwriting data that has changed since the backup.
- `uploaderConfig.deleteExtraFiles`: A boolean flag that controls whether files present in the target volume but absent from the backup should be deleted. **Note:** This setting is *only* applicable to File System restores (PodVolumeBackup or CSI File System Data Move) and has no effect on Block Data Move restores. Furthermore, it is ignored for non-in-place restores (where `existingVolumeDataPolicy` is not set to `full` or `incremental`).
If the target PVC does not exist, Velero will fall back to its default behavior and provision a new PVC for the restore, regardless of whether `existingVolumeDataPolicy` is set to `full` or `incremental`. Furthermore, if `existingVolumeDataPolicy` is set to `incremental` but the underlying storage does not support incremental restores, Velero will automatically fall back to a `full` restore.
The following table summarizes the expected behavior for different combinations of `existingResourcePolicy` and `existingVolumeDataPolicy` when the target PVC already exists:
| `existingResourcePolicy` | `existingVolumeDataPolicy` | PVC Resource Action | Volume Data Restore |
| ------------------------ | -------------------------- | ------------------- | ------------------- |
| `none` | `none` | Untouched | Untouched |
| `none` | `full` | Untouched | Full |
| `none` | `incremental` | Untouched | Incremental |
| `update` | `none` | Patched | Untouched |
| `update` | `full` | Patched | Full |
| `update` | `incremental` | Patched | Incremental |
**DataDownload CRD**
To support incremental restores, the `DataDownload` spec is extended with a new `restoreType` string flag (valid values are `full` and `incremental`) to instruct the data mover to perform an incremental restore. It also introduces a new `csiSnapshot` field, which captures the metadata of a snapshot taken from the existing PVC, acting as the baseline for Changed Block Tracking (CBT) delta calculations during an in-place incremental block restore. Additionally, the `deleteExtraFiles` configuration is passed to the underlying data mover via the existing `dataMoverConfig` map.
```yaml
spec:
restoreType: "incremental"
csiSnapshot:
volumeSnapshot: ""
storageClass: ""
snapshotClass: ""
driver: ""
```
- `restoreType`: A string flag indicating whether the data mover should perform a `full` or `incremental` restore.
- `csiSnapshot`:
- `volumeSnapshot`: the name of the volume snapshot
- `storageClass`: the name of the storage class of the PVC that the volume snapshot is created from
- `snapshotClass`: the name of the snapshot class that the volume snapshot is created with
- `driver`: the driver used by the VolumeSnapshotContent
**PodVolumeRestore CRD**
A new `restoreType` string flag (valid values are `full` and `incremental`) is added to the `PodVolumeRestore` spec to instruct the file system data mover (e.g., Kopia) to perform an incremental restore. Additionally, the `deleteExtraFiles` configuration is passed to the underlying uploader via the existing `uploaderSettings` map.
```yaml
spec:
restoreType: "incremental"
```
- `restoreType`: A string flag indicating whether the data mover should perform a `full` or `incremental` restore.
### CLI
New flags will be added to the `velero restore create` command to support the new policy:
- `--existing-volume-data-policy`: Accepts the values `none`, `full`, or `incremental`, mapping to `existingVolumeDataPolicy`.
- `--delete-extra-files`: A boolean flag mapping to `uploaderConfig.deleteExtraFiles`.
### Workload Management
To ensure data consistency and allow for necessary configuration changes, users must delete any Pods actively using the target volume before initiating an in-place restore. This is required for three primary reasons:
1. **Preventing Data Corruption:** It is critical to prevent the active workload Pods and the temporary restore Pods from writing to the volume simultaneously, which would lead to data corruption.
2. **PVC Recreation:** Velero creates a temporary restore Pod in the Velero namespace to mount the volume and restore the data. Since it cannot directly use the existing PVC located in the workload namespace, Velero must delete the existing PVC, create a temporary restore PVC in the Velero namespace, and bind it to the existing PV. However, Kubernetes' `pvc-protection` finalizer prevents the deletion of any PVC actively used by a running Pod. Consequently, simply pausing the workload is insufficient; the Pods must be completely removed to allow the PVC deletion to proceed.
3. **ReadWriteOncePod Access Mode:** If the volume is configured with the `ReadWriteOncePod` access mode, Kubernetes strictly enforces that the volume can only be mounted by a single Pod at a time. The existing workload Pod must be completely deleted to release the volume, allowing Velero's temporary restore Pod to successfully mount it and perform the data transfer.
Users must manage the lifecycle of their workloads before starting the restore. This applies to various workload types:
- **Standard Controllers (Deployments, StatefulSets, Jobs, CronJobs):** The required action depends on the restore method:
- **For CSI Snapshot Restores:** Users can scale these controllers down to zero replicas to terminate the underlying Pods.
- **For File System Restores (PodVolumeRestore):** Users must completely delete the controllers. Simply scaling down to zero is insufficient because file system restores rely on an init container injected into the restored target Pod to process the data transfer. If the controller is only scaled down, it will immediately terminate the Pod restored by Velero to maintain its zero-replica count. Although the controller may subsequently spawn a new Pod, that new Pod will lack the required restore init container, causing the restore to fail.
- **DaemonSets:** Since Kubernetes lacks a mechanism to scale DaemonSets to zero, users must either delete the DaemonSet entirely or use node selectors/cordoning to evict the Pods.
- **Operator-Managed Pods:** Custom controllers (like ArgoCD) may have fast reconciliation loops that aggressively recreate Pods. These operators must be paused or suspended, and their managed Pods deleted.
- **Out-of-Cluster Clients:** External consumers accessing the storage directly (e.g., via NFS or storage APIs) are invisible to Kubernetes and must be manually disconnected to ensure no external writes occur during the restore.
**Note:** Automating the deletion of these workloads is explicitly out of scope for this feature. It remains the user's responsibility to ensure the volume is not actively consumed and the Pods are removed before triggering the restore.
### Handling Cross-Zone Scheduling (WaitForFirstConsumer)
When performing an in-place restore, Velero deletes the existing target PVC and recreates it. For StorageClasses using the `WaitForFirstConsumer` volume binding mode, this recreation resets the scheduling lifecycle. Even though Velero adds a selector to the PVC spec to ensure it binds exclusively to the original PV, a scheduling issue can still occur. If the target PVC loses its node affinity, the Kubernetes Scheduler might schedule the recreated business Pod to a different availability zone. Because the original PV is physically constrained to its original zone, the Pod will fail to mount the volume and remain stuck in the `ContainerCreating` state with an attachment error.
**Solution**:
During the PVC CSI Restore Item Action (RIA), right before deleting the existing PVC, Velero extracts the `volume.kubernetes.io/selected-node` annotation from that PVC and carries it on the PVC to be restored via a Velero-internal carrier annotation (`restore.velero.io/inplace-restore-selected-node`). After all Restore Item Actions have run, the restore engine translates the carrier back to the `volume.kubernetes.io/selected-node` annotation and strips the carrier so it never lands on the cluster.
A carrier annotation is used instead of the Kubernetes annotation directly because the generic PVC RIA unconditionally strips the `selected-node` annotation during restore, and the execution order of Restore Item Actions is not a documented contract. With the carrier, the behavior is independent of the RIA execution order: the Kubernetes annotation is stripped by default on every path (including when the target PVC does not exist and Velero falls back to provisioning a new PVC), and preservation only happens when the PVC CSI RIA explicitly captured a value from the existing PVC.
By preserving the `selected-node` annotation, the Kubernetes Scheduler is forced to schedule the recreated business Pod to the original node/zone, ensuring it successfully mounts the restored PV.
### Namespace Mapping
When namespace mapping is configured, in-place restores work normally in most scenarios. However, in-place incremental restores using CSI snapshots with a block data mover are not natively supported across different namespaces. Velero cannot use Changed Block Tracking (CBT) to calculate data deltas when the target volume is in a different namespace, as the volumes may belong to different lineages.
Despite this, users can achieve a fast cross-namespace "clone and restore" workflow. For example, to quickly clone a large production workload into a test namespace ((e.g., for debugging, testing, or auditing)), a standard full restore would be too slow. Instead, users can manually take a CSI snapshot of the source PVC and provision a new PVC in the destination namespace from that snapshot.
When a Velero restore is triggered against this new PVC, Velero detects the snapshot and uses CBT to write only the blocks that changed since the backup. This effectively "rolls back" the clone to the backup's state, drastically reducing data transfer and speeding up the restore.
The key requirements for this approach are:
1. **Manual Cloning:** Users must manually snapshot the source PVC and clone it to the destination namespace before the restore. *(Note: Users must manually recreate the `VolumeSnapshotContent` and `VolumeSnapshot` in the destination namespace, or use `CrossNamespaceVolumeDataSource` if supported).*
2. **Workload Management:** Ensure no Pods are mounting the destination PVC during the restore to prevent data corruption.
3. **Snapshot Detection:** Velero inspects the destination PVC's `dataSource`. If it is a `VolumeSnapshot`, Velero uses its `SnapshotHandle` along with the backup's handle to calculate CBT.
4. **One-Shot Operation:** This is a one-time process. To restore a different backup later, users must clean up the destination namespace and repeat the workflow.
5. **Snapshot Cleanup:** Users must manually delete the temporary snapshot after the restore completes.
### Pre-flight Checks
Before initiating an in-place restore for a volume, Velero performs the following pre-flight checks to ensure the operation is safe and valid:
#### 1. PVC is Not Actively Used by a Running Pod
Velero verifies that the target PVC is not currently mounted or consumed by any running Pods in the cluster. If the PVC is in use, Velero will skip the in-place restore for that volume and log an error. This enforces the prerequisite that users must completely delete consuming workloads prior to the restore, which prevents data corruption and avoids deadlocks caused by the Kubernetes `pvc-protection` finalizer during PVC recreation.
#### 2. PVC is Bound to the Original PV
Velero checks whether the existing PVC in the cluster is still bound to the same PersistentVolume (PV) it was bound to at the time of the backup. If the PVC is bound to a different PV, performing an in-place restore (especially an incremental one that relies on Changed Block Tracking) may be unsafe or result in unpredictable behavior. If this check fails, Velero will log an error and skip the in-place restore for that volume.
#### 3. Volume Size Validation
For in-place restores, the target volume must be large enough to accommodate the backed-up data. While the data path performs size checks during the actual restoration (only for block data mover), Velero will fail early to prevent unnecessary operations (such as taking a temporary snapshot).
Before initiating an in-place restore, Velero compares the existing PV's size (`pv.spec.capacity.storage`) against the backup's data size (retrieved from the backup volume info). If the target PV is smaller than the backup data size, Velero will log an error and skip the volume data restoration.
### Error Handling
It is highly recommended that users create a backup (e.g., a CSI snapshot backup without data movement, if possible) before initiating an in-place restore. This ensures that the original state can be recovered in the event of a restore failure.
If an in-place restore fails, Velero will intentionally leave certain temporary resources intact, such as the temporary PVC bound to the existing PV. Velero does not automatically clean up these resources because doing so could inadvertently trigger the deletion of the underlying storage volume. In such failure scenarios, users must manually clean up these temporary resources and, if necessary, use their pre-restore backup to recover the system's state.
### Restore Workflow Update
This section outlines the step-by-step control path and data path workflows for in-place restores. The exact sequence of operations depends on the backup method (CSI snapshot vs. file system backup), the chosen data mover (block vs. file system), and the target volume mode (block vs. file system). The following subsections detail the mechanisms for each supported scenario.
#### In-place Incremental Restore for CSI Snapshot with Block Data Move for Block Volumes
**Control Path**
PVC CSI RIA:
- Capture the `volume.kubernetes.io/selected-node` annotation from the existing PVC into the Velero-internal carrier annotation before deleting the PVC, so the restore engine can re-apply it to the recreated target PVC (see [Handling Cross-Zone Scheduling](#handling-cross-zone-scheduling-waitforfirstconsumer)).
- Create a snapshot of the existing `PVC` to serve as the baseline for CBT delta calculations.
- Patch the existing PV's reclaim policy to `Retain`.
- Delete the existing PVC.
- Create a `DataDownload` resource referencing this snapshot and the existing `PV`, with `restoreType` set to `incremental`.
Restore Exposer:
- Create a temporary restore PVC and bind it to the existing PV.
- Create a temporary restore Pod that mounts the temporary restore PVC.
**Data Path**
Block Uploader:
- The block uploader leverages Changed Block Tracking (CBT) to calculate the delta between the volume's current state and the backup snapshot. By skipping unchanged blocks and exclusively overwriting the modified ones, it significantly reduces I/O operations and accelerates the overall restore process. If the underlying storage system lacks CBT support, Velero will automatically fall back to performing an in-place full restore.
#### In-place Full Restore for CSI Snapshot with Block Data Move for Block Volumes
The workflow is identical to the **In-place Incremental Restore for CSI Snapshot with Block Data Move for Block Volumes**, with the following exceptions:
- No baseline snapshot is taken.
- The uploader does not use CBT to calculate deltas; instead, it overwrites all data on the volume.
#### In-place Incremental Restore for CSI Snapshot with File System Data Move for File System Volumes
**Control Path**
The control path workflow is identical to the **In-place Incremental Restore for CSI Snapshot with Block Data Move for Block Volumes**, with the following exceptions:
- No baseline snapshot is taken.
**Data Path**
Kopia Uploader:
- Set the `incremental` flag to `true` when initiating the restore with the Kopia uploader.
- Pass the `deleteExtraFiles` configuration to the Kopia uploader based on the user's settings.
- Kopia evaluates file metadata (e.g., modification times and sizes) to identify changed files. It skips downloading and overwriting files that are identical to the backup, only restoring those that are modified, missing, or corrupted.
#### In-place Full Restore for CSI Snapshot with File System Data Move for File System Volumes
The workflow is identical to the **In-place Incremental Restore for CSI Snapshot with File System Data Move for File System Volumes**, with the following exceptions:
- The `restoreType` flag set to `full`.
- The Kopia uploader does not evaluate file metadata to skip unchanged files; instead, it overwrites all data on the target volume.
#### In-place Incremental Restore for CSI Snapshot with Block Data Move for File System Volumes
**Control Path**
PVC CSI RIA:
- Capture the `volume.kubernetes.io/selected-node` annotation from the existing PVC into the Velero-internal carrier annotation before deleting the PVC, so the restore engine can re-apply it to the recreated target PVC (see [Handling Cross-Zone Scheduling](#handling-cross-zone-scheduling-waitforfirstconsumer)).
- Create a snapshot of the existing `PVC` to serve as the baseline for CBT delta calculations.
- Patch the existing `PV` to set its `persistentVolumeReclaimPolicy` to `Retain`.
- Delete the existing `PVC`.
- Create a `DataDownload` resource referencing the snapshot and the existing `PV`, with `restoreType` set to `incremental`.
Restore Exposer:
- Delete the existing `PV`.
- Create a temporary restore `PV` with `volumeMode` set to `Block`, using the same volume handle as the original `PV`.
- Reset the bind information of the temporary restore `PV` to ensure it only binds to the temporary restore `PVC`.
- Create a temporary restore `PVC` with `volumeMode` set to `Block`.
- Create a temporary restore Pod that mounts the temporary restore `PVC`.
**Data Path**
Block Uploader:
- The block uploader leverages Changed Block Tracking (CBT) to calculate the delta between the volume's current state and the backup snapshot. By skipping unchanged blocks and exclusively overwriting the modified ones, it significantly reduces I/O operations and accelerates the overall restore process. If the underlying storage system lacks CBT support, Velero will automatically fall back to performing an in-place full restore.
**Control Path (Post-Restore)**
Restore Exposer:
- Delete the temporary restore Pod, `PVC`, and `PV`.
- Recreate the original `PV` with its `volumeMode` set back to `Filesystem`.
- Proceed with the standard process to allow the target `PVC` to bind to the recreated `PV`.
#### In-place Full Restore for CSI Snapshot with Block Data Move for File System Volumes
The workflow is identical to the **In-place Incremental Restore for CSI Snapshot with Block Data Move for File System Volumes**, with the following exceptions:
- No baseline snapshot is taken.
- The uploader does not use CBT to calculate deltas; instead, it overwrites all data on the volume.
#### In-place Incremental Restore for File System Backup for File System Volumes
**Control Path**
- Create a `PodVolumeRestore` resource with `restoreType` set to `incremental`.
**Data Path**
Kopia Uploader:
- Set the `incremental` flag to `true` when initiating the restore with the Kopia uploader.
- Pass the `deleteExtraFiles` configuration to the Kopia uploader based on the user's settings.
- Similar to the CSI File System Data Move, Kopia evaluates file metadata to skip unchanged files and only restores those that are modified or missing.
#### In-place Full Restore for File System Backup for File System Volumes
The workflow is identical to the **In-place Incremental Restore for File System Backup for File System Volumes**, with the following exceptions:
- The `restoreType` flag is set to `full`.
- The Kopia uploader does not evaluate file metadata to skip unchanged files; instead, it overwrites all data on the target volume.
## Installation
No change to Installation.
## Upgrade
No impacts to Upgrade. The new fields in the CRDs are all optional fields and have backwards compatible values.