OnCallReady

Lesson 16.44 · Kubernetes: Networking & Storage · 18 min read

Reclaim policies, protection, expansion, StatefulSet claims

In plain words

Imagine you rent a locker and then cancel the rental. The locker company has a policy written on the locker. "Delete" means: empty the locker and throw everything away as soon as you cancel. "Retain" means: lock it up with your things still inside, and nobody else can rent it until the manager checks it. And the company won't let you cancel while you're still using the locker; the cancellation just waits.

In Kubernetes, the reclaim policy on a PV decides what happens to the storage when its claim is deleted: Delete removes the PV and the disk, Retain keeps both as Released. The pvc-protection finalizer is the "not while you're using it" rule. StatefulSets give each replica its own claim that outlives the pods.

Why this lesson can save your data

kubectl delete namespace shop takes a second to type. With the wrong settings it deletes every claim in the namespace, every volume behind them, and every disk behind those - the database included. This lesson is about what happens to storage when claims and pods go away, how to get a volume back, how to grow one, and how StatefulSets keep one disk per replica.

What you need to know already: PV, PVC, StorageClass, binding, claimRef and volumeName (16.39), access modes and zones (16.41), StatefulSets and ordinals (15.19), JSON patch (16.10), kubectl describe events (15.14).

What happens when the claim goes away

The reclaim policy decides. It is persistentVolumeReclaimPolicy on the PV, copied from the StorageClass's reclaimPolicy when the volume was provisioned:

Delete   the PV AND the backing storage are deleted. The disk is gone.
         Default for dynamically provisioned volumes.
Retain   the PV stays, phase Released, data intact. Nobody can bind it until
         an admin intervenes. Default for manually created PVs.
Recycle  deprecated (rm -rf of the volume). Do not use.

Production data gets Retain. With Delete, kubectl delete namespace deletes every PVC in it, which deletes every PV, which deletes every disk. That is one command away from a very bad day. You can switch an existing PV:

k patch pv pvc-ca5c067a-... -p '{"spec":{"persistentVolumeReclaimPolicy":"Retain"}}'

But not a StorageClass: reclaimPolicy: Forbidden: updates to reclaimPolicy are forbidden. - make a new class.

Getting a Released volume back

$ k get pv
NAME                                       CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS     CLAIM    STORAGECLASS
pvc-ca5c067a-48e0-43ad-bf16-2195742424e1   3Gi        RWO            Retain           Released   shop/d   ceph-rbd

STATUS Released = its claim was deleted, the data is still there. The PV still carries claimRef to the dead claim, including that claim's uid (the unique ID every object gets). A new claim with the same name has a different uid and does not match.

So: remove the reference, the PV becomes Available, then bind a new claim to it by name. A JSON patch remove deletes one field:

k patch pv pvc-ca5c067a-... --type=json -p='[{"op":"remove","path":"/spec/claimRef"}]'
# or keep the name and drop only the uid:
k patch pv pvc-ca5c067a-... --type=json -p='[{"op":"remove","path":"/spec/claimRef/uid"}]'
kind: PersistentVolumeClaim
spec:
  storageClassName: ceph-rbd     # must match the PV's class
  volumeName: pvc-ca5c067a-...   # this exact PV
  accessModes: [ReadWriteOnce]
  resources: {requests: {storage: 3Gi}}

Protection: why a PVC hangs in Terminating

A finalizer is a marker on an object that says "do not really delete this yet - someone must clean up first". Deleting such an object only marks it Terminating; it disappears once every finalizer has been removed by the controller that owns it.

Every PVC gets the finalizer kubernetes.io/pvc-protection, every PV kubernetes.io/pv-protection. Deleting a claim that a pod still uses (any pod not Succeeded/Failed - a crash-looping one counts) only marks it:

# old-data = a claim still mounted by a forgotten pod
k get pvc old-data
NAME       STATUS        VOLUME         CAPACITY   ACCESS MODES   STORAGECLASS
old-data   Terminating   pvc-7d1c...    1Gi        RWO            lab-disk
k describe pvc old-data | grep -E 'Finalizers|Used By'
Finalizers:    [kubernetes.io/pvc-protection]
Used By:       forgotten-debug

Used By names the pod. Delete that pod and the claim finishes deleting. Removing the finalizer by hand "works" and leaves a pod writing to a volume whose claim no longer exists - don't.

Expansion

To grow a claim, raise its requested size:

k patch pvc logs -p '{"spec":{"resources":{"requests":{"storage":"4Gi"}}}}'

Allowed only if the class has allowVolumeExpansion: true and the volume was dynamically provisioned:

Error from server (Forbidden): persistentvolumeclaims "logs" is forbidden: only dynamically provisioned pvc can be resized and the storageclass that provisions the pvc must support resize

The sequence, visible in the claim's events and conditions:

  1. The driver's resizer grows the backend volume (ExternalExpanding, Resizing).
  2. The filesystem on it must grow too (like resize2fs after growing a disk). The claim shows condition FileSystemResizePending until the kubelet does it on a node where the volume is mounted (FileSystemResizeSuccessful).

Most CSI drivers do it while the pod runs.

Shrinking is not possible: spec.resources.requests.storage: Forbidden: field can not be less than status.capacity. To make a volume smaller you copy the data to a new one.

StatefulSet claims

A StatefulSet (15.19) can create one claim per replica from a template, the volumeClaimTemplates:

kind: StatefulSet
spec:
  volumeClaimTemplates:
  - metadata: {name: data}
    spec:
      accessModes: [ReadWriteOnce]
      resources: {requests: {storage: 1Gi}}

The claims are named <template>-<statefulset>-<ordinal>: data-queue-0, data-queue-1... Pod queue-1 always gets data-queue-1, whichever node it lands on (within the volume's zone). The claims outlive the pods and, by default, the StatefulSet:

scale 3 -> 1      data-queue-1 and data-queue-2 stay (Bound, unused)
scale 1 -> 3      queue-1 and queue-2 come back with their old data
delete the sts    all claims stay - reinstall and the data is there

persistentVolumeClaimRetentionPolicy (stable since 1.32) changes that:

spec:
  persistentVolumeClaimRetentionPolicy:
    whenDeleted: Delete     # delete the claims with the StatefulSet (Retain = default)
    whenScaled: Retain      # Delete = remove claims of scaled-away ordinals

It works by making the claims "owned" by the StatefulSet (or by the pod, for whenScaled) through ownerReferences, so deleting the owner deletes them. The claims still follow their PVs' reclaim policy - a Retain PV survives even a deleted claim.

Backups are not optional

None of this is a backup. Retain protects against a deleted object, not a corrupted database or a dead disk. Real backups are snapshots (point-in-time copies of a volume: VolumeSnapshot objects through the CSI driver, or the cloud's disk snapshots) plus copies outside the cluster (a backup tool such as Velero, or the database's own dump tools) - and a restore you have actually tested.

What you can now do:

Why it helps

This is the lesson that prevents the worst storage incident: someone runs kubectl delete namespace on a namespace with a database, the claims are deleted, and with the default Delete policy every disk goes with them. Knowing to use Retain for production data, and how to patch an existing PV, is basic platform hygiene.

It also covers the tickets that follow: a PVC hanging in Terminating (a forgotten pod still uses it), getting a Released volume back after an accidental delete (remove the claimRef and bind by volumeName), a disk that won't grow (allowVolumeExpansion, FileSystemResizePending), and StatefulSet claims that are still there after a scale-down. In hands-on exams, expanding a PVC and binding to a specific PV are common tasks.

FAQ

Why is my PVC stuck in Terminating?

The kubernetes.io/pvc-protection finalizer keeps it until no pod uses it. Any pod that isn't Succeeded or Failed counts, including a crash-looping one. k describe pvc shows "Used By" with the pod name; delete that pod and the claim finishes deleting. Don't remove the finalizer by hand, since it leaves a pod writing to a volume whose claim is gone.

I deleted the claim and the PV is Released. Can I reuse it?

Yes, if the policy was Retain. Released means the PV still has a claimRef pointing at the dead claim, including its uid, so a new claim with the same name doesn't match. Remove the claimRef (or just its uid) with a JSON patch, the PV becomes Available, and bind a new claim to it with volumeName and the same storageClassName.

Can I shrink a PVC?

No. Kubernetes refuses a request smaller than the current capacity ("field can not be less than status.capacity"). To make a volume smaller, create a new, smaller claim and copy the data over, with the app stopped or using its own tooling (a database dump and restore, for instance). Growing is allowed if the class has allowVolumeExpansion: true.

Why does my PVC show FileSystemResizePending after I expanded it?

Expansion has two steps. The CSI resizer first grows the backend volume; then the file system on it has to grow, which the kubelet does on the node where the volume is mounted. Until that happens the claim shows FileSystemResizePending. Most CSI drivers do it online while the pod runs; some need the pod restarted.

Why are there still PVCs after I scaled my StatefulSet down?

By design. StatefulSet claims (data-queue-1, data-queue-2) outlive the pods and, by default, the StatefulSet itself, so scaling back up or reinstalling brings the old data back. persistentVolumeClaimRetentionPolicy (GA in 1.32) changes it: whenScaled: Delete and whenDeleted: Delete remove claims automatically. The PVs still follow their own reclaim policy.

In an interview Junior

What are the PersistentVolume reclaim policies, and what happens when someone deletes a PVC?

The reclaim policy (persistentVolumeReclaimPolicy on the PV, copied from the StorageClass) decides what happens when the claim goes away:

Production data gets Retain: k patch pv NAME -p '{"spec":{"persistentVolumeReclaimPolicy":"Retain"}}' (a StorageClass's policy cannot be changed; make a new class).

To recover a Released volume: remove its stale claimRef (it holds the old claim's uid), so it becomes Available, then create a claim that binds to it by volumeName.

A PVC still used by a pod only goes Terminating - the pvc-protection finalizer waits for the pod. None of this is a backup: snapshots and tested restores are.

Also asked: Someone deleted a PVC holding production data. What can you do? · What happens to StatefulSet volumes when you scale down or delete it? · How do you grow a PVC, and can you shrink one?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.