OnCallReady

Lesson 16.39 · Kubernetes: Networking & Storage · 21 min read

PersistentVolume, PersistentVolumeClaim, StorageClass

In plain words

Think of a library. A PersistentVolume is an actual book on a shelf. A PersistentVolumeClaim is your request slip: "I need a book on space, at least 200 pages". You never go to the shelf yourself; the librarian matches your slip to a book and puts your name on it, and from then on that book is yours only. If there's no fitting book, a StorageClass is like a print-on-demand machine: "for requests of this kind, print a new book with this recipe".

In Kubernetes: pods reference claims, claims bind 1:1 to PVs, and a StorageClass's provisioner (a CSI driver running as pods) creates PVs on demand. A claim stays Pending until a PV is bound, and k describe pvc tells you which of the four reasons it is.

Why claims, volumes and classes

An app team wants "10 GiB of disk that survives my pod". They should not need to know whether the cluster's storage is a cloud disk, a network file share or a storage cluster, nor call its API to create one. Kubernetes splits the job: the app asks (a claim), a recipe says how to make storage of each kind (a class), and the result is a volume object that represents the real storage.

What you need to know already: volumes and the claim idea (16.37), mounts and filesystems (4.18), namespaces and annotations (15.26), Deployments and DaemonSets (15.16, 15.22), sidecars (15.35), kubectl patch (15.40).

Three objects, three owners

PersistentVolume    cluster-scoped. A real piece of storage: an NFS export, a disk,
                    a Ceph image. Capacity, access modes, reclaim policy, source.
PersistentVolumeClaim  namespaced. An app's request: size, access mode, class.
                    Pods reference claims, never volumes.
StorageClass        cluster-scoped. A recipe for creating volumes on demand:
                    which provisioner, which parameters, reclaim, binding mode.

NFS (network file system) is the classic way to share one directory over the network with many machines. Ceph is a storage cluster software that offers disks over the network. The split lets an app team say "10Gi, read-write, fast" without knowing which is behind it.

The lab's storage

sc is short for storageclass:

$ k get sc
NAME                 PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ALLOWVOLUMEEXPANSION   AGE
ceph-rbd             rbd.csi.ceph.com        Delete          Immediate              true                   12d
lab-disk (default)   disk.csi.lab            Delete          WaitForFirstConsumer   true                   12d
lab-disk-immediate   disk.csi.lab            Delete          Immediate              true                   12d
local-path           rancher.io/local-path   Delete          WaitForFirstConsumer   false                  12d
nfs-csi              nfs.csi.k8s.io          Delete          Immediate              true                   12d

The columns:

NAME                   the class name; "(default)" = used when a claim names none
PROVISIONER            the driver that creates volumes of this class
RECLAIMPOLICY          what happens to the volume when its claim is deleted (16.44)
VOLUMEBINDINGMODE      when the volume is created: at once, or when a pod needs it (16.41)
ALLOWVOLUMEEXPANSION   whether claims of this class may grow later (16.44)

The classes:

lab-disk      a zonal block disk (simulator) - behaves like a cloud provider's disk:
              attaches to ONE node, lives in ONE zone
local-path    a directory on one node's disk (Rancher's local-path-provisioner)
ceph-rbd      Ceph RBD block images, reachable from every node (ceph-csi)
nfs-csi       subdirectories of an NFS export (csi-driver-nfs), shared by many nodes

(A zone is one data centre, or an isolated part of one, in a cloud region; 16.41 explains why it matters.)

Most provisioners are CSI drivers. CSI (Container Storage Interface) is the standard plugin interface between Kubernetes and a storage system, as CNI is for networks (16.35). A driver is pods: a controller Deployment with helper sidecars (external-provisioner creates volumes, -attacher attaches them to nodes, -resizer grows them) and a node DaemonSet that mounts volumes on each node. If the provisioner is down, nothing new gets provisioned:

$ k get pods -A | grep -E 'csi|local-path'
kube-system          lab-disk-csi-controller-kjb965mwh5-ffnb4    1/1     Running   0          12d
kube-system          lab-disk-csi-node-lptrr                     2/2     Running   0          12d
local-path-storage   local-path-provisioner-8qvp6tv4n6-gmwmb     1/1     Running   0          12d
...

(-A = all namespaces.)

Dynamic provisioning

Dynamic provisioning: you create only the claim, and the class's provisioner creates a matching volume for it.

apiVersion: v1
kind: PersistentVolumeClaim
metadata: {name: pg, namespace: data}
spec:
  accessModes: [ReadWriteOnce]
  resources:
    requests:
      storage: 500Mi
  # storageClassName omitted: the default class is filled in (admission)

accessModes says how it may be mounted (ReadWriteOnce = by one node at a time, 16.41); resources.requests.storage the size. "Admission" is the API server's step that fills in or checks fields before storing an object.

# pg = the claim above, applied to namespace data (the PVC mission)
k get pvc -n data
NAME   STATUS    VOLUME   CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
pg     Pending                                      lab-disk       <unset>                 3s
k describe pvc pg -n data | tail -3
  Normal  WaitForFirstConsumer  3s    persistentvolume-controller  waiting for first consumer to be created before binding

The columns of get pvc:

STATUS        Pending = no volume yet; Bound = attached to a volume
VOLUME        the PV it is bound to
CAPACITY      the size of that PV
ACCESS MODES  RWO / ROX / RWX / RWOP (16.41)
STORAGECLASS  the class it asked for (or was given)
VOLUMEATTRIBUTESCLASS  an optional performance setting; <unset> here

Pending is correct here: lab-disk waits for a pod (16.41). Once a pod using the claim is scheduled:

k get pvc,pv -n data
NAME                       STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS
persistentvolumeclaim/pg   Bound    pvc-b2a4c203-9afe-482e-944c-36e1c87bb030   1Gi        RWO            lab-disk
NAME                                                        CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS   CLAIM
persistentvolume/pvc-b2a4c203-9afe-482e-944c-36e1c87bb030   1Gi        RWO            Delete           Bound    data/pg
k describe pvc pg -n data | tail -3
  Normal  Provisioning           9s    disk.csi.lab_lab-disk-csi-controller-...  External provisioner is provisioning volume for claim "data/pg"
  Normal  ProvisioningSucceeded  7s    disk.csi.lab_lab-disk-csi-controller-...  Successfully provisioned volume pvc-b2a4c203-9afe-482e-944c-36e1c87bb030

The PV's own columns: RECLAIM POLICY (16.44), STATUS (Available = free, Bound = taken, Released = its claim was deleted, Failed), CLAIM (namespace/name of the claim that has it).

Static provisioning

Static provisioning: an admin creates the PV by hand, typically for storage that already exists:

apiVersion: v1
kind: PersistentVolume
metadata: {name: nfs-archive}
spec:
  capacity: {storage: 5Gi}
  accessModes: [ReadWriteMany]
  persistentVolumeReclaimPolicy: Retain
  storageClassName: ""               # belongs to no class
  nfs:
    server: 10.0.3.60
    path: /exports/archive

A claim binds to an Available PV when all of these match:

The smallest fitting PV wins.

The trap: a claim that omits storageClassName is given the default class at creation. It then never binds to your class-less static PV - it waits for lab-disk instead. To bind to static class-less PVs, write storageClassName: "" explicitly.

Pending: which of the four is it?

k describe pvc <name>                       the Events say which:

no persistent volumes available for         no class set, no default class, and no
this claim and no storage class is set      matching static PV
storageclass.storage.k8s.io "x" not found   a typo in storageClassName
waiting for first consumer to be created    WaitForFirstConsumer: normal until a pod
before binding                              uses it (and then check the POD's events)
Waiting for a volume to be created either   the provisioner is not running
by the external provisioner ...
ProvisioningFailed ... <driver message>     the driver refused: access mode, size, quota

And for the pod, in its events: 0/3 nodes are available: pod has unbound immediate PersistentVolumeClaims (an Immediate claim is not bound yet) or volume node affinity conflict (bound, but to a volume no allowed node can reach, 16.41).

Since Kubernetes 1.28 the default class is also assigned retroactively: claims created without a class while no default existed get it as soon as one is marked default.

Defaults

The default StorageClass is the one with this annotation:

k patch sc lab-disk -p '{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'

Two defaults at once is allowed (the newest wins, with a warning) - and a source of "why did this land on the slow class".

StorageClasses are almost immutable: provisioner, parameters, reclaimPolicy and volumeBindingMode cannot be changed (parameters: Forbidden: updates to parameters are forbidden.); you create a new class.

What you can now do:

Why it helps

A Pending PVC is a daily ticket on a platform team, and the four causes (no default class, a typo in storageClassName, WaitForFirstConsumer doing its job, a provisioner that's down or refusing) each have a distinct event message. Reading the events tells you in seconds whether it's the team's YAML or your CSI driver.

When you migrate existing data (an NFS export, a restored disk), you'll create a static PV and hit the classic trap: the claim omits storageClassName, gets the default class, and never binds. When you review a PR that adds a new StorageClass, you'll know which fields are immutable and why two default classes cause surprises. In hands-on exams, creating a PV, a PVC and a pod that uses it is a standard task.

FAQ

Why did I get 1Gi when I asked for 500Mi?

A claim gets at least what it asked for. Many drivers round up to their allocation unit; block disks usually come in whole GiB. The PVC's CAPACITY shows what you actually got. For static PVs the same rule holds: a claim binds to the smallest Available PV whose capacity is at least the request.

My claim won't bind to the PV I created by hand. Why?

The most common reason: the claim omits storageClassName, so admission fills in the default class, and it now waits for that class instead of your class-less PV. Write storageClassName: "" explicitly. Other mismatches: access modes (the PV must have every mode requested), capacity smaller than the request, volumeMode, a label selector, or a volumeName pin.

Is Pending always an error?

No. With volumeBindingMode: WaitForFirstConsumer a claim stays Pending until a pod that uses it is scheduled, and the event says "waiting for first consumer to be created before binding". That's by design. Look at the pod's events instead; if the pod is Pending too, the scheduler's message explains why.

Who actually creates the disk?

The provisioner named in the StorageClass, today almost always a CSI driver. It runs as pods: a controller Deployment with sidecars (external-provisioner, attacher, resizer) and a node DaemonSet that mounts volumes. If the controller is down, claims stay Pending with "Waiting for a volume to be created either by the external provisioner". Check it with k get pods -A | grep csi.

Can I change a StorageClass's parameters or reclaim policy?

No. Provisioner, parameters, reclaimPolicy and volumeBindingMode are immutable, and the API server refuses the update ("updates to parameters are forbidden"). Create a new class and point new claims at it. Existing PVs keep what they got at provisioning; you can patch an individual PV's reclaim policy directly.

In an interview Junior

Explain PersistentVolume, PersistentVolumeClaim and StorageClass.

Dynamic provisioning: create only the claim; the class's provisioner creates the PV (pvc-<uid>, rounded up in size). Static: an admin creates PVs for existing storage, and a claim binds when class, access modes, size, volume mode and selector all fit.

A claim stuck Pending: k describe pvc events - waiting for a pod (WaitForFirstConsumer, normal), no class or a broken provisioner, no matching static PV, or the default-class trap: a claim without storageClassName gets the default and never binds to a class-less PV - write storageClassName: "".

Also asked: A PVC is stuck in Pending. How do you figure out why? · What is the difference between static and dynamic provisioning? · What is a CSI driver?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.