Why access modes are an architecture decision
"Run three replicas and let them share one volume" sounds like a detail to settle later. It is not: most disks can only be attached to one machine at a time, and a disk lives in one zone. Those two facts decide where your pods may run, why a rollout hangs, and why a database does not come back after maintenance. This lesson is those facts.
What you need to know already: PV, PVC, StorageClass, binding and Pending (16.39), block devices and filesystems (4.18), how the scheduler picks a node and nodeSelector / taints in one line each (15.11, 15.22), Deployments and their rolling update (15.16), StatefulSets (15.19), node labels (15.26).
The four modes
ReadWriteOnce RWO read-write by ONE NODE at a time (any number of pods on it)
ReadOnlyMany ROX read-only by many nodes
ReadWriteMany RWX read-write by many nodes at once
ReadWriteOncePod RWOP read-write by ONE POD in the whole cluster (GA 1.29)
An access mode says how a volume may be mounted. RWO is about nodes, not pods - a common misreading. Two pods of a Deployment on the same node can both mount an RWO volume; on different nodes they cannot. RWOP is the strict version for when you really mean one writer.
A mode is a promise of the backend:
- Block storage - a raw disk device attached to one machine, which then puts a filesystem on it (a cloud provider's disk, Ceph RBD, the lab's disk.csi.lab): RWO / RWOP only.
- Shared file storage - a filesystem served over the network (NFS, a cloud provider's file share service, CephFS): can do RWX.
Asking a block driver for RWX fails at provisioning, in the claim's events:
Warning ProvisioningFailed disk.csi.lab_... failed to provision volume with StorageClass "lab-disk": rpc error: code = InvalidArgument desc = Volume capability not supported
This is an architecture constraint, not a detail. "Three replicas share one volume" rules out cloud disks, full stop. The options are a shared file service (NFS or similar, RWX), or redesigning so each replica has its own disk (a StatefulSet, 16.44) or no disk at all (object storage: files stored through an HTTP API instead of a filesystem).
Later (Ch 23): the specific disk and file-share services of one cloud, and which modes each supports.
Attach, and the Multi-Attach error
For block drivers, a volume is attached to a node (like plugging a disk into that machine) before the kubelet can mount it. A control-plane controller, the attach/detach controller, records each attachment as a VolumeAttachment object:
$ k get volumeattachments
NAME ATTACHER PV NODE ATTACHED AGE
csi-4cef39282797f5562f20a4376a7061e499d0d14b028eb0161c771916ee9584dc disk.csi.lab pvc-b2a4c203-9afe-482e-944c-36e1c87bb030 worker-2 true 8s
ATTACHER = the driver, PV = which volume, NODE = where it is plugged in, ATTACHED = done.
A pod on another node wanting the same RWO volume gets this event:
Warning FailedAttachVolume 9s (x3 over 29s) attachdetach-controller Multi-Attach error for volume "pvc-f67aaaff-..." Volume is already used by pod(s) app-4qcb5hsbw5-cvdw8
It sits in ContainerCreating until the other pod is gone and the volume detaches (after about two minutes the kubelet adds Unable to attach or mount volumes: ... timed out waiting for the condition).
The classic way in: a Deployment with an RWO volume and the default RollingUpdate (15.16) - the new pod starts before the old one stops. If it lands on another node, it waits for the old pod's volume, and the old pod waits for the new one to be ready. Stuck. Use strategy: Recreate (stop the old pod first: a short downtime, but honest), or a StatefulSet.
Zones
A cloud region is split into availability zones: separate data centres, so one fire or power cut does not take out everything. Cloud disks live in one zone and can only attach to machines in it. Nodes carry their zone in the label topology.kubernetes.io/zone (a topology label: it says where the node sits). The PV records its zone as node affinity - a rule saying which nodes may use it:
# the PV provisioned for a claim in zone lab-b
k describe pv pvc-ddb405cc-...
Node Affinity:
Required Terms:
Term 0: topology.kubernetes.io/zone in [lab-b]
The scheduler honours it: a pod using this claim can only run on nodes labelled topology.kubernetes.io/zone=lab-b. In this lab, cp-1 and worker-1 are in lab-a, worker-2 in lab-b. -L LABEL adds a column with that label's value:
$ k get nodes -L topology.kubernetes.io/zone
NAME STATUS ROLES AGE VERSION ZONE
cp-1 Ready control-plane 12d v1.34.1 lab-a
worker-1 Ready <none> 12d v1.34.1 lab-a
worker-2 Ready <none> 12d v1.34.1 lab-b
So RWO limits your scheduling: once a disk exists, every pod that uses it is confined to that disk's zone (and for node-local storage like local-path, to one node). If that zone is full, cordoned (marked "no new pods") or down, the pod is Pending:
0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 node(s) had volume node affinity conflict, 1 node(s) were unschedulable. preemption: 0/3 nodes are available: 3 Preemption is not helpful for scheduling.
The scheduler explains each node: cp-1 has the control-plane taint (15.22), one node has a volume node affinity conflict - this pod's volume cannot be reached from it - and one is unschedulable (cordoned). The "preemption" part says evicting other pods would not help.
volumeBindingMode
A StorageClass's volumeBindingMode decides when the volume is created:
Immediate provision + bind as soon as the claim exists - in some zone the
driver picks, before anyone knows where the pod will run
WaitForFirstConsumer wait until a pod using the claim is scheduled, then provision
in THAT node's zone (the scheduler annotates the claim with
volume.kubernetes.io/selected-node)
With Immediate in a cluster that spans zones, the disk can land in zone B while the pod's other constraints (a nodeSelector, a full zone A) require zone A - the pod can never start. Use WaitForFirstConsumer for zonal storage. Immediate is fine for storage reachable from everywhere (NFS, Ceph RBD, or a cloud "zone-redundant" disk that is copied across zones).
allowedTopologies on a StorageClass restricts where it may provision at all (e.g. only zones 1 and 2).
Drains and RWO
A node drain (kubectl drain) evicts all pods from a node before maintenance. A StatefulSet pod with an RWO disk in zone B can only come back in zone B, so draining the only zone-B node leaves it Pending until the node returns. Plan maintenance per zone and keep capacity in every zone that holds disks.
Later (Ch 17): PodDisruptionBudgets, which make a drain wait instead of taking a database down.
What you can now do:
- pick an access mode, and know which backends can provide it
- explain a Multi-Attach error during a rolling update, and fix it
- read "volume node affinity conflict" and choose WaitForFirstConsumer for zonal disks