OnCallReady

Lesson 16.41 · Kubernetes: Networking & Storage · 16 min read

Access modes, zones, and why RWO constrains scheduling

In plain words

Think of a USB stick versus a shared folder on the school network. The USB stick can be plugged into only one computer at a time. Several kids at that computer can use it, but a kid at another computer can't until it's unplugged and carried over. The shared folder can be opened from every computer at once. Also, the USB stick lives in one classroom building, so you can only use it on computers in that building.

That's access modes and zones. ReadWriteOnce (RWO) means one node at a time, like the USB stick; block storage (a cloud disk, Ceph RBD) can only do this. ReadWriteMany needs a shared filesystem (NFS, a cloud file share). A cloud disk lives in one zone, so its PV carries node affinity, and every pod using it can only run in that zone.

Why access modes are an architecture decision

"Run three replicas and let them share one volume" sounds like a detail to settle later. It is not: most disks can only be attached to one machine at a time, and a disk lives in one zone. Those two facts decide where your pods may run, why a rollout hangs, and why a database does not come back after maintenance. This lesson is those facts.

What you need to know already: PV, PVC, StorageClass, binding and Pending (16.39), block devices and filesystems (4.18), how the scheduler picks a node and nodeSelector / taints in one line each (15.11, 15.22), Deployments and their rolling update (15.16), StatefulSets (15.19), node labels (15.26).

The four modes

ReadWriteOnce    RWO    read-write by ONE NODE at a time (any number of pods on it)
ReadOnlyMany     ROX    read-only by many nodes
ReadWriteMany    RWX    read-write by many nodes at once
ReadWriteOncePod RWOP   read-write by ONE POD in the whole cluster (GA 1.29)

An access mode says how a volume may be mounted. RWO is about nodes, not pods - a common misreading. Two pods of a Deployment on the same node can both mount an RWO volume; on different nodes they cannot. RWOP is the strict version for when you really mean one writer.

A mode is a promise of the backend:

Asking a block driver for RWX fails at provisioning, in the claim's events:

Warning  ProvisioningFailed  disk.csi.lab_...  failed to provision volume with StorageClass "lab-disk": rpc error: code = InvalidArgument desc = Volume capability not supported

This is an architecture constraint, not a detail. "Three replicas share one volume" rules out cloud disks, full stop. The options are a shared file service (NFS or similar, RWX), or redesigning so each replica has its own disk (a StatefulSet, 16.44) or no disk at all (object storage: files stored through an HTTP API instead of a filesystem).

Later (Ch 23): the specific disk and file-share services of one cloud, and which modes each supports.

Attach, and the Multi-Attach error

For block drivers, a volume is attached to a node (like plugging a disk into that machine) before the kubelet can mount it. A control-plane controller, the attach/detach controller, records each attachment as a VolumeAttachment object:

$ k get volumeattachments
NAME                                                                   ATTACHER       PV                                         NODE       ATTACHED   AGE
csi-4cef39282797f5562f20a4376a7061e499d0d14b028eb0161c771916ee9584dc   disk.csi.lab   pvc-b2a4c203-9afe-482e-944c-36e1c87bb030   worker-2   true       8s

ATTACHER = the driver, PV = which volume, NODE = where it is plugged in, ATTACHED = done.

A pod on another node wanting the same RWO volume gets this event:

Warning  FailedAttachVolume  9s (x3 over 29s)  attachdetach-controller  Multi-Attach error for volume "pvc-f67aaaff-..." Volume is already used by pod(s) app-4qcb5hsbw5-cvdw8

It sits in ContainerCreating until the other pod is gone and the volume detaches (after about two minutes the kubelet adds Unable to attach or mount volumes: ... timed out waiting for the condition).

The classic way in: a Deployment with an RWO volume and the default RollingUpdate (15.16) - the new pod starts before the old one stops. If it lands on another node, it waits for the old pod's volume, and the old pod waits for the new one to be ready. Stuck. Use strategy: Recreate (stop the old pod first: a short downtime, but honest), or a StatefulSet.

Zones

A cloud region is split into availability zones: separate data centres, so one fire or power cut does not take out everything. Cloud disks live in one zone and can only attach to machines in it. Nodes carry their zone in the label topology.kubernetes.io/zone (a topology label: it says where the node sits). The PV records its zone as node affinity - a rule saying which nodes may use it:

# the PV provisioned for a claim in zone lab-b
k describe pv pvc-ddb405cc-...
Node Affinity:
  Required Terms:
    Term 0:        topology.kubernetes.io/zone in [lab-b]

The scheduler honours it: a pod using this claim can only run on nodes labelled topology.kubernetes.io/zone=lab-b. In this lab, cp-1 and worker-1 are in lab-a, worker-2 in lab-b. -L LABEL adds a column with that label's value:

$ k get nodes -L topology.kubernetes.io/zone
NAME       STATUS   ROLES           AGE   VERSION   ZONE
cp-1       Ready    control-plane   12d   v1.34.1   lab-a
worker-1   Ready    <none>          12d   v1.34.1   lab-a
worker-2   Ready    <none>          12d   v1.34.1   lab-b

So RWO limits your scheduling: once a disk exists, every pod that uses it is confined to that disk's zone (and for node-local storage like local-path, to one node). If that zone is full, cordoned (marked "no new pods") or down, the pod is Pending:

0/3 nodes are available: 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 1 node(s) had volume node affinity conflict, 1 node(s) were unschedulable. preemption: 0/3 nodes are available: 3 Preemption is not helpful for scheduling.

The scheduler explains each node: cp-1 has the control-plane taint (15.22), one node has a volume node affinity conflict - this pod's volume cannot be reached from it - and one is unschedulable (cordoned). The "preemption" part says evicting other pods would not help.

volumeBindingMode

A StorageClass's volumeBindingMode decides when the volume is created:

Immediate              provision + bind as soon as the claim exists - in some zone the
                       driver picks, before anyone knows where the pod will run
WaitForFirstConsumer   wait until a pod using the claim is scheduled, then provision
                       in THAT node's zone (the scheduler annotates the claim with
                       volume.kubernetes.io/selected-node)

With Immediate in a cluster that spans zones, the disk can land in zone B while the pod's other constraints (a nodeSelector, a full zone A) require zone A - the pod can never start. Use WaitForFirstConsumer for zonal storage. Immediate is fine for storage reachable from everywhere (NFS, Ceph RBD, or a cloud "zone-redundant" disk that is copied across zones).

allowedTopologies on a StorageClass restricts where it may provision at all (e.g. only zones 1 and 2).

Drains and RWO

A node drain (kubectl drain) evicts all pods from a node before maintenance. A StatefulSet pod with an RWO disk in zone B can only come back in zone B, so draining the only zone-B node leaves it Pending until the node returns. Plan maintenance per zone and keep capacity in every zone that holds disks.

Later (Ch 17): PodDisruptionBudgets, which make a drain wait instead of taking a database down.

What you can now do:

Why it helps

This lesson explains three of the most confusing production failures. The Multi-Attach error: a Deployment with an RWO volume and a RollingUpdate, where the new pod lands on another node and waits forever for the old one. "volume node affinity conflict": a StatefulSet pod stuck Pending after its zone's node was drained. And a disk provisioned in zone B while the pod can only run in zone A, because the class used Immediate binding.

When a team says "our three replicas will share one volume", you'll know that rules out cloud disks and what the real options are. When you plan node maintenance, you'll know to keep capacity in every zone that holds disks. Architecture reviews and hands-on storage tasks both lean on this.

FAQ

Does ReadWriteOnce mean only one pod can use the volume?

No, it means one node. Any number of pods on the same node can mount an RWO volume read-write. Pods on a different node can't, and get a Multi-Attach error. If you really need a single writer in the whole cluster, use ReadWriteOncePod (RWOP), GA since Kubernetes 1.29.

Why can't I get ReadWriteMany from a cloud disk?

Because a disk is a block device attached to one VM at a time; the access mode is a promise the backend has to keep. Block drivers only offer RWO and RWOP, and asking for RWX fails at provisioning with "Volume capability not supported". For shared read-write use a file system: NFS, CephFS or a cloud file-share service, or redesign so each replica has its own disk.

Why is my new pod stuck in ContainerCreating with a Multi-Attach error?

The RWO volume is still attached to another node, usually because an old pod there still uses it. The classic path is a Deployment with RollingUpdate: the new pod starts before the old one stops and lands on a different node. Use strategy: Recreate for single-replica Deployments with RWO volumes, or a StatefulSet. It clears once the old pod is gone and the volume detaches.

What does "volume node affinity conflict" mean?

The pod's PV can only be reached from certain nodes, recorded as node affinity on the PV (for example topology.kubernetes.io/zone in [lab-b]), and none of those nodes is available to the pod: they're cordoned, full, tainted, or excluded by the pod's own nodeSelector or affinity. The disk won't move zones; bring capacity back in that zone.

Immediate or WaitForFirstConsumer?

WaitForFirstConsumer for anything zonal or node-local: the volume is created only after the scheduler picks a node, in that node's zone. With Immediate the driver picks a zone before anyone knows where the pod will run, and the pod may never be schedulable there. Immediate is fine for storage reachable from everywhere: NFS, Ceph RBD, zone-redundant disk classes.

In an interview Junior

What are the PersistentVolume access modes, and how do they affect scheduling?

The backend decides what is possible: block storage (cloud disks, Ceph RBD) is RWO/RWOP only; shared file storage (NFS) can do RWX. So "three replicas share one volume" rules out disks.

Effects on scheduling:

Also asked: A Deployment with one replica and an RWO volume hangs during every rollout. Why? · What does WaitForFirstConsumer change? · What happens to a pod with a zonal disk when you drain the only node in that zone?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.