OnCallReady

Lesson 22.31 · Azure I: CLI, Identity & Data Planes · 9 min read

Managed disks (RWO) and Azure Files (RWX) for Kubernetes

In plain words

Imagine a notebook and a whiteboard. A notebook can only be written in by one person at a time: if two people scribble in it at once, the pages become nonsense. A whiteboard on the classroom wall can be used by the whole class at once, but it's slower to walk up to and write on than your own notebook.

An Azure managed disk is the notebook: a block device attached to one VM at a time, so in Kubernetes it's ReadWriteOnce. It's also tied to one availability zone. Azure Files is the whiteboard: a network filesystem, SMB or NFS, that many nodes mount at once, so ReadWriteMany, at the cost of latency. On a Kubernetes cluster in Azure that's managed-csi versus azurefile-csi.

The shop team scales its uploads Deployment from 1 replica to 3. One pod runs; the other two sit in ContainerCreating forever. Nothing is wrong with the cluster - they asked Azure for something a disk physically cannot do. Three pods need a shared volume on a Kubernetes cluster in Azure - what do you use? (A Notion question.) Not an Azure Disk. Here is why, from the Azure side.

What you need to know already: 16.39 (PersistentVolume, PVC, StorageClass), 16.41 (access modes, zones), 16.44 (reclaim policies), 22.29 (storage accounts, redundancy, availability zones).

A managed disk attaches to one VM

A managed disk is Azure's virtual hard disk: a block device (like /dev/vda on your VM, 1.20) that Azure stores and attaches to a VM. When a PVC in the cluster uses a disk StorageClass, the CSI driver creates one of these per claim, named pvc-<uuid>. They live in the cluster's node resource group, the MC_... group Azure creates for the cluster's own machines.

az disk list -g <group> lists disks; this --query shows size, state and managedBy (which VM has it attached):

$ az disk list -g MC_rg-shop-prod_aks-shop-prod_westeurope --query "[].{name:name, gb:diskSizeGb, state:diskState, vm:managedBy}" -o table
Name                                      Gb    State       Vm
----------------------------------------  ----  ----------  -----------------------------------------------
pvc-3c9a51f2-8d1e-4b7a-9f02-6e4d1c8b7a10  128   Attached    /subscriptions/.../virtualMachineScaleSets/aks-user-31415926-vmss/virtualMachines/0
pvc-8b2e7d40-1f6c-4e3a-a5d9-0c7b6e2f1a93  64    Attached    /subscriptions/.../virtualMachineScaleSets/aks-user-31415926-vmss/virtualMachines/1
pvc-e41d9a07-5c2b-4f8e-b316-9a0d7c4e2b58  1024  Unattached

The VMs are instances of a VM scale set (22.9): each node of the cluster is one.

managedBy is a single VM. A disk is a block device; two machines writing a filesystem on one block device corrupt it, so Azure attaches it to one VM at a time. In Kubernetes terms that is ReadWriteOnce: one node at a time (so several pods on the same node can share it, which is rarely what you want).

Disks are also zonal: a Premium SSD created in availability zone 1 can only attach to a VM in zone 1. A pod rescheduled to a node in zone 2 hangs in ContainerCreating with a volume node affinity conflict. StorageClasses with volumeBindingMode: WaitForFirstConsumer (the defaults here) wait until the pod is scheduled and create the disk in that pod's zone. That avoids the first-time problem but not the move-later one. ZRS disks (Premium_ZRS, copies in three zones) remove it.

Azure Files is a network filesystem

An Azure Files share is a network file share served by a storage account (22.29), over SMB (or NFS on Premium). Any number of nodes mount it at once: ReadWriteMany. The price is latency and a throughput limit per share - fine for shared config, uploads, legacy apps that insist on a shared directory; wrong for a database.

The storage classes Azure gives your cluster

managed-csi              Azure Disk, StandardSSD_LRS     RWO   (the default)
managed-csi-premium      Azure Disk, Premium_LRS         RWO
azurefile-csi            Azure Files, Standard           RWX
azurefile-csi-premium    Azure Files, Premium (SSD)      RWX

(RWO = ReadWriteOnce, RWX = ReadWriteMany, 16.41.) A claim for a shared volume:

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: shared-uploads
  namespace: shop
spec:
  accessModes: ["ReadWriteMany"]
  storageClassName: azurefile-csi
  resources:
    requests:
      storage: 100Gi

Ask a disk class for ReadWriteMany and the claim never binds - the provisioner (the component that creates volumes for claims) rejects the access mode and the PVC stays Pending with a ProvisioningFailed event. Nothing is wrong with the cluster; the request is impossible.

So the answer: Azure Files (azurefile-csi, or Premium / NFS for performance), with ReadWriteMany. Or redesign so the pods do not share a filesystem - object storage (blobs, 22.29) is usually the better architecture.

Where the orphaned 1 TB disk came from

Every dynamically provisioned volume has a reclaimPolicy (16.44). The built-in classes use Delete: delete the PVC and the disk goes too. Custom classes (or PVs someone patched) with Retain keep the disk when the PVC is deleted - deliberately, so data survives mistakes. The side effect is disks nobody owns: diskState: Unattached, still billed every hour.

az disk list --query "[?diskState=='Unattached'].{name:name, gb:diskSizeGb, rg:resourceGroup, created:timeCreated}" -o table

The kubernetes.io-created-for-pvc-name tag tells you which claim it belonged to, before you decide whether anyone still needs it.

Snapshots and backup

A snapshot is a point-in-time copy of a disk: az snapshot create --source <disk id>, or, better, the Kubernetes VolumeSnapshot API with the CSI snapshot class. It is crash-consistent: like pulling the power cord, files the app was halfway through writing may be half-written. For anything that matters, use a real backup tool that can restore a whole namespace elsewhere (Azure Backup, or Velero, an open-source Kubernetes backup tool).

What you can now do

Why it helps

"Three pods need a shared volume, what do you use?" is a classic interview and ticket question for Kubernetes on Azure, and asking a disk StorageClass for ReadWriteMany gives a PVC stuck in Pending forever. Knowing the Azure side of storage classes lets you answer immediately and explain why.

Two other real incidents come from here. A pod rescheduled to a node in another zone hangs in ContainerCreating with a volume node affinity conflict, because its disk can't follow. And a Retain reclaim policy leaves unattached disks behind, still billed, like the 1 TB disk this chapter finds. Recognising both saves an outage and money, and knowing when shared storage is the wrong design, and blobs are better, is a senior-level answer.

Commands in this lesson

az

FAQ

What does ReadWriteOnce mean exactly?

The volume can be mounted read-write by one node at a time. Several pods on the same node can technically share it, which is rarely what you want. For Azure Disks that follows from how they work: a block device attached to a single VM, since two machines writing one filesystem would corrupt it. ReadWriteOncePod restricts it further to a single pod. For many nodes at once you need ReadWriteMany, which Azure Files provides.

Why did my pod hang with a volume node affinity conflict?

Azure Disks are zonal: a disk created in zone 1 can only attach to a VM in zone 1. If the pod is rescheduled to a node in zone 2, the disk can't follow and the pod stays in ContainerCreating or Pending. WaitForFirstConsumer on the StorageClass creates the disk where the pod first lands, but doesn't help when it moves later. Options: ZRS disks like Premium_ZRS, node pools per zone, or topology-aware scheduling.

When should I use Azure Files?

When several pods on different nodes genuinely need the same filesystem: shared configuration, user uploads, legacy apps that insist on a shared directory. It's a network filesystem with per-share throughput limits and higher latency than a disk, so it's the wrong choice for databases or anything with heavy small random I/O. Premium shares are SSD-backed and support NFS. Often the better design is object storage in blobs instead.

What does reclaimPolicy Retain do?

When the PVC is deleted, the PersistentVolume and the underlying Azure Disk are kept instead of deleted, so data survives mistakes. The built-in StorageClasses use Delete. The side effect of Retain is disks nobody owns, showing diskState: Unattached and still billed every hour. The kubernetes.io-created-for-pvc-name tag tells you which claim a disk belonged to before you decide whether to delete it.

Is a disk snapshot a backup?

It's a point-in-time, crash-consistent copy, which is useful but not a complete backup strategy. Crash-consistent means it's like pulling the power: databases usually recover, but not always cleanly without application coordination. A real backup also covers the Kubernetes objects that go with the data, retention, and restoring elsewhere. For a cluster, the VolumeSnapshot API with CSI snapshot classes, or Azure Backup or Velero for full backup and restore.

In an interview Mid

What is the difference between Azure Disks and Azure Files for Kubernetes persistent volumes?

Ask a disk class for ReadWriteMany and the PVC stays Pending with ProvisioningFailed - the request is impossible. So three pods sharing a volume = Azure Files, or better, redesign to object storage (blobs).

Housekeeping: classes with Retain leave disks behind when the PVC is deleted - az disk list --query "[?diskState=='Unattached']" finds them, and the created-for-pvc-name tag says whose they were.

Also asked: A pod is stuck in ContainerCreating after being rescheduled to a node in another zone. Why? · How do you find and clean up orphaned disks safely? · Three pods need a shared volume on a Kubernetes cluster in Azure. What do you use and why?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.