The shop team scales its uploads Deployment from 1 replica to 3. One pod runs; the other two sit in ContainerCreating forever. Nothing is wrong with the cluster - they asked Azure for something a disk physically cannot do. Three pods need a shared volume on a Kubernetes cluster in Azure - what do you use? (A Notion question.) Not an Azure Disk. Here is why, from the Azure side.
What you need to know already: 16.39 (PersistentVolume, PVC, StorageClass), 16.41 (access modes, zones), 16.44 (reclaim policies), 22.29 (storage accounts, redundancy, availability zones).
A managed disk attaches to one VM
A managed disk is Azure's virtual hard disk: a block device (like /dev/vda on your VM, 1.20) that Azure stores and attaches to a VM. When a PVC in the cluster uses a disk StorageClass, the CSI driver creates one of these per claim, named pvc-<uuid>. They live in the cluster's node resource group, the MC_... group Azure creates for the cluster's own machines.
az disk list -g <group> lists disks; this --query shows size, state and managedBy (which VM has it attached):
$ az disk list -g MC_rg-shop-prod_aks-shop-prod_westeurope --query "[].{name:name, gb:diskSizeGb, state:diskState, vm:managedBy}" -o table
Name Gb State Vm
---------------------------------------- ---- ---------- -----------------------------------------------
pvc-3c9a51f2-8d1e-4b7a-9f02-6e4d1c8b7a10 128 Attached /subscriptions/.../virtualMachineScaleSets/aks-user-31415926-vmss/virtualMachines/0
pvc-8b2e7d40-1f6c-4e3a-a5d9-0c7b6e2f1a93 64 Attached /subscriptions/.../virtualMachineScaleSets/aks-user-31415926-vmss/virtualMachines/1
pvc-e41d9a07-5c2b-4f8e-b316-9a0d7c4e2b58 1024 Unattached
The VMs are instances of a VM scale set (22.9): each node of the cluster is one.
managedBy is a single VM. A disk is a block device; two machines writing a filesystem on one block device corrupt it, so Azure attaches it to one VM at a time. In Kubernetes terms that is ReadWriteOnce: one node at a time (so several pods on the same node can share it, which is rarely what you want).
Disks are also zonal: a Premium SSD created in availability zone 1 can only attach to a VM in zone 1. A pod rescheduled to a node in zone 2 hangs in ContainerCreating with a volume node affinity conflict. StorageClasses with volumeBindingMode: WaitForFirstConsumer (the defaults here) wait until the pod is scheduled and create the disk in that pod's zone. That avoids the first-time problem but not the move-later one. ZRS disks (Premium_ZRS, copies in three zones) remove it.
Azure Files is a network filesystem
An Azure Files share is a network file share served by a storage account (22.29), over SMB (or NFS on Premium). Any number of nodes mount it at once: ReadWriteMany. The price is latency and a throughput limit per share - fine for shared config, uploads, legacy apps that insist on a shared directory; wrong for a database.
The storage classes Azure gives your cluster
managed-csi Azure Disk, StandardSSD_LRS RWO (the default)
managed-csi-premium Azure Disk, Premium_LRS RWO
azurefile-csi Azure Files, Standard RWX
azurefile-csi-premium Azure Files, Premium (SSD) RWX
(RWO = ReadWriteOnce, RWX = ReadWriteMany, 16.41.) A claim for a shared volume:
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: shared-uploads
namespace: shop
spec:
accessModes: ["ReadWriteMany"]
storageClassName: azurefile-csi
resources:
requests:
storage: 100Gi
Ask a disk class for ReadWriteMany and the claim never binds - the provisioner (the component that creates volumes for claims) rejects the access mode and the PVC stays Pending with a ProvisioningFailed event. Nothing is wrong with the cluster; the request is impossible.
So the answer: Azure Files (azurefile-csi, or Premium / NFS for performance), with ReadWriteMany. Or redesign so the pods do not share a filesystem - object storage (blobs, 22.29) is usually the better architecture.
Where the orphaned 1 TB disk came from
Every dynamically provisioned volume has a reclaimPolicy (16.44). The built-in classes use Delete: delete the PVC and the disk goes too. Custom classes (or PVs someone patched) with Retain keep the disk when the PVC is deleted - deliberately, so data survives mistakes. The side effect is disks nobody owns: diskState: Unattached, still billed every hour.
az disk list --query "[?diskState=='Unattached'].{name:name, gb:diskSizeGb, rg:resourceGroup, created:timeCreated}" -o table
The kubernetes.io-created-for-pvc-name tag tells you which claim it belonged to, before you decide whether anyone still needs it.
Snapshots and backup
A snapshot is a point-in-time copy of a disk: az snapshot create --source <disk id>, or, better, the Kubernetes VolumeSnapshot API with the CSI snapshot class. It is crash-consistent: like pulling the power cord, files the app was halfway through writing may be half-written. For anything that matters, use a real backup tool that can restore a whole namespace elsewhere (Azure Backup, or Velero, an open-source Kubernetes backup tool).
What you can now do
- Explain why an Azure Disk is ReadWriteOnce and zonal, and what that does to pod scheduling.
- Pick the StorageClass for a shared volume, and read a PVC stuck in Pending.
- Find unattached disks and trace them back to the claim they came from.