OnCallReady

Lesson 22.9 · Azure I: CLI, Identity & Data Planes · 17 min read

Managed identities: system-assigned vs user-assigned

In plain words

Imagine a school where each classroom has a key card built into its door frame. The classroom doesn't need to remember a password: when the teacher wants to open the supply cupboard, the door frame quietly asks the school office for a temporary pass, and the office hands one over because it knows which classroom is asking.

A managed identity is that built-in key card for Azure compute. Code asks the local metadata endpoint at 169.254.169.254 for a token, and there's no secret anywhere in its config. A system-assigned identity is born and dies with its resource, like a card built into one door. A user-assigned identity, like id-reports, is a separate card you can create first, give permissions to, and attach to many resources. In a Kubernetes cluster on Azure, pods get their own through workload identity.

A service principal's client secret has to live somewhere: an environment variable, a config file, a pipeline setting. Wherever it lives, it can leak, and it expires on a Friday night. For code that runs on Azure, there is a way to have no secret at all: let Azure itself vouch for the machine.

What you need to know already: 22.8 (service principals, appId vs object ID, federated credentials, OIDC), 17.33 (ServiceAccounts and their tokens), 11.1 (containers), 9.21 (HTTP and curl).

A managed identity (MI) is a service principal whose credential Azure creates, stores and rotates for you. Code running on Azure compute asks a local address on its own machine for a token; there is no secret anywhere in your config.

Azure compute means the Azure services that run your code: a VM (virtual machine, like your UTM VM but in Azure), a VM scale set (a group of identical VMs - the nodes of a Kubernetes cluster on Azure are one), App Service (Azure runs your web app for you), Functions, Container Apps.

How code gets a token

On a VM, a VM scale set node or App Service, the Instance Metadata Service (IMDS) answers on a link-local address - 169.254.169.254, an address that never leaves the machine, so only that machine can reach its own IMDS:

curl -s -H Metadata:true \
  "http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https://vault.azure.net"

-H Metadata:true is a required header (it proves the request did not come through a proxy); resource= says which Azure service the token is for.

{"access_token":"eyJ0eXAi...","expires_in":"86399","resource":"https://vault.azure.net","token_type":"Bearer",...}

access_token is the token itself; expires_in is seconds until it expires. resource is the audience: the service the token is valid for. A token for Key Vault is useless against Storage.

You rarely call IMDS by hand. The Azure SDKs (libraries for calling Azure from Java, Python, ...) do it: DefaultAzureCredential in every language tries environment variables, workload identity, then managed identity, then the developer's az login - which is why the same code works on a laptop and in Azure.

az login --identity does the same for the CLI - and fails on oncall-lab, because it is not an Azure VM and there is no 169.254.169.254 to ask.

System-assigned

A system-assigned identity is switched on as a property of one resource.

az vm identity assign -g rg -n vm-build        # enable on an existing VM

az aks show ... --query identity: just the cluster's own identity block.

$ az aks show -g rg-oncall-lab -n aks-sysop --query identity
{
  "principalId": "8d0b1c5e-...",
  "tenantId": "11111111-2222-3333-4444-555555555555",
  "type": "SystemAssigned",
  "userAssignedIdentities": null
}

User-assigned

A user-assigned identity is a resource of its own that you create and then attach to things. az identity create -g <group> -n <name> creates one.

$ az identity create -g rg-oncall-lab -n id-reports
{
  "clientId": "3f1e...",
  "id": "/subscriptions/.../resourceGroups/rg-oncall-lab/providers/Microsoft.ManagedIdentity/userAssignedIdentities/id-reports",
  "location": "westeurope",
  "name": "id-reports",
  "principalId": "a91c...",
  "resourceGroup": "rg-oncall-lab",
  "tenantId": "11111111-2222-3333-4444-555555555555",
  "type": "Microsoft.ManagedIdentity/userAssignedIdentities"
}

clientId is what the workload uses to say which identity it wants (a VM can carry several user-assigned identities). principalId is what RBAC is granted to. Same rule as for service principals (22.8): clientId to log in, object ID for permissions.

So which one?

questionsystem-assigneduser-assigned
lifecycletied to the resourceindependent
shared by many resourcesnoyes
RBAC before the resource existsnoyes
recreate the resourcenew principal, re-grant everythingnothing changes
cleanupautomaticyou delete it

The practical rule most platform teams use: user-assigned for anything managed by Terraform and anything where access must exist before the first start (which is most workloads); system-assigned for a one-off resource whose access is genuinely its own.

Managed identity or service principal?

The identities of a Kubernetes cluster on Azure

Azure can run a Kubernetes cluster for you (Ch 23 builds aks-sysop). That cluster alone has several identities:

$ az aks show -g rg-oncall-lab -n aks-sysop --query "{cluster:identity.type, kubelet:identityProfile.kubeletidentity.resourceId}"
{
  "cluster": "SystemAssigned",
  "kubelet": "/subscriptions/.../resourceGroups/MC_rg-oncall-lab_aks-sysop_westeurope/providers/Microsoft.ManagedIdentity/userAssignedIdentities/aks-sysop-agentpool"
}
identityused bytypically needs
cluster (control plane)Azure's cluster service itself: load balancers, public IPs, disksNetwork Contributor on a custom VNet subnet
kubelet (<cluster>-agentpool)every node: pulling imagesAcrPull on the registry
add-on identitiese.g. azurekeyvaultsecretsprovider-<cluster>whatever that add-on reads
workload identitiesyour podsexactly what that app needs

The kubelet identity is shared by every pod on every node. Granting it Key Vault access "to make the app work" gives every workload in the cluster that access. Pods get their own identities through workload identity.

Workload identity for the cluster

Workload identity lets one Kubernetes ServiceAccount act as one Azure identity. Pods do not reach IMDS for their own identity. Instead the cluster is an OIDC issuer (it signs its ServiceAccount tokens and publishes the keys to check them), and Entra ID trusts it:

1. The cluster runs with --enable-oidc-issuer --enable-workload-identity
2. You create a user-assigned identity           id-orders
3. You add a federated credential to it:
      issuer   = the cluster's OIDC issuer URL
      subject  = system:serviceaccount:<namespace>:<serviceaccount>
      audience = api://AzureADTokenExchange
4. The ServiceAccount carries   azure.workload.identity/client-id: <clientId of id-orders>
5. The pod carries the label    azure.workload.identity/use: "true"
6. A mutating webhook injects   AZURE_CLIENT_ID, AZURE_TENANT_ID,
                                AZURE_FEDERATED_TOKEN_FILE, AZURE_AUTHORITY_HOST
7. The SDK exchanges the projected service-account token for an Entra token
8. Azure RBAC decides what id-orders may do

Two new words in there:

Every link has to match exactly. The failures you will actually see:

(pod-managed identity, the older aad-pod-identity approach, is deprecated. Workload identity replaced it.)

An identity can have at most 20 federated credentials, so one identity per application, not one for the whole cluster.

What you can now do

Why it helps

Managed identities remove the most common source of Azure credential leaks: secrets in config. Every workload you run on an Azure cluster or VM should use one, and you'll be the person wiring them up in Terraform and debugging them when they fail.

The failures are specific and recognisable once you know the chain: AADSTS70021 when a workload identity's federated subject doesn't match the pod's ServiceAccount, a 403 when the identity is fine but the role is missing, or ten minutes of failed starts because a system-assigned identity got its role only after the resource existed. You'll also recognise the dangerous shortcut: granting the kubelet identity Key Vault access, which gives every pod in the cluster that access. That's a common finding in cluster security reviews.

Commands in this lesson

az

FAQ

Should I use system-assigned or user-assigned?

User-assigned for most platform workloads: it has its own lifecycle, can be shared by several resources, and, crucially, can receive its role assignments before the workload exists, so the app works on its first start. Recreating the workload doesn't change it. System-assigned suits a one-off resource whose access is truly its own and should vanish with it. Recreating a resource with a system-assigned identity gives it a new principal, and all its role assignments must be made again.

What are clientId and principalId on an identity?

clientId is what the workload uses to say which identity it wants, since a VM or pod can have several user-assigned identities. It goes into SDK configuration, the ServiceAccount annotation or a SecretProviderClass. principalId is the object ID of the underlying service principal, and it's what you grant RBAC roles to. Same rule as for app registrations: client ID to sign in, object ID for permissions.

Why doesn't az login --identity work on oncall-lab?

Because oncall-lab is not an Azure VM. Managed identity tokens come from the Instance Metadata Service at 169.254.169.254, a link-local address that exists only on Azure compute and answers only for that machine's identities. On your laptop or a VM in UTM there's nothing to answer. That's also why DefaultAzureCredential falls through to your az login locally: the managed identity step finds nothing.

Why not grant the kubelet identity access to Key Vault?

The kubelet identity is shared by every node in the pool and so, in effect, by every pod on those nodes. Granting it Key Vault Secrets User to make one app work gives every workload in the cluster, including anything compromised, access to those secrets. Its job is pulling images, so it needs AcrPull on the registry. Each app should get its own user-assigned identity through workload identity, with roles only on what that app needs.

What does AADSTS70021 mean with workload identity?

"No matching federated identity record found for presented assertion": Entra ID received a ServiceAccount token that doesn't match any federated credential on the identity. The usual causes are a subject mismatch, where the credential says system:serviceaccount:orders:orders but the pod uses orders-api, or an issuer mismatch after the cluster was recreated with a new OIDC issuer URL. It's an authentication failure, so no role assignment will fix it.

In an interview Mid

What is a managed identity, and when do you use system-assigned versus user-assigned?

A managed identity is a service principal whose credential Azure creates, stores and rotates. Code on Azure compute asks the Instance Metadata Service (169.254.169.254, reachable only from that machine) for a token - the SDKs' DefaultAzureCredential does it - so there is no secret in any config.

Rule most platform teams use: user-assigned for anything Terraform manages; system-assigned for a one-off resource. clientId says which identity the workload wants; principalId is what RBAC is granted to.

On Azure, managed identity always; outside Azure, a federated credential. And never grant app access to a cluster's kubelet identity - every pod on every node shares it; pods get their own through workload identity.

Also asked: How does workload identity let a pod authenticate to Azure without a secret? · When would you use a service principal instead of a managed identity? · Why should you not give the cluster's kubelet identity access to application secrets?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.