A service principal's client secret has to live somewhere: an environment variable, a config file, a pipeline setting. Wherever it lives, it can leak, and it expires on a Friday night. For code that runs on Azure, there is a way to have no secret at all: let Azure itself vouch for the machine.
What you need to know already: 22.8 (service principals, appId vs object ID, federated credentials, OIDC), 17.33 (ServiceAccounts and their tokens), 11.1 (containers), 9.21 (HTTP and curl).
A managed identity (MI) is a service principal whose credential Azure creates, stores and rotates for you. Code running on Azure compute asks a local address on its own machine for a token; there is no secret anywhere in your config.
Azure compute means the Azure services that run your code: a VM (virtual machine, like your UTM VM but in Azure), a VM scale set (a group of identical VMs - the nodes of a Kubernetes cluster on Azure are one), App Service (Azure runs your web app for you), Functions, Container Apps.
How code gets a token
On a VM, a VM scale set node or App Service, the Instance Metadata Service (IMDS) answers on a link-local address - 169.254.169.254, an address that never leaves the machine, so only that machine can reach its own IMDS:
curl -s -H Metadata:true \
"http://169.254.169.254/metadata/identity/oauth2/token?api-version=2018-02-01&resource=https://vault.azure.net"
-H Metadata:true is a required header (it proves the request did not come through a proxy); resource= says which Azure service the token is for.
{"access_token":"eyJ0eXAi...","expires_in":"86399","resource":"https://vault.azure.net","token_type":"Bearer",...}
access_token is the token itself; expires_in is seconds until it expires. resource is the audience: the service the token is valid for. A token for Key Vault is useless against Storage.
You rarely call IMDS by hand. The Azure SDKs (libraries for calling Azure from Java, Python, ...) do it: DefaultAzureCredential in every language tries environment variables, workload identity, then managed identity, then the developer's az login - which is why the same code works on a laptop and in Azure.
az login --identity does the same for the CLI - and fails on oncall-lab, because it is not an Azure VM and there is no 169.254.169.254 to ask.
System-assigned
A system-assigned identity is switched on as a property of one resource.
az vm identity assign -g rg -n vm-build # enable on an existing VM
- Created with the resource, deleted with the resource. One-to-one.
- Its object ID appears as
identity.principalIdon the resource. - Delete and recreate the VM and you get a new principal, so every role assignment has to be made again. The old assignments stay behind, pointing at an object that no longer exists (the portal shows them as Identity not found).
az aks show ... --query identity: just the cluster's own identity block.
$ az aks show -g rg-oncall-lab -n aks-sysop --query identity
{
"principalId": "8d0b1c5e-...",
"tenantId": "11111111-2222-3333-4444-555555555555",
"type": "SystemAssigned",
"userAssignedIdentities": null
}
User-assigned
A user-assigned identity is a resource of its own that you create and then attach to things. az identity create -g <group> -n <name> creates one.
$ az identity create -g rg-oncall-lab -n id-reports
{
"clientId": "3f1e...",
"id": "/subscriptions/.../resourceGroups/rg-oncall-lab/providers/Microsoft.ManagedIdentity/userAssignedIdentities/id-reports",
"location": "westeurope",
"name": "id-reports",
"principalId": "a91c...",
"resourceGroup": "rg-oncall-lab",
"tenantId": "11111111-2222-3333-4444-555555555555",
"type": "Microsoft.ManagedIdentity/userAssignedIdentities"
}
- A standalone Azure resource with its own lifecycle.
- Can be attached to many resources (all the nodes of a pool, several apps that need the same access).
- You can grant it roles before the workload exists - so the Terraform that creates the identity and the role assignment can run first, and the app gets its access on the very first start instead of failing for ten minutes while RBAC propagates (spreads through Azure; a new assignment can take minutes to take effect).
- Recreating the workload does not change the identity; its role assignments survive.
clientId is what the workload uses to say which identity it wants (a VM can carry several user-assigned identities). principalId is what RBAC is granted to. Same rule as for service principals (22.8): clientId to log in, object ID for permissions.
So which one?
| question | system-assigned | user-assigned |
|---|---|---|
| lifecycle | tied to the resource | independent |
| shared by many resources | no | yes |
| RBAC before the resource exists | no | yes |
| recreate the resource | new principal, re-grant everything | nothing changes |
| cleanup | automatic | you delete it |
The practical rule most platform teams use: user-assigned for anything managed by Terraform and anything where access must exist before the first start (which is most workloads); system-assigned for a one-off resource whose access is genuinely its own.
Managed identity or service principal?
- The code runs on Azure (VM, a Kubernetes cluster, App Service, Functions, Container Apps): managed identity, always. There is no secret to leak or rotate.
- The code runs outside Azure (a GitHub workflow, build machines you do not host, your own data centre): an app registration with a federated credential if the platform issues OIDC tokens, otherwise a certificate. A client secret is the last resort.
The identities of a Kubernetes cluster on Azure
Azure can run a Kubernetes cluster for you (Ch 23 builds aks-sysop). That cluster alone has several identities:
$ az aks show -g rg-oncall-lab -n aks-sysop --query "{cluster:identity.type, kubelet:identityProfile.kubeletidentity.resourceId}"
{
"cluster": "SystemAssigned",
"kubelet": "/subscriptions/.../resourceGroups/MC_rg-oncall-lab_aks-sysop_westeurope/providers/Microsoft.ManagedIdentity/userAssignedIdentities/aks-sysop-agentpool"
}
| identity | used by | typically needs |
|---|---|---|
| cluster (control plane) | Azure's cluster service itself: load balancers, public IPs, disks | Network Contributor on a custom VNet subnet |
kubelet (<cluster>-agentpool) | every node: pulling images | AcrPull on the registry |
| add-on identities | e.g. azurekeyvaultsecretsprovider-<cluster> | whatever that add-on reads |
| workload identities | your pods | exactly what that app needs |
- The kubelet identity is the one the node's kubelet (Ch 15) uses, mainly to pull images from ACR (Azure Container Registry, Azure's image registry, Ch 10). AcrPull is the role that allows pulling.
- An add-on is an optional cluster feature Microsoft runs for you (like the Key Vault driver, lesson 22.27); each gets its own identity.
The kubelet identity is shared by every pod on every node. Granting it Key Vault access "to make the app work" gives every workload in the cluster that access. Pods get their own identities through workload identity.
Workload identity for the cluster
Workload identity lets one Kubernetes ServiceAccount act as one Azure identity. Pods do not reach IMDS for their own identity. Instead the cluster is an OIDC issuer (it signs its ServiceAccount tokens and publishes the keys to check them), and Entra ID trusts it:
1. The cluster runs with --enable-oidc-issuer --enable-workload-identity
2. You create a user-assigned identity id-orders
3. You add a federated credential to it:
issuer = the cluster's OIDC issuer URL
subject = system:serviceaccount:<namespace>:<serviceaccount>
audience = api://AzureADTokenExchange
4. The ServiceAccount carries azure.workload.identity/client-id: <clientId of id-orders>
5. The pod carries the label azure.workload.identity/use: "true"
6. A mutating webhook injects AZURE_CLIENT_ID, AZURE_TENANT_ID,
AZURE_FEDERATED_TOKEN_FILE, AZURE_AUTHORITY_HOST
7. The SDK exchanges the projected service-account token for an Entra token
8. Azure RBAC decides what id-orders may do
Two new words in there:
- A mutating webhook is a hook in the Kubernetes API server that may change a pod as it is created. Here it adds the environment variables and the token file to every pod with the label.
- The projected service-account token is the ServiceAccount token (17.33) the kubelet writes into the pod as a file; the SDK trades it for an Azure token (token exchange).
Every link has to match exactly. The failures you will actually see:
- subject mismatch: the federated credential says
system:serviceaccount:orders:orders, the pod runs asorders-api. Token exchange fails withAADSTS70021: No matching federated identity record found for presented assertion. - wrong issuer: the cluster was recreated and has a new issuer URL; the credential still trusts the old one. Same AADSTS70021.
- token exchange works, access denied: identity fine, RBAC missing or at the wrong scope - a 403 (HTTP Forbidden) from the service. The incident at the end of this chapter is this one.
(pod-managed identity, the older aad-pod-identity approach, is deprecated. Workload identity replaced it.)
An identity can have at most 20 federated credentials, so one identity per application, not one for the whole cluster.
What you can now do
- Choose between system-assigned and user-assigned, and between a managed identity and a service principal.
- Name the identities of a Kubernetes cluster on Azure and why the kubelet one must not get app access.
- Trace the eight links of workload identity, and which failure breaks which link.