A control loop over the replica count
The problem. At 09:00 traffic triples; at 03:00 it is almost zero. A fixed replica count is either too few in the morning or wasted money at night. The HorizontalPodAutoscaler (HPA) changes the replica count for you, based on measured load. ("Horizontal" = more or fewer pods; "vertical" would be bigger pods.)
What you need to know already: requests (17.1), metrics-server and kubectl top (17.1), the reconciliation loop and kube-controller-manager (15.5, 15.9), Deployments (15.16).
The HPA controller (a loop inside kube-controller-manager) runs every 15 seconds. For each HPA it reads the metric of the target's pods from the metrics API (metrics-server for CPU and memory), computes a desired replica count and writes it into the target's spec.replicas through the scale subresource (the small part of the API that changes only the replica count - what kubectl scale uses too). That is all: the Deployment does the rest. (ceil below = round up to the next whole number.)
desiredReplicas = ceil( currentReplicas x currentMetricValue / targetMetricValue )
With 3 pods at 90% CPU utilization and a 50% target: ceil(3 x 90 / 50) = ceil(5.4) = 6. If the ratio is within 10% of 1 (the default tolerance), nothing changes - that stops it chasing noise.
Utilization is a percentage of the pods' requests. 50% of a 200m request is 100m. So:
An HPA on CPU utilization cannot work without CPU requests. No request = no denominator.
# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl get hpa -n shop
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
web Deployment/web cpu: <unknown>/50% 1 10 1 3m
kubectl describe hpa web -n shop | tail -3
Warning FailedGetResourceMetric 15s (x12 over 3m) horizontal-pod-autoscaler failed to get cpu utilization: missing request for cpu in container web of Pod web-5d8f7c9b4d-2kq9x
Warning FailedComputeMetricsReplicas 15s (x12 over 3m) horizontal-pod-autoscaler invalid metrics (1 invalid out of 1), first error is: failed to get cpu resource metric value: failed to get cpu utilization: missing request for cpu in container web of Pod web-5d8f7c9b4d-2kq9x
<unknown> is also what you see in the first ~30 seconds (no metrics yet) and when metrics-server is down (unable to fetch metrics from resource metrics API). The sidecar version of the trap: every container in the pod needs the request.
Creating one
# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl autoscale deployment php-apache --cpu=50% --min=1 --max=10
horizontalpodautoscaler.autoscaling/php-apache autoscaled
kubectl autoscale deployment NAME creates an HPA for that Deployment: --cpu=50% = keep average CPU at 50% of the request, --min=1 / --max=10 = never fewer / more replicas than that.
(--cpu-percent=50 still works but is deprecated since 1.34; --cpu=500m would target an average value instead of utilization.) The object it creates is autoscaling/v2:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: php-apache
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: php-apache
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
Metric types: Resource (cpu/memory of the pods), ContainerResource (one container only - useful with sidecars), Pods and Object (custom metrics - numbers your app reports, such as requests per second, served to the HPA by an extra add-on called a metrics adapter), External (a number from outside the cluster, such as the length of a message queue). With several metrics, the HPA computes a count for each and takes the highest.
Reading it
# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl get hpa php-apache
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
php-apache Deployment/php-apache cpu: 47%/50% 1 10 5 4m
kubectl describe hpa php-apache
Metrics: ( current / target )
resource cpu on pods (as a percentage of request): 47% (94m) / 50%
Min replicas: 1
Max replicas: 10
Deployment pods: 5 current / 5 desired
Conditions:
Type Status Reason Message
---- ------ ------ -------
AbleToScale True ReadyForNewScale recommended size matches current size
ScalingActive True ValidMetricFound the HPA was able to successfully calculate a replica count from cpu resource utilization (percentage of request)
ScalingLimited False DesiredWithinRange the desired count is within the acceptable range
Events:
Normal SuccessfulRescale 74s horizontal-pod-autoscaler New size: 5; reason: cpu resource utilization (percentage of request) above target
The three conditions are the diagnosis: AbleToScale (can it read/write the target's scale?), ScalingActive (does it have valid metrics? False = look at the reason), ScalingLimited (is it pinned at min or max? TooManyReplicas = raise max or fix the app).
behavior: stopping the flapping
Load is spiky; scaling on every spike is expensive and scaling down too eagerly means scaling up again a minute later. behavior controls both directions:
spec:
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # react immediately (default)
policies: # default: +100% or +4 pods per 15s, whichever is more
- {type: Percent, value: 100, periodSeconds: 15}
- {type: Pods, value: 4, periodSeconds: 15}
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 300 # default: use the HIGHEST recommendation of the last 5 min
policies:
- {type: Percent, value: 100, periodSeconds: 15}
The scale-down window is why load disappears and replicas stay up for five minutes - by design. The condition says so: AbleToScale True ScaleDownStabilized recent recommendations were higher than current one, applying the highest recent recommendation. Shorten it for bursty batch, lengthen it for services with slow startup (a JVM that takes 90s to become Ready should not be scaled down and up every two minutes).
Living with an HPA
- Do not set
replicasin the manifest you apply once an HPA owns the Deployment - everykubectl applyresets the count, the HPA fixes it 15s later, and you get a sawtooth. Remove the field from the YAML you apply. kubectl scalefights the HPA the same way: the next sync overwrites you.- Scaling up adds pods, which need capacity: the HPA creates Pending pods if the nodes are full; the cluster autoscaler adds nodes. Both have to be tuned together.
- CPU is a proxy. For request-driven services, requests per second or queue depth (custom metrics) often scale better; for JVMs, CPU at startup spikes and can trigger scale-ups that are pure noise - another reason for sensible behavior settings.
- VPA (VerticalPodAutoscaler - an add-on that recommends or sets requests instead of replica counts); do not let VPA and HPA both act on CPU for the same workload.
What you can now do
- Create an HPA (
kubectl autoscale), and compute the replica count it will pick. - Diagnose
<unknown>targets (missing requests, no metrics) and read the three conditions. - Tune
behaviorso it scales down faster or slower on purpose.