OnCallReady

Lesson 21.11 · Spring Boot Runtime, Resilience & Python Ops · 11 min read

Micrometer and /actuator/prometheus

In plain words

Imagine a bus driver with a tally counter and a stopwatch. Every passenger who gets on, click. Every trip, the stopwatch adds up the time. At the end of the day anyone can read the counter and the total time and work out the average trip. A separate little board shows how many people are standing right now.

Micrometer is the counter and stopwatch inside a Spring app. Counters only go up, like requests served; timers record count, total time and a recent maximum, like http_server_requests_seconds; gauges show a current value, like hikaricp_connections_pending. Actuator shows all of them at /actuator/prometheus in a plain text format monitoring tools collect, and a pool that is full shows up as pending connections rising while CPU stays idle.

Micrometer

The problem. "Is it slow, and why?" needs numbers from inside the app: request counts and times, connections in use, threads busy. The app keeps those numbers itself; you only have to know where to read them and what they mean.

What you need to know already: the golden signals and percentiles (0.1, 0.2), Actuator (21.1), grep and awk (7.1, 7.8), jq (7.11), the JVM heap and GC (20.1, 20.12).

A metric is a number the app keeps about itself; a meter is one named metric inside the app. Micrometer is the Java library Spring Boot uses to record them - the app records a Counter (only goes up: requests served), a Timer (how long things take), a Gauge (a current value: connections in use) or a DistributionSummary (sizes), and a registry turns them into the format some monitoring system wants. With micrometer-registry-prometheus on the classpath, Actuator serves them as plain text at /actuator/prometheus - one line per number, readable with curl and grep, and the format most monitoring tools can collect. (Prometheus is the monitoring system the format is named after.) Spring, Tomcat, HikariCP, the JVM and Resilience4j all register meters automatically.

Later (Ch 27): Prometheus collects this page from every pod and lets you query the history.

Reading the text format

# with prometheus exposed (the next mission's setup)
curl -s localhost:8080/actuator/prometheus | grep -A3 'hikaricp_connections_active'
# HELP hikaricp_connections_active Active connections
# TYPE hikaricp_connections_active gauge
hikaricp_connections_active{pool="HikariPool-1"} 4.0

A Timer: http_server_requests

http_server_requests_seconds_count{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 500
http_server_requests_seconds_sum{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 2.001521
http_server_requests_seconds_max{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 0.004798

Averages hide the tail (0.2). For percentiles you need a histogram (counts of requests per latency bucket: how many under 0.1 s, under 0.5 s...):

management.metrics.distribution.percentiles-histogram.http.server.requests=true

which adds http_server_requests_seconds_bucket{le="0.1",...} series (le = "less than or equal": requests that took at most 0.1 s). A monitoring system turns those bucket counts into a p99 for you.

The meters that matter for this chapter

http_server_requests_seconds_*     latency and rate per endpoint and status
hikaricp_connections_active        connections in use          (gauge)
hikaricp_connections_pending       threads WAITING for one      <- pool exhaustion, the early warning
hikaricp_connections_max           the configured pool size     <- what Hikari really uses
hikaricp_connections_timeout_total borrow timeouts              (counter)
hikaricp_connections_acquire_seconds_*   time to get a connection
tomcat_threads_busy_threads        busy request threads (needs server.tomcat.mbeanregistry.enabled=true)
executor_active_threads / executor_queued_tasks   an async executor's saturation
jvm_memory_used_bytes{area="heap",id="G1 Old Gen"}   heap by pool
jvm_gc_pause_seconds_*             GC pauses by cause
jvm_threads_live_threads           thread count: alert on growth (thread leaks)
process_cpu_usage                  0..1
resilience4j_circuitbreaker_state{state="open"}   1 when open

tomcat_threads_* are missing by default: Boot disables Tomcat's MBean registry to save memory, and the meters come from it. Set server.tomcat.mbeanregistry.enabled=true if you want them.

Reading them from the shell

curl -s localhost:8080/actuator/prometheus | grep -E '^hikaricp_connections_(active|pending|max)'
curl -s localhost:8080/actuator/prometheus | grep '^http_server_requests_seconds_count' | awk '{s+=$2} END {print s}'
curl -s 'localhost:8080/actuator/metrics/hikaricp.connections.pending' | jq '.measurements'
curl -s 'localhost:8080/actuator/metrics/http.server.requests?tag=uri:/api/orders&tag=status:200' | jq .

The /actuator/metrics/{name} form aggregates across series and lists the availableTags you can drill into with ?tag=key:value.

Pool exhaustion in the numbers

hikaricp_connections_active{pool="HikariPool-1"}   20.0     = max
hikaricp_connections_pending{pool="HikariPool-1"} 143.0     threads queued for a connection
http_server_requests_seconds_max{...,uri="/api/orders"} 27.9  and no 5xx series growing
process_cpu_usage 0.02                                      nothing is computing

p99 climbing, error rate flat, CPU normal: the signature of a pool that is full. The next lessons are about why that happens and how to stop it.

What you can now do

Why it helps

Metrics are how you'll see most Java incidents before anyone runs a thread dump. hikaricp_connections_pending above zero, executor_queued_tasks climbing, jvm_threads_live_threads growing, or a circuit breaker state flipping to open each tell you what kind of problem you have from a dashboard. "p99 up, errors flat, CPU idle" is the signature of pool exhaustion, and you'll recognise it.

It also matters for dashboards and alerts you build: average latency is rate(sum)/rate(count), percentiles need histogram buckets turned on, and _max is a decaying window, not all-time. Getting those wrong produces dashboards that hide the problem. And knowing tomcat_threads_* needs the MBean registry enabled saves an afternoon of "why is this metric missing?".

FAQ

Why are metric names different at /actuator/metrics and /actuator/prometheus?

Micrometer uses dotted names like hikaricp.connections.active, which is what you query at /actuator/metrics/{name}. The text-format registry converts dots to underscores and adds conventional suffixes: _seconds or _bytes for base units, and _total for counters. So the same meter is hikaricp_connections_active at /actuator/prometheus. A timer produces _count, _sum and _max series, and _bucket if histograms are enabled.

How do I get p99 latency from Spring metrics?

Enable histogram buckets: management.metrics.distribution.percentiles-histogram.http.server.requests=true. That adds http_server_requests_seconds_bucket{le="..."} series (counts of requests faster than each limit), from which a monitoring system can compute the p99 across all pods. Without buckets you only have count, sum and max, which gives average latency and a recent maximum but no percentiles that can be aggregated across pods.

What does _max actually measure?

The largest value recorded in a decaying window, about two minutes by default, not since the application started. It tells you "the slowest request recently" and drops back once slow requests stop. That makes it good for spotting spikes and GC pauses, but it cannot be averaged or turned into a percentile across pods, and a single slow request moves it. Use histograms for SLOs.

Why is the uri label a template like /api/orders/{id}?

To keep cardinality bounded. Every unique combination of label values is a separate time series (one stored line of numbers) in the monitoring system. If the raw path /api/orders/123456 were the label, every order ID would create a new series, and the monitoring system's memory would explode. Spring uses the route template instead. When you add your own tags, the same rule applies: never use user IDs, request IDs or other unbounded values as labels.

Why are the tomcat_threads metrics missing?

Spring Boot disables Tomcat's MBean registry by default to save memory, and Micrometer's Tomcat thread metrics come from it. Set server.tomcat.mbeanregistry.enabled=true and restart to get tomcat_threads_busy_threads and friends. Other metrics like http_server_requests_seconds and hikaricp_* do not depend on it and are there by default.

In an interview Mid

Explain counters, gauges and timers or histograms, with an example of each from a Java service.

Each label combination is its own series - which is why uri is the route template (/api/orders/{id}), not the raw path: otherwise cardinality explodes. Micrometer records them; /actuator/prometheus serves them as text.

Also asked: Which metrics would you put on a dashboard for a Spring Boot service, and why? · What is metric cardinality and why does it matter? · Why is average latency a poor signal on its own?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.