Micrometer
The problem. "Is it slow, and why?" needs numbers from inside the app: request counts and times, connections in use, threads busy. The app keeps those numbers itself; you only have to know where to read them and what they mean.
What you need to know already: the golden signals and percentiles (0.1, 0.2), Actuator (21.1), grep and awk (7.1, 7.8), jq (7.11), the JVM heap and GC (20.1, 20.12).
A metric is a number the app keeps about itself; a meter is one named metric inside the app. Micrometer is the Java library Spring Boot uses to record them - the app records a Counter (only goes up: requests served), a Timer (how long things take), a Gauge (a current value: connections in use) or a DistributionSummary (sizes), and a registry turns them into the format some monitoring system wants. With micrometer-registry-prometheus on the classpath, Actuator serves them as plain text at /actuator/prometheus - one line per number, readable with curl and grep, and the format most monitoring tools can collect. (Prometheus is the monitoring system the format is named after.) Spring, Tomcat, HikariCP, the JVM and Resilience4j all register meters automatically.
Later (Ch 27): Prometheus collects this page from every pod and lets you query the history.
Reading the text format
# with prometheus exposed (the next mission's setup)
curl -s localhost:8080/actuator/prometheus | grep -A3 'hikaricp_connections_active'
# HELP hikaricp_connections_active Active connections
# TYPE hikaricp_connections_active gauge
hikaricp_connections_active{pool="HikariPool-1"} 4.0
# HELP- description.# TYPE-counter,gauge,summary,histogram.- the sample:
name{label="value",...} number. A label is a key=value tag on the number (which pool, which URL); each distinct label combination is its own series. - Micrometer names use dots (
hikaricp.connections.active); the text-format registry converts them to underscores and adds unit suffixes (_seconds,_bytes) and_totalfor counters.
A Timer: http_server_requests
http_server_requests_seconds_count{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 500
http_server_requests_seconds_sum{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 2.001521
http_server_requests_seconds_max{error="none",exception="none",method="GET",outcome="SUCCESS",status="200",uri="/api/orders"} 0.004798
_count- requests since start (a counter)._sum- total seconds spent in them. Average latency = sum / count (here 4 ms). For a recent average, take two readings a minute apart and divide the change in sum by the change in count - the same trick ascpu.statin 17.6._max- the maximum in a decaying window (about 2 minutes), not since start.- One series per combination of labels: method, status, outcome, exception, and
uri- the route template (/api/orders/{id}), not the raw path, otherwise cardinality (the number of different series) would explode.
Averages hide the tail (0.2). For percentiles you need a histogram (counts of requests per latency bucket: how many under 0.1 s, under 0.5 s...):
management.metrics.distribution.percentiles-histogram.http.server.requests=true
which adds http_server_requests_seconds_bucket{le="0.1",...} series (le = "less than or equal": requests that took at most 0.1 s). A monitoring system turns those bucket counts into a p99 for you.
The meters that matter for this chapter
http_server_requests_seconds_* latency and rate per endpoint and status
hikaricp_connections_active connections in use (gauge)
hikaricp_connections_pending threads WAITING for one <- pool exhaustion, the early warning
hikaricp_connections_max the configured pool size <- what Hikari really uses
hikaricp_connections_timeout_total borrow timeouts (counter)
hikaricp_connections_acquire_seconds_* time to get a connection
tomcat_threads_busy_threads busy request threads (needs server.tomcat.mbeanregistry.enabled=true)
executor_active_threads / executor_queued_tasks an async executor's saturation
jvm_memory_used_bytes{area="heap",id="G1 Old Gen"} heap by pool
jvm_gc_pause_seconds_* GC pauses by cause
jvm_threads_live_threads thread count: alert on growth (thread leaks)
process_cpu_usage 0..1
resilience4j_circuitbreaker_state{state="open"} 1 when open
tomcat_threads_* are missing by default: Boot disables Tomcat's MBean registry to save memory, and the meters come from it. Set server.tomcat.mbeanregistry.enabled=true if you want them.
Reading them from the shell
curl -s localhost:8080/actuator/prometheus | grep -E '^hikaricp_connections_(active|pending|max)'
curl -s localhost:8080/actuator/prometheus | grep '^http_server_requests_seconds_count' | awk '{s+=$2} END {print s}'
curl -s 'localhost:8080/actuator/metrics/hikaricp.connections.pending' | jq '.measurements'
curl -s 'localhost:8080/actuator/metrics/http.server.requests?tag=uri:/api/orders&tag=status:200' | jq .
The /actuator/metrics/{name} form aggregates across series and lists the availableTags you can drill into with ?tag=key:value.
Pool exhaustion in the numbers
hikaricp_connections_active{pool="HikariPool-1"} 20.0 = max
hikaricp_connections_pending{pool="HikariPool-1"} 143.0 threads queued for a connection
http_server_requests_seconds_max{...,uri="/api/orders"} 27.9 and no 5xx series growing
process_cpu_usage 0.02 nothing is computing
p99 climbing, error rate flat, CPU normal: the signature of a pool that is full. The next lessons are about why that happens and how to stop it.
What you can now do
- Read the plain-text metrics page: HELP, TYPE, samples, labels.
- Compute average latency from a Timer's _sum and _count, and know why you need a histogram for p99.
- Name the meters that show pool, thread and executor saturation.