Why you need more than ssh and grep
So far you debugged one box by logging in and looking: top, journalctl, grep on a log. That stops working when the app runs as 20 copies on 20 machines and the question is "checkout got slow at 20:02 yesterday - why?". Nobody was logged in at 20:02, and you cannot grep 20 machines by hand.
What you need is data the system records about itself all the time, kept in one place, that you can ask questions later. That is this block of chapters.
What you need to know already: the golden signals - latency, traffic, errors, saturation (0.1); percentiles such as p99 (0.2); SLIs and SLOs (0.8); journalctl (2.30); grep and jq (7.1, 7.11); HTTP requests, headers and status codes (9.21); the orders Spring Boot app and its Actuator endpoints (21.1, 21.11).
The words you need first
- Observability - being able to answer a question about a running system from the outside, from data it already produces, without shipping new code to ask it.
- Telemetry - that data: what a program reports about itself while it runs. It comes in three kinds, often called the three signals (or pillars): metrics, logs and traces.
- Metric - a number that is measured again and again over time: requests served so far, bytes of memory in use, CPU seconds spent.
- Log - a line of text a program writes when something happens: "order 88121 failed: timeout". You have read plenty of them in the journal (2.30).
- Trace - the record of one request's path through several services, with how long each step took.
Three instruments, three questions
They are not three copies of the same data. Each answers a different question, and each gets expensive in a different way.
| metrics | logs | traces | |
|---|---|---|---|
| what it is | a number per time series, sampled | a line per event | a tree of timed steps (spans) per request |
| answers | how much / how often / how fast, overall | what exactly happened, to this one thing | where did the time go, across services |
| cost grows with | number of series (label combinations) | number of events (traffic) | number of requests kept |
| typical tool | Prometheus (this chapter) | a log store | a trace store |
| you alert on it | yes | rarely | almost never |
(An alert is an automatic message - a page to the on-call phone, a chat message - sent when a condition on the data becomes true. Chapter 0 (0.21) covered when to page; building the alerts is the next chapter.)
The rule of thumb that holds up in interviews:
- Metrics tell you THAT something is wrong. Error ratio went from 0.05% to 8%. p99 latency tripled. They are pre-counted numbers, cheap to keep for months and fast to query, so dashboards and alerts are built on them.
- Logs tell you WHAT happened. The exact exception, the order id, the SQL that timed out. Expensive per event, so you keep days or weeks, not years.
- Traces tell you WHERE. A checkout took 30 s: 29.9 s of it was waiting for a database connection inside the orders service, not the database itself.
A real incident uses all three in that order: an alert on a metric pages you, a dashboard (a web page of graphs drawn from metrics) narrows it to one endpoint on one instance, the logs for that endpoint show the exception, and one trace shows which hop ate the time.
Later (Ch 29): alerts get the next chapter (Alertmanager delivers them); then dashboards you build yourself in Grafana, a log store you query (Loki), trace waterfalls, and OpenTelemetry - the vendor-neutral standard libraries and Java agent that produce all three signals. This chapter is metrics, the foundation the others lean on.
What a metric actually is
The tool this chapter uses is Prometheus: a server that collects metrics from programs every few seconds and stores them in a time-series database (TSDB) on its own disk.
A time series is one metric, for one specific thing, over time. It has a name, a set of labels (key="value" pairs saying which thing), and a list of samples - each sample is a (timestamp, number) pair:
http_server_requests_seconds_count{instance="localhost:8080", job="orders", method="GET", status="200", uri="/api/orders"}
@ 20:00:00 84906
@ 20:00:15 85112
@ 20:00:30 85319
Reading it: the metric name is http_server_requests_seconds_count (how many HTTP requests the server has handled). The labels say which requests: the orders job, on the instance localhost:8080, GET requests to /api/orders that answered 200. Then three samples, 15 seconds apart: the count keeps growing (84906, 85112, 85319) because requests keep arriving.
Every distinct combination of label values is its own series, with its own memory in the TSDB. That single fact drives most of what goes wrong with Prometheus:
labels series
method (4) x status (8) x uri (30) 960 fine
... x instance (20) 19,200 fine
... x user_id (50,000) 960,000,000 the server falls over
The number of series you get from a label combination is its cardinality. Multiply the number of possible values of each label: 4 methods x 8 statuses x 30 uris = 960. Add a label with a value per user and it explodes. Anything unbounded (user ids, order ids, full URLs with ids in them, request ids, email addresses, timestamps) must never be a label. It belongs in a log line. Incident 27.6 is a service that does exactly this.
Structured logging
A log line is only as useful as your ability to filter it. Compare:
2026-09-22 20:02:24 ERROR Request processing failed for order 88121 after 30004ms
{"@timestamp":"2026-09-22T20:02:24.000Z","level":"ERROR","logger_name":"o.a.c.c.C.[.[.[/].[dispatcherServlet]","message":"Request processing failed: org.springframework.jdbc.CannotGetJdbcConnectionException: Failed to obtain JDBC Connection","traceId":"4db9ca9a4eb9cc2d4bb9c7744cb9c907","spanId":"4db9ca9a4eb9cc2d","http.method":"POST","http.uri":"/api/checkout"}
The first is free text: a sentence for humans. To find "all errors on /api/checkout" you need a regex (7.3) per question, and it breaks the day someone rewords the message.
The second is a structured log: one JSON object (7.11) per line, every fact in its own named field. It is the orders app on this box (Spring Boot with logging.structured.format.console=logstash, a setting that switches its log format to JSON). A log system can parse it once and let you filter level="ERROR", group by http.uri, and - the important bit - jump from the line to the trace of that request by its traceId field. You can already do this by hand: jq 'select(.level=="ERROR") | .["http.uri"]' on such lines.
What goes in a structured log line: timestamp, level, logger (which part of the code wrote it), message, the request's trace and span ids, and a few stable business fields. What does not: secrets, full request bodies, personal data you are not allowed to keep.
Traces, at concept level
A trace is the tree of work one request caused. Each node is a span: one timed step - a name ("POST /api/checkout", "SELECT orders"), a start time, a duration, a status and a few attributes. A span can have child spans: the checkout span contains a database span and a call-to-payments span.
Every trace has a trace id (a long random hex number) and every span a span id. Spans in different programs are linked because the caller passes the trace id to the callee in an HTTP header - that is context propagation. The standard header is called traceparent; it carries the trace id and the id of the calling span, so the callee can say "my span is a child of that one". If one service in the middle drops the header, the trace breaks into two unrelated halves - the most common "our traces are useless" cause.
When is each the right tool?
"Is checkout slower than yesterday?" metrics (p99 over time)
"Page me when users are failing." metrics (error-ratio alert)
"Why did order 88121 fail?" logs (search by order id)
"Which of our 6 services made this request slow?" traces (one request, every hop)
"How many 5xx did each endpoint serve last hour?" metrics (count per uri)
"What did the payment provider send back?" logs
"Is the new version slower for the same requests?" metrics first, then traces
A trace is the right tool when the question is about one request crossing several components and where its time went. A metric cannot tell you that (it has no per-request identity: it is already added up); a log line can only tell you about its own process.
What this box runs
Everything is a systemd unit (2.1) listening on a port:
prometheus.service :9090 the TSDB, the query engine, an HTTP API
prometheus-node-exporter.service :9100 machine metrics (CPU, memory, disks, network)
orders.service :8080 the Spring Boot app, metrics at /actuator/prometheus
oncall-lab-exporter.service :9199 (simulator) fresh data for every drill
Three more units run on :9093, :3100 and :5001; they are parts of the stack the next two chapters configure, and you can ignore them here.
An exporter is a small program whose only job is to publish metrics about something that cannot publish its own - the node exporter reads /proc and /sys (4.1) and publishes CPU, memory and disk numbers. The orders app needs no exporter: Actuator publishes its metrics itself (21.11).
Prometheus and the node exporter are the Ubuntu 26.04 packages (prometheus 2.53.5, prometheus-node-exporter 1.10.2), so the paths, the unit files and the default config are what sudo apt install prometheus gives you on your real VM. Upstream Prometheus is on 3.x; the differences that matter for this chapter are pointed out where they come up.
What you can now do
- Say which signal answers a question: metrics for that, logs for what, traces for where.
- Read a time series: metric name, labels, samples - and explain why an unbounded label (a user id) must never be one.
- Tell free-text logs from structured ones, and a trace from a span.