OnCallReady

Lesson 27.1 · Observability I: Prometheus & PromQL · 17 min read

Metrics, logs and traces

In plain words

Imagine a big school. The head teacher sees numbers on a board: how many kids are in class, how many are late, how many went to the nurse today. The numbers show that something is wrong: ten kids at the nurse is unusual. The nurse's diary says what happened to each kid: "Ana, 10:02, fell in the yard". And if you follow one kid's whole day, you see where the time went: stuck 40 minutes in the lunch queue.

Metrics are the board, logs are the diary, traces are following one request through every service. On oncall-lab, Prometheus holds the numbers from the orders app, the app writes JSON log lines, and each line carries a traceId so you can jump from a log to the trace of that request.

Why you need more than ssh and grep

So far you debugged one box by logging in and looking: top, journalctl, grep on a log. That stops working when the app runs as 20 copies on 20 machines and the question is "checkout got slow at 20:02 yesterday - why?". Nobody was logged in at 20:02, and you cannot grep 20 machines by hand.

What you need is data the system records about itself all the time, kept in one place, that you can ask questions later. That is this block of chapters.

What you need to know already: the golden signals - latency, traffic, errors, saturation (0.1); percentiles such as p99 (0.2); SLIs and SLOs (0.8); journalctl (2.30); grep and jq (7.1, 7.11); HTTP requests, headers and status codes (9.21); the orders Spring Boot app and its Actuator endpoints (21.1, 21.11).

The words you need first

Three instruments, three questions

They are not three copies of the same data. Each answers a different question, and each gets expensive in a different way.

metricslogstraces
what it isa number per time series, sampleda line per eventa tree of timed steps (spans) per request
answershow much / how often / how fast, overallwhat exactly happened, to this one thingwhere did the time go, across services
cost grows withnumber of series (label combinations)number of events (traffic)number of requests kept
typical toolPrometheus (this chapter)a log storea trace store
you alert on ityesrarelyalmost never

(An alert is an automatic message - a page to the on-call phone, a chat message - sent when a condition on the data becomes true. Chapter 0 (0.21) covered when to page; building the alerts is the next chapter.)

The rule of thumb that holds up in interviews:

A real incident uses all three in that order: an alert on a metric pages you, a dashboard (a web page of graphs drawn from metrics) narrows it to one endpoint on one instance, the logs for that endpoint show the exception, and one trace shows which hop ate the time.

Later (Ch 29): alerts get the next chapter (Alertmanager delivers them); then dashboards you build yourself in Grafana, a log store you query (Loki), trace waterfalls, and OpenTelemetry - the vendor-neutral standard libraries and Java agent that produce all three signals. This chapter is metrics, the foundation the others lean on.

What a metric actually is

The tool this chapter uses is Prometheus: a server that collects metrics from programs every few seconds and stores them in a time-series database (TSDB) on its own disk.

A time series is one metric, for one specific thing, over time. It has a name, a set of labels (key="value" pairs saying which thing), and a list of samples - each sample is a (timestamp, number) pair:

http_server_requests_seconds_count{instance="localhost:8080", job="orders", method="GET", status="200", uri="/api/orders"}
  @ 20:00:00  84906
  @ 20:00:15  85112
  @ 20:00:30  85319

Reading it: the metric name is http_server_requests_seconds_count (how many HTTP requests the server has handled). The labels say which requests: the orders job, on the instance localhost:8080, GET requests to /api/orders that answered 200. Then three samples, 15 seconds apart: the count keeps growing (84906, 85112, 85319) because requests keep arriving.

Every distinct combination of label values is its own series, with its own memory in the TSDB. That single fact drives most of what goes wrong with Prometheus:

labels                                  series
method (4) x status (8) x uri (30)      960        fine
... x instance (20)                     19,200     fine
... x user_id (50,000)                  960,000,000   the server falls over

The number of series you get from a label combination is its cardinality. Multiply the number of possible values of each label: 4 methods x 8 statuses x 30 uris = 960. Add a label with a value per user and it explodes. Anything unbounded (user ids, order ids, full URLs with ids in them, request ids, email addresses, timestamps) must never be a label. It belongs in a log line. Incident 27.6 is a service that does exactly this.

Structured logging

A log line is only as useful as your ability to filter it. Compare:

2026-09-22 20:02:24 ERROR Request processing failed for order 88121 after 30004ms
{"@timestamp":"2026-09-22T20:02:24.000Z","level":"ERROR","logger_name":"o.a.c.c.C.[.[.[/].[dispatcherServlet]","message":"Request processing failed: org.springframework.jdbc.CannotGetJdbcConnectionException: Failed to obtain JDBC Connection","traceId":"4db9ca9a4eb9cc2d4bb9c7744cb9c907","spanId":"4db9ca9a4eb9cc2d","http.method":"POST","http.uri":"/api/checkout"}

The first is free text: a sentence for humans. To find "all errors on /api/checkout" you need a regex (7.3) per question, and it breaks the day someone rewords the message.

The second is a structured log: one JSON object (7.11) per line, every fact in its own named field. It is the orders app on this box (Spring Boot with logging.structured.format.console=logstash, a setting that switches its log format to JSON). A log system can parse it once and let you filter level="ERROR", group by http.uri, and - the important bit - jump from the line to the trace of that request by its traceId field. You can already do this by hand: jq 'select(.level=="ERROR") | .["http.uri"]' on such lines.

What goes in a structured log line: timestamp, level, logger (which part of the code wrote it), message, the request's trace and span ids, and a few stable business fields. What does not: secrets, full request bodies, personal data you are not allowed to keep.

Traces, at concept level

A trace is the tree of work one request caused. Each node is a span: one timed step - a name ("POST /api/checkout", "SELECT orders"), a start time, a duration, a status and a few attributes. A span can have child spans: the checkout span contains a database span and a call-to-payments span.

Every trace has a trace id (a long random hex number) and every span a span id. Spans in different programs are linked because the caller passes the trace id to the callee in an HTTP header - that is context propagation. The standard header is called traceparent; it carries the trace id and the id of the calling span, so the callee can say "my span is a child of that one". If one service in the middle drops the header, the trace breaks into two unrelated halves - the most common "our traces are useless" cause.

When is each the right tool?

"Is checkout slower than yesterday?"                 metrics  (p99 over time)
"Page me when users are failing."                    metrics  (error-ratio alert)
"Why did order 88121 fail?"                          logs     (search by order id)
"Which of our 6 services made this request slow?"    traces   (one request, every hop)
"How many 5xx did each endpoint serve last hour?"    metrics  (count per uri)
"What did the payment provider send back?"           logs
"Is the new version slower for the same requests?"   metrics first, then traces

A trace is the right tool when the question is about one request crossing several components and where its time went. A metric cannot tell you that (it has no per-request identity: it is already added up); a log line can only tell you about its own process.

What this box runs

Everything is a systemd unit (2.1) listening on a port:

prometheus.service                 :9090  the TSDB, the query engine, an HTTP API
prometheus-node-exporter.service   :9100  machine metrics (CPU, memory, disks, network)
orders.service                     :8080  the Spring Boot app, metrics at /actuator/prometheus
oncall-lab-exporter.service         :9199  (simulator) fresh data for every drill

Three more units run on :9093, :3100 and :5001; they are parts of the stack the next two chapters configure, and you can ignore them here.

An exporter is a small program whose only job is to publish metrics about something that cannot publish its own - the node exporter reads /proc and /sys (4.1) and publishes CPU, memory and disk numbers. The orders app needs no exporter: Actuator publishes its metrics itself (21.11).

Prometheus and the node exporter are the Ubuntu 26.04 packages (prometheus 2.53.5, prometheus-node-exporter 1.10.2), so the paths, the unit files and the default config are what sudo apt install prometheus gives you on your real VM. Upstream Prometheus is on 3.x; the differences that matter for this chapter are pointed out where they come up.

What you can now do

Why it helps

The first minutes of every incident are about picking the right instrument. The alert on an error ratio tells you that checkout is failing; sum by (uri, instance) on a dashboard tells you it is one endpoint on one node; the logs show CannotGetJdbcConnectionException; one trace shows 29.9 of 30 seconds spent waiting for a connection pool, not in the database. Knowing which question each signal answers stops you grepping logs for something a metric shows in one query.

This lesson also plants the cardinality rule you will enforce in code reviews: no user or order ids as labels. And it explains why traces "break into halves" when one hop drops the traceparent header, a common complaint when a platform team rolls out tracing. "Metrics vs logs vs traces" is a standard opening interview question.

FAQ

When should I use a metric versus a log?

Use a metric when the question is about totals, rates or distributions over many events: requests per second, error ratio, p99 latency, queue depth. Metrics are cheap and fast, so you alert on them. Use a log when you need the details of one event: which order failed, the exception, what the upstream returned. Logs are expensive per event, so you do not alert on them directly in most setups.

What is structured logging and why does it matter?

Writing each log line as fields, usually JSON, instead of free text: timestamp, level, logger, message, trace and span ids, and a few stable business fields. The log system can then filter by level="ERROR", group by http.uri and jump to a trace by traceId, without a fragile regex per question. Spring Boot 3.4 and later supports it natively with logging.structured.format.console.

What is the traceparent header?

The W3C Trace Context header that carries a trace between services: 00-<trace-id>-<parent-span-id>-<flags>. The first service or proxy creates the 16-byte trace id; every hop keeps it and puts its own span id as the parent for the next call. Flags 01 mean sampled. If any hop drops the header, for example an uninstrumented HTTP client or queue, the trace splits into unrelated pieces.

What is the difference between a trace and a span?

A trace is everything one request caused, across all services. A span is one timed step inside it: "POST /api/checkout" in orders, "SELECT orders" in the database client, a call to payments. Spans have a name, start time, duration, status and attributes, and point to their parent span, so together they form a tree. The trace id is shared by every span of the request; each span has its own span id.

Why not just use logs for everything?

Logs cost per event: every request writes lines that must be shipped, stored and searched, so you keep them days or weeks, and counting "errors per second per endpoint" from them means scanning millions of lines. A metric is already a count, so the same question is one cheap query over months of data. Use logs for the details of one event, and metrics for totals, rates and alerts.

In an interview Mid

What is the difference between metrics, logs and traces, and when do you use each?

An incident uses them in that order: a metric alert pages you, a dashboard narrows it to one endpoint and instance, the logs show the error, one trace shows which hop was slow.

The trap that matters most: every distinct label combination is its own series, so an unbounded label (user id, order id, full URL) explodes cardinality - those values belong in logs.

Also asked: Your traces show one request split into two unrelated traces. What is going on? · How do you avoid high-cardinality problems when instrumenting a service? · What makes a log line useful during an incident?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.