OnCallReady

Chapter 27 Observability I: Prometheus & PromQL

Metrics vs logs vs traces, the pull model, scrape configs and relabelling, and PromQL properly: rate, increase, aggregation, histograms, joins, recording rules.

In plain words

Imagine a car dashboard. The speedometer and the fuel gauge are numbers you glance at all the time: they tell you something is wrong ("fuel almost empty") but not why. The trip log in the glovebox records exactly what happened on each drive. And a GPS replay of one trip shows where you lost time: 20 minutes stuck at one junction.

Those are the three signals: metrics (the gauges), logs (the trip log) and traces (the replay of one request). This chapter is about the gauges. Prometheus visits every service every 15 seconds, reads its /metrics page, and stores each number as a time series identified by labels. PromQL is the language for asking questions like "error ratio per endpoint" or "p99 latency of checkout", and recording rules precompute the expensive answers.

Why it matters on call

Everything on-call is built on metrics: the alert that pages you, the dashboard you open first, the SLO your team is judged by. When the page says "orders error ratio above 5%", you need to write sum by (uri) (rate(...{status=~"5.."}[5m])) / sum by (uri) (rate(...[5m])) without looking it up, and know why [15s] returns nothing and why averaging p99s lies.

On a platform team you will also run Prometheus itself: add scrape jobs, debug a target that is up == 0 with context deadline exceeded, stop a service from exploding cardinality with user ids in labels, and reload config safely. PromQL is tested in SRE interviews directly ("write the p99 by endpoint", "why rate before sum"), and it is the base for the next two chapters: alerting, SLO burn rates, dashboards and incidents.

Lessons

  1. Metrics, logs and traces
  2. Prometheus: the pull model, exposition, the four metric types
  3. Service discovery and relabelling
  4. PromQL I: selectors, rate, increase, aggregation
  5. PromQL II: histograms, joins, absent, predict_linear
  6. Recording rules

22 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.

Questions people ask

What is the difference between monitoring and observability?

Monitoring is watching known failure modes with predefined dashboards and alerts: CPU, error ratio, disk space. Observability is the property of a system that lets you answer new questions from the outside, without shipping new code, using rich telemetry: metrics with useful labels, structured logs, and traces. In practice you need both. Alerts come from monitoring; debugging an incident you did not anticipate needs observability.

Why does Prometheus pull instead of receiving pushed metrics?

Pull makes target health free: a failed scrape records up 0, so you can tell "no errors" from "the app is dead". Prometheus owns the target list through service discovery, and apps do not need to know where monitoring lives. The cost is that Prometheus must reach every target, which is why it runs inside each cluster or network. Short-lived batch jobs push to the Pushgateway, the one sanctioned exception.

Is Prometheus a good place to store logs or per-request data?

No. Prometheus stores numeric time series, and every unique label combination is a new series with its own memory. Per-request identifiers like order ids, user ids or trace ids as labels explode cardinality and can take the server down. Put those in log lines or trace attributes. Metrics answer "how much, how often, how fast overall"; logs and traces answer questions about one event or request.

Is the Ubuntu Prometheus package the same as upstream?

It is Prometheus, but an older major version: the lab has the 26.04 package, 2.53.5, while upstream is on 3.x. Paths and units follow Debian conventions: /etc/prometheus/prometheus.yml, /etc/default/prometheus for arguments, data in /var/lib/prometheus/metrics2. The lifecycle API is off by default, so you reload with systemctl reload. In Kubernetes you would normally run upstream images via the Prometheus Operator instead.

What is an exporter and when do I need one?

An exporter is a small program that publishes metrics, in Prometheus's text format, for something that cannot publish its own. The node exporter reads /proc and /sys and serves CPU, memory, disk and network numbers on :9100. Apps you write, like the orders Spring Boot service, publish their own metrics (Actuator's /actuator/prometheus), so they need no exporter. Databases, proxies and the operating system usually need one.