OnCallReady

Lesson 29.1 · Observability III: Dashboards, Logs, Traces & Incidents · 15 min read

Dashboards you build yourself

In plain words

Imagine the dashboard of a car you built yourself. You chose the gauges: speed first, then fuel, then engine temperature, and you know what "normal" looks like on each. Now imagine a borrowed car with 60 dials in a language you don't read. When a warning light flashes on the motorway, the first car tells you what is wrong in a second; the second one just makes you panic.

A Grafana dashboard is a JSON file of panels, each with a PromQL query. For orders you build it top to bottom: SLO burn and budget, then the four golden signals by uri and instance, then dependencies like the Hikari connection pool, then CPU and memory. Every rate uses $__rate_interval, so panels stay correct when you zoom.

Why build a dashboard yourself

You are paged at 3 a.m. and open a dashboard someone imported from the internet: 60 panels, labels your services do not have, red lines at thresholds nobody on your team chose. You stare at it and learn nothing. A dashboard you built yourself answers questions you asked, with queries you can read, and you know what "normal" looks like on it.

What you need to know already: the four golden signals (0.1) and percentiles (0.2); SLOs, error budgets and burn rate (0.8, 0.17); PromQL selectors, rate() and sum by (27.8); histogram_quantile and the le label (27.15); the Dashboards console you used in 27.26; recording rules (27.27) and the SLO recording rules job:slo_errors_per_request:ratio_rate1h (28.14); the HikariCP connection pool (21.15); JSON (7.11) and quoted heredocs (6.18).

The words you need first

Import community dashboards to learn how others write their queries, never to use during an incident.

What goes on a service dashboard

Top to bottom, in the order you read it during an incident:

  1. Is the SLO OK? Burn rate (1h and 6h) and budget remaining, as big numbers (a stat panel shows one big number).
  2. The four golden signals (0.1): traffic, errors, latency, saturation - one row, each broken down by the label that tells you where the problem is (uri = which endpoint, instance = which server).
  3. Dependencies: the database pool, latency and errors of calls to other services.
  4. Resources: CPU, memory, garbage collection, threads - the causes, placed below the symptoms.

For the orders service, one query per signal:

Traffic      sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[$__rate_interval]))
Errors       sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[$__rate_interval]))
               / sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[$__rate_interval]))
Latency      histogram_quantile(0.5,  sum by (le) (rate(http_server_requests_seconds_bucket{job="orders"}[$__rate_interval])))
             histogram_quantile(0.95, ...)   histogram_quantile(0.99, ...)
Saturation   max by (instance) (hikaricp_connections_active{job="orders"}) / max by (instance) (hikaricp_connections_max{job="orders"})
Burn rate    job:slo_errors_per_request:ratio_rate1h / 0.001

Reading them with what 27.8 and 27.15 taught:

$__rate_interval in the brackets is a Grafana variable, explained two sections down. Where Prometheus would see [5m], Grafana fills in a window for you.

For a service like this one, saturation is the connection pool, not CPU: every checkout needs a database connection, and when all 20 are busy the 21st request waits (up to db.pool.timeout.ms, 21.18). This chapter's capstone incident is exactly that. Pick the saturation signal by asking "what runs out first?".

$__rate_interval

The problem: rate(x[1m]) looks at the last minute of samples at each point. On a 3-hour graph the step is ~3 minutes, so each point sees one minute out of three and the other two are never looked at - a short spike can fall between points and vanish. Zoom out to a week (step ~1 hour) and [1m] shows noise.

A Grafana variable is a $name placeholder that Grafana replaces before sending the query. $__rate_interval (two underscores) is a built-in one that picks the rate() window for you:

$__rate_interval = max($__interval + scrape interval, 4 x scrape interval)

With a 15 s scrape it is never less than 1 minute and it grows with the step, so every sample is used. Use it in every rate on a dashboard. In alerts and recording rules write a fixed window ([5m]): there the meaning must not change when someone zooms.

The dashboard JSON

A Grafana dashboard is a JSON document (7.11). You write one by hand once, so you know what is in it; after that you can edit in the Grafana UI and keep the JSON in git. Grafana's provisioning - loading dashboards from files at startup - then makes the git copy the source of truth, so nobody can break the on-call dashboard with an unreviewed click.

{
  "title": "orders - golden signals",
  "time": { "from": "now-3h", "to": "now" },
  "panels": [
    {
      "type": "timeseries",
      "title": "Latency p50 / p95 / p99",
      "fieldConfig": { "defaults": { "unit": "s" } },
      "targets": [
        { "expr": "histogram_quantile(0.99, sum by (le) (rate(http_server_requests_seconds_bucket{job=\"orders\"}[$__rate_interval])))", "legendFormat": "p99" }
      ]
    }
  ]
}

Field by field:

This box has no browser, so grafana-preview (simulator) stands in for Grafana: it reads a dashboard file, runs every panel's query against the real Prometheus and draws each result as a row of bars (a sparkline):

# orders.json is the dashboard you write in the next mission
grafana-preview ~/oncall-lab/labs/4d-observability/dashboards/orders.json
(simulator) orders - golden signals   last 3h, step 3m, $__rate_interval=3m15s

Latency p50 / p95 / p99  [timeseries]  unit=s
  ▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂  p99  last 912ms  max 1.96s

Line by line: the header says the range (3h), the step Grafana would use (3m) and the value it substituted for $__rate_interval (3m + 15 s = 3m15s). Then one block per panel: title, type and unit, and one sparkline per legend entry with its last and highest value. A panel with no data or a broken query says so instead of drawing bars. --range 6h changes the range. The JSON itself is the real dashboard model: the file imports into a real Grafana unchanged.

Variables

A template variable is a drop-down at the top of a dashboard, filled by a query:

label_values(up{job="orders"}, instance)

label_values(<series>, <label>) (a Grafana function, not PromQL) lists every value the instance label has on the up series of the orders job. A variable $instance filled from it lets one dashboard serve every server: {job="orders", instance=~"$instance"}. Use the regex matcher =~: when you pick "All" or several values, Grafana turns them into a|b|c, which only a regex matches. Do not add a variable per label "just in case": each one runs a query every time the dashboard loads.

What belongs on a dashboard and what in an alert

A good dashboard makes the first five minutes of an incident boring: the SLO row says how bad, the golden-signal row says where (which uri, which instance), and the row below says which resource ran out.

What you can now do

Why it helps

The dashboard you build is the first thing you open when paged, and it decides whether the first five minutes are calm or chaotic. A good one answers "how bad" (SLO row), "where" (errors and latency by uri and instance) and "what ran out" (the connection pool, not CPU, for a service like orders). The capstone incident is found in minutes precisely because saturation is the pool.

The details matter in real work. A panel with rate(x[1m]) at a 3-minute step silently drops two thirds of the data; $__rate_interval fixes it. Units make a graph readable at a glance. Keeping dashboard JSON in git and provisioning it means a colleague cannot silently break the on-call dashboard. And "what would you put on a service dashboard?" is a common interview question with an expected structure.

FAQ

What is $__rate_interval and why use it?

A Grafana variable that picks a safe rate window: the larger of the panel step plus one scrape interval, and four scrape intervals. A fixed [1m] on a 3-hour graph with a 3-minute step skips samples between points, and on a week-long view shows noise. $__rate_interval grows with the step so every sample contributes and never drops below the four-scrape minimum. Use it in dashboards, and fixed windows in alerts and recording rules.

What are the four golden signals?

From the Google SRE book: traffic (demand, such as requests per second), errors (rate or ratio of failed requests), latency (how long requests take, separately for successes and failures, as percentiles), and saturation (how full the most constrained resource is). The RED method is a service-focused subset: rate, errors, duration. USE is for resources: utilisation, saturation, errors.

How do I choose the saturation signal for a service?

Ask what runs out first under load. For a database-backed API it is often the connection pool: active connections against the maximum and threads waiting. For a queue consumer, backlog and consumer lag. For a CPU-bound service, CPU against its limit and throttling. For a JVM, heap after GC. CPU is often the wrong answer, since services can fail with plenty of CPU left.

Why use =~ with dashboard variables?

When a variable allows multiple values or "All", Grafana expands it into a regex alternation such as (a|b|c) or .*. Only the regex matcher =~ handles that; with = the query looks for the literal string and returns nothing. So write {job="orders", instance=~"$instance"}. Keep variables few, since each one runs a query every time the dashboard loads.

How do I keep dashboards under version control?

Store the dashboard JSON in git and load it with Grafana provisioning: a provider config points at a directory, and Grafana loads the files at startup and on change. In Kubernetes, the kube-prometheus-stack Helm chart (Prometheus, Alertmanager and Grafana in one install) runs a small helper container next to Grafana that loads dashboards from ConfigMaps with a specific label. Edit in the UI if you like, export the JSON, and commit it. Grafonnet (a library that generates dashboard JSON from code) or the Grafana Terraform provider write dashboards as code.

In an interview Mid

What should a good service dashboard show?

Laid out in the order you read it during an incident:

  1. Is the SLO OK? Burn rate (1h and 6h) and budget remaining as big stat numbers.
  2. The four golden signals, each broken down by the label that says where: traffic (sum by (uri) (rate(...))), errors as a ratio, latency p50/p99 from histogram_quantile, and saturation - whatever runs out first (for orders, connection pool in use / pool size, not CPU).
  3. Dependencies: the database pool, latency and errors of calls to other services.
  4. Resources below the symptoms: CPU, memory, GC, threads - the causes.

How to make it trustworthy: build it yourself (an imported 60-panel dashboard teaches nothing at 3 a.m.), units on every panel, readable legends ({{uri}}), $__rate_interval in every rate so a zoom never skips samples, a variable like $instance matched with =~, and the JSON in git via provisioning.

And the split: alerts for what needs a human now; the dashboard for the human already looking. Nobody watches a dashboard at 3 a.m.

Also asked: What is the difference between a dashboard and an alert, and what belongs where? · The numbers on a dashboard look wrong during an incident. What could be misleading you? · Why use $__rate_interval instead of a fixed window in Grafana?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.