Why build a dashboard yourself
You are paged at 3 a.m. and open a dashboard someone imported from the internet: 60 panels, labels your services do not have, red lines at thresholds nobody on your team chose. You stare at it and learn nothing. A dashboard you built yourself answers questions you asked, with queries you can read, and you know what "normal" looks like on it.
What you need to know already: the four golden signals (0.1) and percentiles (0.2); SLOs, error budgets and burn rate (0.8, 0.17); PromQL selectors, rate() and sum by (27.8); histogram_quantile and the le label (27.15); the Dashboards console you used in 27.26; recording rules (27.27) and the SLO recording rules job:slo_errors_per_request:ratio_rate1h (28.14); the HikariCP connection pool (21.15); JSON (7.11) and quoted heredocs (6.18).
The words you need first
- Grafana - the web app that draws graphs from Prometheus (and, later in this chapter, from a log database too). It does not store data itself: every graph is a query it sends to a data source (a database it reads from, here Prometheus).
- Dashboard - one Grafana page: a title, a time range, and a list of panels.
- Panel - one box on the dashboard: a chart, a big number or a table. Each panel has one or more queries - Grafana calls each one a panel target - here PromQL.
- Panel type - how the panel draws its data:
timeseries(a line graph over time),stat(one big number),gauge,table,heatmap. - Step - the time between two points on a graph. Grafana picks it from the time range and the panel's width: 3 hours drawn with ~56 points means one point every ~3 minutes.
Import community dashboards to learn how others write their queries, never to use during an incident.
What goes on a service dashboard
Top to bottom, in the order you read it during an incident:
- Is the SLO OK? Burn rate (1h and 6h) and budget remaining, as big numbers (a stat panel shows one big number).
- The four golden signals (0.1): traffic, errors, latency, saturation - one row, each broken down by the label that tells you where the problem is (
uri= which endpoint,instance= which server). - Dependencies: the database pool, latency and errors of calls to other services.
- Resources: CPU, memory, garbage collection, threads - the causes, placed below the symptoms.
For the orders service, one query per signal:
Traffic sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[$__rate_interval]))
Errors sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[$__rate_interval]))
/ sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[$__rate_interval]))
Latency histogram_quantile(0.5, sum by (le) (rate(http_server_requests_seconds_bucket{job="orders"}[$__rate_interval])))
histogram_quantile(0.95, ...) histogram_quantile(0.99, ...)
Saturation max by (instance) (hikaricp_connections_active{job="orders"}) / max by (instance) (hikaricp_connections_max{job="orders"})
Burn rate job:slo_errors_per_request:ratio_rate1h / 0.001
Reading them with what 27.8 and 27.15 taught:
- Traffic:
http_server_requests_seconds_countis the counter Micrometer (21.11) keeps of finished requests;rate(...)turns it into requests per second;sum by (uri)adds the instances together, one line per endpoint. - Errors: the same, with
status=~"5.."(a regex matcher: any status starting with 5), divided by all requests. The result is a ratio, 0 to 1. - Latency:
sum by (le)keeps the bucket boundaries, andhistogram_quantile(0.99, ...)estimates the 99th percentile from them. - Saturation: connections in use divided by the pool size, per instance. 1.0 means the pool is full.
- Burn rate: the recorded 1-hour error ratio divided by the allowed ratio (0.001 for a 99.9% SLO). 1 means "burning exactly on budget".
$__rate_interval in the brackets is a Grafana variable, explained two sections down. Where Prometheus would see [5m], Grafana fills in a window for you.
For a service like this one, saturation is the connection pool, not CPU: every checkout needs a database connection, and when all 20 are busy the 21st request waits (up to db.pool.timeout.ms, 21.18). This chapter's capstone incident is exactly that. Pick the saturation signal by asking "what runs out first?".
$__rate_interval
The problem: rate(x[1m]) looks at the last minute of samples at each point. On a 3-hour graph the step is ~3 minutes, so each point sees one minute out of three and the other two are never looked at - a short spike can fall between points and vanish. Zoom out to a week (step ~1 hour) and [1m] shows noise.
A Grafana variable is a $name placeholder that Grafana replaces before sending the query. $__rate_interval (two underscores) is a built-in one that picks the rate() window for you:
$__rate_interval = max($__interval + scrape interval, 4 x scrape interval)
$__interval- the graph's current step (~3m on a 3-hour graph).- scrape interval - how often Prometheus collects samples: 15 s on this box (27.2).
- 4 x scrape interval - the floor: a window must hold at least two samples for
rate()to work (27.8), and four gives room for a missed scrape.
With a 15 s scrape it is never less than 1 minute and it grows with the step, so every sample is used. Use it in every rate on a dashboard. In alerts and recording rules write a fixed window ([5m]): there the meaning must not change when someone zooms.
The dashboard JSON
A Grafana dashboard is a JSON document (7.11). You write one by hand once, so you know what is in it; after that you can edit in the Grafana UI and keep the JSON in git. Grafana's provisioning - loading dashboards from files at startup - then makes the git copy the source of truth, so nobody can break the on-call dashboard with an unreviewed click.
{
"title": "orders - golden signals",
"time": { "from": "now-3h", "to": "now" },
"panels": [
{
"type": "timeseries",
"title": "Latency p50 / p95 / p99",
"fieldConfig": { "defaults": { "unit": "s" } },
"targets": [
{ "expr": "histogram_quantile(0.99, sum by (le) (rate(http_server_requests_seconds_bucket{job=\"orders\"}[$__rate_interval])))", "legendFormat": "p99" }
]
}
]
}
Field by field:
title- the dashboard's name.time- the default range: from 3 hours ago (now-3h) to now.panels- the list of panels, drawn in order.type- the panel type:timeseriesfor graphs,statfor one big number, alsogauge,table,heatmap(the right way to draw a histogram over time).targets[].expr- the PromQL query. Because it sits inside a JSON string, its own double quotes are escaped:{job=\"orders\"}.legendFormat- the name each line gets in the legend.{{uri}}is replaced by the value of theurilabel, so one query drawing five endpoints gets five readable names instead of{uri="/api/orders", ...}.fieldConfig.defaults.unit- what the numbers mean:s(seconds),percentunit(a 0-1 ratio shown as %),bytes,reqps(requests per second). Units are what make a graph readable at a glance.
This box has no browser, so grafana-preview (simulator) stands in for Grafana: it reads a dashboard file, runs every panel's query against the real Prometheus and draws each result as a row of bars (a sparkline):
# orders.json is the dashboard you write in the next mission
grafana-preview ~/oncall-lab/labs/4d-observability/dashboards/orders.json
(simulator) orders - golden signals last 3h, step 3m, $__rate_interval=3m15s
Latency p50 / p95 / p99 [timeseries] unit=s
▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▃▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂▂ p99 last 912ms max 1.96s
Line by line: the header says the range (3h), the step Grafana would use (3m) and the value it substituted for $__rate_interval (3m + 15 s = 3m15s). Then one block per panel: title, type and unit, and one sparkline per legend entry with its last and highest value. A panel with no data or a broken query says so instead of drawing bars. --range 6h changes the range. The JSON itself is the real dashboard model: the file imports into a real Grafana unchanged.
Variables
A template variable is a drop-down at the top of a dashboard, filled by a query:
label_values(up{job="orders"}, instance)
label_values(<series>, <label>) (a Grafana function, not PromQL) lists every value the instance label has on the up series of the orders job. A variable $instance filled from it lets one dashboard serve every server: {job="orders", instance=~"$instance"}. Use the regex matcher =~: when you pick "All" or several values, Grafana turns them into a|b|c, which only a regex matches. Do not add a variable per label "just in case": each one runs a query every time the dashboard loads.
What belongs on a dashboard and what in an alert
- Alert (28.1): a condition that needs a human now or this week (SLO burn, a target gone). Few, reviewed, tested.
- Dashboard: everything that helps a human who is already looking - breakdowns, resources, dependencies, the causes. Many panels are fine.
- Never an alert for something nobody would act on; never a dashboard as the only way to notice an outage (nobody watches it at 3 a.m.).
A good dashboard makes the first five minutes of an incident boring: the SLO row says how bad, the golden-signal row says where (which uri, which instance), and the row below says which resource ran out.
What you can now do
- Lay out a service dashboard: SLO row, golden signals broken down by uri and instance, dependencies, then resources.
- Write a Grafana dashboard JSON by hand, with units, legends and
$__rate_intervalin every rate, and check it withgrafana-preview. - Decide whether a signal belongs in an alert or on a dashboard.