Chapter 29 Observability III: Dashboards, Logs, Traces & Incidents
Build the golden-signals dashboard yourself, query logs with LogQL, read traces, and run an incident end to end: severity, command, mitigation, timeline and postmortem.
In plain words
Imagine a hospital emergency. The monitor above the bed shows heart rate and oxygen: it tells the doctors how bad things are. The patient's chart records what happened and when. A scan shows exactly where the problem is inside. And one senior doctor stands at the end of the bed, not operating, but deciding who does what, while a nurse tells the family what is going on.
This chapter is that room. Grafana dashboards are the monitor: SLO row, golden signals, then resources. Loki and LogQL are the chart: {unit="orders.service"} | json | level="ERROR". Traces with OpenTelemetry are the scan: which span ate the 30 seconds. And incident command is the senior doctor: severity, roles, updates, a timeline, and a postmortem afterwards.
Why it matters on call
This is where the observability block turns into on-call skill. The capstone incident, a checkout outage caused by a saturated connection pool, is solved only by using all of it: the dashboard shows which endpoint and instance, Loki shows HikariPool timeouts, one trace shows 30 seconds in getConnection with no database span, so the database is not the problem. Metrics alone would have looked identical for a slow database.
The incident process matters as much as the tools. Declaring early, mitigating before diagnosing, updating on a cadence and keeping a UTC timeline are what separate a calm 20-minute SEV2 from a chaotic two-hour one. And "walk me through an incident you handled" is the most common behavioural question in SRE interviews; this chapter gives you a real, well-documented one to tell.
Lessons
- Dashboards you build yourself
- Loki and LogQL
- Traces and OpenTelemetry
- Running an incident: severity, command, comms, timeline
15 hands-on labs (missions, incidents and drills) run in the terminal: Open this chapter in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.
Questions people ask
Should I import community Grafana dashboards?
Import them to learn how others query things, not to use them in incidents. A 60-panel import shows labels your services may not have, thresholds someone else chose, and panels you cannot interpret at 3am. A dashboard you built yourself answers questions you asked, with queries you can read, and you know what normal looks like on it. Keep your dashboards as JSON in git and provision them.
What is the difference between Loki and Elasticsearch?
Elasticsearch, a search engine often used for logs, builds a full-text index of every log line, so arbitrary content searches are fast but storage and memory costs are high. Loki indexes only the labels of each stream and stores the lines compressed, scanning content at query time. Loki is much cheaper to run and fits Prometheus-style workflows, but content-heavy searches over wide time ranges are slower, so you must narrow by labels first.
Is Grafana Tempo the same as Jaeger?
Both store and show traces. Jaeger is the older open-source project (under the CNCF, the foundation that also hosts Kubernetes and Prometheus) with its own UI, storing traces in databases such as Elasticsearch or Cassandra. Tempo is Grafana's trace store that keeps traces in cheap object storage (a blob store such as an Azure storage account), indexes little, and integrates with Grafana for exploring and for links from logs and metrics. Both accept OpenTelemetry data, so applications instrumented with OTel can switch backends.
What does an Incident Commander actually do?
The IC owns the incident, not the fix. They set severity, assign roles such as responders, communications and scribe, keep the timeline of decisions, decide on escalation and rollback, set the update cadence, and declare the end. They deliberately do not debug, because someone must steer. On a small team at night one person may wear several hats, but should say out loud which one they are wearing.
What is a blameless postmortem?
A written review after an incident that focuses on how the system and process allowed the failure, not on who made a mistake. It has a summary, impact, a UTC timeline, root cause and contributing factors, what went well and badly, and specific action items with owners and dates. Blamelessness matters because people only share the real details, which are needed to fix causes, when they are not punished for them.