OnCallReady

Observability III: Dashboards, Logs, Traces & Incidents: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 29 of the course.

Walk me through how you would debug a latency spike using metrics, logs and traces. Mid

Each signal answers one question, in this order:

  1. Metrics - that and where. The burn-rate page fires. The service dashboard: the SLO row says how bad, the golden-signal row (histogram_quantile(0.99, sum by (le, uri) (...)), errors by uri and instance) says which endpoint and instance, and the saturation panel says what ran out - here the connection pool at 20/20.
  2. Logs - what. In Loki, labels first, then filters: {unit="orders.service"} | json | level="ERROR" shows CannotGetJdbcConnectionException; line_format "{{.traceId}} {{.message}}" gives the trace ids. topk over count_over_time finds the endpoints logging most errors.
  3. Traces - where the time went. Open one trace: the widest span with no child covering it is HikariPool.getConnection, 30 s, with no database span under it - requests waited for a connection and never reached the database. A slow database would show a wide SELECT span instead.

Then mitigate first (stop the load, roll back), keep a UTC timeline, and fix the cause afterwards. Never compute percentiles from traces - they are sampled examples.

Also asked: What should a good service dashboard show? · How do you run a major incident from page to postmortem? · What is distributed tracing and how does it work?

What should a good service dashboard show? Mid

Laid out in the order you read it during an incident:

  1. Is the SLO OK? Burn rate (1h and 6h) and budget remaining as big stat numbers.
  2. The four golden signals, each broken down by the label that says where: traffic (sum by (uri) (rate(...))), errors as a ratio, latency p50/p99 from histogram_quantile, and saturation - whatever runs out first (for orders, connection pool in use / pool size, not CPU).
  3. Dependencies: the database pool, latency and errors of calls to other services.
  4. Resources below the symptoms: CPU, memory, GC, threads - the causes.

How to make it trustworthy: build it yourself (an imported 60-panel dashboard teaches nothing at 3 a.m.), units on every panel, readable legends ({{uri}}), $__rate_interval in every rate so a zoom never skips samples, a variable like $instance matched with =~, and the JSON in git via provisioning.

And the split: alerts for what needs a human now; the dashboard for the human already looking. Nobody watches a dashboard at 3 a.m.

Also asked: What is the difference between a dashboard and an alert, and what belongs where? · The numbers on a dashboard look wrong during an incident. What could be misleading you? · Why use $__rate_interval instead of a fixed window in Grafana?

Learn it: 29.1 Dashboards you build yourself

What is Loki and how is it different from other log systems? Mid

Loki is a log database that indexes only labels, not the text of the lines. Lines are grouped into streams (all lines with the same label set, e.g. {job="systemd-journal", unit="orders.service", host="oncall-lab"}), compressed into chunks, and only read when a query needs them. A shipper (Grafana Alloy) sends them from each machine.

Consequences:

LogQL looks like PromQL: {unit="orders.service"} |= "ERROR", then a parser (| json, | logfmt) that turns fields into labels for that query, label filters (| level="ERROR"), line_format. Wrap it in count_over_time(...[5m]) or rate and it returns numbers you can sum by - but prefer a real metric when there is one.

Also asked: Write a LogQL query to find the endpoints producing the most errors in the last hour, and explain its cost. · Why should a trace id never be a Loki label? · How do you get from a log line to the trace of that request?

Learn it: 29.6 Loki and LogQL

What is distributed tracing and how does it work? Mid

A trace is the tree of work one request caused across services. Each node is a span: a name, start, duration, status and attributes, with its own span id and the parent span id of its caller; every span shares the request's trace id.

Context propagation links the services: each hop reads the incoming traceparent header (00-<trace id>-<parent span id>-<flags>, the W3C Trace Context), starts its span as a child of it, and sends a new traceparent with its own span id on every outgoing call. A proxy that strips the header, async work that loses the context, or a queue without headers breaks the trace into unrelated halves.

OpenTelemetry produces the spans: an SDK or, for a JVM, the auto-instrumentation agent (-javaagent:opentelemetry-javaagent.jar); apps send OTLP to a local Collector (receivers, processors, exporters) that forwards to Tempo or Jaeger.

Reading one: in the waterfall, the widest span with no child covering its time is where the time went. Sampling (head or tail) keeps only some traces, so never compute rates or percentiles from them - metrics do numbers, traces explain one request.

Also asked: Compare head-based and tail-based sampling and when you would use each. · A trace ends abruptly at one service and another trace starts downstream. What is wrong? · What is an exemplar and why is it useful?

Learn it: 29.11 Traces and OpenTelemetry

What are the roles in incident response, and why separate them? Mid

Why separate: without it, three people debug the same log, nobody tells support, nobody decides to roll back, and a manager asks for updates every four minutes. Separation means someone always owns the decisions, the users' impact and the record. On a two-person night shift you wear several hats - say which one out loud.

The mechanics around it: declare early by user impact (SEV1, SEV2...), mitigate before you diagnose, update even when nothing changed, explicit handoffs, and end it when monitoring confirms impact is over - then a blameless postmortem with owned, dated action items.

Also asked: Walk me through an incident you handled. · What do you do in the first ten minutes after being paged? · What makes a good postmortem action item?

Learn it: 29.15 Running an incident: severity, command, comms, timeline

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.