Observability I: Prometheus & PromQL: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 27 of the course.
How would you measure a service's request rate, error ratio and p99 latency with Prometheus? Mid
From the service's request counter and latency histogram (the orders app exposes http_server_requests_seconds_*):
sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))
sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))
/ sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))
histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))
The rules behind them:
- Counters are graphed as rates, never raw;
ratefirst, thensum- rate must see each series alone to handle restarts. - The window holds at least four scrape intervals (
[5m]with 15 s scrapes), or the result is silently empty. - A ratio needs both sides aggregated to the same labels.
- For p99, sum the buckets (keeping
le) first, then take the quantile - never average per-instance percentiles. - Keep label cardinality bounded (route templates, not user ids), and record the expensive ones with recording rules.
Also asked: What are the three pillars of observability and when do you use each? · What is cardinality in Prometheus and why does it matter? · A Prometheus target is down. How do you troubleshoot it?
What is the difference between metrics, logs and traces, and when do you use each? Mid
- Metrics - numbers measured over time, per series (a name plus labels). Pre-aggregated, cheap to keep for months, fast to query: they tell you THAT something is wrong. Dashboards and alerts are built on them. Cost grows with the number of series.
- Logs - a line per event. They tell you WHAT happened to one thing: the exception, the order id, the SQL that timed out. Cost grows with traffic, so they are kept for days or weeks. Structured (JSON, one field per fact, with the trace id) beats free text.
- Traces - the tree of timed spans one request caused across services, linked by a trace id passed in the
traceparentheader. They tell you WHERE the time went.
An incident uses them in that order: a metric alert pages you, a dashboard narrows it to one endpoint and instance, the logs show the error, one trace shows which hop was slow.
The trap that matters most: every distinct label combination is its own series, so an unbounded label (user id, order id, full URL) explodes cardinality - those values belong in logs.
Also asked: Your traces show one request split into two unrelated traces. What is going on? · How do you avoid high-cardinality problems when instrumenting a service? · What makes a log line useful during an incident?
Learn it: 27.1 Metrics, logs and traces
Explain the four Prometheus metric types. Mid
- Counter - only goes up, back to 0 when the process restarts: requests served, errors, CPU seconds (
_total). Never graphed raw - you take itsrate. - Gauge - goes up and down: memory in use, queue depth, connections,
up. Graphed as it is. - Histogram - observations sorted into buckets counted in the app:
_bucket{le="0.25"}= requests that took at most 0.25 s (buckets are cumulative;+Inf=_count), plus_sumand_count(average = sum / count). Percentiles are computed in Prometheus, and histograms from several instances can be added. Put a bucket boundary at your SLO threshold. - Summary - percentiles computed inside the app (
quantile="0.99"). Cheap to read but cannot be combined across instances - prefer histograms for anything with more than one copy.
How the numbers arrive: Prometheus pulls - it scrapes each target's /metrics every scrape_interval and stores the samples, adding job and instance. A failed scrape writes up 0, so "is it up?" comes for free; the target's lastError says why.
Also asked: A Prometheus target is down. How do you troubleshoot it? · Why does Prometheus pull metrics instead of having applications push them? · How do you choose histogram buckets?
Learn it: 27.2 Prometheus: the pull model, exposition, the four metric types
What is service discovery in Prometheus, and what is relabelling for? Mid
Service discovery = Prometheus asks something else for its list of targets instead of a static list: kubernetes_sd_configs (pods), azure_sd_configs (VMs), dns_sd_configs, or file_sd_configs (a JSON/YAML file another tool writes; Prometheus watches it, no reload needed). New pods and VMs are monitored without anyone editing prometheus.yml, and removed ones disappear.
Each discovered target comes with hidden labels: __address__, __metrics_path__, __scheme__, and __meta_* (for Kubernetes: pod, namespace, every label and annotation). Relabelling turns those into the target you want:
relabel_configs- per target, before the scrape:keep/droptargets,replaceto build labels (regex anchored,$1capture groups,${1}_xwhen text follows),labelmap. Afterwards__*labels are dropped.metric_relabel_configs- per series, after the scrape: a stopgap guard, e.g. drop a metric or a series with an id in itspath. The real fix is in the app.
Debug by comparing discoveredLabels with labels in /api/v1/targets (and droppedTargets). A ServiceMonitor is the same relabelling underneath.
Also asked: A team deployed a new version and Prometheus memory is climbing fast. What do you do? · What is the difference between relabel_configs and metric_relabel_configs? · A ServiceMonitor seems to do nothing. How do you debug it?
Learn it: 27.4 Service discovery and relabelling
How would you calculate the request rate and error rate of a service in PromQL? Mid
sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))
sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))
/ sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))
Inside out: a selector (metric + label matchers), a range [5m], rate = per-second average increase, sum by (uri) = add up instances and statuses, one result per endpoint.
What makes it correct:
- rate, then sum.
ratedetects counter resets (restarts) per series and extrapolates to the window edges;rate(sum(x)[5m])is not even valid. - The window holds at least four scrape intervals - with 15 s scrapes,
[15s]returns nothing at all, silently. - Regex matchers are anchored:
status=~"5..", not"5"or"5xx"(which match nothing). increase(x[1h])=rate(x[1h]) * 3600for counts over a period;irateonly looks at the last two samples - not for alerts.rateon a gauge returns numbers that mean nothing.
Also asked: Explain how rate() works and its edge cases. · Write the PromQL for CPU utilisation percentage per node and explain each part. · When would you use increase() instead of rate()?
Learn it: 27.8 PromQL I: selectors, rate, increase, aggregation
How do you calculate p99 latency per endpoint in Prometheus? Mid
histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))
rate(..._bucket[5m])- each bucket is a counter.sum by (le, uri)- add up all instances, methods and statuses, keepingle(forget it and the result is empty) and the label you want one answer per.histogram_quantile(0.99, ...)- find the bucket the 99th percentile falls in and interpolate inside it. The answer is in seconds.
What to say about it:
- Precision is the bucket width - the true value can be anywhere in that bucket; in
+Infyou only learn "more than the last bucket". - Never average per-instance p99s: 217 ms and 1035 ms average to 626 ms, while the real service p99 (buckets summed first) is 912 ms. Sum the buckets, then take the quantile.
- For an SLO you often want the exact fraction under the threshold:
le="0.5"bucket rate /_countrate. - Average latency = rate of
_sum/ rate of_count- useful next to p99, never instead.
Also asked: How would you alert before a disk fills up? · Why is it wrong to average p99 latencies across instances? · How do you detect that a whole job has disappeared from Prometheus?
Learn it: 27.15 PromQL II: histograms, joins, absent, predict_linear
What are recording rules in Prometheus and why would you use them? Mid
A recording rule is a query Prometheus evaluates on its own every evaluation_interval and stores as a new time series:
groups:
- name: orders-recording
rules:
- record: job_uri:http_server_requests_errors:ratio_rate5m
expr: |
sum by (job, uri) (rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
/ sum by (job, uri) (rate(http_server_requests_seconds_count[5m]))
Why:
- Cost - an expensive query (a p99 over every bucket series) is computed once per interval instead of on every dashboard refresh for every viewer.
- Building blocks - alerts and SLOs reuse one consistently defined ratio.
Naming: level:metric:operations - aggregated to job_uri, from that metric, ratio_rate5m done to it. Colons mark recorded series.
Operating them: promtool check rules (and check config) before systemctl reload prometheus; /api/v1/rules shows each rule's health. And the gotcha: a recorded series has no past - it starts at the first evaluation after loading, so anything using [1h] of it is unreliable for the first hour.
Also asked: You deployed an alert based on a recorded series and it does not fire during testing. Why? · How should recording rules be named, and why? · How do you check that a rule file is valid and that its rules are running?
Learn it: 27.27 Recording rules
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.