OnCallReady

Observability I: Prometheus & PromQL: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 27 of the course.

How would you measure a service's request rate, error ratio and p99 latency with Prometheus? Mid

From the service's request counter and latency histogram (the orders app exposes http_server_requests_seconds_*):

sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))
  / sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))

The rules behind them:

Also asked: What are the three pillars of observability and when do you use each? · What is cardinality in Prometheus and why does it matter? · A Prometheus target is down. How do you troubleshoot it?

What is the difference between metrics, logs and traces, and when do you use each? Mid

An incident uses them in that order: a metric alert pages you, a dashboard narrows it to one endpoint and instance, the logs show the error, one trace shows which hop was slow.

The trap that matters most: every distinct label combination is its own series, so an unbounded label (user id, order id, full URL) explodes cardinality - those values belong in logs.

Also asked: Your traces show one request split into two unrelated traces. What is going on? · How do you avoid high-cardinality problems when instrumenting a service? · What makes a log line useful during an incident?

Learn it: 27.1 Metrics, logs and traces

Explain the four Prometheus metric types. Mid

How the numbers arrive: Prometheus pulls - it scrapes each target's /metrics every scrape_interval and stores the samples, adding job and instance. A failed scrape writes up 0, so "is it up?" comes for free; the target's lastError says why.

Also asked: A Prometheus target is down. How do you troubleshoot it? · Why does Prometheus pull metrics instead of having applications push them? · How do you choose histogram buckets?

Learn it: 27.2 Prometheus: the pull model, exposition, the four metric types

What is service discovery in Prometheus, and what is relabelling for? Mid

Service discovery = Prometheus asks something else for its list of targets instead of a static list: kubernetes_sd_configs (pods), azure_sd_configs (VMs), dns_sd_configs, or file_sd_configs (a JSON/YAML file another tool writes; Prometheus watches it, no reload needed). New pods and VMs are monitored without anyone editing prometheus.yml, and removed ones disappear.

Each discovered target comes with hidden labels: __address__, __metrics_path__, __scheme__, and __meta_* (for Kubernetes: pod, namespace, every label and annotation). Relabelling turns those into the target you want:

Debug by comparing discoveredLabels with labels in /api/v1/targets (and droppedTargets). A ServiceMonitor is the same relabelling underneath.

Also asked: A team deployed a new version and Prometheus memory is climbing fast. What do you do? · What is the difference between relabel_configs and metric_relabel_configs? · A ServiceMonitor seems to do nothing. How do you debug it?

Learn it: 27.4 Service discovery and relabelling

How would you calculate the request rate and error rate of a service in PromQL? Mid

sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

sum by (uri) (rate(http_server_requests_seconds_count{job="orders",status=~"5.."}[5m]))
  / sum by (uri) (rate(http_server_requests_seconds_count{job="orders"}[5m]))

Inside out: a selector (metric + label matchers), a range [5m], rate = per-second average increase, sum by (uri) = add up instances and statuses, one result per endpoint.

What makes it correct:

Also asked: Explain how rate() works and its edge cases. · Write the PromQL for CPU utilisation percentage per node and explain each part. · When would you use increase() instead of rate()?

Learn it: 27.8 PromQL I: selectors, rate, increase, aggregation

How do you calculate p99 latency per endpoint in Prometheus? Mid

histogram_quantile(0.99, sum by (le, uri) (rate(http_server_requests_seconds_bucket{job="orders"}[5m])))
  1. rate(..._bucket[5m]) - each bucket is a counter.
  2. sum by (le, uri) - add up all instances, methods and statuses, keeping le (forget it and the result is empty) and the label you want one answer per.
  3. histogram_quantile(0.99, ...) - find the bucket the 99th percentile falls in and interpolate inside it. The answer is in seconds.

What to say about it:

Also asked: How would you alert before a disk fills up? · Why is it wrong to average p99 latencies across instances? · How do you detect that a whole job has disappeared from Prometheus?

Learn it: 27.15 PromQL II: histograms, joins, absent, predict_linear

What are recording rules in Prometheus and why would you use them? Mid

A recording rule is a query Prometheus evaluates on its own every evaluation_interval and stores as a new time series:

groups:
  - name: orders-recording
    rules:
      - record: job_uri:http_server_requests_errors:ratio_rate5m
        expr: |
          sum by (job, uri) (rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
            / sum by (job, uri) (rate(http_server_requests_seconds_count[5m]))

Why:

  1. Cost - an expensive query (a p99 over every bucket series) is computed once per interval instead of on every dashboard refresh for every viewer.
  2. Building blocks - alerts and SLOs reuse one consistently defined ratio.

Naming: level:metric:operations - aggregated to job_uri, from that metric, ratio_rate5m done to it. Colons mark recorded series.

Operating them: promtool check rules (and check config) before systemctl reload prometheus; /api/v1/rules shows each rule's health. And the gotcha: a recorded series has no past - it starts at the first evaluation after loading, so anything using [1h] of it is unreliable for the first hour.

Also asked: You deployed an alert based on a recorded series and it does not fire during testing. Why? · How should recording rules be named, and why? · How do you check that a rule file is valid and that its rules are running?

Learn it: 27.27 Recording rules

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.