OnCallReady

Lesson 29.11 · Observability III: Dashboards, Logs, Traces & Incidents · 13 min read

Traces and OpenTelemetry

In plain words

Imagine tracking a parcel. At every depot, the parcel's barcode is scanned: "arrived 10:00, left 10:05". The barcode stays the same all the way, and each scan says which depot handed it over. Afterwards you can see the whole journey on one timeline and spot that it sat for two days in one warehouse.

A trace is that timeline for one request. The trace id is the barcode; each span is a depot scan with a start, a duration and a parent span. The traceparent header carries the barcode between services. OpenTelemetry does the scanning, often with a Java agent and no code changes, and the Collector decides which parcels' histories to keep. On orders, the waterfall shows 30 seconds inside HikariPool.getConnection and no database span below it.

Why one request needs its own picture

The checkout graph says p99 is 30 seconds. Two very different problems produce that same graph: a slow database, or requests queueing for a database connection and never reaching the database at all. One means "call the DBA", the other "stop whatever is flooding the pool". Metrics cannot tell them apart. A trace of one slow request can, in one look.

What you need to know already: traces, spans, traceparent, OpenTelemetry, the Collector, and head vs tail sampling at concept level (27.1); histograms and buckets (27.2, 27.15); a trace id in each JSON log line and how to pull it out with LogQL (29.6); the HikariCP pool and its timeout (21.15, 21.18).

A trace, span by span

27.1 said a trace is a tree of spans that share a trace id. Here is everything one span carries:

trace_id        4db9ca9a4eb9cc2d4bb9c7744cb9c907     same for every span of the request
span_id         00f067aa0ba902b7                     unique per span
parent_span_id  (the caller's span id)               builds the tree
name            "POST /api/checkout", "HikariPool.getConnection", "SELECT orders"
start, duration
status          OK / ERROR (+ message)
attributes      http.method, http.route, http.status_code, db.system, db.statement, ...
events          timestamped points inside the span (an exception, a retry)

Reading a waterfall

A trace is drawn as a waterfall: one bar per span, time running left to right, children indented under their parent. trace (simulator) draws one; on a real platform the same view is Grafana Tempo or Jaeger, the two common trace databases (27.1).

$ trace 4db9ca9a4eb9cc2d4bb9c7744cb9c907
(simulator) trace 4db9ca9a4eb9cc2d4bb9c7744cb9c907   30.01s   5 spans   services: nginx, orders

nginx POST /api/checkout                          ▕██████████████████████████████████████████████▏  30.01s  502 upstream
  orders POST /api/checkout                       ▕██████████████████████████████████████████████▏  30.01s  status=500
    orders CheckoutController.checkout            ▕██████████████████████████████████████████████▏  30.00s
      orders CheckoutService.placeOrder           ▕██████████████████████████████████████████████▏  30.00s
        orders HikariPool.getConnection           ▕██████████████████████████████████████████████▏  30.00s  ERROR SQLTransientConnectionException

The header: trace id, total time, number of spans, services involved. Each row: service, span name, a bar for when it ran, its duration, and its status.

How to read it: find the widest span that has no child covering its time - the work that was really happening, not just a parent waiting for its children. Here that is HikariPool.getConnection, 30 s, with no database span below it: the request never reached the database; it waited for a free connection until the pool's timeout. A slow database would show a wide SELECT span instead. The chain of spans that decides the total time (here every row) is called the critical path; speeding up anything off it does not make the request faster.

Context propagation, in detail

27.1 showed the traceparent header: 00-<trace id>-<parent span id>-<flags>. Here is what each hop does with it:

client -> nginx -> orders -> payments
          traceparent: 00-<trace>-<nginx span>-01
                      traceparent: 00-<trace>-<orders span>-01
  1. Read the incoming traceparent (if there is none, start a new trace id).
  2. Start its own span with the incoming span id as the parent.
  3. On every outgoing call, send a new traceparent: same trace id, its own span id in the parent slot. The last field, 01, means "sampled: record this".

A second header, tracestate, carries vendor-specific extras. This standard format is the W3C Trace Context.

Propagation breaks at anything that does not pass the header on: a proxy that strips unknown headers, work handed to a thread pool that loses the context (async code, 21.25), a message queue (the context must travel in the message's headers). The symptom: a trace that ends abruptly at one service, and a second, unrelated trace starting from nowhere downstream.

OpenTelemetry, the parts that matter to an SRE

OpenTelemetry (OTel, 27.1) is the standard way to produce traces. Its parts:

A tail-sampling processor in the Collector's YAML config:

processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow
        type: latency
        latency: { threshold_ms: 1000 }
      - name: baseline
        type: probabilistic
        probabilistic: { sampling_percentage: 5 }

Result: every error, every request over a second, and 5% of the rest.

Sampling, and what it does to your conclusions

Sampling means keeping only some traces, because keeping all of them costs too much. Head sampling decides at the first span (for example parentbased_traceidratio at 10%: keep 10% of trace ids, and every later service follows its parent's decision). Cheap and consistent, but blind: 90% of your slowest requests are dropped before anyone knew they would be slow. Tail sampling (the config above) keeps the interesting ones, but every span of a trace must reach the same Collector instance to be judged together.

Either way: never compute rates or percentiles from traces. The kept traces are a biased handful of examples. Numbers come from metrics; traces explain individual requests.

Exemplars: from a spike on a graph to one trace

An exemplar (27.1) is one trace id stored next to a histogram bucket: "here is a request that landed in this bucket". Prometheus keeps a few per series when started with --enable-feature=exemplar-storage, Grafana draws them as dots on the latency panel, and a click opens the trace. It is the fastest route from "p99 spiked at 20:02" to "this request, this span". The app must expose metrics in the OpenMetrics format (the standardised successor of Prometheus's text format); Micrometer does, with a tracing library on the classpath.

When a trace is the right tool

Not the right tool: "how many", "how often", "is it getting worse" - that is metrics.

What you can now do

Why it helps

Traces answer the question metrics cannot: which hop ate the time. In the capstone, metrics look the same for "the database is slow" and "the pool is exhausted"; the trace tells them apart instantly, which decides whether you page the DBA or stop a load test. In a microservice estate, traces also expose retries (three identical child spans) and N+1 queries (a comb of tiny spans).

On a platform team you will roll out OpenTelemetry: agents for the JVM, resource attributes like service.name, a Collector per node or cluster, and tail sampling policies. You will debug broken propagation, the most common "our traces are useless" complaint. And knowing that you never compute rates or percentiles from sampled traces keeps you from drawing wrong conclusions in an incident review. Tracing and OTel come up in most modern SRE interviews.

FAQ

What is the difference between a trace and a span?

A span is one unit of work: a name like POST /api/checkout or SELECT orders, a start time, a duration, a status, attributes and events, with its own span id and a parent span id. A trace is the whole tree of spans that share one trace id, representing everything one request caused across services. The root span has no parent; the tree is built from parent ids.

Why send telemetry to a Collector instead of straight to the backend?

The Collector decouples applications from the backend. Apps send OTLP to a local Collector; it batches, limits memory, adds or removes attributes, applies tail sampling, and exports to Tempo, Jaeger or a vendor. Sampling and routing become platform decisions you can change without redeploying apps, credentials for the backend live in one place, and switching vendors does not touch application code.

Can I compute latency percentiles from traces?

Not reliably. Traces are sampled, and head sampling drops most requests at random, while tail sampling deliberately keeps errors and slow requests, which biases any statistics towards bad cases. Use metrics, histograms in Prometheus, for rates and percentiles, and use traces to explain individual requests. Some backends generate span metrics from all spans before sampling in the Collector, which is a valid exception.

How do I read a trace waterfall?

Time runs left to right and children are indented under their parents. Look for the widest span that has no child covering its time: that is where the time was actually spent, rather than just passed through. If a request span is 30 seconds and its child getConnection is also 30 seconds with nothing below, the wait is for a connection. Also watch for gaps between children, repeated identical spans and many tiny ones.

What does tail sampling need to work correctly?

All spans of a trace must reach the same Collector instance, because the decision is made when the trace is complete. With several Collectors, a load-balancing exporter in front routes spans by trace id to a consistent instance. The Collector must also buffer traces for decision_wait, which costs memory proportional to traffic. Late spans arriving after the decision may be dropped or handled separately.

In an interview Mid

What is distributed tracing and how does it work?

A trace is the tree of work one request caused across services. Each node is a span: a name, start, duration, status and attributes, with its own span id and the parent span id of its caller; every span shares the request's trace id.

Context propagation links the services: each hop reads the incoming traceparent header (00-<trace id>-<parent span id>-<flags>, the W3C Trace Context), starts its span as a child of it, and sends a new traceparent with its own span id on every outgoing call. A proxy that strips the header, async work that loses the context, or a queue without headers breaks the trace into unrelated halves.

OpenTelemetry produces the spans: an SDK or, for a JVM, the auto-instrumentation agent (-javaagent:opentelemetry-javaagent.jar); apps send OTLP to a local Collector (receivers, processors, exporters) that forwards to Tempo or Jaeger.

Reading one: in the waterfall, the widest span with no child covering its time is where the time went. Sampling (head or tail) keeps only some traces, so never compute rates or percentiles from them - metrics do numbers, traces explain one request.

Also asked: Compare head-based and tail-based sampling and when you would use each. · A trace ends abruptly at one service and another trace starts downstream. What is wrong? · What is an exemplar and why is it useful?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.