Why one request needs its own picture
The checkout graph says p99 is 30 seconds. Two very different problems produce that same graph: a slow database, or requests queueing for a database connection and never reaching the database at all. One means "call the DBA", the other "stop whatever is flooding the pool". Metrics cannot tell them apart. A trace of one slow request can, in one look.
What you need to know already: traces, spans, traceparent, OpenTelemetry, the Collector, and head vs tail sampling at concept level (27.1); histograms and buckets (27.2, 27.15); a trace id in each JSON log line and how to pull it out with LogQL (29.6); the HikariCP pool and its timeout (21.15, 21.18).
A trace, span by span
27.1 said a trace is a tree of spans that share a trace id. Here is everything one span carries:
trace_id 4db9ca9a4eb9cc2d4bb9c7744cb9c907 same for every span of the request
span_id 00f067aa0ba902b7 unique per span
parent_span_id (the caller's span id) builds the tree
name "POST /api/checkout", "HikariPool.getConnection", "SELECT orders"
start, duration
status OK / ERROR (+ message)
attributes http.method, http.route, http.status_code, db.system, db.statement, ...
events timestamped points inside the span (an exception, a retry)
- trace id - 32 hex characters (16 bytes), created once per request and copied into every span.
- span id - 16 hex characters, new for each span.
- parent span id - the span that called this one. The root span (the first one) has none. Following parent ids gives the tree.
- attributes - key/value details about the work: which HTTP method and route, which SQL statement. Named by a shared convention (OpenTelemetry's semantic conventions) so every tool understands
http.status_code. - events (span events) - moments inside a span, such as "exception thrown here".
Reading a waterfall
A trace is drawn as a waterfall: one bar per span, time running left to right, children indented under their parent. trace (simulator) draws one; on a real platform the same view is Grafana Tempo or Jaeger, the two common trace databases (27.1).
$ trace 4db9ca9a4eb9cc2d4bb9c7744cb9c907
(simulator) trace 4db9ca9a4eb9cc2d4bb9c7744cb9c907 30.01s 5 spans services: nginx, orders
nginx POST /api/checkout ▕██████████████████████████████████████████████▏ 30.01s 502 upstream
orders POST /api/checkout ▕██████████████████████████████████████████████▏ 30.01s status=500
orders CheckoutController.checkout ▕██████████████████████████████████████████████▏ 30.00s
orders CheckoutService.placeOrder ▕██████████████████████████████████████████████▏ 30.00s
orders HikariPool.getConnection ▕██████████████████████████████████████████████▏ 30.00s ERROR SQLTransientConnectionException
The header: trace id, total time, number of spans, services involved. Each row: service, span name, a bar for when it ran, its duration, and its status.
How to read it: find the widest span that has no child covering its time - the work that was really happening, not just a parent waiting for its children. Here that is HikariPool.getConnection, 30 s, with no database span below it: the request never reached the database; it waited for a free connection until the pool's timeout. A slow database would show a wide SELECT span instead. The chain of spans that decides the total time (here every row) is called the critical path; speeding up anything off it does not make the request faster.
Context propagation, in detail
27.1 showed the traceparent header: 00-<trace id>-<parent span id>-<flags>. Here is what each hop does with it:
client -> nginx -> orders -> payments
traceparent: 00-<trace>-<nginx span>-01
traceparent: 00-<trace>-<orders span>-01
- Read the incoming
traceparent(if there is none, start a new trace id). - Start its own span with the incoming span id as the parent.
- On every outgoing call, send a new
traceparent: same trace id, its own span id in the parent slot. The last field,01, means "sampled: record this".
A second header, tracestate, carries vendor-specific extras. This standard format is the W3C Trace Context.
Propagation breaks at anything that does not pass the header on: a proxy that strips unknown headers, work handed to a thread pool that loses the context (async code, 21.25), a message queue (the context must travel in the message's headers). The symptom: a trace that ends abruptly at one service, and a second, unrelated trace starting from nowhere downstream.
OpenTelemetry, the parts that matter to an SRE
OpenTelemetry (OTel, 27.1) is the standard way to produce traces. Its parts:
- API and SDK per language - the library an app uses to create spans. Auto-instrumentation for a JVM needs no code changes, only a flag:
java -javaagent:opentelemetry-javaagent.jar -Dotel.service.name=orders -jar orders.jar.-javaagent:loads an agent - a jar that hooks into the JVM (20.2) as it starts - which then creates spans for Spring MVC, JDBC, HTTP clients and Kafka on its own;-Dotel.service.namesets a system property with the service's name. - Resource attributes - facts about the process, set once and attached to every span:
service.name,service.version,deployment.environment,k8s.pod.name. - OTLP (OpenTelemetry Protocol) - the format apps use to send telemetry, over gRPC on port 4317 or HTTP on 4318, for traces, metrics and logs alike.
- The Collector (27.1) - a program that receives telemetry and forwards it. Its config is a Collector pipeline of receivers (what it accepts: OTLP, Jaeger, Prometheus), processors (what it does in between: batch, memory_limiter, tail_sampling, attributes) and exporters (where it sends: Tempo, Jaeger, a vendor). Apps send to a local Collector, never straight to the backend, so sampling and routing stay platform decisions.
A tail-sampling processor in the Collector's YAML config:
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow
type: latency
latency: { threshold_ms: 1000 }
- name: baseline
type: probabilistic
probabilistic: { sampling_percentage: 5 }
decision_wait: 10s- hold each trace's spans for 10 s, until it is (probably) complete, then decide.policies- keep the trace if any policy says yes:errors(any span with status ERROR),slow(took over 1000 ms),baseline(a random 5% of everything else).
Result: every error, every request over a second, and 5% of the rest.
Sampling, and what it does to your conclusions
Sampling means keeping only some traces, because keeping all of them costs too much. Head sampling decides at the first span (for example parentbased_traceidratio at 10%: keep 10% of trace ids, and every later service follows its parent's decision). Cheap and consistent, but blind: 90% of your slowest requests are dropped before anyone knew they would be slow. Tail sampling (the config above) keeps the interesting ones, but every span of a trace must reach the same Collector instance to be judged together.
Either way: never compute rates or percentiles from traces. The kept traces are a biased handful of examples. Numbers come from metrics; traces explain individual requests.
Exemplars: from a spike on a graph to one trace
An exemplar (27.1) is one trace id stored next to a histogram bucket: "here is a request that landed in this bucket". Prometheus keeps a few per series when started with --enable-feature=exemplar-storage, Grafana draws them as dots on the latency panel, and a click opens the trace. It is the fastest route from "p99 spiked at 20:02" to "this request, this span". The app must expose metrics in the OpenMetrics format (the standardised successor of Prometheus's text format); Micrometer does, with a tracing library on the classpath.
When a trace is the right tool
- The latency of one endpoint rose and it calls several things: which one?
- Errors appear in service C, but the cause might be upstream: follow the tree.
- Retries: three identical child spans where there should be one, the first ones in ERROR.
- Fan-out (one request calling many things): 200 database calls for one page - the N+1 query pattern, one query for a list plus one per item - shows up as a comb of tiny spans.
Not the right tool: "how many", "how often", "is it getting worse" - that is metrics.
What you can now do
- Read a waterfall: find the span where the time really went, and tell a slow dependency from a wait in front of it.
- Follow a
traceparenthop by hop and say where propagation broke. - Explain head vs tail sampling, and why percentiles never come from traces.