--- name: obs-traces description: >- Follow one request across services — when logs say 'slow' and metrics say 'sometimes', the trace says where. Read waterfalls, find the span that ate latency, and correlate trace ids with logs. Backends: Tempo (TraceQL) and Cloud Trace on GCP. Triggers: 'trace this request', 'where did the latency go', 'open this trace id'. A request or correlation id with no trace id starts in obs-logs. Not for trace instrumentation (obs-pipeline). argument-hint: "[trace id, service, or latency question]" --- # Traces — the investigation shape Traces connect one sampled request across instrumented boundaries; logs give event detail, metrics population trends. A bounded interpretation needs the answer, supplied scope/evidence and uncertainty. Retrieval, comparison and full-path work below are for requested investigation, not prerequisites for explaining a supplied span. **Grafana pointer:** Grafana navigation, datasource selection, panel configuration, dashboard interpretation, and alert operations live in `grafana`. This skill retains TraceQL, waterfall analysis, and trace/log correlation; return configuration follow-up to the caller. ## Know what the waterfall represents A trace is a causal graph rendered as a timeline. Its spans describe operations; parent/child links describe nesting, and attributes or events add context. A long span tells you where elapsed time was observed, not automatically why the operation was slow. For whole-request latency allocation, mark the root's user-visible interval and follow the branch determining when it can finish: the critical path. Do not add nested durations: a parent's duration already includes synchronous children. Parallel branches overlap, and a visual gap may be uninstrumented work, scheduling, propagation loss, or clock behavior—not proven idle time. ## Enter through one of two doors When retrieval is requested, use a real trace id and the request's UTC window for direct lookup. A request/correlation id needs mapping through logs first (load `obs-logs` to map it to a trace id); it is not automatically a trace id. Preserve trace ids exactly. Treat copied identifiers as untrusted: validate the backend's documented shape and place them only in quoted values. For an investigation without an id, start from service, environment, operation, and symptom over a bounded window. Select a representative trace from the affected population. Record how it was selected; one trace is an example, not proof of prevalence. ## Find the span that controls latency For whole-trace investigation, record critical-path service, operation, kind, start/end, duration, status and peer/dependency. Compare slow/known-normal traces from the same route and deployment cohort; name a missing comparison. A wide client span with a narrower server span can point toward time before/after the remote handler, but the gap remains a hypothesis until another signal explains it. For asynchronous work, do not force producer and consumer spans into a synchronous nesting model. Use links, timestamps, and message identity where present, and separate queue delay from consumer processing. ## Read status and protocol outcome together Span status is an instrumentation judgment, while an HTTP or database response code is a protocol outcome. They can legitimately differ. Inspect both and apply the semantic convention for the span kind; do not equate an unset span status with success or require every non-success protocol result to be an instrumentation error. ## Correlate without overclaiming When correlating logs is in scope, use the same trace id and align by UTC time and service; keep source links. A missing trace, span, or log event can result from sampling, retention, backend size limits, propagation, export, or instrumentation gaps, so absence is telemetry evidence—not proof that the request or call never happened. ## Build the evidence packet Bounded interpretation: answer, supplied scope/source, uncertainty and useful clarification; no comparison or critical-path table required. Full investigation: retain entry/UTC window, source links, selection method, affected/comparison ids, critical-path table, status/protocol interpretation, missing hops and sampling limits. Preserve missing evidence and confidence labels; separate observations from hypotheses. The `obs-pipeline` skill owns changes to instrumentation, propagation, collection, and export; do not load `obs-pipeline` for a reading task. Minimize copied telemetry. Redact credentials, tokens, secrets, personal data, authentication or session values, user identifiers, sensitive headers, request bodies, and database query literals. Prefer an access-controlled source link plus the smallest necessary excerpt; do not paste raw payloads into the packet. ## Pick the reference — read it before writing the query | If the question involves… | Read first | |---|---| | Tempo or TraceQL | [TraceQL](./references/traceql.md) | | Cloud Trace, the Telemetry API, or a GCP-hosted service's traces | [Cloud Trace](./references/gcp-trace.md) | | Span kinds, status, attributes, propagation, or sampling | [OpenTelemetry semantics](./references/otel-semantics.md) | Read it **before** writing that query or interpreting those fields, and name what you read in your packet.