--- name: obs-logs description: >- The answer is in the logs — find error spikes, read them over time, correlate one request across services, compare before/after a deploy. Backends: Splunk (SPL), Loki (LogQL), and Cloud Logging on GCP — the reference teaches the dialect. Triggers: 'search the logs', 'why are there 500s', 'write a log query', 'follow this correlation id'. Ownership map only: obs-metrics owns metrics, obs-dashboards owns dashboard design, grafana owns Grafana operations, and obs-alerting owns alert design. Deciding what a live page means or what to do next belongs to incident-investigation, which routes query work here. argument-hint: "[service, symptom, or log question]" --- # Logs — the investigation shape Match the task: interpret supplied logs, write a query, or investigate. An explanation needs no new query or baseline; apply methods only to the requested question. **Grafana pointer:** dashboard/Explore access, datasource selection, panel configuration, and alert operations live in `grafana`. This skill retains LogQL/SPL and log interpretation. Return a need for Grafana configuration to the caller; a log query does not automatically start another workflow. ## Start narrow For queries, bind environment, service, source and window before symptom filters. `error` can miss failed structured access events. Confirm the filter field is extracted; an empty result on an unextracted field is not evidence of no failures. Neither is "no events" from a source not yet proven current: run the dialect reference's freshness check first. Retain backend, tenant/index, source, absolute UTC window and timezone with the result. Widen one boundary at a time and say why. ## Read it over time For spike investigation, build a complete fixed-width timeline. Count failures inside each bucket; filtering failures first can erase quiet zero buckets and bias the baseline toward error periods. Mark the first anomalous bucket, the last known-normal bucket, and any deploy or configuration event. Keep the current bucket out of a trailing baseline so the spike cannot raise the threshold used to judge itself. ## Find the top offenders When ranking offenders, group by stable service, route, status, error type or host. Keep high-cardinality IDs for one-request work. Show count and traffic share; traffic growth alone is not a worsening error rate. ## Correlate one request across services When following one request, use its request/correlation/trace id within a tight window. Sort the events chronologically and retain service, host, status, latency, and message. If a hop emits no common identifier, record that as a telemetry gap and recommend to the caller that `software-engineer` add it. Treat identifiers copied from tickets or logs as untrusted data. Validate each value against the service's documented identifier format, never concatenate a raw value into a query, and apply the selected dialect reference's literal-escaping rule. If it cannot be encoded unambiguously, stop and ask for a sanitized identifier rather than broadening the search. ## Compare before vs after a deploy For a deploy comparison, compare **rates**, not raw counts: differing traffic can explain counts. Use equal-duration windows and the same query scope. Take the exact deploy time from the platform record — Apps Manager → **Events** (`cf events `) on PCF, or when traffic moved to the new revision on Cloud Run — not from a dashboard annotation. ## Build the evidence packet Bounded interpretation: answer, supplied source/target/window, limits and a useful next check if needed; missing metadata stays unknown. Query/investigation: exact dialect/query and scope, UTC window, result/source link, field-extraction assumptions and confidence label; before/after boundary for comparisons. Separate observations from interpretations. Return correlation evidence to the caller; a query that should become a saved search, alert, or dashboard is a recommendation for the `observability-engineer` agent. Minimize copied telemetry. Redact credentials, tokens, secrets, personal data, authentication or session values, user identifiers, sensitive headers, request bodies, and database query literals. Prefer an access-controlled source link plus the smallest necessary excerpt; do not paste raw payloads into the packet. ## Pick your dialect — read the reference before writing the query | If the question involves… | Read first | |---|---| | Splunk or SPL | [SPL](./references/spl.md) | | Loki or LogQL | [LogQL](./references/logql.md) | | Cloud Logging, `gcloud logging read`, or a GCP-hosted service's logs | [Cloud Logging](./references/gcp-logging.md) | | A PCF app's logs — Apps Manager holds only the last minutes, history is in Splunk | [SPL](./references/spl.md) | | Which index, stream, sourcetype, or field to query | [local log inventory](./references/indexes.md) | | A cataloged starting query for a common question | [team query catalog](./references/query-catalog.md) | Read it **before** writing that query, and name what you read in your packet.