--- name: akamai-edge description: >- Akamai edge work in three lanes — triage (edge vs origin, Reference # error strings, cache status, WAF denials, DataStream 2), delivery config (Property Manager versions, staging-first activation, fast fallback), and mPulse RUM (network-side vs app-side slowdowns). Triggers: 'is it the CDN or the origin', 'an Akamai error page with Reference #...', 'is the WAF blocking real users', 'mPulse shows slow pages'. Not for backend log queries (obs-logs) or a firing alert (incident-investigation). compatibility: Requires Akamai Control Center access; DataStream 2 queries run in the configured log backend argument-hint: "[the edge, CDN, WAF, or RUM problem]" --- # Akamai edge — triage, delivery config, RUM We are Akamai customers, not Akamai operators: our lane is our properties' configuration, our security policies, and the evidence the portal and DataStream 2 expose. The edge network itself is Akamai's platform — a suspected platform-wide Akamai problem is an escalation to Akamai support with evidence, not something to debug past the portal. ## The first question: which side of the edge? Establish which parts of the request path were exercised: client/DNS/TLS, edge/cache/WAF, and origin. Not every request reaches the origin; a cache hit can be served at the edge. Use evidence to distinguish the candidate causes: 1. **An Akamai error page with `Reference #…`** → decode it in Edge Diagnostics' Translate Error String **promptly** — the logs behind a reference number survive roughly 6–24 hours [sourced: techdocs.akamai.com/edge-diagnostics/docs/translate-error-string]. 2. **No reference number** → Edge Diagnostics' Get Error Statistics splits errors into the client→edge and edge→origin legs per URL/CP code; URL Health Check bundles grep + dig + curl + MTR for one URL in one job [sourced: techdocs.akamai.com/edge-diagnostics/docs/get-error-statistics, …/url-health-check]. 3. **Cache behavior in question** → read the cache-status response headers via the debug-header mechanism the property actually supports (Enhanced Debug vs legacy Pragma — the reference explains which and why it changed). 4. **Sustained, fleet-wide or regional questions** → DataStream 2 fields (`statusCode`, `cacheStatus`, `turnAroundTimeMSec`, `errorCode`, `country`) in the log backend `stack-profile` records, with `obs-logs` writing the query, and Traffic by Hostname for offload trends. Edge-side evidence (WAF deny, cache misconfiguration, edge 5xx) stays in this skill's lanes. Route origin-side findings through [Handoffs](#handoffs). ## Three lanes, three authority postures - **Triage is read-only, portal-first.** Edge Diagnostics, Web Security Analytics, Reporting, and DataStream 2 queries change nothing. Production debug-header requests stay **recommend-for-human** whatever tools the invoking lane holds; the human runs them with their own debug token. Other observations, including selected DNS reads, require the invoking agent's granted, scoped, output-protected read path. - **Delivery config is change-managed work.** A property version edit is Tier 1 (prepare); any activation — staging included — is a live change with an approval gate and a named, proven rollback or recovery path. Production activation additionally runs through `production-change-gate` with a human release owner. - **WAF policy changes are security changes.** Evidence of a false positive goes to the human security policy owner with the sampled requests attached; this fleet never loosens a protection itself. ## Read the reference before acting | Need | Reference | |---|---| | Edge vs origin evidence: Edge Diagnostics, reference numbers, cache status, debug headers, DataStream 2, offload reports, WAF events | [Edge triage](./references/edge-triage.md) | | Property Manager change flow: versions, staging/production activation, fast fallback, PAPI/Terraform/CLI, Sandbox | [Property config](./references/property-config.md) | | Cache purge, a live change: invalidate vs delete, scope, network, origin load | [Cache purge](./references/property-config.md#cache-purge) | | Real-user monitoring: beacons, Core Web Vitals, back-end vs front-end time, slicing a regression | [mPulse RUM](./references/mpulse-rum.md) | ## Handoffs The responder with `incident-investigation` retains the overall live investigation. A dispatched `sre-assistant` may follow origin-side evidence (edge→origin errors, healthy edge with slow turnaround) within its assigned question, targets and access; return the leg finding, exact source, timestamps, alternatives and gaps to its invoking caller. A related lead does not grant new access or transfer incident ownership. Origin leg outside an incident: `pcf-ops` or `gcp-ops` for the origin app, `obs-traces` or `obs-metrics` for origin latency. A recurring query, missing alert, or detection gap goes to `observability-engineer`. For a proposed property change, send the human release owner the prepared version and [property change packet](./references/property-config.md#property-change-packet). New durable operational facts route to `scribe` through the operational-learning disposition path.