--- name: service-lifecycle description: >- Take a service through its operational life on this stack: audit an existing service's readiness read-only, onboard a new or materially changed service, and retire one, each as a checklist a human executes under production-change-gate. Triggers: 'audit this service', 'onboard this service', 'retire this service', 'decommission this application'. Onboard and retire support draft planning; live execution requires an approved plan. argument-hint: "[audit|onboard|retire] [environment]" --- # Service lifecycle Every mode reads evidence and produces a checklist with owners; a human release owner or separately approved protected automation executes every live step under `production-change-gate` and returns its receipt. This skill changes nothing live and never requests a credential-bearing read; a prohibited read is recorded as a gap. ## Pick the mode | Mode | Runs when | Otherwise | |---|---|---| | **Audit** (default) | Any readiness question about an existing service | — | | **Onboard** | The caller requests onboarding or its plan | Inventory the authorized scope and prepare a `draft — unapproved` checklist; identify missing service/environment, owner, source commit, authoritative definitions, and approval | | **Retire** | The caller requests retirement or its plan | Inventory the authorized scope and prepare a `draft — unapproved` checklist; identify missing consumers, dependencies, data-retention obligations, recovery evidence, and approval | Unknown service or environment identities bound discovery: inspect supplied repository evidence, and request only the target-scoped reads the caller authorized. Ask for the scope before live reads when it is unknown; never invent a target, owner, consumer, or approval. Discovery and draft planning require no production approval. Before an onboard or retire plan is ready for live execution, require its approved plan, target, owner, source commit, definitions, executor, and current production gate. Retirement also requires the change record, known consumers and dependencies, data-retention obligations, proven recovery path, and approval expiry. Missing ownership, unknown consumers, unclassified data, or unproven recovery blocks removal, not the inventory that resolves it. Load `stack-profile` before interpreting platform, runtime, or backend evidence. Load an owning skill from the table below only for the surface you are on; it supplies expected evidence, not authority. Before any live step in onboard or retire, enter `production-change-gate` with the exact target and change. ## The surfaces Work the rows that apply. Audit inspects; onboard produces; retire dispositions each surface as `remove`, `retain`, `transfer`, `retire-record`, or `BLOCKED`, with an owner. Silence is not done. | Surface | Audit inspects | Onboard produces | Retire | Owning skill | |---|---|---|---|---| | Ownership and boundary | service owner, on-call, runtime and environment, platform escalation boundary | the same, recorded | ownership of retained shared resources transferred | `stack-profile` | | Runtime and health | versioned deployment definition, health check, instance or revision health, crash history | `manifest.yml` or pinned Cloud Run config; a workload-appropriate health check; justified instance target, scale-to-zero preserved unless the SLO needs minimums | workloads, routes, schedules, bindings removed only after quiescence is verified | `pcf-ops` or `gcp-ops` | | Delivery and recovery | CI promotion controls, exact rollback or roll-forward path, last recovery evidence | build and deploy via Actions with promotion gates on | deployment workflows and environment bindings disabled; access removal routed to the identity owner | `ci-actions` | | Telemetry pipeline | structured logs, RED metrics, traces, collector path, one arrival query per signal | OTel SDK wired, cardinality reviewed, collector routing to the selected destinations, arrival proven with one quoted query per signal | collection stopped after the evidence needed to verify the decommission is retained; shared collectors left in place are listed | `obs-pipeline` | | Dashboards | service health overview, drill-downs, owner, verification evidence | the service page: health at the top, drill-down below | removed only as authorized; shared dashboards listed | `obs-dashboards` for design; `grafana` for implementation | | Alerts and SLOs | SLI formula, target and window, symptom-based paging, saturation coverage, runbook link, notification-path evidence | a burn-rate alert on the SLI for request-based services or a freshness, completion, or failure alert for scheduled work, a saturation alert where the signal exists, each linked to a runbook; SLI formula, target, and window recorded where the team keeps them | paging alerts and SLO evaluation retired first, in dependency order | `obs-alerting` for design; `grafana` for Grafana operations | | Operations knowledge | current runbook, service and alert records, escalation procedure | a check, restart, recover runbook on-call can find | records marked `retired`, never deleted; active indexes updated | `runbook` | | Dependencies and capacity | critical dependencies, failure behaviour, limits, headroom, expiry risks | recorded on the service card | inbound callers migrated or stopped, producers and consumers drained, nothing still expects the service; unknown or conflicting evidence stops removal | owning skill | | Data, backup, restore | backup scope plus a dated restore rehearsal; existence alone is not restore evidence | recovery path recorded | data deletion, credential revocation, DNS and certificate removal, and access-path changes stay Tier 3, each with its own recovery evidence and human executor | `database-reliability` | Retire works the rows in a fixed order with stop rules, from traffic exit through independent verification; read [retirement order](./references/retirement-order.md) before drafting the retirement plan. Its execution stop rules identify draft gaps and still block live removal. ## Audit findings - **P0** exposed without required authentication, or stateful with no usable backup or recovery path. - **P1** a current high-impact failure, or a missing control likely to prevent safe detection, mitigation, rollback, or recovery. - **P2** a material gap with a workaround or limited blast radius. - **P3** hygiene, maintainability, or evidence freshness. A finding needs a cited file, record, or minimal sanitized output; causal claims stay `[unverified]`. A control that exists but is absent from the service's record is a documentation gap; one absent from both is a readiness gap. Report them separately, and report onboarding as unverified when no record exists at all. ## Knowledge closeout For a change, remediation, freshness review, or retirement, read [record transitions](./references/record-transitions.md) to name who retains ownership until the reviewed record is read back, and when evidence becomes stale. Reuse the existing follow-up record. After onboarding execution—or after retirement's live-resource receipts but before final record verification—send `scribe` an evidence-bound handoff for affected cards, index entries, and runbooks. Include the authorizing record, exact repository revision, caller's `[verified]` checkout binding, receipts, retained evidence labels, and non-actions. Audit findings use the same route. This skill never loads `operational-learning` or writes records. Drafts describe proposed knowledge changes; missing records use its templates. A retirement handoff remains pending until reviewed records and final verification are supplied. ## Return Lead with the conclusion, dated in UTC, with the age of the oldest load-bearing evidence. Then, by mode: **audit** returns up to three validated fixes in priority order, severity-ranked findings with evidence and owner, checks that passed, gaps and prohibited reads not run, and a plain statement that nothing was changed; **onboard** and **retire** drafts return the proposed checklist, `draft — unapproved`, and unresolved identities, approvals, and recovery evidence. Executed work returns each surface's row with its result or `UNKNOWN`, every gate verdict and receipt, unresolved dependencies, the `scribe` handoff, and what was not done. Never report an effect complete while any surface is `BLOCKED`, `UNKNOWN`, or merely planned, and close an onboarding by recommending its own audit as owed verification. ## Optional resolved context A caller with a compatible resolver may resolve [this skill's context requirements](./context-requirements.yaml) under the SRE operational-context contract ADR. Resolved context is routing input only: it never supplies a plan, an approval, or a credential, and missing context is a gap, not a guess. Before relying on freshness, read the same record-transitions reference: catalog `lastVerified` is not execution-backed `last_verified`.