# AI Operations Runbook Toolkit schema: `PAISEH-TK-1.0` Source artifact: Chapter 11 **AI Operations Runbook** Complete `../common-artifact-header.md` first. This runbook must let an operator move from signal to evidence validation, command, containment, investigation, verified recovery, communication, and learning without inventing authority or semantics. ## 1. Service Charter and Effective State - Critical workflows, outcomes, and prohibited outcomes: - Populations, regions, tenants, and consequence classes: - Effective release, topology, configuration, and route receipts: - Service, domain, on-call, support, and escalation owners: - Critical dependencies and authoritative records: - Accepted degraded modes and scope ceilings: - Tool Permission Matrix, Evaluation Plan, AI Threat Model, and AI Release Manifest identities, versions, and accepted decision references: - Reference Architecture Catalog handoff: operational invariants, capacity limits, degraded modes, containment, and recovery proof: ## 2. Objectives and Evidence Health | Objective | Layer/population | Indicator/window/boundary | Delay/minimum evidence | Evidence-health dependency | Owner/response | | --- | --- | --- | --- | --- | --- | | | Behavioral / execution / economic / evidence | | | | | | Authoritative query/view | Decision served | Dimensions/drill-down | Freshness/health | Change overlays | Known blind spots | | --- | --- | --- | --- | --- | --- | | | | | | | | - Schema, producer, causal identity, sampling, and join-health checks: - Missing or indeterminate evidence behavior: - Protected payload, access, retention, and replay constraints: ## 3. Detection and Declaration | Rule | Protected objective/population | Window/threshold/volume/completeness | Indeterminate behavior | Severity/owner/destination | Automated authority/expiry | Test/last exercise | | --- | --- | --- | --- | --- | --- | --- | | | | | | | | | - Declaration channel and provisional classification: - Severity dimensions: - Incident commander, operations lead, domain lead, communications lead, and scribe: - Security, privacy, legal, action, release, and business handoff triggers: - First scope queries and decision-log location: ## 4. Containment and Degraded Modes | Failure condition | Scope | Action/procedure reference | Authority | Preconditions | Expected effect | Expiry/reversal | Verification query/receipt | | --- | --- | --- | --- | --- | --- | --- | --- | | | | | | | | | | Cover observation, increased sampling, feature/cohort restriction, review-only mode, admission stop, queue pause, state isolation, release transition, cancellation, and reconciliation as applicable. ## 5. Investigation Protocol - Incident identity, scope hypotheses, sources, clock offsets, and gaps: - Alerts, acknowledgements, automated actions, commands, approvals, and communications: - Effective-state and dependency history: - Causal graphs, action receipts, queue/state references: - Versioned queries, snapshots, denominators, sampling, and evidence health: - Protected evidence references, redaction, access, and retention holds: - Hypotheses, disconfirming evidence, observations, decisions, and uncertainty: - Reproduction type, result, and limitation: ## 6. Recovery, Verification, and Communication | Dimension | Required recovery claim | Evidence/query | Observation window | Owner | Exit criterion | | --- | --- | --- | --- | --- | --- | | New traffic/infrastructure | | | | | | | Queued/in-flight work | | | | | | | State/caches/indexes | | | | | | | External actions/effects | | | | | | | Behavioral outcomes/slices | | | | | | | User correction/remediation | | | | | | | Telemetry/evidence health | | | | | | | Delayed outcomes/uncertainty | | | | | | - Residual degraded mode or uncertainty and accepting authority: - Audiences, message owner, cadence, and constraints: - Recovery and final-status receipts: ## 7. Learning and Closure - Event, impact, population, duration, and uncertainty: - Detection path, missed signals, and time to credible awareness: - Response actions, authority delays, and effects: - Contributing system and organizational conditions: - Failed or absent prevention, detection, containment, recovery, and assurance controls: - Feasible counterfactuals and residual risk: | Action | Class | Owning chapter/system | Owner/due date | Verification and release/policy path | Closure authority | Recurrence signal | | --- | --- | --- | --- | --- | --- | --- | | | Immediate repair / prevention / detection / recovery / evidence / accepted risk | | | | | |