--- name: qa description: "General-purpose QA verdict for any artifact type" --- # /ouroboros:qa Standalone quality assessment for any artifact — code, documents, API responses, test output, or custom content. Unlike `ooo evaluate` (3-stage formal verification pipeline), `ooo qa` is a fast single-pass verdict with actionable suggestions. ## Usage ``` ooo qa [file_path | artifact_text] ooo qa # evaluate recent execution output /ouroboros:qa [file_path | artifact_text] # plugin mode ``` **Trigger keywords:** "ooo qa", "qa check", "quality check" ## How It Works The QA Judge evaluates an artifact against a quality bar and returns a structured verdict: 1. **Parse the Quality Bar** — What EXACTLY must be true to pass? 2. **Assess Dimensions** — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific 3. **Render Verdict** — Score (0.0-1.0) with PASS / REVISE / FAIL 4. **Determine Loop Action** — `done` (pass), `continue` (revise), `escalate` (fail) ### Verdict Thresholds | Score Range | Verdict | Loop Action | |--------------|---------|-------------| | >= 0.80 | PASS | done | | 0.40 - 0.79 | REVISE | continue | | < 0.40 | FAIL | escalate | ## Instructions When the user invokes this skill: ### Step 0: Determine execution mode This skill works in two modes. Determine which one **before** attempting any tool calls: - **MCP mode** — If the QA MCP tool is available (already exposed, or loadable via discovery), use it: ``` tool discovery query: "+ouroboros qa" ``` If found (typically named `mcp__plugin_ouroboros_ouroboros__ouroboros_qa`), proceed with **QA Steps** below. - **Fallback mode** — Only if the QA MCP tool is genuinely absent (no Ouroboros MCP server) skip to the **Fallback** section; an empty discovery result for an already-exposed tool is expected — call it directly rather than falling back. This skill is designed to work without MCP setup. ### QA Steps (MCP mode) 1. **Determine the artifact to evaluate:** - If user provides a file path: Read the file with Read tool - If user provides inline text: Use that directly - If no artifact specified: Look for the most recent execution output in conversation context - Ask user if unclear what to evaluate 2. **Determine the quality bar:** - If a seed YAML is available in context: Extract acceptance criteria from it - If user specifies a quality bar: Use that - If neither: Ask the user "What does 'good' mean for this artifact?" 3. **Determine artifact type:** - `code` — source code files - `test_output` — test results, CI output - `document` — specs, docs, READMEs - `api_response` — API responses, JSON payloads - `screenshot` — visual artifacts - `custom` — anything else 3.5. **Acting verification fan-out — probe in parallel, then judge (do not skip for behaviour-bearing artifacts):** A text judge can be fooled by a hopeful log line. When the artifact actually *does* something (code, an app, an API, a UI), fan out empirical probes using the host's native parallel sub-agent primitive — one probe sub-agent per acting modality the runtime actually exposes, all spawned **in the same message so they run concurrently**: - **process probe** (`Bash`/shell): run the command / start the app / run the declared smoke commands with bounded timeouts; capture exit codes and real output. - **browser probe** (browser-use tools, when the artifact serves HTTP or is a web UI): load it, click the primary flows, capture what actually renders and any console/network errors. - **computer-use probe** (desktop computer-use tools, when the artifact is a GUI/TUI): drive it like a user, screenshot the observed states. - **artifact probe** (file reads): verify declared files/paths exist with real content, not placeholders. Each probe returns structured evidence only — commands run, observed effects, screenshots/paths, pass/fail per probed behaviour. Every probe also hits the applicable adversarial classes (the QA tool lists them): `misleading_output` (claimed success vs. real effect), `hung_command` (bounded timeout?), `malformed_input`, `stale_state`, `dirty_worktree`. Skip a modality only when its tools are absent or the artifact type makes it meaningless — and say which modalities were skipped and why. Await all probes, then pass the merged evidence into the judge as `reference` (prefer observed behaviour over source text as the `artifact` when they disagree). **Empirical evidence outranks the judge**: if the judge scores PASS but any probe observed the behaviour failing, present the verdict as REVISE/FAIL on that evidence and say so explicitly — a score contradicted by observation is not a pass. If no acting tools are available at all, judge on the text alone but flag that behaviour was not observed. 4. **Call the `ouroboros_qa` MCP tool:** ``` Tool: ouroboros_qa Arguments: artifact: quality_bar: artifact_type: "code" (or other type) reference: pass_threshold: 0.80 (adjustable) seed_content: ``` 5. **Present results clearly:** - Show the score and verdict prominently - List dimension scores - Highlight specific differences found - Show actionable suggestions - End with next step guidance based on verdict: - **PASS (done)**: `Next: Your artifact meets the quality bar. Proceed with confidence.` - **REVISE (continue)**: `Next: Address the suggestions above, then run ooo qa again to re-check.` - **FAIL (escalate)**: `Next: Fundamental issues detected. Consider ooo interview to re-examine requirements, or ooo unstuck to challenge assumptions.` ### Iterative QA Loop For iterative usage, track the `qa_session_id` and `iteration_history` from the response meta: 1. First call returns `qa_session_id` and `iteration_entry` in meta 2. On subsequent calls, pass `qa_session_id` and accumulated `iteration_history` 3. Continue until verdict is `pass` or `fail` In fallback mode, generate a `qa-` session ID on the first run and maintain iteration count in conversation context to preserve the same iterative contract. ## Fallback (No MCP Server) If the MCP server is not available, adopt the `ouroboros:qa-judge` agent role directly: 1. Read the canonical agent definition: `/src/ouroboros/agents/qa-judge.md` (This is the same prompt used by the MCP QA tool, ensuring consistent verdicts.) 2. Run the same acting-verification fan-out as step 3.5 (parallel probe sub-agents per available modality; empirical evidence outranks the judge) before judging behaviour-bearing artifacts. 3. Follow the QA Judge framework to evaluate the artifact 4. Output the verdict in the standard format (must match MCP output shape): ``` QA Verdict [Iteration N] ======================== Session: qa- Score: X.XX / 1.00 [PASS/REVISE/FAIL] Verdict: pass/revise/fail Threshold: 0.80 Dimensions: Correctness: X.XX Completeness: X.XX Quality: X.XX Intent Alignment: X.XX Domain-Specific: X.XX Differences: - Suggestions: - Reasoning: <1-3 sentence summary> Loop Action: done/continue/escalate ``` ## Example ``` User: ooo qa src/main.py QA Verdict [Iteration 1] ============================================================ Session: qa-a1b2c3d4 Score: 0.72 / 1.00 [REVISE] Verdict: revise Threshold: 0.80 Dimensions: Correctness: 0.85 Completeness: 0.60 Quality: 0.75 Intent Alignment: 0.80 Domain-Specific: 0.60 Differences: - Missing error handling for network timeout in fetch_data() - No input validation on user_id parameter - Type hints missing on 3 public functions Suggestions: - Add try/except with TimeoutError in fetch_data() (line 42) - Add isinstance check for user_id at function entry - Add return type annotations to get_user(), fetch_data(), process_result() Reasoning: Core logic is correct but lacks defensive programming patterns expected for production code. Loop Action: continue Next: Address the suggestions above, then run `ooo qa` again to re-check. ``` ## RFC #1392 State Breadcrumb Footer Your final response MUST end with exactly one breadcrumb footer line: ``` ◆ → next: ``` Derive `` from live session state via `ouroboros_session_status` when that MCP projection is available; otherwise derive it from this skill's actual outcome. Never use a linear `Step N of M` footer because Ouroboros is an evolutionary loop. When the next action is genuinely a choice, list 2-3 honest options in the `next:` clause. The breadcrumb line must be the last line of the response.