--- name: design-iteration-loop description: | 반복적 디자인 품질 개선을 위한 빌더-평가자 GAN 루프 워크플로. 스프린트 컨트랙트 협상, 4차원 평가(디자인 품질·오리지널리티·완결성·기능성), 정체 감지, 에스컬레이션 프로토콜을 구현합니다. design.yaml에서 파라미터를 읽습니다. Use for the Builder-Evaluator GAN loop: iterative design-quality improvement via Sprint Contract negotiation, 4-dimension scoring (Design Quality, Originality, Completeness, Functionality), stagnation detection, and escalation. metadata: invocation-scope: "workflow" version: "1.1.3" --- > MoAI-ADK 프로젝트에서는 `.moai/config/sections/design.yaml`과 기존 스프린트 산출물을 읽는다. Claude Cowork·ChatGPT Work 데스크톱에서는 사용자가 제공한 브리프와 실제 시안으로 같은 평가 항목을 점검할 수 있다. 설정 파일이나 검사 도구가 없으면 그 항목을 미검증으로 기록하고 점수나 PASS를 만들어 내지 않는다. # design-iteration-loop Implements the Builder-Evaluator GAN loop for iterative design quality improvement. Absorbed from the retired v2.x design constitution Section 11 and Section 12 (per the design constitution absorption policy). Integrates Sprint Contract Protocol, 4-dimension scoring, stagnation detection, and Evaluator Leniency Prevention. In a MoAI-ADK project, read loop parameters from `.moai/config/sections/design.yaml`; the values below describe the current project configuration and are not portable defaults. In a desktop conversation without that file, agree on criteria with the user and report observations without inventing thresholds. --- ## Quick Reference ### Loop Parameters (from design.yaml) ``` design.gan_loop: max_iterations: 5 # Maximum Builder-Evaluator cycles pass_threshold: 0.75 # Score >= this value to exit loop escalation_after: 3 # Escalate to user after N iterations without passing improvement_threshold: 0.05 # Minimum score delta per iteration strict_mode: false # If true, each must-pass criterion must pass individually sprint_contract: enabled: true required_harness_levels: [thorough] optional_harness_levels: [standard] artifact_dir: ".moai/sprints" max_negotiation_rounds: 2 ``` ### 4-Dimension Evaluation | Dimension | Description | | --- | --- | | Design Quality | Visual consistency, supplied brand token compliance, measured contrast | | Originality | Brand-specific expression against the supplied brief | | Completeness | Requested BRIEF sections present, copy matches the agreed text | | Functionality | Observed responsive behavior and tested interactions | The current `design.yaml` does not define weights for these dimensions. Do not present 30/25/25/20 or any other split as configured. The ADK evaluator owns the overall score; cite its actual rubric and result. In a standalone desktop review, report each observed dimension and gaps without a synthetic overall score. In the ADK loop, pass requires the evaluator's measured overall score to reach `pass_threshold` and every contracted must-pass criterion to pass. Strict mode adds the constitution's two-iteration minimum. A standalone desktop review without a configured evaluator does not issue this numeric PASS. --- ## Implementation Guide ### GAN Loop Execution Flow **Phase 1: Sprint Contract (when required by harness level)** Required when `harness_level == thorough`. Optional when `harness_level == standard` and user opts in. Skipped when `harness_level == minimal`. Sprint Contract generation: 1. Evaluator analyzes the BRIEF document and current iteration scope. 2. Evaluator produces the Sprint Contract document: - `acceptance_checklist`: concrete, testable criteria for this iteration - `priority_dimension`: which of the 4 dimensions to focus on - `test_scenarios`: specific verification steps - `pass_conditions`: minimum score per criterion 3. Builder reviews the contract: - Accept: proceed with implementation - Request adjustment: propose alternatives (max `max_negotiation_rounds` rounds) 4. Contract is saved to `design.gan_loop.sprint_contract.artifact_dir/sprint-N.json` Constraint: Evaluator must not score on criteria outside the Sprint Contract. Builder must not claim criteria as met without evidence. **Phase 2: Builder Execution** Builder implements based on: - Accepted Sprint Contract (if present) - BRIEF document - Copy JSON from `design-copywriting` - Design tokens from `design-brand-system` or `design-workflow` (Path A handler) Builder outputs: code files, rendered previews (if Playwright available), implementation notes. **Phase 3: Evaluator Scoring** Evaluator scores against the 4 dimensions using the Evaluator Leniency Prevention mechanisms: 1. **Rubric Anchoring**: Score each dimension against the rubric (0.25 increments) with explicit justification. Scores without rubric reference are invalid. 2. **Evidence-Only Verdicts**: No PASS without concrete evidence (screenshot, test output, code reference). 3. **Anti-Pattern Cross-check**: Check known anti-patterns before finalizing. Any detected anti-pattern caps the relevant dimension score at 0.50. 4. **Must-Pass Firewall**: Copy integrity, mobile viewport, and WCAG AA are must-pass criteria. Failure in any must-pass = overall FAIL regardless of other scores. Output: `evaluation-report-N.json` in `sprint_contract.artifact_dir`. **Phase 4: Loop Decision** ``` if overall_score >= pass_threshold: EXIT LOOP → proceed to next phase elif iteration >= max_iterations: ESCALATE → present failure report to user elif stagnation_detected: ESCALATE → present stagnation options else: ITERATE → pass feedback to Builder, increment N ``` **Phase 5: Iteration Feedback** If looping back: 1. Evaluator generates targeted feedback per failed criterion. 2. Builder receives the feedback and previous Sprint Contract. 3. Previously passed criteria carry forward (no regression allowed). 4. New Sprint Contract is generated for failed criteria only. --- ### Stagnation Detection Stagnation is detected when the score improvement between consecutive iterations is below `improvement_threshold` for 2 or more iterations. Tracking: - After each iteration, record `{iteration: N, score: X}` in the sprint artifact. - Calculate `delta = score[N] - score[N-1]`. - If `delta < improvement_threshold` for the last 2 iterations, flag stagnation. When stagnation is detected, present the findings through the available user question channel with three options: 1. Continue with current approach (Evaluator tries a different dimension focus) 2. Adjust criteria (user provides guidance or relaxes constraints) 3. Abort loop (accept current output as-is) The escalation trigger at `escalation_after` iterations applies independently: if 3 iterations pass without a PASS score, escalate regardless of stagnation state. --- ### Evaluator Leniency Prevention Mechanisms The following 5 mechanisms prevent score inflation and must be applied on every evaluation: **Mechanism 1: Rubric Anchoring** Score descriptions for each dimension: - 0.25: Major defects, fails most criteria - 0.50: Partial compliance, notable issues remain - 0.75: Solid compliance, minor issues only - 1.00: Full compliance, no issues found Always state which rubric level applies and why before assigning a numeric score. **Mechanism 2: Must-Pass Firewall** The following conditions cause immediate FAIL regardless of other scores: - Copy text differs from the original `copy.json` or BRIEF copy section - AI slop detected: purple gradient (#8B5CF6-#6D28D9) as primary visual element with generic white cards - Mobile viewport broken at 375px width (content overflow, unreadable text) - Any interactive element returns 404 or broken state - An agreed accessibility criterion fails in a tool result; when Lighthouse is used and the brief sets a threshold, report its measured score **Mechanism 3: Anti-Pattern Penalty** Known anti-patterns that cap dimension score at 0.50: - Generic icon set without brand customization (Originality capped) - Hard-coded spacing values outside the design token scale (Design Quality capped) - Missing `alt` attributes on non-decorative images (Functionality capped) - Section copy that does not match the contracted copy (Completeness capped) **Mechanism 4: Evidence Requirement** Each dimension score must cite specific evidence: - Design Quality: Reference supplied tokens and a measured contrast ratio where available - Originality: Describe what makes the design non-generic - Completeness: List each BRIEF section and its implementation status - Functionality: Reference actual test or browser observation; otherwise mark unverified **Mechanism 5: Regression Baseline** If a previous iteration passed a criterion, the current iteration must maintain that criterion. Regression from a previously passed criterion triggers an automatic score reduction in the relevant dimension. --- ### Sprint Contract Structure Sprint Contract document format (`sprint-N.json`): ```json { "sprint_id": "sprint-N", "iteration": N, "priority_dimension": "Design Quality | Originality | Completeness | Functionality", "acceptance_checklist": [ { "id": "AC-01", "criterion": "Hero headline contrast ratio >= 4.5:1", "verification": "Check color pair with contrast calculator", "status": "pending | passed | failed" } ], "test_scenarios": [ { "id": "TS-01", "description": "Mobile viewport renders without horizontal scroll", "tool": "Playwright | visual inspection", "command": "playwright test --viewport 375x667" } ], "pass_conditions": { "Design Quality": 0.75, "Originality": 0.70, "Completeness": 0.80, "Functionality": 0.75 }, "negotiation_history": [], "created_at": "ISO-8601" } ``` --- ## Advanced Patterns ### Strict Mode When `strict_mode: true` in `design.yaml`: - Each contracted must-pass criterion must individually pass, as constitution §11 requires. - An overall score cannot compensate for a failed must-pass criterion. - Minimum 2 iterations required even if the first iteration achieves a passing weighted average. - Strict mode is recommended for client-facing deliverables. ### Independent Re-evaluation Every 5th project triggers an independent re-evaluation: - The same build is scored twice with independent prompts. - If scores diverge by more than 0.10, a calibration warning is logged. - Calibration results are stored in `sprint_contract.artifact_dir/calibration-log.json`. ### Playwright Integration When a browser or Playwright is available, the Evaluator may run the relevant checks: - Desktop screenshot (1280x720): full page - Mobile screenshot (375x667): full page - Interaction test: click all CTAs, verify no 404 - Accessibility scan: record the tool and its findings; an automated scan alone does not establish full WCAG conformance When testing tools are unavailable, record code observations and mark browser behavior, accessibility and interaction criteria unverified. Do not score an unobserved criterion as passed. --- ## Works Well With - `design-brand-system`: Provides design tokens that Evaluator validates in Design Quality dimension - `design-copywriting`: Copy JSON is the reference for Completeness dimension - 자체 Evaluator: GAN 루프는 매 스코어링 패스마다 본 스킬의 4-dimension scoring으로 평가합니다. MoAI harness(`moai`)의 `sync-auditor`가 함께 설치된 환경에서는 해당 agent로 평가를 보강할 수 있습니다. - `design-workflow`: Extracted tokens (Path A) serve as the design reference baseline --- Source: Absorbed from the retired v2.x design constitution per the design constitution absorption policy (Section 11 GAN Loop Contract, Section 12 Evaluator Leniency Prevention). REQ coverage: (internal provenance omitted) Version: 0.1.0