--- name: blend-automation-test description: Execute supplied BLEND Test Spec cases against an authorized web target with the current Playwright MCP, capture reviewed evidence of the necessary view, and update one XLSX or native Google Sheets report. Supports scoped reruns and stop/resume; does not redesign cases or migrate reports. --- # Blend Automation Test Read [shared workflow](../../shared/workflow.md), [artifact formats](../../shared/artifact-formats.md) and [automation contract](../../shared/automation-testing.md). Resolve the user's checkouts and verified feature; source IDs, frozen design revision and one authoritative report determine execution identity. Use the actual supplied Test Spec/Test Data and report, including supported legacy formats. Never replace them with a fresh design or blank report to make execution easier. ## Intake Resolve the Test Cases, scope-and-approach, Test Data, target URL, requested scope, report language/location and auth/fixture preconditions. A run request authorizes the requested test actions and result entry in that report; it does not authorize database seeding, app fixes, tool/config updates, report migration, skill installation or publication. An authenticated/open tab supplies neither identity nor extra permissions. Existing session authorization persists; do not ask again for actions already authorized. Read [execution](references/execution.md) before running, [MCP evidence](references/mcp-evidence.md) before capture, [output](references/output.md) for run creation/resume/archive, and [report writeback](references/report-writeback.md) before reading or updating results. Load only required business/context references and preparations; whole-codebase research is not a default execution step. ## Run and evidence 1. Verify source/report revision and actual TC/Variant bindings. Inventory **every variant** in the requested scope before the first action. Full-suite requests include all variants; partial/rerun requests select the stated subset. Keep expected authority, preparation readiness and observed execution independent. A permitted Draft observation run does not approve its oracle or eligibility. 2. Discover currently callable project Playwright MCP tools and their real schemas. Use desktop viewport **1920×1080**, except a case's explicit viewport variant. Take a snapshot before interactions and use current refs, one action per call. `ai-workflow:th-using-playwright-mcp` and `ai-workflow:th-execute-automation-tests` may supplement available browser/execution guidance; this skill's bundled references provide the complete required fallback and BLEND report semantics. Do not adopt their different output/status defaults or database-seeding examples. 3. Independently prepare each variant, follow its actions/deltas and evaluate its concrete overall/final Expected plus required checkpoints, then perform its one real final Reset. Keep full material source context/fixtures/checkpoint oracles and preserve supported older source semantics; report2.4 core criteria and per-variant Actual/Status remain visible. Meaningful English is required for text the agent creates during testing and for screenshot notes; preserve UI labels, identifiers and required test payloads. An earlier failed case does not skip an independent later case. 4. Default both raw and annotated evidence to **`browser_take_screenshot(full_page=false)` explicitly**; use `true` only when the assertion needs the entire layout. Record the actual scope and keep raw/annotated viewport, scroll/state and scope unchanged. Each report image needs minimal precise highlights without touching/overlap, a short English observed-state note at the top-right in verified free space, and actual pixel inspection. Adjust the view or use a safe annotation gutter when no free space exists; never cover UI/data. Capture additional views only when needed; follow the bounded correction/cleanup protocol. Full-page does not prove hidden/internal/horizontal/virtualized content was included. 5. Archive each actual returned file immediately into the feature/Run-ID folder, before another capture. Verify digests and preserve raw/rejected evidence. Accept an EvidenceRecord only after opening pixels. Actual app-export files need independent verification; a browser screenshot or `browser_pdf` cannot substitute for them. 6. Update only the target variant after evaluating its mandatory assertions and reviewed evidence. TC, Actual, localized Status and images stay together in the existing two-sheet report. Customer2.4 presents each direct embedded image below its human-readable title, stacked vertically with original pixels/aspect ratio, and one initially collapsed evidence group per TC; core content always stays visible. No image URL is required. Without images show only the localized no-image line. Use returned bindings for image capacity/placement; keep supplied legacy layouts unchanged. Do not create an Evidence tab, migrate history, auto-sync two reports or reinterpret an evidence error as an app failure. Read back and inspect changed layout/images; native Sheets grouping/image behavior needs separate live verification. ## Stop, resume and handoff User stop ends further testing and external writes. Preserve the exact resume point, pending observations and cleanup state in the sole local run ledger; untouched report variants remain localized NOT RUN. Resume rereads frozen source, saved report and ledger, verifies their identities, reconciles uncertain writes and continues only incomplete requested work. Do not erase earlier failures with retries or rerun the entire suite automatically. Return the authoritative report location, requested/executed/remaining inventory, observed failures/blockers/evidence gaps, run output and cleanup/resume point. Distinguish eligible formal results from Draft observations, package/structural checks, pixel review and verified export/native Sheets proof. A successful tool call or checker is not a completed automation result.