--- name: mirroir-onboard description: Onboard a consumer web app to mirroir's .mirroir/ dotfile by EXPLORING the running app (chrome-devtools-mcp) — derive real selectors from the accessibility tree, exercise each surface's primary action, emit the .mirroir/ tree, and validate by LIVE REPLAY with a self-heal loop. Reject shallow "page renders" coverage. allowed-tools: - "mcp__chrome-devtools__*" - "Read" - "Write" - "Edit" - "Bash" - "Glob" - "Grep" - "TaskCreate" - "TaskUpdate" - "TaskList" --- # mirroir-onboard You are the **web explorer**. You drive a consumer's running web app the way an agent drives an iPhone in `generate_skill action=explore`: navigate, read the live accessibility tree, find each surface's *primary action*, derive the exact selector that resolves it, execute it, capture the resulting state change, and emit a replayable `.mirroir/` suite. Then you **prove it by live replay** and **self-heal** any selector that doesn't resolve. The output is a `.mirroir/` tree at the consumer root that goes **green under `mirroir-run`** (a real boot + real Playwright replay against the real backend), not merely one that compiles. Compile-clean is necessary but not sufficient — selectors that pass `--emit playwright` still fail at replay (wrong element, strict-mode, or never resolves). The replay is the gate. You do not hand-author from spec files. Every selector comes from a DOM you observed via `mcp__chrome-devtools__take_snapshot`. You are mechanizing the human loop, not transcribing a test plan. --- ## What makes this worth doing (read before you start) A mocked Playwright/Cypress suite (`page.route(...)`) verifies the **frontend in isolation** — it stays green even when the backend wiring, DB seed, auth gate, dev proxy, or LLM path is broken. mirroir's only job is to catch exactly those: **"the real stack boots and the headline journeys actually work."** If your scenarios don't exercise real primary actions against the real backend, you've rebuilt the mocked suite slower — delete them and stop. So: **complementary, never a replacement.** Run `find -name '*.spec.ts' -o -name '*.cy.ts' | grep -v node_modules | wc -l` first; if non-zero, the `mirroir.yaml` description must say "complementary" and name that count. Assume an existing suite exists until proven otherwise. --- ## Phase 0 — Get the app to a usable, logged-in-able state This is the phase that silently eats hours. A web app being "up on :PORT" does **not** mean a real user can log in and reach an authenticated surface. Confirm each link of the chain before exploring, because every one of these has bitten real onboarding: 1. **Boot a genuinely fresh stack — kill stale processes first.** If a previous stack is still listening on the frontend port, a fresh boot's "wait for port" can race the old stack's teardown and you'll drive a dying server. Stop all services, confirm the port is **down**, then boot. 2. **Beware test-mode env flags — they often disable the real backend.** Many apps gate a dev proxy / mock layer on a flag like `E2E_TEST`, `CI`, `MOCK_API`. Their *own* mocked suite sets it so the frontend serves canned responses. For mirroir that is poison: with the proxy off, the login POST (e.g. `/oauth/token`) 404s and every authed scenario fails. **Do not set these flags.** If login fails, open the network panel / check the auth endpoint's status — a `404`/`502` on the auth POST means the proxy is off. 3. **IPv4 vs IPv6 loopback.** Dev servers (Vite's default `localhost`) often bind `[::1]` only, so `http://127.0.0.1:PORT` is refused while `http://localhost:PORT` works. Use `localhost` in scenario URLs, and know that any port-readiness probe must check both `127.0.0.1` and `[::1]`. 4. **The auth/onboarding gate may demand seeded state.** A user can authenticate yet be trapped on a "connect a provider / finish onboarding" wall that blocks every real surface — by design, and often *not* satisfied by the default seed. If a freshly-seeded user lands on such a gate, you cannot explore the app as them. Find which seeded user clears the gate (read the seeder), or surface to the user that the seed needs to make one demo user fully usable. Do not fake your way past it — that's no longer a real-stack test. 5. **Backend readiness ≠ frontend readiness.** The frontend may bind its port before the backend has finished a slow startup (migrations, a remote content sync). With the proxy on, the app's first calls hit a not-ready backend. Give the boot real headroom and confirm an authed API call returns `200` before exploring. Record the working facts as a discovery task: boot command, frontend port, the **usable** user credentials (and which gate they clear), and any flag you had to *avoid*. You will reference these throughout. > You generally do **not** boot the user's stack for them — it's their > environment. But you must *verify* the chain above against the running app > before exploring, and tell the user precisely which link is broken if one is. --- ## Phase 1 — Explore (one primary action per surface) Confirm the app is up (`list_pages`; navigate to the consumer URL if needed). Then walk it like a graph. ### 1a. Unauthenticated landing Clear storage + cookies for the origin, navigate to `/`, `take_snapshot`. Capture the form field selectors and a **unique-on-page** heading string (avoid "Sign in" if it appears as both an `

` and a button). → one `unauthenticated-landing` scenario asserting the form is present. ### 1b. Login per role, map the landing For each role the seeder creates (admin, regular/demo user): fill credentials, submit, wait for an authenticated signal, capture `location` and one sidebar item **unique to that role** (e.g. an admin-only nav button). → one login scenario per role asserting a real post-login element (not just "logged in"). ### 1c. Per-surface deep walk — THE anti-shallow rule For **every** top-level nav surface, find and **execute its primary action** — the thing a user comes to that surface to *do*, not the fact that it rendered: | Surface kind | Primary action (examples) | |---|---| | Chat / assistant | type a message → send → wait for the real reply | | List/store of items | open an item → assert its detail view | | Empty-state collection | click "Create …" → assert the create modal/form | | Feed | trigger the share/react/adapt action → assert the result | | Admin table | Approve/Suspend/Delete a row → assert the count/row change | A read-only surface with no action is **not** a deep flow — don't manufacture a shallow scenario for it (that's the anti-pattern); either find its real action or leave it to the mocked suite. One scenario per (surface × primary action); never collapse two distinct actions into one. --- ## Phase 2 — Derive selectors from the accessibility tree mirroir compiles `tap:`/`wait_for:`/`assert_visible:` labels to Playwright via a helper that resolves a label as follows: - starts with `[ # . : > *` → raw CSS / locator passthrough (`page.locator`) - starts with `role= text= xpath= css= id= data-testid=` → Playwright locator-engine passthrough (`page.locator('role=button[name="X"]')`, etc.) - otherwise → `[data-test="