--- name: verify-omnigent description: Drive Omnigent the way a user does and prove a behavior with recorded evidence, using an isolated server, runner, and mock model plus a feature map of every user entry point. Load before reproducing a user-facing bug, before claiming a fix works, or when reviewing whether a change covered every surface (session vs. new-session composer, terminal strip vs. agent terminal, all twelve native harnesses). Covers the web UI, native harness terminals, and the CLI. --- # Verify Omnigent Use this skill to see a behavior happen in the real app and to prove a change, not to reason about it from code. It has two parts: - **An isolated instance.** [`scripts/verify-env`](scripts/verify-env) wraps `python -m dev.repro_env`: a server, runner, and mock model server on private ports, with their own config, data, Claude, and Codex directories. It never touches `~/.omnigent`, a running host daemon, or another developer server. - **A feature map.** [feature map](../../README.md) lists each user-facing feature's entry points, the tests that drive them, and the traps. A fix is verified only when every entry point listed for its feature has proof. Run all commands from the repository root. Put `scripts/` on your path or call `feature-map/skills/verify-omnigent/scripts/verify-env` directly. ## Launch 1. Install dependencies, Chromium, and build the web UI for the checkout: ```sh uv sync --frozen --extra all --group test uv run --no-sync playwright install --with-deps chromium pnpm install --frozen-lockfile --filter web && pnpm --filter web run build ``` 2. Start an instance and load its paths: ```sh verify-env start # waits until the runner is online, about 10 seconds eval "$(verify-env paths)" ``` `start` writes `VERIFY_ROOT` to `.omnigent/verify/current`, so later shells find the same instance. The instance stops itself after its lease (default 90 minutes; `--lease SECONDS`, 60 to 21600). Ready means the server answers, the runner reports online, and the mock model server answers. `start` fails with the environment's error and log path otherwise. Ready does not check browser launch. Use the Chromium version installed by this checkout's Playwright package; if the environment provides it through `PLAYWRIGHT_BROWSERS_PATH`, keep that path available to the test process. `--ui-skip-build` reuses the existing bundle, so rebuild after frontend edits. Native-terminal journeys also need the relevant CLI and terminal prerequisites (such as tmux); a browser installation alone does not provide them. In CI, the repro workflow already runs this environment. Use `python -m dev.repro_env exec -- ...` there instead of starting another. ## Doctor Run `verify-env doctor` before the first drive, after any failed drive, and whenever something looks off. It checks that the instance is ready, that its supervisor is alive, that the runner and model server answer, and it warns when the checkout has moved since launch. A warning about a moved checkout means the instance runs old code: stop it and start a new one. ## Drive 1. Open the matching file in [feature map](../../README.md) and list every entry point for the behavior in question. 2. For each entry point, run the named test through the instance, recording on: ```sh verify-env run -- python -m pytest \ --ui-skip-build --video=on --screenshot=on \ --output="$VERIFY_EVIDENCE/" ``` Tests under `tests/browser_ui/` need no instance: `uv run pytest --browser-ui-skip-build --video=on --output=...`. Tests the feature file marks "own environment" also run with plain `uv run pytest`. Server/transport/component tests use plain pytest without browser recording or build flags. 3. For an entry point with no test, hand-drive it with Playwright against `OMNIGENT_REPRO_SERVER_URL` inside `verify-env run -- python