# Repo conventions This file describes the conventions in place across the devtools monorepo — how code is organized, how packages relate to each other, how tests are structured, and what the coding style looks like. It's the companion to [ARCHITECTURE.md](./ARCHITECTURE.md): that file says where the pieces are; this one says why they're shaped the way they are and what to look for when adding or changing code. Anyone working in the repo, human or AI agent, can use this as the source of truth for "how do we do things here." --- ## What this repo is A devtools dashboard for end-to-end browser tests. Three test frameworks (WebdriverIO, Nightwatch, Selenium) push the same normalized event stream through a single backend into a single Lit-based browser UI. The adapters are deliberately thin — they translate framework hooks into calls on a shared core capture/reporting library and own only the framework-specific glue. Package map and data flow are in [ARCHITECTURE.md](./ARCHITECTURE.md). The summary: `shared` for types and contracts, `trace` for the event→zip transforms, `core` for framework-agnostic capture, three adapters (`service`, `nightwatch-devtools`, `selenium-devtools`) for framework glue, `backend` for the server, `app` for the UI, `script` for the page-injected runtime. --- ## Commands Run from repo root unless noted. | Command | What it does | |---|---| | `pnpm install` | Install workspace dependencies. | | `pnpm build` | Build all packages (`pnpm -r build`). | | `pnpm test` | Run vitest suite once. | | `pnpm test:watch` | Run vitest in watch mode. | | `pnpm test:coverage` | Run vitest with v8 coverage. The thresholds in `vitest.config.ts` are aspirational, not a gate: that file states CI does not run this and the suite is currently below all four. CI runs `test`/`lint`/`test:ui`. | | `pnpm lint` | Lint all packages in parallel. Includes `eslint-plugin-security` for a subset of CodeQL findings; deeper taint-flow checks surface on the PR's CodeQL scan. | | `pnpm demo:wdio` / `pnpm demo:nightwatch` / `pnpm demo:selenium` | Run the per-framework example projects. Useful for manual verification of UI or runtime changes. **No package pins a browser driver**, deliberately: a pinned `chromedriver` rots against whatever Chrome the developer has, and because pnpm puts the package's `node_modules/.bin` on `PATH`, Selenium Manager finds it there, prefers it over resolving, and on a major mismatch only *warns* before returning it anyway. The session then fails to start. With nothing on `PATH` both adapters resolve a matching driver (Nightwatch's Chrome service builder sets `requiresDriverBinary: false` and passes an unset `server_path` through to the same resolver), and nightwatch declares `chromedriver` an optional peer. | | `pnpm demo:wdio:mobile` / `:selenium:mobile` / `:nightwatch:mobile` / `:python:mobile` | The same, against Appium. All four build the same capability bag and drive the Clock app that ships with every Android system image — starting a timer, pausing it, clearing it — so a native example needs no `.apk`. `DEVTOOLS_MOBILE_PLATFORM=ios` runs the iOS spec instead — all four adapters — which drives Settings because the simulator ships no Clock, from a separate spec per platform rather than a branch. The simulator is chosen by udid and defaults to whichever is already booted, and an unmatched `IOS_DEVICE_NAME` is refused: naming one that does not exist makes the XCUITest driver create and boot it, every run, rather than fail. That refusal and the booted-simulator preflight are **local** policy — `xcrun simctl` enumerates local simulators and nothing else — so `IOS_UDID` and a non-local `APPIUM_HOST` both bypass them, or a real device and a cloud grid would be refused a run they were correctly configured for. `DEVTOOLS_MOBILE=web` drives the device's own browser on both platforms — Chrome on Android, Safari on iOS, which the XCUITest driver serves without a chromedriver. `examples/MOBILE.md` holds the prerequisites and the `DEVTOOLS_MOBILE` / `APPIUM_APP` switches; `DEVTOOLS_MODE=trace` flips any demo to trace mode. | | `pnpm dev` | Run all packages in parallel dev mode. | | `python3 packages/selenium-devtools-py/scripts/changes.py next-version` | The version a Python-adapter release would publish, from the fragments pending in `changes/`. `check --base ` is the CI gate; `apply` is what the release runs. | `selenium-devtools` exposes per-runner variants of its example via `pnpm --filter @wdio/selenium-devtools example:mocha` / `:mocha:allure` / `:jest` / `:cucumber`. --- ## Path aliases Defined in root `tsconfig.json`: | Alias | Resolves to | |---|---| | `@/*` | `packages/app/src/*` | | `@components/*` | `packages/app/src/components/*` | | `@core/*` | `packages/app/src/core/*` (app-internal — not the framework-agnostic `packages/core`) | | `@wdio/devtools-backend` / `*` | `packages/backend/src/...` | | `@wdio/devtools-script` / `*` | `packages/script/src/...` | | `@wdio/devtools-service` / `*` | `packages/service/src/...` | | `@wdio/selenium-devtools` / `*` | `packages/selenium-devtools/src/...` | | `@wdio/devtools-shared` / `*` | `packages/shared/src/...` | | `@wdio/devtools-core` / `*` | `packages/core/src/...` | | `@wdio/devtools-trace` / `*` | `packages/trace/src/...` | | `@wdio/elements` / `*` | `packages/elements/src/...` | These exist so imports stay short and grep-able. Long relative paths (`../../../components/…`) aren't used. The `@core/*` name is a historical alias for app-internal helpers and predates `packages/core`. They don't collide because they resolve to different roots, but the names are confusable. --- ## Conventions ### One source of truth per concept Every shared type, constant, enum, schema, and HTTP/WS contract lives in `packages/shared`. Adapter packages and the app never re-declare a concept that already exists upstream — they re-export shared definitions when a local consumer name needs to stay stable (e.g. nightwatch's `TEST_FILE_PATTERN` is `export { SPEC_FILE_RE as TEST_FILE_PATTERN } from '@wdio/devtools-shared'`). When a duplicate is discovered, the next change that touches either copy consolidates them into shared. ### Framework-agnostic logic lives in `core` Anything that captures, parses, normalizes, formats, or transports test-event data and doesn't depend on a specific framework's API lives in `packages/core`. Adapters call into core; they don't reimplement. If the same logical change would land in two or more adapters, the logic belongs in core. This rule produced the current `SessionCapturerBase`, `TestReporterBase`, `ScreencastRecorderBase`, `resolveAdapterOutputDir`, and the pure helpers around console capture, error serialization, UID generation, stack-trace parsing, BiDi attachment, and screencast finalization. Some helpers are framework-agnostic by nature but used in only one adapter today (e.g. nightwatch's `parseNetworkFromPerfLogs` for CDP perf-log parsing, selenium's `detectRunner`/`captureLaunchCommand`). They stay in their adapter until a second consumer appears; at that point they move to core. ### Trace-format transforms live in `trace`, one layer below `core` `packages/trace` holds the pure transforms that turn captured events into trace-zip content — the zip writer, action events, group paths, frame snapshots, mutations, HAR, sources, transcript. `core` keeps the adapter-side *orchestration and policy* that calls them: `trace-finalizer`, `spec-trace-helpers`, `trace-retention`. The split is not aesthetic. **`backend` may not import `core`** (it would pull framework-adapter logic into the server), but it does need to build a trace on behalf of an adapter that can't — the Python adapter ships no Node. `trace` is the layer both can reach, so it may import `shared` and nothing else; a single import of `core` from it re-creates the cycle the split exists to remove, and ESLint enforces that. The test for a new helper: *would the backend ever need this to build a zip?* If yes, `trace`. If it needs a driver, a framework hook, or a capture session, `core`. ### One DOM snapshot per action, taken before it The per-action trace snapshot is captured in `beforeCommand`, stamped at the previous action's end, so an action's result IS the next action's "before" and every row resolves to a state the driver was idle for. `afterCommand` captures nothing in trace mode. An eager post-action capture — added and removed twice — lands while the screen is still moving, which is the only reason `waitForActionResult` and the `__wdioSnapMark` document tag existed; both are deleted, and both should stay that way. The last action has no successor, so `#finalizePerScenario` supplies it and names its capture after that action — `lastRenderedScreenshot` skips a capture named `__final__`, which is reserved for a session that ran no action at all. What that one capture waits on is `settleAfterLastAction`, and the wait is **gated, not timed**: the drain that runs immediately before it (`captureTrace(browser, true)`) anchors each document once, so `SessionCapturer.replacedDocumentInLastDrain` says whether the last action navigated to a document the session had not seen. No → return at once (the app has been at rest since the last action, so the test's own teardown is the gap that let the paint land). Yes → `waitUntil(document.readyState === 'complete')`, which is only meaningful once you know the document being described is the incoming one: ungated, the OUTGOING document already reports 'complete' right after a click, so a blind poll returns instantly and captures the page the test just left, and the old `body.childElementCount > 0` clause that papered over that made a legitimately blank destination a guaranteed 8 s timeout per test. Native pauses 250 ms instead (no document to poll; a capture at a 0 s gap measures 359–476 KB against 1.87 MB settled on Appium). Residual: a navigation that has not *committed* when the anchor is read still reads as "no navigation" — the same blind spot the deleted document tag had. Don't reintroduce a poll without a gate, or a tag. ### Adapters are thin and isolated Adapter packages own only: - Framework-specific hook registration and lifecycle binding. - Framework-specific driver/browser patching. - Framework-specific config and capabilities. They import from `shared` and `core`, never from each other. They aren't imported by `backend` or `app`. ### Backend and app are framework-agnostic `backend` and `app` import from `shared` only (for contracts) and from each other via the WS/HTTP boundary. Neither imports an adapter package. Framework-specific behavior in the backend is contained in two files: `runner.ts` and `framework-filters.ts`. Both branch on a typed `TestRunnerId` from shared, never on a magic string. The `framework-filters` dispatch is a `switch` over `TestRunnerId` (not a table lookup) so CodeQL's `unvalidated-dynamic-method-call` query trusts the call site. ### Boundaries have typed contracts Every `fetch(...)` and `ws.send(...)` has a typed request/response shape in shared. `SocketMessage` is the canonical WS wire format — receivers narrow on `scope` to get the exact payload type per branch. No `any` crosses a package boundary. When a framework API forces a loosely-typed value (Nightwatch's `currentTest`, Selenium's BiDi events, raw HTTP payloads), the `any` is cast to a typed shape immediately at the boundary, with the cast site documenting why. ### Workspace-internal packages stay bundled `packages/shared`, `packages/trace` and `packages/core` are `"private": true` and never published. Each consumer inlines their code into its own `dist/` at build time. - All three are listed in `devDependencies` with `workspace:^`, never in `dependencies`. Vite and tsup both externalize anything in `dependencies` by default; `devDependencies` is what gets inlined. - None of them is added to a bundler's `external` config. Vite's `external` callback receives both the bare package name *and* the resolved absolute path (e.g. `/Users/.../packages/core/src/index.ts`); a check for only one form silently externalizes the other. - That callback enumerates the private packages, so **adding a fourth one means editing it** — a package missing from the list falls through to the default and is externalized silently, producing a dist that dies at install with `ERR_MODULE_NOT_FOUND`. It is a `PRIVATE_WORKSPACE_PACKAGES` array rather than a chain of `||`s for exactly that reason. Adding a workspace package also means adding it to `pnpm-workspace.yaml`, whose `packages:` list is explicit rather than a glob. - The same callback receives bare relative imports (`./utils.js`, `../constants.js`). A check that allows only `./` will externalize `../`-style imports from subfolders and the dist crashes with `ERR_MODULE_NOT_FOUND` at install time. - `packages/service/vite.config.ts` is the canonical pattern for getting both right. - After any change to a bundler config or build script, `grep -nE "(from|require\()\s*['\"](@wdio/devtools-(core|shared|trace)|.*/packages/(core|shared|trace)/)" packages//dist/*.js` should return nothing. That's how you catch the absolute-path leak. Match on the `from`/`require(` prefix, not the bare package name: `LIBRARY_NAME = "@wdio/devtools-core"` (written into the trace's `context-options`) and `Symbol.for("@wdio/devtools-core/assert-patched")` are inlined string *values* that legitimately survive bundling, so a bare-name grep always reports a false leak. - A **CJS-only dependency must be externalized, not inlined**, or esbuild rewrites its `require` into a shim that throws `Dynamic require of "fs" is not supported` the moment the module loads. Declaring it in `dependencies` is what externalizes it; that is why all three adapters — and now `backend` — list `yazl` there rather than in `devDependencies`. This is the opposite of the workspace-internal rule above, and for the same underlying reason: `dependencies` is externalized, `devDependencies` is inlined. Neither `pnpm build` nor `pnpm test` nor the leak grep notices — every one of them passes on a dist that dies on first import — so `packages/backend/tests/dist-bundling.test.ts` asserts the shim is absent. Bundlers in use: **vite** for `app`, `service`, `script`; **tsup** for `backend`, `nightwatch-devtools`, `selenium-devtools`. - **A published package that ships a BUNDLE declares its build libraries as `devDependencies`.** `app` and `script` each publish a vite build with everything inlined — the app's dist carries no bare import of lit, preact or codemirror, the script's none of htm, parse5 or preact — so listing those under `dependencies` installed a whole toolchain on every consumer that needed none of it. `script` also listed `vite-plugin-singlefile`, which pulls vite, rolldown and lightningcss; `app` listed `@wdio/devtools-service`, which it never imports and which pulls webdriverio. Measured against the registry: installing `@wdio/devtools-backend` cost **338 packages / 264 MB**, against **~85 / ~27 MB** once both were moved. Every adapter paid that; the Python one pays it hardest, since it fetches the backend at runtime. The test is the same grep as above — a bare import surviving in `dist/` means the dependency is real and belongs in `dependencies`; nothing surviving means it was build-time. Note this is the opposite default from the workspace-internal rule: there `devDependencies` is chosen so code is *inlined*, here it is chosen because the code *already* is. ### Separation of concerns within a file Files own one concern: - UI components render. They don't `fetch`, manage WebSocket state, or run business logic. - Controllers and services own I/O and state. They don't render. - Backend route handlers wire requests to services. They don't contain business logic inline. - Reporters report. They don't also resolve sourcemaps, read files, and generate step UIDs in the same module. Mixed-concern files are split as they're touched. The app-side helpers like `contextUpdates.ts`, `runnerCapabilities.ts`, `renderDetailBlock.ts`, `compareUtils.ts`, `suite-merge.ts`, `mark-running.ts`, `run-detection.ts`, and `stepResolution.ts` are all extractions from larger god-files. ### TypeScript - `strict: true` is on (root `tsconfig.json`). - No `any`. If a framework or library forces it, the `any` is isolated at the boundary and cast to a typed shape with a one-line comment explaining why. As of writing, there are no `no-explicit-any` warnings repo-wide. - No `as unknown as X` double-casts unless the reason is documented inline. - `type` for unions, `interface` for object shapes that may be extended. - Names exported from `shared` and `core` are public API of those packages — renames are breaking changes for downstream consumers. ### Naming - One name per concept across the whole repo. The canonical test-status name is `TestStatus` in shared; the sidebar `TestState` is a value-only enum-style accessor over the same string union. - Constants are `SCREAMING_SNAKE_CASE`. Types are `PascalCase`. Functions and variables are `camelCase`. Files are `kebab-case.ts` unless they match a class name (`SessionCapturer.ts`). ### File and function size Soft caps (warnings in `pnpm lint`, not errors): - **File**: 500 logic lines (blank lines and comments excluded). Files growing toward this cap are split as their sections are edited. - **Function**: 50 logic lines. A few declarative blocks (`#getInternals` accessor bags in the adapter plugins) exceed the function cap intentionally — splitting them artificially hurts readability. Those are marked with an inline `eslint-disable-next-line max-lines-per-function` plus a one-line justification. ### Comments - Default to no comments. Names should explain *what*. - A comment is written only when the *why* is non-obvious: a hidden constraint, a workaround for a specific bug, a subtle invariant, behavior that would surprise a reader. - `// TODO`, `// added for X`, `// removed Y`, `// keep in sync` aren't used — the first three belong in git history; the fourth means a single source of truth is missing. - One line max. Multi-paragraph docstrings aren't used. ### Error handling - Validation happens at boundaries (HTTP input, WS messages, framework callbacks). Internal code is trusted. - Errors aren't swallowed silently. `catch` only adds context, then rethrows or logs with enough detail to debug. Empty catches don't appear in production code. ### Dead code Unused exports, unused imports, commented-out blocks, and `_unused` parameters get deleted when discovered. Git history is the safety net for "in case we need it later" code. --- ## Testing The repo uses **vitest** at the root. The current state: 1776 tests across 139 files; thresholds at `vitest.config.ts` enforce a floor of 85/77/86/85 (statements/branches/functions/lines). Coverage is ratcheted upward as gaps close, never downward. ### What gets tested - **`shared` and `core`**: unit tests for every exported function and type guard. These are the foundation; regressions cascade. - **Bug fixes (any package)**: a regression test that fails before the fix and passes after. When a real test is genuinely impossible (e.g. requires a live browser the infra doesn't have), the PR description says so. - **New HTTP/WS contracts**: a test that exercises the contract end-to-end at least once. ### Adapter and backend logic Non-trivial parsing or transformation logic in adapters has unit tests. Hook wiring is verified manually via `examples//`. `backend` and `app` test their non-UI logic (parsers, transforms, state reducers); UI verification is manual. ### Manual verification For UI or runtime changes, `examples//` is the verification harness. Type-checks and unit tests verify code correctness, not feature correctness — claiming a UI change works on the basis of `tsc --noEmit` alone misses the point. When CI can't run an example (no real browser), the PR description says so explicitly. ### Skipping tests that depend on workspace-internal build artifacts A handful of tests need `@wdio/devtools-script` to be built first (the browser-injected bundle). CI test jobs sometimes run before that build step; those tests gate on `it.skipIf` after probing `createRequire(import.meta.url).resolve('@wdio/devtools-script')`. Locally they run normally. --- ## Workflow ### When adding code The decision tree from [ARCHITECTURE.md "Where things live"](./ARCHITECTURE.md#where-things-live) is the starting point. The general shape: - Shared concept → `shared`. - Pure transform producing trace-zip content → `trace`. - Framework-agnostic capture/reporting logic → `core`. - Framework-specific glue → the matching adapter. - Server route/WS handler → `backend` (contract in `shared` first). - UI → `app`. - Code that runs in the browser under test → `script`. When the right place is ambiguous (something between `shared` and `core`, or between `core` and an adapter), the question that resolves it is: *who else would want this?* If the answer is "any future adapter would," it's `core`. If "only the framework with X-specific API does," it's the adapter. If "the backend would, to build a zip for an adapter that can't," it's `trace`. ### While editing - Boy-scout rule applies: when touching a file or section, leave it more aligned with these conventions than it was found. Touch a duplicated type, consolidate it into shared. Touch a section of a god-file, split that section out. Touch a magic-string framework check, replace it with `TestRunnerId`. The cleanup scope matches the change scope — don't rewrite the whole file, but don't leave a clear convention violation in lines just touched. - New code doesn't introduce violations to match existing style. Where existing style violates these conventions, that's documented debt (§ Known debt), not a template. ### Before pushing - `pnpm build`, `pnpm test`, `pnpm lint`. Don't push red. - A changeset for a published npm package, or a `changes/` fragment for the Python adapter — see § Releasing a change. - For UI or runtime changes: verify in `examples//`. - Deeper security findings (taint flow, polynomial-redos with adjacent quantifiers) surface on the PR's CodeQL scan; review and fix those before merge. ### Releasing a change Two mechanisms, and the Python one exists because the npm one cannot reach it. Changesets discovers packages through the pnpm workspace and identifies them by `package.json`; `packages/selenium-devtools-py` is in neither, so a changeset naming `selenium-devtools-py` does not degrade — it raises "not in the workspace", fails `changeset version`, and takes the npm release for every other package down with it. - **Published npm package changed** → `pnpm changeset`, committed as `.changeset/*.md`. - **`packages/selenium-devtools-py/src/` changed** → a fragment under `packages/selenium-devtools-py/changes/`, frontmatter carrying the bump level alone (`patch`/`minor`/`major`). `python.yml` refuses a branch that changes `src/` and documents nothing. A direct `CHANGELOG.md` edit satisfies it **only until the first release** — before one there is nothing to bump from and the pending entry IS the changelog section; after one (detected by a `py-v*` tag existing) it documents the change but bumps nothing, so the release would find no fragment and republish a version the index already holds. Neither is hand-versioned: both assemble the version and the changelog at release. The Python release additionally consumes its fragments, bumps `__version__` (the single source — `pyproject.toml` reads it via `dynamic = ["version"]`), and tags `py-v` **after** a successful publish, so the tag is an output pointing at the published tree rather than an input naming a version nothing has computed yet. `BACKEND_NPM_VERSION` is the backend a `pip install` user actually runs, so the npm release goes first; `release.yml` opens the pin bump as a PR, and `scripts/check_backend_pin.py` refuses a PyPI publish whose pinned backend cannot serve the contract. That PR carries its own `changes/` fragment, because `python.yml` refuses a branch that changes `src/` and documents nothing — a pin-only PR would fail its own CI. It is a PR rather than a push because the pin is a claim about a *published* artifact and `check_backend_pin.py` is what adjudicates it; merging one queues a `patch` for the next PyPI release rather than bumping `__version__` there and then. Raised with `GITHUB_TOKEN` it arrives with **no checks at all** — GitHub suppresses workflow runs for events its own token raises — so either set `PIN_BUMP_TOKEN` or close/reopen the PR to get CI onto it. ### Commits - Small, focused. Don't bundle unrelated changes. - Imperative mood. The commit message explains *why*; the diff shows *what*. - New commits, not amends to pushed/shared commits. - No `--no-verify` to skip hooks. If a hook fails, the underlying issue gets fixed. ### PRs - One concern per PR. A refactor and a feature are two PRs. - A PR touching more than one adapter package answers in its description: *why isn't this in `core`?* ### Documentation - User-facing docs live in two places that must stay in sync: this repo's `README.md` (+ per-package READMEs) and the **WebdriverIO devtools webpage** (`website/docs/devtools/**` in the `webdriverio/webdriverio` repo — e.g. `wdio/TraceMode.md`). When a change adds, removes, or alters user-facing behavior (a new option, CLI, flag, output, or workflow), update the README here **and** mirror it to the matching webpage doc in the same change. A docs PR that updates only one side isn't complete. --- ## Known debt Documented divergences from the conventions above. They exist today as debt to be paid down, not exceptions to the rules. Each change reduces this list; new violations don't get added. ### Architecture - `replaceCommand` has two semantics — Selenium mutates in place (preserves `_id`/`id` for chained calls); Nightwatch splices and reissues. Both call the same `core/suite-helpers` factories; the storage strategy stays adapter-specific because runner integrations differ. Could be unified by parameterizing the policy if the divergence ever causes a real problem. - `patchNodeAssert` (via `core/assert-patcher`) is now wired in all three adapters, default-on behind each adapter's `captureAssertions` option (opt out with `captureAssertions: false`). Framework matcher libraries differ: Service taps expect-webdriverio's `beforeAssertion`/`afterAssertion` hooks so passing+failing matchers render as `expect.*` actions (mechanism in the assert-capture entry below); Nightwatch native `assert`/`verify` and Selenium's `node:assert` also surface passing+failing rows via their reconcile/patch paths. The remaining gap is Selenium's jest-style `expect()` (and chai): jest/vitest expose no pass+fail assertion hook, so only failing matchers surface there. - BiDi is auto-attached in Service and Selenium; Nightwatch is opt-in via `bidi: true` and requires `webSocketUrl: true` in capabilities. - Retry-aware trace policies share one mechanism: adapters feed a per-attempt **outcome ledger** (`core/attempt-tracker.ts` `TestAttemptTracker.recordStart(uid, specFile)` + `recordOutcome(uid, state, attempt)`) keyed by the **retry-stable** uid, and the finalizer reads the scoped views (`all`/`forSpec`/`forTest`) so `trace-retention.ts` evaluates **group-by-test** — `retain-on-failure` keys on each test's *final* attempt (no over-retaining a fail-then-pass), `retain-on-first-failure` on *attempt 0*. `recordStart` on a second attempt stamps the prior attempt `failed` (a retry only follows a failure), which corrects runners that swallow the intermediate failure — e.g. Mocha via a `--require` plugin never surfaces the retried attempt's failure, so the ledger would otherwise see `[passed, passed]`. An empty scoped view falls back to `testMetadata` (never fail-open-retains). - **WDIO + Selenium: verified end-to-end** (manual fail-then-pass runs: retained under `retain-on-first-failure`, dropped under `retain-on-failure`). - **Nightwatch: `retain-on-failure` works; the other retry-aware policies degrade.** Its `--retries` re-runs the testcase *internally* without re-firing the plugin's per-test hooks, and the per-testcase results carry no attempt/retry field (retries live only in undocumented, version-varying Nightwatch internals — `suiteRetries.testRetriesCount` / `reporter.testResults.retryTest`), so the ledger sees only the final attempt for the `describe/it` and exports-object interfaces. Cucumber scenarios expose per-scenario hooks, so the feed captures their attempts. Not cleanly fixable without depending on those internals. - WDIO `specFileRetries` spawns a fresh worker per retry, so cross-process attempts aren't in the (process-scoped) ledger. - Run identity across worker sockets is env-propagated. `core/run-id.ts` `resolveRunId()` publishes `DEVTOOLS_RUN_ID` (`RUNNER_ENV.RUN_ID`) and every worker socket carries it as `?runId=` (`WORKER_WS_QUERY`), so the backend keeps accumulated run state when the *next spec's* worker connects and wipes it only for a genuinely new run. Without it every connect read as a new run: Preserve & Rerun 409'd for every spec except the last one that ran, and a dashboard opened mid-run replayed only the current spec. The WDIO service stamps it in the launcher's `onPrepare`, before workers fork, so all workers of one run agree; single-process adapters self-stamp on first use. **Gap: multi-process parallel runs in Selenium/Nightwatch** (jest/vitest workers, nightwatch `test_workers`) load the plugin per worker with no launcher-side hook to stamp first, so each worker generates its own id and still reads as a new run — the pre-fix behaviour, not a regression. Deriving the fallback from `process.ppid` would group those siblings, but would also make two sequential single-process runs share an id and inherit each other's state against a standalone dashboard, so the per-process fallback stands. - **A rerun template selects EITHER by name pattern or by exact id, and the two cannot share a slot.** `shared/src/runner.ts` `RERUN_SLOT` names both, and `backend/src/runner.ts` `#resolveGenericCommand` branches on which one the adapter's template carries: `{{testName}}` is filled from `label`/`fullTitle` through `escapeFilterRegex` because mocha `--grep`, jest `--testNamePattern` and cucumber `--name` all match by regex; `{{testId}}` is filled from `uid`, shell-quoted and **never escaped**, because pytest selects by nodeid and matches it literally. Measured: `pytest 'test_thing\.py::test_a'` collects nothing and exits 0, so the escaped form fails as a rerun that appears to have run and passed. Which slot a payload can service also decides `isTargetedRerun` — an id template needs a `uid`, not a label. - A pytest nodeid addresses a file, a class or one test in one syntax (`file.py`, `file.py::Class`, `file.py::Class::test`), and the Python adapter's uids already *are* nodeids at all three levels, so one slot covers every row the tree offers — no per-level filter flag and no cucumber-style feature special case. Verified end-to-end: substituting a test nodeid collects 1, a file nodeid collects 3. - `selenium-devtools-py/src/selenium_devtools/rerun.py` derives both commands from pytest's own view of its invocation (`config.invocation_params.args` plus `config.args` for which of them were positional) rather than parsing argv itself: inferring positionals needs a table of every option that takes a value, and dropping a value while keeping its option makes that option swallow the appended id. Capabilities are **derived from which commands got built**, never declared — the backend's fallback for a rerun it was given no command for is the wdio binary, so an advertised-but-unserviceable control is worse than an absent one. A plain script publishes a launch command only and advertises Run-all alone. - Selectors are stripped from a targeted rerun (`-k`, `-m`, `--deselect`, `--lf`/`--ff`/`--sw` family, `-n`/`--numprocesses`/`--dist`): the rerun already names its test, so a surviving filter can only narrow further — usually to nothing, which pytest reports as a clean exit. Positionals go too, or a rerun's own child would union the inherited nodeid with the next one and each generation would run one test more. The xdist flags also go because each worker would connect under its own run id. - The rerun spawns in pytest's **rootdir** (`RUNNER_ENV.RUNNER_CWD`, stamped before the backend is launched so its process inherits it): a nodeid is reported relative to rootdir while a positional path resolves against the process's cwd, so anywhere else makes every nodeid a path that does not exist. The launch command's positionals are absolutised for the same reason. The variable is *replaced* on a second `enable()` in one process but only while it still holds **the value we wrote** (tracked, and deliberately surviving `reset()`): our leftover would otherwise spawn the next run's reruns in the previous project, while a value someone exported since is an instruction. A boolean "we wrote it once" cannot serve both — it says nothing about whether the current value is still ours. The remaining ambiguity is accepted and untouchable: a caller who exports the *same* path we already stamped is byte-identical to our leftover in the only channel there is, so that override is replaced; pinning a directory across runs works by exporting it before the first `enable()`, which is never claimed as ours. Residual: an option carrying a *relative* path (`-c`, `--junitxml`) resolves against rootdir on a rerun, and an already-running dashboard keeps the directory it was started in. - **Preserve & Rerun needed no adapter code at all — the blocker was the capability gate.** The button renders on `hasFailed && !runDisabled`, so an adapter advertising no run capabilities never showed it; `baselineStore` snapshots from the stream every adapter already sends, and `toMs` accepts the ISO strings Python puts on `SuiteStats.start`/`end`. Verified end-to-end against a live backend with frames built by the adapter's own `frames`/`SessionCapturer`: both attempts carried their commands, console and network with distinct windows and correct states. What was missing was *coverage*: `preserveBaseline` appeared in no test in the repo, and it is the one flag deciding whether the request that exists to compare wipes what it means to compare against. Preserved attempts live in `#baselines`, outside the `#activeRun` accumulator, which is why a rerun's new run id resets capture without losing the snapshot — and why the order matters: preserving *after* a new run connects is a deliberate 409. - **A rerun does not travel down the worker socket.** `POST /api/tests/run` spawns a fresh process; the socket carries only `clientConnected`/`clientDisconnected`. So the single-`workerSocket` limitation is about which process the *dashboard state* belongs to under `pytest -n`, not about routing the rerun. - **A spawned rerun must be pointed back at the backend that asked for it, or it reports into a dashboard nobody is looking at.** `REUSE_ENV` (`DEVTOOLS_APP_REUSE`/`_HOST`/`_PORT`) is how the backend does that, and an adapter that ignores it launches a *second* backend and a *second* window: measured on the Python adapter, a rerun opened a new dashboard carrying the rerun's data while the window the user pressed Rerun in stayed as it was — which reads as a rerun that captured nothing. `backend.py` `reuse_target()` now attaches to it **ahead of `DEVTOOLS_PORT`** (that variable is an ambient preference inherited from the parent; the handshake names the backend that requested *this* run), and the window gate lives in `lifecycle.auto_open_enabled()` rather than at the `enable()` call site so it is directly testable. An incomplete handshake deliberately still opens a window — no usable target means the child launched its own backend, and then the window is the only way to see it. - A plain script's tree is one synthetic suite holding one synthetic test, and both denote the whole run, so its launch command doubles as its rerun template (no slot — the backend substitutes nothing) and all three controls are honest. Refusing the row-scoped ones instead would disable the button beside the only row the tree has. - **Two unrelated events share the `clearExecutionData` scope, and the receiver cannot tell them apart from the uid.** A run STARTING (`backend/src/index.ts` `handleTestRun`, one per `POST /api/tests/run`) and ONE ENTRY resetting inside a run already in flight (`nightwatch-devtools/src/cucumber-lifecycle.ts`, which re-emits a scenario suite and must not wipe its siblings) arrive under the same scope with the same shape. The app inferred the difference by comparing the uid against `rerunState.activeRerunSuiteUid` — a latch that outlived its rerun, so the *next* run start at a different scope read as a child clear of the last one and **skipped its wipe entirely**: rerun a suite, then the file or Tests, and the Actions/Console/Network tabs kept the previous run's rows and grew with each rerun. `ClearExecutionDataWsPayload.runStart` now states it on the wire (it has to be on the wire, not local to the clicking window — popouts see only WS events), and the app clears both latches when it is set. A backend test asserts the flag actually ships: the app-side fix reads it, so dropping it would restore the bug with every app test still green. - Still open, same class: `app/src/components/browser/snapshot.ts` `#videos` is only ever pushed to, so the screencast "Recording N" dropdown accumulates every session of every run for the life of the page (observed at 17). That component listens only to the `screencast-ready` window event and never learns a run started. - **A rerun's process collects a SUBSET, so anything it derives from "this collection" is wrong for the tree it merges into.** Three bugs of that one shape, all found by rerunning a single pytest test. The third is the one that shows the rule has a limit: (c) the **launch command** — what Run-all spawns — was built from the child's own invocation, which the backend had narrowed to a single nodeid, so one targeted rerun rescoped Run-all to that test permanently and the tree kept showing only what that child collected. Unlike (a) and (b) this is not recoverable inside the child: its arguments no longer mention what it was narrowed from. The original travels down instead, as `REUSE_ENV.LAUNCH_COMMAND` beside the rest of the reuse handshake; the *rerun template* stays locally derived, since only this process can say how its own interpreter selects a test. **Fixed in the Python adapter only** — `selenium-devtools/src/rerunManager.ts` still derives its launch command from `captureLaunchCommand()`, i.e. from the child's own argv, so a mocha/jest rerun carrying an inherited `--grep` republishes that as Run-all. It already strips those filters out of the *template* for this reason; the getter is what is left. A second consumer makes the inherit-or-derive resolution itself core's, with `captureLaunchCommand` staying adapter-local. The other two: (a) `SuiteStats.order` — which `test-entry-state.ts` `orderedChildren` sorts a suite's tests and child suites by — was pytest's `enumerate(session.items)` index, so a rerun restamped its one test as position 0 and the row jumped above the class it was written below. It is now the item's **source line**, a property of the test rather than of the collection; within a module pytest collects in definition order, so the two agree wherever both are meaningful (a plugin that reorders collection is the exception, and there the line is the more stable answer anyway). (b) `suite-merge.ts` `resetStaleChildrenOnRerun` flipped every settled child *suite* to `pending` whenever an incoming suite arrived `pending` — but a single-test rerun re-emits the parent as `pending` carrying only the one test it collected, so a sibling class suite was set spinning and never reported again, keeping the spinner for the rest of the session with all of its own tests still green. `mergeTests` already froze sibling *tests* on `activeRerunTestUid`; that guard now covers child suites too. A suite on the path to the target is unaffected either way — it re-reports its own state. - **A trace archive is a full recording of the page, and there is no redaction policy anywhere in capture.** Whatever the run put on screen or typed is in the zip, usually several times over: measured on the Python login example, its demo credential appears ~103 times across six places — the page's own displayed text (90x, the-internet prints it), the DOM mutation stream, the `Element.fill` command args, the transcript, the captured test source, and `*-elements.json`. `shared/element-scripts.ts` blanks an `` value, which is worth having because nothing downstream reads that field, but it removes **2 of those ~103** and closes nothing on its own. `buildElementScripts` now projects a captured record down to what is actually read (`selector` + `boundingBox` + context), so `value` and `href` leave the archive entirely — justified as dead data, **not** as a redaction: the same archive still carries 15 hrefs in `trace.mutations` independent of `elements.json`, and 29 `value` attribute mutations recording a typed string keystroke by keystroke (`t`, `to`, `tom`, ...). `@wdio/elements` keeps returning the full `BrowserElementInfo` from its own live call, which is its documented API. A real policy has to act at the collector and the command-arg serializer — a masking-selector or `maskInputs` option — not at one resource. Until then, treat a trace zip as sensitive as the run that produced it. - **An `ActionSnapshot` carries no session identity, in any adapter.** `shared`'s type has never had one and `core/action-snapshot.ts` records none, so per-action captures from two concurrently-driven sessions land in one list and are resolved purely by the command's completion timestamp. `claimAfter` is an exact keyed lookup, so the window is narrow — two commands completing in the **same millisecond**, where `trace-frame-snapshots.ts` breaks the tie by "keep the richest capture" (largest screenshot), which is session-blind — plus the documented `latestAtOrBefore` fallback for a command that took no capture of its own. Python is not worse than the JS adapters here and leans on that fallback less, since it stamps each snapshot with its own command's `row["timestamp"]`; reaching the failure at all needs threaded drivers in one process (pytest's function-scoped fixtures are sequential, and `-n` is multi-process). Fixing it is a shared-contract change: `sessionId` on the snapshot, and an index keyed by the pair. - **Chrome discards all WebDriver-synthesized input to a tab after a breached credential is submitted.** The first time a test types a `(username, password)` pair that Chrome's password-leak check finds in a breach corpus into an `` and submits a form whose destination no longer shows that login form, Chrome queries `passwordsleakcheck-pa.googleapis.com` and ~0.3-0.9 s later stops delivering **all** synthesized input — mouse *and* keyboard — to that tab. chromedriver returns HTTP 200 for every subsequent Element Click / Send Keys; nothing reaches the page. Untrusted JS (`element.click()`) still works and direct CDP `Input.dispatchMouseEvent`/`dispatchKeyEvent` are equally dead, so this is Chrome, not chromedriver and not our capture. `tomsmith` / `SuperSecretPassword!` — the-internet's demo credential — triggers it; changing only the *username* does not, nor does a random password. - **Workaround: add `--host-resolver-rules=MAP passwordsleakcheck-pa.googleapis.com 127.0.0.1` to the browser args.** Every example that submits the demo credential carries it — WDIO, Nightwatch, and both Python ones (`login.py` was missing it and its logout click silently did nothing, which is exactly the symptom). Verified 3/3 on the WDIO mocha example and on the Nightwatch example, where it also fixes the **within-one-test** logout click that a session reset never could. `--guest` also works (3/3); `--incognito` works at the raw-WebDriver level but WebdriverIO rejects it at session creation; disabling the password manager via `prefs` does **not** (6/6 still fail). - **Not a version regression, not headless-specific, not the site, not "the Nth navigation".** Measured identically on Chrome 149.0.7827.155 / 150.0.7871.124 / 151.0.7922.77 / 152.0.7977.30 with matched chromedrivers (5/5 each), headless and headed, and on a purely local two-page static form. It fires **once per browser profile** on a wall clock — a liveness probe that never navigates again goes dead 904 ms after the submit — so the historical ~25% intermittency was the race between the next input command and that round trip. Do **not** pin `browserVersion` to 149; every part of the earlier "Chrome 150 regression, fixed in 151" attribution is contradicted. - Minimal reproduction (own HTTP server, raw `fetch` to chromedriver, no repo, no client library, no framework) is in the session scratchpad as `minimal-repro.mjs`; it is what an upstream chromedriver bug report needs. If a session is already stuck, navigating away and back or opening a new tab restores input (4/4 each); `refresh()`, ESC, JS focus/blur and a 10 s wait do not (0/4 each). - **Live mode has no per-action DOM snapshot, so its replay is only as fresh as the last drain.** Per-action snapshots cost two injected scripts plus a screenshot and stay trace-only; all three adapters instead drain the collector after a command that could have moved the page. Service: `#drainAfterLiveCommand`. Selenium: `commandPostActions.ts` `warrantsLiveDrain` + `SessionCapturer.drainAfterLiveCommand`, the same deny-list shape over its own command vocabulary plus `mapAssertCommand` (a node:assert row never reaches the browser) — the predicate is *not* in core because the vocabularies are per-framework and only two `includes` calls would be shared. Without it Selenium drained only at navigation, and that hook is deferred behind an injection and a 500 ms settle: measured on the login example, **2 mutation entries and 2 anchors for a 16-row run**, with the page test 1 spent most of its life on never anchored, so all 11 of its rows replayed the page the test *ended* on (2 → 24 entries, 2 → 3 anchors, 0 → 21 field-state mutations after the fix). Selenium's drain is serialized on a tail because the driver patcher does not await `onCommand`, and the app scans the mutation stream in order and stops at the first entry past a row's window — an overtaken batch strands every row after it. - **A live client receives commands in ARRIVAL order, and one consumer assumed timeline order.** Nightwatch withholds native asserts until their outcome is known and flushes them in one batch at test-end (BDD fires `afterEach` once per *module*, so a whole module's asserts arrive after every driver row). The display list already sorts and `utils/elapsed.ts` already treats capture order as untrusted, but `app/src/components/browser/mutation-at-command.ts` bounded a row's DOM by `commands[idx + 1]` — the *array* neighbour — so a row was bounded by a time before it ran: measured, the run's last `waitForElementVisible('#username')` took its bound from an assert that had run 7.6 s earlier and replayed `/secure`. It now orders by `(startTime ?? timestamp, sequence ?? 0, array index)` — the key `buildActionEvents` uses, with the index last so a chronologically ordered array (every trace) resolves to its own successor. Measured: live **5/21 → 1/21** rows on the wrong document, trace **0/21 with 0/21 selections differing**. Live-mode anchoring itself is not the gap: `processTracePayload` sends mutations upstream unconditionally, and a live run streams one anchor per document visited. - Residual: a submit click whose end, destination birth and next command's start land in the **same millisecond** still shows its pre-navigation page; and the app deliberately leaves the **last** row unbounded, rendering the newest DOM. - **The Nightwatch filmstrip was never losing frames across `browser.end()`.** Instrumented: 155 poll ticks, 129 frames appended, the 26 skipped only in the null-session gap, and the login-page image present once per session in the export. It works because Nightwatch mutates `sessionId` in place on one `browser` object and the screenshot probe reads it fresh. An observed 17→6 drop was `thinScreencastFrames`' byte-identical dedup meeting a different failure profile — 10 of the 17 were a **blinking text caret** captured while a `waitForElementVisible` sat 5 s on a focused form. Don't read a low polling-mode filmstrip count as frame loss without a sha1 histogram of the emitted events. What *was* real: `#emitTestArtifacts` read `recorder?.frames ?? filmstripFrames`, so once a recorder existed it dropped every frame from before a session change. - Eager per-test trace slice (Nightwatch + Selenium) can drop an action snapshot whose fire-and-forget capture hasn't resolved by `afterEach` / scenario end — the slice is written from whatever snapshots exist at flush time. The WDIO service is immune because it awaits each snapshot inline before flushing. - **A command row is stamped at COMPLETION, and the DOM anchor carries the document's own birth time.** These two together are what make the replay line up; both adapters got them wrong in the same way and the fix is symmetric. (a) `selenium-devtools/src/driverPatcher.ts` and `nightwatch-devtools/src/helpers/browserProxy.ts` both ran their capture at completion but stamped `timestamp` with the *invocation* clock, keeping the invocation time as `startTime` only after this fix. The page-side mutation stream is on real time, so an invocation-stamped row ended before its own effect landed and replayed the page from before it — the `#username` fill rendered an empty field, the `#password` fill rendered only the username, and a navigation row rendered the page it had just left. Rows also now span their real duration instead of a synthetic 1 ms. (b) `collector.captureCurrentDom` (the only producer of a mutation with a `url`) stamps `performance.timeOrigin`, not the drain clock. A drain is forced from Node whenever a collector might be fresh, which is always after the navigation — a round trip at best, a whole page load at worst — so drain-stamping put the anchor after several later actions (measured: 9/15 Selenium and 8/15 Nightwatch rows on the wrong DOM). With both in place a navigation row ends after its destination document was born, so the anchor needs no repositioning at all. - `core/trace-mutations.ts` `reattributeDomAnchors` remains as a narrow backstop for the one case the stamps can't cover: an anchor born *after* the last logged command, i.e. a click whose navigation commits once the click has already returned. It snaps such an anchor to the newest logged command, but **only when no logged command completed after it** — if one did, that command's row already resolves the anchor and pulling it earlier mis-credits it to a preceding action and steals the new page's DOM from rows still on the old one (measured: a 206 ms pull moved `/login` onto two rows that were on `/add_remove_elements`). Anchors are only pulled earlier, never past the newest timestamp already in the stream, or replay would apply the outgoing document's refs to the incoming tree. - Residual, accepted: Nightwatch's `click` resolves *before* its navigation commits (measured 5 ms), so a submit-click row can still show its pre-navigation page. Selenium is immune — its click waits for page load. Not worth another heuristic; every heuristic tried here regressed a different row. - **A pushed screencast needs bounding at both ends, and the obvious bound biases toward the end of the run.** Per-command capture is self-limiting — one frame per command, so the test's own length caps it — which is why `selenium-devtools-py` had no frame cap at all. Chrome's `Page.startScreencast` removes that property: `cdp_screencast.py` subscribes over a websocket of its OWN, which is a *different connection* from the session's command channel and therefore safe where the poll thread the module's docstring warns about was not. Every frame must be acked (Chrome sends nothing after an unacknowledged one, so a missed ack ends the recording rather than degrading it), and the rate is thinned at the source with `every_nth_frame` rather than buffered and discarded here. - The buffer cap then needs care. Halving the buffer and keeping first/last — core's documented `maxBufferFrames` shape — drifts toward the run's end, because each decimation thins what is already held while new frames keep arriving unthinned: measured on a 40-frame run at a cap of 6, it kept frames 0, 1, 35, 37, 38, 39, i.e. the last moments and nothing from the middle. `_buffer` therefore thins the INCOMING frames by the same factor it has halved the buffer (`_stride` doubles per decimation), giving 0, 1, 11, 23, 31 for the same run and, at the real 2000 cap over 12000 frames, 1503 frames with 751 from the middle half. Thinning then costs the END of the run, because the last frame offered is only kept when the run happens to stop on a stride position — 41 frames at a cap of 6 ended on frame 31, eight frames stale, and 12000 at 2000 kept its last only because the two aligned. The newest skipped frame is therefore HELD rather than dropped and folded in by `_keep_tail` when the recorder stops (`finalize` stops before reading the buffer), so the video always ends where the run did — which is the part a failure is inspected for. Asserting only the endpoints does not catch this — tail truncation also leaves frame 0 plus whatever arrived since the last decimation, so the test has to assert something from the *middle* survives. - **`driver.start_devtools()` cannot be used for this, because selenium caches ONE `_websocket_connection` per driver and hands it to whichever of BiDi or CDP asks first.** The adapter attaches BiDi before arming the screencast, so `start_devtools()` returned the BiDi socket and `Page.startScreencast` reached a BiDi endpoint: measured, `unknown command: Unknown command 'Page.startScreencast'`, followed by `BiDi command has no 'params' of type dictionary: {"method": "Page.stopScreencast"}` — that second line being the proof of which endpoint it was, and coming from a `stop` the failed start should never have sent. BiDi carries console and network, so it keeps the shared connection and the screencast opens its own. - Resolving that endpoint needs BOTH routes. `se:cdp` is a **Grid** capability and is absent for a locally started chromedriver — the common case, and the one the demo runs — so without the `debuggerAddress` → `/json/version` → `webSocketDebuggerUrl` lookup that selenium's own `_get_cdp_details` performs, push mode would decline on every local run and the feature would be dead code. Done with stdlib urllib rather than that private method, since selenium moving internals is what broke network capture in #293. An `se:cdp` equal to `webSocketUrl` is rejected as the BiDi socket. - **Performance timings ride on the command ROW, not a scope of their own** (`CommandLog.performance`, plus `cookies`/`documentInfo`/`result`), so the row is sent when the command completes and sent again under `replaceCommand` once the page has answered — which is why `capture_command` returns the row it built and `send_replace_command` keys on its `timestamp` rather than the per-process `id` counter. Python does **not** sleep before reading, where the JS adapters wait 500 ms: their navigation command can resolve before the load event, selenium's `get()` returns after it, and a sleep on this thread would be a real delay in the user's test rather than a detached await. A read that lands early anyway carries no `navigation` entry and is discarded rather than replacing a good row with an empty one. The read goes through `_guarded_execute_script`, or it lands in the same `execute` hook it was called from and shows up as an `executeScript` row beside every navigation — a fake driver whose `execute_script` does not route through `execute` cannot catch that, and did not. - Per-command screenshots keep being taken for the command ROWS while a stream is live, but stop feeding the video: the pushed frames already cover the timeline and interleaving would duplicate one of them a few milliseconds off. - **A drain must anchor the document it reads, and the flag for that has only ever had one value.** `core/script-loader.ts` `collectorDrainExpression(forceAnchor)` prepends `captureCurrentDom()` so a freshly injected collector's *async* initial anchor is not lost: the collector schedules it after `waitForBody`, so a drain issued right after a navigation beats it, reads an empty buffer, and the destination's buffer then dies with the page — leaving the navigating action with no DOM. Every production caller in both JS adapters passes `true` (selenium's `drainAfterLiveCommand`, its re-inject-after-navigation and teardown paths; nightwatch's five sites), so the `false` default is vestigial. Python's drain read `getTraceData()` with no anchor at all, which is the same missing backstop the preload does not cover; `selenium-devtools-py/src/selenium_devtools/snapshot.py` `_DRAIN_SCRIPT` now forces it **unconditionally and carries no flag** — one setting is not a knob. Forcing is free after the first anchor of a document because `packages/script` guards `captureCurrentDom` with an `#anchored` flag that deliberately survives its `reset()`, which is why selenium anchors on every live command and still emits ~3 anchors across a 16-row run rather than 16. - **Document-start injection is what removes the whole race class; everything else is reconstruction.** `