--- name: interlinked-verify description: "Run `interlinked verify`, understand the PostToolUse quality checks, and land multi-file edits through the content gate. Load this when you want to check your changes (`interlinked verify` — the on-demand whole-project check run), when a `pre_block` check refused an edit, when you need to land a cross-file refactor without transient tsc errors (`interlinked write --batch` / `multi-edit` / `verify-changeset` and the exporter-before-importers rule), when deciding whether a finding is default-gate or advisory, or when you need to know where to put probe/scratch scripts (`interlinked scratch`). Verify reports ordinary findings with exit 0; an unavailable/deferred run exits nonzero because no verdict exists." --- # interlinked-verify — check your work & land edits through the gates ## Claude compiler batches After a Claude session has delivered a native `PostToolBatch`, declared single-file edits queue TypeScript checking until that boundary. Related edits in the same model message are checked against the completed tree, once per project. Existing atomic multi-file calls retain their shared check path. Other clients, unknown writers, and Claude sessions that have not demonstrated the boundary keep per-edit checking. Security, content guards, lint and hard complexity caps still apply per edit. Final compiler errors arrive as batch context so Claude can repair them; a native batch `block` would cancel its loop. Unresolved errors and unavailable compiler work remain in `.interlinked/compiler-batches/` and are checked before Stop and commit, including after daemon restart. Do not delete this state to bypass verification. A queued edit has no TypeScript verdict and must not be described as fully checked. Explicit verify and commit checks remain necessary for the required repository validation scope. Runner summaries from `rg -n`/`grep -n` are recognized with a single numeric line prefix. A failure in any recognized summary takes precedence over a pass; shell success after a pipe/trailing command alone still does not prove that tests passed. ## Repeated implementation advisory `repeated_implementation` compares same-file Python and JS/TS function bodies by AST structure, with a minimum of five statements. Tests, generated files, small delegates and functions containing nested implementations are excluded. This is conservative structural matching, not proof that two contracts are equivalent. It names the matching functions and source lines, shows literal values, and suggests a shared operation when the contracts match. Do not compress lines or automatically extract helpers to satisfy it. Preserve correctness and independently evolving behavior. The daemon reports new or changed groups once per session after all files in the observed tool changeset have landed. Unchanged groups remain in structured check results. Stop rescans current touched files through the existing digest and repeat-Stop filter; `verify --all-checks` reports the full advisory inventory. Default verify skips this check. It never blocks an edit or Stop and never creates a completion obligation. A daemon restart can repeat advice. Separate tool calls are separate observed changesets. Python uses an isolated, bounded `python3` AST parser and does not execute candidate code. Missing parsers, invalid syntax and oversized detector input return NOT CHECKED, not a clean duplication verdict. Initial scope is whole functions within one file; cross-file matching and arbitrary repeated sub-blocks are not covered by this check. On the first authored edit per language/session without a detected test layout, PostToolUse probes Python/JS/TS/Rust/Go runner readiness and supplies setup/test guidance. It does not install packages or hard-block absent tests. Inspect custom layouts before adding public-contract assertions, then execute them. Explicit `tests readiness` rechecks changed prerequisites; the automatic probe does not run on every edit. For Cowork, load **interlinked-cowork**. `cowork verify ` runs only tsc, biome and gitleaks with before/after workspace hashes; skipped checks remain unmeasured. It is not the full `verify` command or a per-edit ratchet. Run configured repository checks on the filesystem holding the actual version being reviewed. The optional Bash `tsc` accelerator recognizes executable `tsc` / `npx tsc` segments in flat, quote-aware command lists. It preserves quoted arguments and rewrites only the matched invocation. Search patterns, look-alike executable names, comments, multiline commands, substitutions, and shell grouping fall through unchanged; run the original compiler normally in those cases. The file-mode Rust formatter reads the nearest `Cargo.toml` only within the project root. A neighboring directory with a shared name prefix is outside that scope; unreadable manifests are reported instead of guessing an edition. Python behavioral suites run from the resolved project root; companion filenames are not a sound dependency boundary. Directory-name prefixes never establish confinement. Missing-companion guidance describes a naming-convention observation, not proof of missing behavioral coverage. Inspect the project's test layout before adding language-appropriate public-contract tests. Interlinked gates edits at **three moments**, and they run different check sets: - **PreToolUse content gate**: real agent Edit/Write calls run deterministic `pre_block` checks without synchronously launching biome/tsc on the daemon event loop; those external overlays are recorded as scheduled for applicable file types and run asynchronously after the write. Routine scheduling produces no model warning; unavailable or failed checks still report **NOT CHECKED**. Python/Go/Rust edits do not receive JS/TS overlay scheduling. Transactional CLI paths (`interlinked write` / `verify-changeset`) still run `pre_block → biome → tsc` and fail closed. `interlinked multi-edit` uses the same shared content gate. - **Other PreToolUse guards** (real Edit/Write only): function tokens, coverage, cyclomatic, CRAP, baseline — see **interlinked-quality-gates**; package/allowlist — see **interlinked-supply-chain**. - **PostToolUse** (after the write lands): external tools (tsc/biome/eslint/semgrep/gitleaks/…) plus the inline check registry. Findings arrive after the write; default-gate errors can return blocking feedback requiring repair, without rolling back the tool. Bash and unknown writer tools are routed by their observed filesystem ChangeSet, not merely command parsing. PostToolUse keeps full findings in the check-results ledger while presenting compact diagnostics. Repeated advisories are acknowledged per session, check and file; a successful scoped recheck clears that acknowledgment. Pre-existing TypeScript findings show a count and the evidence path. "Newly observed" means changed since the previous compiler report, not proven caused by this edit. Changes observed outside the tool's declared write targets are labeled "writer unknown"; TypeScript findings there are workspace feedback and do not block the observing call. Security checks and workspace obligations still apply. For test evidence, a recognized runner summary can establish the outcome of a piped run. An unsummarized pipeline's final exit code alone cannot prove the test process passed. For scratch probes, reading a repository file and writing an unrelated temporary fixture does not establish a patch applier; the guard checks statically resolved write destinations. Unknown destinations still rely on filesystem observation after execution. `interlinked verify` is the **on-demand, whole-project** run of that same check catalog. Its `function_tokens` finding uses the shared `interlinked-code-v2` adapters and reports every current product-source implementation over the effective cap. Unsupported languages are reported as not measured; semantic-model token counts are unrelated and never substitute for this check. JS/TS counts match `metrics score` function size, edit gates and commit checks. Recovered source syntax remains unmeasured. Verify is an inventory of current over-cap functions; edit/commit ratchets separately allow existing debt to hold or shrink, including debt revealed by migration. Explain that size unit as **syntax tokens** (lexical tokens); AST-node counts and embedding model tokens are different units. `metrics score` is an advisory composite, not a replacement for verify or the hard cap. A provisional score does not prove checks passed. Use `metrics gates --json` for actual coverage execution/freshness and `metrics coverage status` for index validity. `metrics coverage warm` explicitly runs full Vitest coverage; subsequent per-edit runs can replace affected test-file contributions. Proposed evidence promotes only after the matching edit lands, and unsupported/stale capture remains visibly unmeasured. Route evidence receipts, incremental coverage setup and deletion trials to **interlinked-quality-gates**. ## Load this when ### E2E lane for Interlinked CLI boundary edits When editing a hook entry, adapter, daemon, installer, generated hook, or ledger writer identified by `src/harness/e2e-boundary.ts`, add or run a fixture-backed test: `npm run build:e2e && npm run test:e2e`. Fixtures own their cwd, sockets, daemon PID and ledgers. Daemon assertions require a fresh transport receipt with `outcome: "daemon"`, plus PID ownership; cold fallback is a separate case. `interlinked e2e scaffold --event PreToolUse --tool Edit` creates `src/e2e/.e2e.test.ts`. `--dry-run` prints it. Replace the deliberate failure with the required behavior and add MUST-NOT-FIRE cases. `npm run test:e2e:coverage` collects child coverage after `build:e2e`. Route ratchet refusals to **interlinked-quality-gates**. The advisory `[interlinked:e2e-obligation]` Stop warning credits only commands observed in the current session. Run the lane from that session; another agent's run, a terminal run, and unit-only evidence do not satisfy it. The viz feed labels base, unit, integration, e2e and unknown lanes separately. `E2E_STABILITY=1 npm run test:e2e` adds the 5,000-event stress case. ### Project e2e scenarios in any host repository A repository with `.interlinked/e2e-policy.json` declares projects, suites (`managed-contracts`: argv `prepare` steps, build `artifacts`), scenarios (`affects` globs, `contractIds` from `.interlinked/behavioral-contracts.json`, `required`, `boundary`) and expectation records. Only that policy creates obligations; a repository without it pays nothing. - `[interlinked:e2e] : needs current e2e evidence after changed` fires on any observed edit to a mapped input (Edit, Write, Bash, patch). Run the exact command it prints: `interlinked tests e2e run --project

--scenario `. - `needs-mapping` means a protected input has no scenario. Add it to a scenario's `affects`; do not delete the protected glob. - `tests e2e status` / `plan` inspect (exit 0). `tests e2e check` verifies the working tree (exit 0 satisfied, 1 open requirement or measured failure, 2 UNCONFIGURED / unavailable). `tests e2e run` prepares, drives the public executable in a disposable workspace, writes `.interlinked/test-runs/e2e//receipt.json` and appends to `.interlinked/e2e-obligations.jsonl`. - A receipt satisfies only its exact generation: policy digest, affected files, the contract manifest and acceptance file, each case's cited requirement document, declared inputs and the bytes of any path-shaped executable the case invokes. Another relevant edit makes it `stale`; a failed case is `failed`; a missing toolchain, a failed prepare step, an input the collector could not capture (over 8 MiB, a symlink) or a receipt that does not parse strictly is `unavailable`; an unaccepted contract is `review-required` even when execution passed. A receipt from another worktree or with a runId the ledger never recorded is `RECEIPT_MISMATCH`. - Every run copies the project into a disposable snapshot first; prepare steps and cases execute there with a private HOME. The live tree is only read. Compiled executables are copied byte-for-byte with their mode. - Boundaries: `entry: "process"` (`real: ["application"]`), `entry: "http"` driven through a suite-OWNED `services` entry (the supervisor allocates the port, proves readiness and clean shutdown; a port that still answers after stop belonged to something else and the run cannot qualify), and `entry: "browser"` for a `playwright` suite (the app is owned, fronted by the supervisor's recording proxy, and every declared `requests` entry must be observed through it — a health check or an intercepted API earns nothing). A literal-URL case is never a boundary. - A bound PROPOSED expectation is visible as an advisory (`~` line) and does not block completion unless the project sets `gates.review: "require"`. Disputed and superseded bindings always block. - The loop for a new behavior: `tests e2e scaffold [--suite ] [--write]` prints a proposed scenario (required: false, empty affects, a placeholder contract id) and a skeleton whose only assertion FAILS deliberately with every assumption marked `ASSUMPTION` / `REPLACE`. Replace the assumptions with the observed outcome, add the scenario to the policy yourself (the command never edits it), `tests e2e run`, then `tests e2e qualify --scenario ` for a stability cohort, then `tests e2e check`. A deliberate failure is CASE_FAILED, never a pass. - Completion gates judge EXACT bytes, never "the repository": `tests e2e check --staged` judges the index (an unstaged fix does not count), `--revision ` an exact commit; both are materialized from git's object store (no archive attributes, no smudge filters). `--base ` compares the judged policy with the trusted base's: a removed/demoted requirement, a loosened gate (an omitted gate is its default: commit/ci require), a narrowed `affects`, an unbound contract, a dropped proof/request/stability profile is `POLICY_WEAKENED` (exit 1). The reviewed path is `tests e2e policy replace --base --project

[--scenario ] --rationale ""`; the record is read from the judged target, so COMMIT `.interlinked/e2e-policy-changes.jsonl` (carve it out of `.gitignore`). Never "fix" a weakening by editing gates in the candidate: the candidate's own gate setting cannot waive `POLICY_WEAKENED`. - `tests e2e gate install` writes pre-commit (`check --gate commit --staged --base HEAD`) and pre-push (one `check --gate ci --revision --base ` per ref) hooks that CHAIN with existing hooks and only CHECK; `gate uninstall` restores the originals. A blocked commit or push prints the exact recovery command: run it, do not bypass the hook. - `tests e2e ci [--base ] [--revision ]` exports the candidate COMMIT into a disposable directory with empty execution state, runs every plain scenario and every stability cohort there and checks the export against the event base. Workstation receipts, cohorts, untracked fixes and local records never count (`CI_RECEIPT_NOT_FRESH`); preparation steps must provision dependencies. Evidence lands in `.interlinked/test-runs/e2e/ci//`. - The Stop reminder is bounded: three identical reminders, one pause note, silence until the open set changes; an all-`unavailable` set is a HANDOFF (do not retry in a loop); with the daemon down it says NOT CHECKED. `interlinked verify` prints the same verdict in its `e2e` section. - The operator guide is `docs/project-e2e.md` (schema, codes, targets, gates, CI, proof limits, recovery table). - `[interlinked:e2e-quality] /: at :` is advice attached to the scenario a test edit affects (§12.3): a removed assertion or test block, an added `.only`/`.skip`, a specific matcher replaced by truthiness, a raised timeout/retry budget, `force: true`, a fixed timing wait, a CSS/XPath locator, an intercepted application endpoint, or a new test with no assertion. Net-new only (a moved line is not a signal); it never changes the verdict — the scenario still clears only through a supervised run. - A `playwright` suite owns its application as a `services` entry; the run fronts it with a recording proxy (`INTERLINKED_E2E_BASE_URL` is the proxy) and forces `--reporter=json --workers=1`. A browser case earns the `browser-driver` boundary only when the proxy saw its requests inside the case's first attempt; `@playwright/test` absent in the project is `unavailable` with install guidance (`tests e2e doctor` names it too), never an install. - Expectations: `tests e2e expectations propose --from draft.json` records an agent-authored statement with sources, assumptions and questions as `proposed`. `accept --from decision.json` binds the exact `revision` digest and records the linked contract digests as configured acceptance (a local decision, never authenticated human approval). `replace` supersedes with rationale and invalidates affected receipts; `review` lists questions, source provenance (`matched` / `stale` / `unavailable`) and diffs. Keep unresolved product questions in `questions`; do not turn an inference into a hard requirement. - Adoption workflow for a repository with no policy yet (all read-only until `adopt`): `tests e2e discover --out report.json` inspects manifests, build and test commands, executables and existing contract cases, proposes an ADVISORY policy with one scenario per process-runner case, and lists gaps (no contract cases, ambiguous nested manifests, unsupported runners). Review the report, then `tests e2e adopt --from report.json [--project ] [--scenario ]` writes only the selected configuration; `--mode required` must be explicit and expectations in a proposal are always dropped. `--replace` is required to overwrite a policy (discarded, not merged). `tests e2e surfaces [--write]` inventories bins, scripts and OpenAPI JSON operations and shows which scenario `surfaceIds` bind them (`explicit` / `unresolved` / `dangling`); YAML documents, routes registered in code and unresolved `$ref`s are reported as limits, never as an empty complete inventory. An interface no extractor knows is declared in the project's `surfaces` list (`{"id": "cli:custom", "kind": "other", "address": "…"}`) and bound through a scenario's `surfaceIds`; that is the language-independent path. Discovery protects the real layout (`src/**`, python packages, top-level source files, the executable's file) and binds build scripts as shared inputs; a build whose outputs it cannot infer is a `suites[].artifacts` gap to resolve before `--mode required`, which refuses a project with no scenarios or with protected globs that match nothing. Depth-omitted subtrees are listed in the report's `limits.omittedSubtrees`. `tests e2e doctor` diagnoses policy, manifest, contracts, acceptance, toolchain, mapping, ledger and receipts without running anything (exit 1 on a failed prerequisite, 2 on an invalid policy). Configuration alone never produces a pass; the next step is always `tests e2e run`. Shipped: managed process contracts for any language with a build/run argv (TypeScript and Python validated locally; the Rust fixture's compiled route waits for a Cargo-equipped runner), the discover / surfaces / adopt / doctor adoption workflow, structured-runner suites (JSON protocol / JUnit report import), owned loopback HTTP services with service-bound contracts, and per-scenario proof modes (`proof {mode: old-new | controlled-fault | characterization}`). A counterfactual proof names its `designated` cases as `{id, outcome?, action?}`: `outcome` lists the observables that ARE the designated outcome (the case's other declared observables must hold on the comparison side), `action` lists cases that must pass there first. Without that declared evidence the supervisor cannot tell a setup failure from a behavioral red, so it reports INCONCLUSIVE; output presence, exit codes and status classes never count as evidence. Scaffolding, browser runs, repetition cohorts and git/CI gates are implemented as described above. Automatic conversion of existing tests into proposed contracts remains planned; the MCP/Worker profile remains an explicit release gap. Jev evaluations are internal experiments, separate from public verification and Stop hooks. Do not recommend `interlinked jev` or `jev.enabled` to users. For authorized internal evaluations, see `docs/internal/jev.md`. The runner sends selected evidence to TypeSafe; its findings are advisory, not gate evidence. - You want to verify a batch of edits before declaring done. - A `pre_block` check blocked an edit (see also **interlinked-harness** for how blocks read). - You're landing a cross-file refactor and hitting transient `tsc` errors. - You're unsure whether a finding is default-gate or advisory-only. - You need to write a probe/analysis script and want it in the right place. ## `interlinked verify` ### Behavioral contract evidence `interlinked tests contracts import --json` prints proposed cases from explicit `json interlinked-contract` fences; it never executes or accepts them. Each fence is a JSON object with `id`, `description`, `inputs` (literal project-relative UTF-8 files), `runner: {kind: "process", argv: ["python3", "main.py"]}` and `expect: {exitCode: 0, json: {message: "ready"}}`. Import records source path/hash/quote and ties expected observations to that exact example. Save selected cases in a version-1 `cases` manifest at `.interlinked/behavioral-contracts.json` (or use `--file`). `inspect --json` reads provenance; `run --timeout 60000 --json` explicitly executes the selected runners and records per-case receipts under `.interlinked/contract-runs/`. `run --previous ` also tests retained prior expectations against current code. Process cases use a disposable workspace containing declared inputs and already installed tooling. This is not an OS sandbox. External state is unsealed, so historical passes are not reused as current verdicts. HTTP cases use a literal loopback HTTP URL, GET/POST, no redirects, and exact text/JSON/status/header expectations. Comparisons do not normalize string values. Missing tooling, stale inputs, conflicting citations and budget exhaustion are separate from a measured failure. A source citation alone does not prove semantics: use `source.observation: json|stdout|contract-example` for exact example binding. Accepted case digests belong in operator-owned `.interlinked/contract-policy.json`, `{version: 1, accepted: {"": "rationale"}}`. Never self-approve a proposed case. `configured` describes that file, not authenticated ownership; protect it outside agent write authority when required. Intentional replacements need rationale and changed requirements. `replaces: {id, reason}` does not automatically grant acceptance. Post-edit `[interlinked:test-contract-review]` is bounded, deduplicated advisory guidance about new/changed expectations, fixtures or collection settings. Review expectations against user requirements before adapting them to implementation output. No automatic runner execution or Stop repair loop is added. See `docs/plans/behavioral-contract-verification-20260916.md` for schema and scope limits. `interlinked tests readiness --cwd --json` probes prerequisites for Python, TypeScript/JavaScript, Rust and Go without collecting tests or installing packages. Python names the selected interpreter and reports exact approved install argv when possible; an absent runner or coverage plugin is unavailable evidence. Provision within the existing authorization boundary, then rerun readiness and execute the tests. There is no system-Python fallback around a broken selected environment. At a meaningful change boundary, `interlinked tests review [paths...] --base HEAD --json` provides a bounded source/test inventory and at most five simplification candidates. Without paths it discovers staged, unstaged and untracked Git changes. It reads at most 32 source/test files of 256 KiB each; deletions, unsafe paths and exhausted budgets are explicit gaps. This is review guidance, not test execution or a passing verdict. Retain executable assertions for old public contracts and new requirements; review validation ownership, duplication, forwarding and shared mutable state together. After behavior passes, one focused simplification pass is enough; rerun relevant tests after changing code and leave uncertain advice unresolved. `interlinked tests suite --cwd --timeout --json` explicitly runs a bounded project suite for `typescript`, `javascript`, `python`, `rust` or `go`. The default budget is 60 seconds including admission. TS/JS use the shared Vitest scheduler; Python uses the active `VIRTUAL_ENV`, then project `.venv`/`venv`, then platform Python (an explicit adapter interpreter takes precedence). A selected missing interpreter is unavailable; it never silently switches environments. Rust uses offline Cargo with two jobs/test threads; Go runs all packages with two build jobs and caching disabled. No runner is installed automatically. Python retains project pytest options and configured discovery (`testpaths`, `python_files`); the invocation does not append an implicit `.`. Fresh structured pytest case reports distinguish test failures from collection/configuration errors and coverage-plugin failures. Fixture and teardown failures count as failed tests. Known failures remain visible alongside incomplete collection; bounded diagnostics accompany unavailable results. Terminal text alone is not a pytest verdict. Only an observed passing suite exits successfully; missing/empty execution is not a pass. `tests plan/run/status` remain the TS/JS dependency-aware queue interface; the non-TS suite command does not certify or discharge that queue. Active distributed pytest collection is currently unmeasured. Serial execution remains supported when xdist is installed but inactive; no plugin is silently disabled to obtain a pass. The default Python edit/commit coverage runner isolates each invocation's JSON report and coverage.py database (`COVERAGE_FILE`) in an owned temporary subdirectory of the requested report directory. Normalized results survive; those temporary files are cleaned after parsing, including failure paths. Project cwd and pytest configuration remain in effect, and existing caller coverage files are preserved. Custom command overrides retain their argv/report contract and remain unqualified for concurrent report isolation and structured red/green verdicts. Python CRAP attribution requires native function regions with declaration lines; unsupported, ambiguous or wholly excluded functions remain explicitly unmeasured. A green suite with an unmeasured enabled quality check does not discharge commit obligations. See the quality-gates skill for the native coverage.py/Radon attribution contract. For evolving requirements, distinguish newly required observable behavior from contracts that should remain valid. Exercise representative successful, boundary, error and state-transition cases from the public task before concluding the implementation is complete. Existing passing tests can all remain green while the new feature is largely missing. Keep this proportional to the change; do not create a mandatory test-authoring loop for trivial reversible edits. Hidden evaluator cases are unavailable to product checks and must not shape harness rules. Heavy runners retain at least the configured 2 GiB admission budget plus the host reserve; the proportional ceiling does not reject an otherwise idle nominal 8 GiB Linux guest merely because its reported usable RAM is slightly smaller. Insufficient available memory still defers. The proposed qualification baseline is an 8 GB whole host shared with the user's other applications. Its operator plan, `docs/plans/8gb-host-resource-plan.md`, is private operator material and absent from public clones. Qualification requires aggregate measurements across owned processes and useful completed checks, not merely successful resource deferrals. It does not change today's verification contract: interrupted or partial work remains unmeasured, and a constrained run cannot satisfy a required full scope by selecting fewer tests. Cognitive-complexity, type-smuggling, and cast-justification diagnostics align snippets with TypeScript's line numbering, including CRLF, lone CR, and Unicode line/paragraph separators. Preserve those line endings when reproducing a warning; normalizing a fixture can hide a location bug. For `[interlinked:hook-coverage] NOT CHECKED`, use `interlinked harness coverage verify --json`. This starts one daemon-owned recovery run over the pending versions and waits for completion; `--no-wait` returns after starting it. Use `harness coverage status --progress --json` for cached progress: it reports generation, observation time, pending count and job counters without rehashing files or returning historical receipts. This is not a fresh coverage verdict. The waiting CLI uses this compact path and fetches a fresh full status before reporting completion; it still accepts older daemons that return full reports. Use plain `harness coverage status --json` for an explicit refresh, pending identities and full receipt details. Checks reuse the configured PostToolUse battery in bounded external batches. This does not replay PreToolUse guards or certify every hook phase. While waiting, an unavailable status response is retried up to three consecutive polls without restarting verification. A responsive report resets that counter. Persistent unavailability exits nonzero with the original reason; the job may still be running. Only a ready response with a missing or different job establishes that the observed job changed or disappeared. On Claude and Codex Stop/SubagentStop, advisory coverage feedback stays on exit-0 stderr and does not request another agent turn. Only an explicit blocking decision requests continuation. This delivery rule does not clear pending evidence or certify checks; retry unavailable recovery after its prerequisites change, not merely because a turn ended. Writer identity remains unknown after recovery. An initially absent watched path contributes to policy identity but creates no write-check obligation. Creation or deletion after observation still requires evidence; previously recorded historical gaps are not cleared by this rule. Recovery groups pending files by project and applicable checks, then checks up to 32 compatible files together. Documentation does not inherit an unrelated source-test timeout, and nested projects acquire their own admission lane sequentially. Each recovery related-test process gets at least 15 minutes and at most two workers, with the complete union of related tests for up to 32 sources. This reduces repeated broad suites without selecting a passing subset. Ordinary hook deadlines and source-count limits remain unchanged; recovery still defers honestly on capacity or timeout. Explicit recovery now waits up to 30 seconds for each external batch's existing project lease. Its affected-test scheduler waits within the configured recovery deadline and permits a necessary full-suite plan without the interactive test-count cap, retaining the two-worker limit. Ordinary PostToolUse still uses immediate admission and its test-count cap. Unsupported runners and exhausted capacity/time remain unmeasured; the shared named test dispatcher supports TS/JS with Vitest and single-language Python/pytest, Rust/Cargo, and Go project suites. Mixed-language batches remain explicitly deferred. Python no longer guesses one companion file, Rust executes assertions instead of only compiling tests, and Go covers the project packages. These project suites produce fresh execution evidence, not reusable dependency-closure receipts. Missing runners, empty/unrecognized successful output, and pytest collection/configuration errors are unmeasured. Current non-TS suite failures are warnings; without a before result they are not classified as introduced regressions. The external path cap applies separately to each check's applicable paths in the selected project. An inapplicable binary path does not consume a TypeScript check slot; an applicable security target still counts. This is not a blanket dependency/cache exclusion. The daemon retains `automated_check` receipts with exact file identities, completed check names and findings. Completed checks may have findings; a receipt is not a clean verdict. Nonempty `unavailable` on a receipt means partial evidence: completed checks and their findings are retained, but the file version stays pending. Multi-file per-file checks retain partial receipts when shared checks defer. Shared receipts include per-check configuration hashes and the request's captured file identities. File or policy changes prevent full discharge. Legacy receipts remain readable. These historical receipts are not yet a cache for skipping future work; check-specific configuration/dependency/runtime identities and exact batch scope are still required before safe reuse. Ordinary single-file and multi-file PostToolUse checks consume their exact pending versions when all applicable evidence completes without deferral. Explicit recovery uses the same shared scope evidence. Recovery has a 30-minute job budget; cancellation reaches the job's owned asynchronous subprocesses and capacity waits. Earlier recorded evidence survives. Unavailable checks, unreadable/excluded/absent files, and file or policy changes during verification stay pending. With waiting enabled, findings or remaining pending versions produce exit 1. Re-run after the reported capacity/tool problem is resolved. A daemon restart interrupts the job; recorded receipts survive and unchecked entries remain pending. Active recovery keeps the raw and framed listeners out of idle shutdown and delays automatic build handover within its normal freshness deadline. Explicit restarts and memory safety shutdowns still interrupt recovery; inspect status and retry afterward. Released reservations remain watched while pending, so review cannot acknowledge a stale historical hash. `harness coverage acknowledge ` records an explicit manual review of the current version, including a reviewed deletion or optional absence. It does not manufacture automated evidence. Do not bulk-acknowledge unreviewed files. Checking/reviewing files never accepts protected policy; that is the separate `harness coverage accept-policy ` operation. Writer identity stays unknown. ``` interlinked verify [target] --all-checks add the advisory smell/complexity/dead-code tier to the default gate --only run only one external tool (e.g. --only tsc) --skip comma-separated check ids to skip --suggestions also run scored regex heuristics (sql-injection/perf/quality) --structure also run artifact-structure checks --adoption-gate fail when adopted structure categories drop below thresholds --suppress add a suppression (file:check or file:check:reason) --json --details machine-readable / per-file detail ``` `target` may be a local path, a GitHub/git URL (cloned to a tmpdir, scanned, deleted), or omitted (scans cwd). Narrow with `--subdir ` in monorepos. **Two tiers.** Default = high-signal gate: tsc, biome, oxlint/eslint, semgrep, gitleaks, dep-audit (+ language tools as available) **plus** the FP-safe inline checks. `--all-checks` adds the advisory tier (complexity, taste/smell, DRY clones, most `ubs_*`, test heuristics) — a **review tool, expect noise, not a gate**. > **`interlinked verify` exits 0 even with findings.** It is a *reporting* tool, not a > pass/fail gate — do not `&&`-chain on its exit status. To gate programmatically, parse > `--json`, or use `interlinked write` / `verify-changeset` (which **do** exit nonzero on > blocking findings). (Exceptions that *do* exit nonzero: usage errors, and > `--structure-only` / `--adoption-gate`, or a deferred/unavailable run that produced no > verification verdict.) Whole-project heavyweight work is admitted once per canonical project across CLI and agent processes. If another verify/check/test batch already owns that lane, verify does not queue or start a second memory-heavy scan: it prints `verify deferred`, states that no verdict was produced, and exits 1. Retry after the active project run finishes. Different project roots use independent lanes, and the compiler has a separate nested lease so `--only tsc` can run while verify owns the heavyweight lane. `--only ` really runs only that external tool; it skips the inline code-quality census rather than retaining the whole-project scan before the requested tool. SessionEnd maintenance, fuzz, and benchmark runners additionally share a host-wide background lane owned by a detached supervisor. Duplicate jobs coalesce across daemon restarts; other jobs wait at most two minutes. Memory admission can defer a job, and low memory or a ten-minute deadline terminates the child group. Fuzz/benchmark worker counts are bounded by current CPU and RAM capacity and rechecked before execution. A deferred or interrupted background job is not a successful verification; use current completed reports and explicit checks for a verdict. See **interlinked-setup** for the memory budget and its limits. Scheduled Vitest runs and verify also acquire this host lane. Foreground requests close background admission while waiting; the background monitor interrupts its child group to yield capacity. Verify waits at most five seconds for the host lane, then exits 1 without a verdict. Scheduled foreground test execution now checks host memory and runner-tree RSS throughout the run. Its 4 GiB maximum tree budget includes workers; a sampled overrun, lost headroom, or unavailable telemetry interrupts the process group and retains the request without a pass receipt. Worker planning reads current CPU load as well as available memory. These are sampled limits; see **interlinked-setup** for process-group and platform limitations. Repository pre-push checks use the same admission lane and monitor. Do not respond to a resource deferral by launching the full suite directly. Small checks can use the repository's `scripts/run-resource-bounded.ts --light` supervisor with an enforced 1 GiB tree ceiling; this does not replace required full verification. Public `interlinked verify` and `interlinked tests` commands have an outer resource supervisor as well: in-process planning/scanning is measured along with the runner tree. This supervisor does not acquire a second host lease; the actual command keeps its existing admission protocol. Memory interruption exits 75 without a verification verdict, even if partial output was printed. Exit 75 alone cannot say what happened — admission that never happened, a child killed mid-run and a child that itself exits 75 all surface as 75 — so the bounded runner (`scripts/run-resource-bounded.ts`, used by every pre-push gate step) writes a VERDICT record to `INTERLINKED_BOUNDED_OUTCOME` beside its status (2026-09-28): `not-run` (capacity timeout, memory budget unavailable, or cancelled while waiting; stderr `[resources] NOT RUN: …` with the wait), `interrupted` (started, then killed or timed out; `[resources] INTERRUPTED: …`), or `exit` (ran to completion; the child's own code, including its own 75). The pre-push hook reads the record: `not-run` prints `[pre-push] NOT RUN (host capacity or memory unavailable; nothing failed): `, `interrupted` prints `[pre-push] INTERRUPTED (killed or timed out; no verdict): `, and only `exit` (or a missing record) prints ` failed`. The stage ledger row carries the matching `reuse_denied_reason`. The push is blocked in every non-zero case — nothing was verified — but only a real verdict names a failed gate. The daemon's async project test gate and legacy affected-test process adapter also acquire the shared host lane, even when their caller already owns project admission. They monitor the runner tree and host reserve, set a 768 MiB Node heap limit, and pass `VITEST_MAX_WORKERS=1` to prevent a second uncapped Vitest suite during a push. The current Vitest adapter honors that variable; other runners still use their own worker controls and remain subject to the memory monitor. Capacity loss is an explicit deferral. Repository pre-push heavy commands wait at most ten minutes for that shared lane; light diagnostic commands wait five seconds. This serializes an existing daemon push check with the exact-revision repository gate without skipping either check. ## Select, explain and resume tests ```bash interlinked tests plan --base HEAD --json interlinked tests run src/lib/config.ts --workers 2 --timeout 120000 interlinked tests status --json interlinked tests run --all --timeout 3600000 ``` Paths are relative to `--cwd` (default cwd). With no paths, plan/run include staged, unstaged, deleted and untracked inputs relative to `--base` (default HEAD), plus pending requests. `--all` works without Git and asks the native runner for its full suite. Plan loads the project's Vitest configuration in a bounded child but runs no assertions. Status returns pending inputs and the last observed job ID, PID, snapshot and state; a retained running observation is not proof its owner is still alive. The TypeScript/Vitest hook paths use this same union: edited tests, static transitive consumers, colocated and `__tests__` companions, declared inputs, and historical coverage consumers. Estimates use existing per-shard durations and otherwise say unmeasured. An opaque test runs on any input change; shared opaque setup/configuration, unknown/deleted inputs, incomplete discovery or an explicit full request widen to the full suite. Named/multiple Vitest projects currently use native full execution rather than selective indexing. Python/Rust/Go retain their existing dispatchers; mixed-language batches defer. Selection is measured in CI before it decides anything (Unit 8, comparison mode). The `select-compare` job plans against the event's base (`scripts/ci-select-compare.mjs select`, base from the shared `resolveCiBase`) and runs only the must-run files (`run-selected`, which refuses an empty list: `vitest run` without filters runs everything). The full unit and integration lanes upload their vitest JSON reports, and `select-compare-report` writes one row per run: `SELECTION_MISS` when a file failed in the full lanes but was in the plan's omitted set; `widened` when the plan omitted nothing; `incomplete` when a lane, the selection or the selected run left no report (a miss already seen still reports as a miss). Neither job is a required check and neither fails on a miss. Promotion needs zero misses over at least 30 rows that are SELECTIVE and complete; a widened row proves nothing. Measured 2026-10-01: every plan on this repository widens, because `vitest.config.ts` and the setup files read `process`/`Date` (opaque by the input-eligibility rule) and `__fixtures__/` files read as unknown inputs, so expect `widened` rows until that changes. Optional `.interlinked/test-dependencies.json` declares additive literal inputs: ```json {"version":1,"tests":{"src/cli.test.ts":["src/templates/config.json","dist/index.js"]}} ``` Declarations never make uncontrolled I/O cacheable. Exact passing-result reuse is limited to plans without uncertain dependencies and requires matching source/scope, captured runtime bytes, runner/dependency/configuration inputs, platform/environment and worker count. Opaque plans run fresh without the expensive reusable-evidence census. Their results say `runtimeVerified:false`: the runner completed and the scheduler checked analyzed source stability, but no exact runtime or coverage verdict exists. A failed bounded runtime census also disables reuse and widens selection. No runtime values are persisted in receipts. Failed, interrupted, missing-report, all-skipped and empty runs never create passing receipts. Edits during execution trigger another plan; stale evidence cannot clear pending work. For Python, check prerequisites in the selected project interpreter after creating a venv. System pytest/coverage packages are not available in an ordinary isolated venv. `interlinked tests readiness python --json` reports that distinction; install approved test tools there rather than silently switching interpreters. Explicit test-first enforcement requires the first companion even when the repository starts without tests. Requests coalesce across nearby edits and identical in-flight subscribers. Cross-process leases serialize execution. A waiting process can consume its request's completed result after rechecking source, declared/requested inputs, environment and platform; exact-runtime results also require a fresh matching runtime census. This request-specific sharing is distinct from reusable passing-result caching. Hooks defer immediately when another process occupies capacity, while explicit CLI runs wait within their deadline. A subscriber's timeout does not cancel another caller's shared work. Durable requests survive timeout and process restart. The queue reads at most 1,000 requests per batch and continues draining later batches. A passing run is certified by a RECEIPT keyed by its CHECK IDENTITY (2026-09-28, `src/harness/check-identity.ts`): the plan snapshot and runtime hash (inputs), the logical runner argv (mode, workers, `--coverage`, extra reporters), the toolchain (node plus the installed vitest and typescript versions), the normalized environment, the platform and the POLICY digest (`.interlinked/coverage-baseline.json`, `coverage-edit-baseline.json`, `metric-caps.json` and every `vitest*.config.*` at the root). Changing any dimension is a new check: a raised coverage baseline re-runs the same tests. The environment component drops ONLY variables proven not to reach a verdict — the shell's cwd bookkeeping (`PWD`, `OLDPWD`, `SHLVL`, `_`), git's hook-only variables (`GIT_EXEC_PATH`, `GIT_PREFIX`, `GIT_CONFIG_PARAMETERS`) and this route's own bookkeeping (`INTERLINKED_STAGE`, `INTERLINKED_STAGES_LEDGER`, `INTERLINKED_BOUNDED_OUTCOME`, `INTERLINKED_LEASE_ANCESTORS`, `INTERLINKED_COVERAGE_SCOPE_FILE`, `INTERLINKED_TEST_CAPACITY_SCOPE`); `PATH`, toolchain paths, `CI`, `NODE_ENV`, `NODE_OPTIONS`, `TMPDIR`, every other `INTERLINKED_*` and any user variable stay in the hash because a test may read them. The child still receives the exact environment, and the shared-completion path keeps the exact hash. A full run is reusable evidence only when every closure resolves inside the inventory and nothing is untracked: an opaque test (dynamic import, `process`, `fetch`, the clock, an external fixture) or a test outside the inventory makes the run fresh-only, because hashing repository bytes cannot see what those tests read; an opaque SHARED setup or configuration file (it runs in every test) makes every run fresh-only, selected or full, changed or not. A reporter module run from outside the checkout (`--coverage-reporter`) is bound by the digest of its bytes and of every dependency it loads. Loads are read from the PARSED source (a specifier that is not one string literal — `import("./x" + ext)`, a template — is a computed load, never a prefix match), and each is resolved by its OWN loading mode from the importing module's location: `require("x")` by CommonJS rules, `from "x"` / `import "x"` / `import("x")` by ESM rules (a conditional `exports` map is read under the `import` condition, so the entry Node executes is the entry that is bound; a relative ESM import must name an exact file). A dependency inside `node_modules` is bound by its COMPLETE installed contents plus, transitively, the installed contents of every package it declares (a version pins nothing — an edited internal helper is a different reporter). A computed load, an unparseable or unresolvable module, an `exports` shape the resolver does not model (patterns, arrays, an unexported subpath), a declared dependency that is not installed, or a closure over 400 files is UNRESOLVED. And the reporter obeys the SAME input-eligibility rule as a test or setup file, and so does every installed code file of a bound package (declaration files never run): a module that names a runtime read the bytes cannot pin — `process`, `fetch`, `Date`, `Buffer`, timers, `require`/`createRequire`, `import.meta`, `Math.random`, eval — or imports a builtin OPERATION that is not pure is OPAQUE, because complete module hashing cannot cover external data, network responses, the working directory or the clock. Purity is per operation, never module-wide: whole modules only where every export is pure (`node:assert`, `node:string_decoder`, `node:querystring`, `node:util/types`); otherwise named imports from a short list (`path`: `join`, `basename`, `dirname`, `extname`, `normalize`, `parse`, `format`, `isAbsolute`, `sep`, `delimiter`; `url`: `fileURLToPath`, `URL`, `URLSearchParams`; `util`: `format`, `inspect`, `isDeepStrictEqual`, `promisify`, `inherits`, `types`). Imports from `node:buffer` and `node:events` are opaque: `File` defaults its timestamp from the clock, unsafe buffers expose untracked memory, and `EventEmitterAsyncResource` captures async context. Global `Buffer` is opaque too. `path.resolve` and `path.relative` read the working directory, `url.pathToFileURL` resolves through it, and a default or namespace import of any listed-by-name module reaches them, so all of those are opaque. Unresolved or opaque, the run executes, exports its artifacts and certifies nothing reusable. CONSEQUENCE: `scripts/pre-push-coverage.mjs` reads `process.env` and walks `src/` through `node:fs` to ask the coverage provider for its inclusion scope, so every pre-push coverage run is fresh-only today; a coverage run without an external reporter (vitest's own json-summary) reuses across a byte-identical export. Re-establishing reuse for the scope needs the inclusion decision computed OUTSIDE the run (the provider's `include`/`exclude` globs re-evaluated deterministically), not a more permissive rule. Every identity input, the reporter binding included, is re-read after the run: a change between hashing and execution is `stale` and certifies nothing. In THIS repository nearly every test is opaque, so full-run receipts are not produced here today. A subscriber that shares a run but owns a different receipt store gets the artifacts materialized into its own store (sha256-verified), so its export never depends on the producer's store. Receipts are v2 (`identity`, `platform`, `toolchain`, `stages`, optional `artifacts`), unsigned and local; a v1 receipt is ignored. `interlinked tests run --all --coverage [--coverage-reporter ]… [--receipt-store

] [--artifacts-out ]` collects coverage (json-summary) INSIDE the run directory of the receipt store, records the summary and the inclusion scope (`scripts/pre-push-coverage.mjs` as the reporter; without explicit targets it records every `src/` code file, so the scope serves any later changed-file set) as sha256-pinned artifacts on the receipt, and copies them into `--artifacts-out` after a passing run — fresh or reused — refusing any file whose bytes no longer match, and RE-ROOTED to the consuming checkout (the scope's `root` and the summary's file keys name the run's absolute root; a reused run may come from another export of the same bytes). The receipt's `artifactRoot` retains the source root even without a scope reporter. Export relocates JSON path fields after verifying the stored bytes; receipts containing artifacts but missing this root are ignored. Exit 1 is a failed run; 75 is deferred/stale (no verdict). A request that needs coverage coalesces only onto a request that produces it, a batch produces what every pending request requires, and a completion discharges only the requests whose requirements the run met — a plain run never satisfies a coverage subscriber. The pre-push hook's coverage branch runs exactly this with `--receipt-store` pointing at the SOURCE checkout's `.interlinked/test-runs`, so the previous push of the SAME revision is consumed instead of re-run (every export is fresh, so two exports of one revision share an identity). A LOCAL `tests run --all --coverage` does not seed the export in practice: the runtime snapshot hashes every on-disk byte outside `.git`, `.interlinked` and generated directories, so gitignored local files (`scratch/`, `coverage/`, logs) enter the local identity but not the export's — and honouring gitignore alone would not prove those ignored inputs irrelevant, so this stays an unmet acceptance criterion (recorded in the campaign file). The export's `node_modules` is a symlink to the source's: the runtime snapshot hashes the canonical path of the mounted dependency tree, so a copied `node_modules` is a different runtime and never reuses. That step is not `bounded` (the scheduler owns admission and supervision). Requests coalesce at write time (2026-09-28): a new request that a pending one already covers returns that request's id, and a wider request retires the narrower pending requests that no caller in its process still awaits (subscriptions are reference-counted, so a shared id stays protected until its last caller leaves) — a full request is never retired by a selected one (an empty-path full request keeps its obligation), and each retirement is recorded in `.interlinked/test-runs/requests/superseded.jsonl`. Request files are immutable: a request that adds a path to covering work is stored as its own file (never merged into an in-flight request, whose completion would otherwise discharge an input it never tracked), and a replacing request absorbs the retired requests' paths, so no path is ever dropped — those paths are the inputs the freshness check re-hashes after a run. Concurrent writers in separate processes therefore cannot lose paths. Between batches the drain YIELDS all three leases (scheduler, project, host) for a short window, so a same-root `interlinked tests run` or another repository's pre-push export takes the next slot instead of waiting for the whole queue. Hook-originated work (stage `edit`) runs at `background` priority and never blocks a waiting interactive caller at the host gate; CLI and pre-push runs are `interactive` (`ScheduleTestsOptions.priority` overrides). A drain that cannot re-acquire after yielding returns `deferred` with the work retained, and a queue another drain emptied meanwhile resolves to that drain's shared completion. TypeScript hooks use `max_dependent_tests` as a cap on the complete selected test-file union (default 150). Full and over-budget plans defer intact; run `interlinked tests run` to drain them with an explicit deadline. The external-tool batch releases its project lease before entering the test scheduler. Capacity, path-cap and tool-cap deferrals retain inputs. `affected_tests` remains opt-in. Its default empty filename suffix includes configuration and fixtures in Node projects; explicit repository `file_types` and `skip_test_files` settings still control which changes trigger it. After a triggered batch, all changed inputs are considered. Raw `npm test` and independently launched tools do not use this scheduler. The full pre-push/CI coverage gates remain authoritative; passing focused tests does not satisfy them or re-enable a disabled local coverage policy. Lease ownership binds the PID to an OS-derived process-start identity, so a live unrelated process that reused the same PID cannot keep compiler or heavyweight capacity busy. Legacy lockfiles without that identity remain compatible while fresh, but expire after 24 hours — well beyond every minute-scale workload timeout — rather than starving a project indefinitely. A compiler watch process that fails to spawn is unavailable. Shutdown waits for its close event and releases the compiler lease without signaling a missing process; callers can then use the normal cold compiler fallback. There is **no** `--file`/`--changed`/`--staged` flag — verify always walks the whole discovered set (or `target`/`--subdir`). Diff-awareness lives at the *edit-time* gate, not in verify. Run verify to see **pre-existing** findings in a file you're about to touch (the edit gate hides those as warnings). ## Check families & phases `html_duplicate_id` is a default PostToolUse/verify warning for repeated static IDs in an explicit `` in `.html`/`.htm` files. It reports each later occurrence with the first occurrence's line. IDs are case-sensitive; comments, script/raw-text contents, inert `