--- name: ship description: >- Autonomous PR-to-merge loop. Normally adopts the early PR that /quality opened, with one CI and review round already harvested and fixed, so ship mostly closes it out. Polls CI and review bots, fixes failures, rebases only on real conflicts, and lands the PR on main. Soft cap of 5 normal iterations plus one force-finalize iteration that bypasses review and fixes only CI. Pure loop — it does not replace the baseline /quality or /test runs; run those first. It does revalidate quality after any ship-loop mutation so the final result is bound to the exact reviewed PR head and content tree. Opt-in --stack-ready runs the same loop for one layer of a coordinator-owned stack: it fixes and pushes its own layer but stops at ready-stacked instead of merging, rebasing, or running any gh stack command. Full phase logic lives in docs/playbooks/ship-lane.md. --- # Ship Skill — Autonomous Merge Loop Drive the current lane from "work is ready" to "merged on main" without manual shepherding. **Pure loop:** `/ship` assumes you already ran `/quality` and `/test` — it does not bundle them. `/quality` opened the PR when it started, and `/quality` and `/test` each harvested and fixed whatever CI and the bots reported (playbook: **Early PR and harvests**). So `/ship` normally adopts an open PR in state `prepping`, skips PR creation and re-review, and closes the PR out. It polls, fixes CI + review, rebases only when there's a real conflict, and merges. It does not exit until the PR is merged or the merge is genuinely blocked by repo policy. Print a compact status line each iteration (no banner): ``` ship · iter 2/5 · PR #184 · POLL → DECIDE → FIX → MERGE · FIXING CI (test-desktop 3) + 2 comments ``` Stack mode prints the layer and its terminal instead of `MERGE`: ``` ship · stack 12 layer 2/5 · iter 2/5 · PR #1007 · POLL → DECIDE → FIX → READY · FIXING CI (windows-foundation) + 1 comment ``` **Invocation:** `/ship` (auto-detect state), `/ship `, or the opt-in `/ship --stack-ready [] --base `. ### Stack-ready mode (opt-in only) `--stack-ready` drives one layer of a coordinator-owned stack to *ready*, not to *merged*. It is the same loop — Phase 0 through Phase 5, the same poll/fix machinery, the same 5-iteration budget — with merging and every stack-wide operation removed. Resolve the direct parent from `--base`, an existing PR's `baseRefName`, then non-interactive `gh stack view --json`; normalize it with the `/quality` rules. Persist `mode: "stack"` plus the complete stack binding: stack number, size, position, expected parent branch, validated head SHA, base SHA, content-tree SHA, test-evidence SHA, required and deferred proof scenarios, proof links, and quality/test status. **The lane owns its layer; the coordinator owns the stack.** The lane commits and pushes its own layer branch, opens its PR against the resolved direct parent when none exists, polls CI and review bots, fixes red CI and verified findings on its own layer, reruns commit-bound quality revalidation against the exact resulting head, and repeats until the layer is genuinely clean. A red check on its own code is work to do, not a reason to stop. The lane never merges, never enables auto-merge, never deletes a branch, never rebases or restacks (`git rebase`, `gh stack sync --remote origin`, `gh stack rebase --upstack --remote origin`, `gh stack push --remote origin`, and `gh stack submit --auto --remote origin` are all coordinator-only), never retargets a PR base, never touches another layer's branch or files, and never enters force-finalize or any bypass-review path. Before any cap, force-finalize, rebase, merge, or branch-deletion decision, branch on `mode == "stack"`. Escalate only what the lane genuinely cannot do, with exact evidence: `stack-coordinator-sync-required` (a restack or base retarget is needed — the parent moved, a lower layer changed, or the PR base is not the direct parent), `stack-coordinator-fix-required` (the fix belongs to a lower layer, or the iteration budget is spent and the layer is still red), `stack-coordinator-pr-required` (the parent branch is missing on `origin`, or PR creation failed on auth or an unusable base ref), and `stack-coordinator-merged` (the coordinator already landed it). The playbook's **Stack escalation states** table is authoritative. None of them is a general stop at the first red check. Write `status: "ready-stacked"` only when the exact head is green, review-terminal, quality-clean, test-clean, and every mandatory proof scenario either has a current evidence link bound to the validated head or is recorded in `deferredProofScenarios` against a named higher layer that exists in this stack. The top layer defers nothing, and cumulative clean-host/cross-client/release scenarios never masquerade as lower-layer evidence. A known-missing mandatory scenario is `blocked` with the scenario ids listed — never `ready-stacked` with a caveat. Missing or ambiguous stack metadata is `blocked`, not a fallback to `main`. Without `--stack-ready`, every existing `/ship` default and merge behavior is unchanged: the base is `main`, green work proceeds through Phase 3c, and the terminal success state is `done-clean` only after merge confirmation. --- ## Source of truth **Follow `docs/playbooks/ship-lane.md`** — all phase logic, the state schema, commands, decision rules, and bot-ping rules live there. This skill is the runtime-neutral entrypoint and the ADE-specific deltas below. If re-invoked by a scheduled wake, read the state file first; if `status == running`, skip Phase 0 and go to Phase 1. If `status == ready-stacked`, revalidate the complete binding first: when it holds, print the persisted coordinator handoff and exit without scheduling or mutating anything; when it is stale, external movement exits `stack-coordinator-sync-required` and this lane's own newer head re-enters the loop at Phase 1. The playbook's Phase 0 is **checkpoint → commit-bound quality revalidation → push → open PR** when no early PR exists. With state `prepping`, it runs only **0.6 Adopt the early PR**: keep a current binding, mark held commits as fix work, and go to Phase 1. Every revalidation reviews only the delta since `qualityReviewedSha`; the full review ran once, in `/quality`. Baseline test generation and the local-CI gate are NOT part of ship — that's `/test` (and optionally `/finalize`) before you reach this skill. ## Precondition: `/quality` must be empty and bound to the final tree Before Phase 0, require a completed `/quality` result with an empty gate. Before Phase 3c, run the playbook's single canonical **Validate the current quality binding** procedure. It binds the reviewed head, content tree, and base so GitHub's squash/merge/rebase result has the reviewed tree. Green CI on a later head or base does not preserve this binding. A non-empty gate **blocks the merge** — every row in it is a finding that was verified as real and left unfixed, and by `/quality`'s contract the only three things that may be there are a product decision the author owes, a behavior change this branch was not asked to make, or a capability whose Windows parity is not achievable. All three need the author. - Gate rows exist → do not merge. Surface them, state the decision needed, and stop with `blocked`. Do not merge and mention them afterwards. - If `/quality` was never run on this lane, or its final gate result is not available in the lane handoff, stop with `blocked`; unknown is not empty. - Any base movement, rebase, conflict resolution, Phase 3b edit, or force-finalize edit clears all three quality binding fields. In stack mode, this lane's own Phase 3b edit clears them and is rebound by revalidation on the head it then pushes; external movement of the parent, base, or head instead clears the complete stack binding and returns `stack-coordinator-sync-required` without rebasing. Run the playbook's single canonical **Commit-bound quality revalidation** procedure before pushing that mutation. - Never enter Phase 3c with a missing or mismatched binding. Revalidate first; do not merge and disclose stale quality evidence afterwards. - Bind every normal or admin merge attempt with `--match-head-commit "$QUALITY_VALIDATED_SHA"`. Persistent auto-merge is not allowed because a later push can replace the validated head while it remains armed. - GitHub creates a new commit for squash/merge/rebase. The validation claim is deliberately about its exact content tree, not its not-yet-created commit OID. After merge, run the playbook's canonical **Confirm the validated merge result** procedure; a mismatch is never `done-clean`. Severity is irrelevant here: a Medium in the gate blocks exactly as hard as a Blocker, because presence in the gate means it needed a human, not that it was minor. --- ## Execution Mode: Autonomous Runs end-to-end without user interaction. Do NOT ask to confirm/choose/approve, pause between phases, or ask whether to apply a fix — apply, verify, commit. The only user-visible output is the per-iteration status line and the final summary. --- ## Repo facts (ADE) - **Package manager:** `npm`; each app under `apps/` has its own `node_modules` + `package-lock.json` (no workspaces). Node 22 (`.nvmrc`). - **CI:** `.github/workflows/ci.yml` — desktop tests shard **8-way** (`npx vitest run --shard=/8`) plus `test-ade-cli`; a `ci-pass`/`ci-status` gate aggregates required jobs. Discover the required-check list live via `gh pr checks` / the `ade-pr-workflows` skill — do not hardcode. - **PR creation:** prefer the `ade` CLI (registers the PR in ADE's tracking — lane ↔ PR link, check/comment inventory). `gh pr create --base main --head --fill` is the ordinary fallback; stack mode substitutes the persisted direct parent for `main`. See the playbook's discovery protocol. - **State file:** `.ade/shipLane/.json`. `status`: `prepping` (early PR open, before ship) | `running` | `ready-stacked` | `done-clean` | `done-max` | `blocked`; it also records `mode` and the complete stack binding. Rebase rebates the iteration counter by 2 (floor 0). **Windows parity gate.** Windows parity is a default requirement for all new code, so this gate runs on **every** lane, not only Windows-labelled ones. Before Phase 3c (or `ready-stacked`), confirm the branch's behavior works on the Windows build. If any capability on this branch cannot work on Windows and the human has not already chosen what to do about it, **stop with `blocked`** — this is the one product decision `/ship` never makes for itself, and force-finalize does not clear it. `/ship` is autonomous, so "ask" here means exit blocked with the question stated, not pause mid-loop. The blocked summary must give, per capability: 1. the exact capability that cannot work and the OS-level reason; 2. whether macOS/Linux keep it; 3. **hidden** (absent on Windows) vs **disabled with a reason shown** (visible, inert, explained) vs **removed** (deleted from the Windows build), with your recommendation. Hidden and disabled are different user experiences; the author picks per item. A `/quality` gate row with reason three carries this content already — surface it verbatim rather than re-deriving it. Never merge a lane whose Windows behavior is unknown; unknown is not parity. Failure classes and the canonical helpers: `../quality/references/windows-quirks.md`. **Windows proof gate.** For a Windows-relevant stack entry, require the native Windows foundation check to be terminal-green on the bound head. Require the packaged Windows check when packaging or native bundle contents changed. Computer Use evidence is capability-specific: native OS capture/control may be explicitly blocked while App Control and proof ingestion remain supported and tested. Clean-host Stable/Beta coexistence, second-account pipe denial, restart/reboot, installed-update, and GUI artifacts remain named external proof blockers until captured; never mark them proven from simulated tests. A stack entry cannot reach `ready-stacked` while any of them is required at its position and still uncaptured — record it as `blocked` with the scenario id, or defer it to a named higher layer in `deferredProofScenarios`. --- ## ADE deltas to the playbook **Poll with one command, every time, and read all of it.** The poll is `node scripts/ship-poll.mjs --pr ` (add `--text` for a readable summary). Run it on every wake-up, before every push (fixes, rebases, held commits, after rerunning a flaky job), and immediately before the merge. It reads CI, review bots, every open review thread whatever its age, findings that bots put in a review body outside any thread, new comments, bot notices and base movement in one call. Do not hand-write a poll: a CI-only or time-filtered check is how PR #1469 nearly merged over a CodeRabbit finding and seven unanswered threads. - `next: fix` — fix every item: CI failures, open threads not yet addressed, `reviewFindings`, `newComments`. Record each handled thread or comment id in `addressedCommentIds` and each review-body finding key in `addressedReviewFindings`. - `next: resolve-threads` — the code is fixed but the thread is unanswered. Reply with the fix commit and the test that pins it, or the reason it is rejected, then resolve the thread. - `next: merge` — the only state in which Phase 3c may merge. - `botNotices` — a bot that could not run reviewed nothing. Name it in the final summary; never report it as clean. **Every push restarts the review bots — Greptile especially.** Pushing a new commit re-triggers Greptile and Codex from scratch; an in-progress Greptile review (`Greptile Review` status stuck `pending`/`IN_PROGRESS`, often 15-25 min) is *cancelled and restarted* by the next push, so a rapid fix-every-iteration cadence means Greptile never actually lands a re-review. Consequences: - **Batch all fixes for an iteration into ONE push**, then genuinely wait for Greptile to reach a terminal state before pushing again. Do not push a follow-up while its status check is still `pending` — you'll just reset its ~20-min clock. - Codex re-reviews fast (~3-5 min) and tends to surface the *next* instance of a bug class each round (e.g. you pinned 2 of 4 cleanups → it flags the other 2). **Sweep the whole class in one iteration** (every cleanup pinned, every mutating call guarded) so you don't trade N fast Codex rounds for N Greptile restarts. - When deciding to merge: a perpetually-restarted Greptile that never completed on the latest commit is not a "still reviewing" signal to wait on forever — it's a signal you pushed too often. Once Codex is clean and CI is green on a commit you have NOT pushed over, let Greptile finish that commit, then merge. **No bot signal means off, not pending.** A review bot blocks Phase 1 only when there is positive evidence that it started for the current head: a queued/pending/in-progress check, a review, a trigger acknowledgement, or a current-head comment. After one full 12-minute post-push grace window, if every available ADE and GitHub surface shows no check, review, acknowledgement, or comment for that bot, classify it as `inactive` / `not-triggered`, treat that as terminal-neutral, and continue. Record it under `inactiveReviewBots`, never `pendingReviewBots`. Do not schedule a second wait for a bot with zero evidence. If branch protection requires an absent check, Phase 3c will surface that as a merge-policy block. **Rebase only on real conflicts or a stale quality base.** `behindBase` alone does not normally trigger a rebase. The one safety exception is base movement after quality validation: the final tree is no longer the reviewed head tree, so ordinary merge mode rebases and reruns the canonical quality procedure even when GitHub reports a clean merge. Stack mode instead invalidates the current and upstack bindings and returns `stack-coordinator-sync-required`; it never rebases and never pushes another layer, though it does push its own layer branch in Phase 0 and Phase 3b. Otherwise, skip needless rebases. **Bot pings by iteration.** Never ping GitHub Copilot and never ping `@codex` — neither is an expected review signal here, and Copilot quota exhaustion otherwise leaves the loop waiting forever. No push, initial or fix-iteration, gets a direct review ping. For a >250-file diff only, ping `@greptile` and `@coderabbit` (separate comments). Phase 1 still waits for the expected review signals to settle before fixing. This is the playbook's Phase 4 rule — defer to it for exact bodies. **Merge needs admin.** `main` is ruleset-guarded — `gh pr merge --squash --match-head-commit "$QUALITY_VALIDATED_SHA"` will show BLOCKED. Retry with `gh pr merge --admin --squash --match-head-commit "$QUALITY_VALIDATED_SHA"`; the ruleset's non-linear-history rule can still reject `--admin`. Do not fall back to a locally-created commit: it would not be the reviewed and CI-tested PR merge result. Exit blocked if both direct `gh` paths fail. After a successful merge, run **Confirm the validated merge result**. Do NOT pass `--delete-branch` (it fails from a worktree); delete the head ref server-side via `gh api -X DELETE "repos/{owner}/{repo}/git/refs/heads/"`. **Fix discipline (every fix agent must follow):** (1) Fix CI and review together in one push, only after BOTH signals are terminal — review fixes routinely cause new CI failures. (2) Never run a full vitest shard or the whole suite inside the loop; run only the failing test file or the touched package's check. This is the playbook's "Fix discipline" block — cite it to every ci-fix / review-fix agent. **Worktree path discipline.** Every `Edit`/`Write` MUST target the lane worktree path (`.ade/worktrees//...`). After edits and before commit, `git status` from the worktree — if it's empty but you "just edited," you wrote to the wrong tree. --- ## Concurrency Use `TeamCreate` if available (one team, reused across iterations: ci-fix / review-fix / rebase / conflict-resolver agents spawned on demand); per the global git-worktrees policy, do **not** pass worktree isolation. Fallback to parallel `Agent` calls. The lead runs the poll itself — `scripts/ship-poll.mjs` prints the structured summary — and never reads raw CI logs or full threads instead; fix agents edit, the lead commits. --- ## Scheduling wake-ups (harness-dependent — pick one, stick with it) **Claude Code CLI (interactive terminal):** `ScheduleWakeup` is honored — the scheduler re-invokes `/ship $ARGUMENTS` later. Use it at the end of each iteration with the playbook cadence (270s just-pushed / 720s CI or bots running / 1800s waiting on human review). **ADE Work chat (Claude Agent SDK):** Work confidently inside the current turn, but treat `ScheduleWakeup` as unavailable in this harness. It does not start a later turn by itself, and `run_in_background` notifications are not a reliable self-resume signal. Poll synchronously inside the current turn (one bounded foreground `until ... ; do sleep N; done`), then fix/merge/exit. If the turn must end while CI or review is still in flight, arm `ade chat scheduled-work create` as in the next section before ending. Asking the user to re-ping is only for a harness that is not an ADE Work chat and has no scheduler. **No native wake in this harness:** Claude Code has `ScheduleWakeup`. Codex in a terminal can `sleep`. Any other ADE Work chat has no scheduler of its own that starts a later turn. Before ending the turn while CI or review is still in flight, arm ADE's scheduler: ``` ade chat scheduled-work create --in 12m --prompt "/ship" ``` Use `--in 4m` just after a push, `--in 12m` while CI or review is running, and `--in 30m` when only a person is left. Pass the same `/ship` arguments this run was given. Ending the turn without that command, or without an in-turn sleep, leaves the lane idle. --- ## The loop (summary — full detail in the playbook) - **Phase 0 (first run):** with state `prepping`, adopt the early PR (0.6) and go to Phase 1. Otherwise: safety rails (clean tree, GitHub origin, refuse `main`) → checkpoint → canonical commit-bound quality revalidation (delta since `qualityReviewedSha`) → push → open PR (`ade`, gh fallback) → verify the provisional binding → write state → schedule first wake. - **Phase 1 — Poll:** run `node scripts/ship-poll.mjs --pr `. Wait while it reports CI or a review bot running. After one 12-minute grace window, classify bots with zero evidence as inactive/terminal-neutral. Its `next` field is the Phase 2 route. Don't fix on a partial signal. - **Phase 2 — Decide:** merged → run **Confirm the validated merge result** and only then set `done-clean`. Real conflict → Phase 3a rebase (rebate). CI or bots running → reschedule. Both terminal, no work → 3c merge. Both terminal, work exists, `iter < 5` → 3b fix. `iter >= 5`, not merged → 3d force-finalize. - **Phase 3a Rebase / 3b Fix / 3c Merge / 3d Force-finalize** — per the playbook. Force-finalize runs once: ignore review comments (bookkeep their IDs), fix only CI, never delete/skip tests or weaken lint/tsconfig, then merge on green. - **Phase 4/5:** post only the >250-file bot pings (Phase 4 sends no ping for an ordinary push), update state, schedule the next wake (or stop per harness above). - **Stack mode:** the same phases run, minus 3a, 3c.1–3c.5, and 3d. Phase 2 routes remaining fix work to 3b and terminal-green to 3c.0 (`ready-stacked`); the spent iteration budget escalates via `stack-coordinator-fix-required` instead of forcing. --- ## Exit states | Status | Meaning | |--------|---------| | `ready-stacked` | Opt-in stacked layer is green, review-terminal, quality/test-clean, and every mandatory proof scenario is linked to the validated head or validly deferred to a named higher layer. The lane fixed its own layer; the coordinator owns restacking, base retargeting, submission, and landing | | `done-clean` | PR merged on main | | `done-max` | 5 normal + 1 force-finalize exhausted, merge genuinely blocked | | `blocked` | Unrecoverable conflict, gate failure, API error, force-finalize CI failed, a non-empty `/quality` gate awaiting an author decision, an unresolved Windows parity decision (hide / disable / remove), a missing mandatory proof scenario, or a `stack-coordinator-*` escalation | Always print the final summary (PR, branch, iterations, status, reason, per-iteration log, unaddressed items) on exit. Do NOT schedule a wake when `status` is `ready-stacked` / `done-clean` / `done-max` / `blocked`.