--- name: codex-pr-review description: >- Run an independent local Codex (gpt-6.1-sol, medium) review of a GitHub pull request and relay its verdict. Use for every PR an agent opens or updates, and whenever the user asks to review a PR "with Codex". The review is posted on the PR as approve / request-changes; it never edits, commits or merges. --- # Local Codex PR review One command does everything. Run it in the background and wait for it; do not re-verify the Codex binary, the scripts, the token minter or the gh login first. The script fails loudly if anything is missing. ```bash nohup scripts/codex-review/codex-pr-review.sh [extra-prompt-file] > 2>&1 & ``` Then wait on the log (`until grep -q 'codex-pr-review: done\|FAILED' ; do sleep 15; done`) and keep working. Every exit path prints exactly one of those markers, including preflight failures. A run normally takes 5-20 minutes. The watchdog kills an attempt, with every process it spawned, after 8 minutes without an event from the Codex process, or 30 minutes total, and retries once unless the verdict was already posted. Never launch a bare `codex exec` for a review. ## What the script does - Fetches the PR head into the sibling worktree `-pr-`, so the main checkout is never touched. If GitHub has not advertised the PR's Git pull ref yet, it fetches the exact head from the PR API and rechecks that head before starting the review. It never substitutes a moving branch tip. The tool marks worktrees it created with a `-pr-.codex-review-owned` sentinel; on reruns it resets and cleans only those (ignored files such as `node_modules` are kept) and refuses a path at that location it did not create. A per-PR lock refuses a second concurrent run; a lock whose owner is dead is reclaimed through a short-lived `.reclaim` mutex, so two reclaimers cannot clobber each other. Cancelling the script (Ctrl-C, SIGTERM) stops the runner and every process of the attempt before the lock is released; the runner pid is kept in the lock so a runner that outlives a hard-killed wrapper still counts as live. A worktree created by an earlier version of the script has no sentinel: add the file by hand (any content) or remove the worktree with `git worktree remove`. - Chooses the posting account. GitHub refuses a formal review from the PR author, so a PR authored by the logged-in gh user posts through a `mentra-release-coordinator` GitHub App token minted by `mentra-release-coordinator-token.mjs`; anyone else's PR posts from the logged-in account. Override with `GH_ACCOUNT=app|own`. - Writes the standard prompt: read the diff against the base, read every existing comment (bots included) and judge each on its merits, favour low complexity and a coherent design over patchy fixes, no edits or pushes, and finish with `gh pr review --approve` or `--request-changes` whose body starts with `Reviewed by local Codex (gpt-6.1-sol, medium).` and lists what was checked. If GitHub refuses the formal review it falls back to `--comment` with `Approve.` or `Request changes.` as the first line. - Always adds an "Is this change needed?" step: Codex must judge whether the change is needed or desirable for the project at all (real problem, existing or planned alternative, surface area and maintenance cost versus benefit, smaller change possible) and let that weigh in the verdict, not only implementation quality. - Runs `codex-review.sh` (watchdog + one retry) and prints Codex's final message. The runner starts `codex exec --json` by default, or the small stdio app-server adapter for a configured desktop project. Both use their own server event stream as the heartbeat, so other Codex sessions on the machine can never be mistaken for this one; every process of the attempt carries a unique `CODEX_REVIEW_ATTEMPT` marker so it can be terminated even if Codex exits first. Success means a review carrying the marker was found on the head commit after the run started, checked through the GitHub API by `review-receipt.sh`; the runner performs the same check before any retry so a verdict posted just before a crash is never duplicated. Artifacts: `$CODEX_REVIEW_HOME/-pr--/{prompt.txt,last-message.txt,runner.log,events-.jsonl}` (default `~/.codex-reviews`). ## Configuration ### Use a desktop Review project On the host that runs reviews, create a folder and add it as a local Codex project named **Review**. Save its absolute path once: ```bash mkdir -p "$HOME/dev/Codex-Reviews" "${CODEX_REVIEW_HOME:-$HOME/.codex-reviews}" printf '%s\n' "$HOME/dev/Codex-Reviews" > "${CODEX_REVIEW_HOME:-$HOME/.codex-reviews}/project-directory" ``` Future skill invocations use `codex app-server --stdio` and the documented `initialize` → `thread/start` → `turn/start` protocol, with that folder as the saved session's starting directory. The server creates a persistent interactive session and the adapter gives it a ` # · review` title. This matters: `codex exec` has a separate session source that default history lists exclude; changing its cwd alone does not establish desktop visibility. `--thread-source` is an analytics label and does not change that history source. The reviewer receives the separate PR worktree path for all code inspection and tests; worktree ownership, locks, watchdogs and verdict checks still apply. Existing sessions are not moved. This is configured per host: a Mini review needs the folder registered as a project on the Mini's connection. Do not use `--ephemeral` or a separate `CODEX_HOME`, which would hide sessions from the usual desktop account. For durable project assignment, initialize the same host's `codex app-server --stdio` and call its documented `project/list` method. Follow pagination and match the project's `roots[].path` to the exact configured directory. Save that returned ID in `$CODEX_REVIEW_HOME/project-id` (default `~/.codex-reviews/project-id`) or set `CODEX_REVIEW_PROJECT_ID`. Desktop-tool `list_projects` IDs can use a different namespace: do not copy those IDs into this configuration. The adapter passes the app-server ID as `thread/start.projectId` and verifies the response. It never reads or edits Codex's database. Without an ID, grouping depends on the registered folder; verify visibility on the next actual review rather than creating a duplicate review as a probe. Override the saved path with `CODEX_REVIEW_PROJECT_DIR=/absolute/path`, or set it to an empty string for one invocation to start in the PR worktree instead. Without a saved path or override, the existing `exec` behavior is unchanged. A configured missing directory is an error, so reviews cannot silently appear somewhere else. An explicit directory override ignores the saved project ID; also pass `CODEX_REVIEW_PROJECT_ID` if that different folder needs a durable assignment. Project reviews require Node.js and a Codex version supporting the app-server protocol. Incompatible protocol, failed/interrupted turns and missing final text fail the attempt; they never produce a successful review merely from progress. Protocol reference: [Codex App Server](https://learn.chatgpt.com/docs/app-server), including its `sourceKinds` history filter. The adapter uses the server's normal session source; it does not relabel existing sessions or spoof a source. | Variable | Default | Purpose | | ------------------------------------------------------------- | ------------------------------------------------------ | --------------------------------------------------------------------------------------------- | | `CODEX_BIN` | `codex` | Codex CLI binary. Point it at a specific install if the shim on PATH does not support `exec`. | | `CODEX_REVIEW_MODEL` / `CODEX_REVIEW_EFFORT` | `gpt-6.1-sol` / `medium` | Model and reasoning effort. | | `CODEX_REVIEW_HOME` | `~/.codex-reviews` | Where prompts, final messages and runner logs are kept. | | `CODEX_REVIEW_PROJECT_DIR` | Saved `project-directory` file, otherwise PR worktree | Starting directory for future sessions; register that folder in Codex to group reviews. | | `CODEX_REVIEW_PROJECT_ID` | Saved `project-id` file, otherwise unset | Optional durable project assignment for the configured project on this host. | | `MENTRA_RELEASE_COORDINATOR_KEY` | `~/.config/mentra-release-coordinator/private-key.pem` | App private key (0600) for posting on your own PRs. | | `GH_ACCOUNT` | auto | `app` or `own`, see above. | | `STALL_SECONDS` / `MAX_SECONDS` / `ATTEMPTS` / `POLL_SECONDS` | 480 / 1800 / 2 / 5 | Watchdog limits and poll interval. | | `LOCK_STALE_SECONDS` | 60 | Age after which a lock directory without a pid file is reclaimed. | The App must be installed on the target repository or the mint fails with HTTP 422 and the run falls back to a comment review from the logged-in account. With an already configured GitHub App installation credential, set `GH_ACCOUNT=own` to retain that credential. An explicit account skips the user-only `/user` lookup; it does not grant additional access or change the review receipt requirements. The usual self-review comment fallback still applies. Controllers can inspect `scripts/codex-review/codex-pr-review.sh --capabilities` before admitting this mode. ## Tests `bun test scripts/codex-review` runs the lifecycle suite with fake `gh` and `codex` binaries: markers on every exit path, worktree ownership, lock behaviour, receipt-before-retry, and stall handling. Project transport tests cover the real JSON protocol, completion ordering, final-text validation, cancellation and detached-child cleanup. They do not claim a live desktop sidebar check. CI runs them on changes under `scripts/codex-review/`. ## Extra prompt file Only when the review needs context Codex cannot get from the PR itself (a design decision made in chat, a known-flaky test to ignore). Do not use it to steer the verdict on a specific comment; let Codex weigh comments itself. ## After the run - Relay the verdict and Codex's list to the user. Do not merge. - If the PR is the agent's own and the verdict is request-changes: read the bot comments too, fix real defects with a coherent design change, say which comments were ignored and why, push, and rerun the script on the new head. - If the PR belongs to someone else, report only; do not rework unless asked. - On `codex-pr-review: FAILED`, read `runner.log`, report the failure, and do not claim a review was posted. - Remove completed review worktrees first, before finished fixer worktrees, once the canonical outcome and needed reproduction material are preserved outside the checkout. Wait for the wrapper, runner and their children to exit; keep a checkout still needed by active or saved review follow-up. - Before removal, check process use, incoming dependency symlinks and Git common-directory consumers. Preserve source commits/refs, review receipts, logs, evidence and shared dependencies. Use `git worktree remove`, never `--force`; if the checkout is still needed or removal refuses, keep it and report why. Store shared dependencies outside disposable review worktrees.