--- name: run-integ description: Run integration tests (deploy + destroy) against real AWS. Use when you need to verify cdkd works end-to-end with actual AWS resources. argument-hint: " [--synth-only] [--no-destroy]" --- # Integration Test Runner Run integration tests against a real AWS account: deploy actual resources, verify, clean up. ## Arguments - `test-name`: which test (see `ls tests/integration/`). If unspecified, ask via `AskUserQuestion` showing the options. - `all`: run all tests - `--synth-only`: synthesis only, skip deploy/destroy - `--no-destroy`: deploy but don't destroy (debugging) - `--deploy-args ""`: forward extra args to the `cdkd deploy` invocation verbatim (destroy is unaffected). ## Steps 1. **Rebase, then build**: `git fetch origin` and rebase onto current `origin/main` (merge it when a force push is denied) BEFORE the run — a stale base verifies code other than what will merge, and nothing warns. Then `vp run build` so `dist/` is current. **Never build beside a live fixture in this tree**: any build (`/check`, `/verify-pr`, `vp run verify`, `vp run runtime:smoke`, a building fixture, …) rewrites `dist/`, and the fixture's next `node dist/cli.js` (its cleanup trap's too) dies `ERR_MODULE_NOT_FOUND`. In a set, run each building fixture (`grep -lE '^\s*\(cd [^)]*&& vp run build[ )]' /*.sh`; a bare `vp run build` also matches error text) alone, BEFORE the others. After a build beside a live run, re-scan per steps 6-7, record that fixture `FAIL`, and re-run it. 2. **List available tests**: `ls tests/integration/` — never a hardcoded list. 3. **Determine state bucket**: account via `aws sts get-caller-identity --query Account --output text`, then `cdkd-state-{accountId}` (region-free default). If absent, fall back to the legacy `cdkd-state-{accountId}-us-east-1` and note the deprecation. 4. **Pre-flight orphan scan** (mandatory): a prior run killed mid-deploy leaves orphans matching the stack about to deploy, and cdkd's diff does not see them, so the deploy attempts CREATE and collides. **Pick the region first**: `us-east-1`, unless the fixture's `verify.sh` header names a constraint; its last ledger note shows a region that passed. In a parallel set, each fixture whose header has a `SAFETY NOTE` (bootstrap-marker or asset-storage ownership, #4063) gets its OWN region, probed for what that header needs absent or present, passed via `AWS_REGION` or the header's region variable. A fixture that runs `cdkd gc` runs ALONE, after the set: gc refuses while ANY stack holds a lock, its batch-mates' included. Every resource-scan `--region`, `AWS_REGION` and synth/deploy/destroy `--region` in steps 4-7 then uses the fixture's region. Synth first (for the stack name and resource types), then scan: ```bash # Always: # deployments/** is retained history, not an orphan (below). aws s3 ls s3:///cdkd// --recursive --region us-east-1 | grep -v '/deployments/' aws iam list-roles --query 'Roles[?contains(RoleName, ``)].RoleName' --output text aws lambda list-functions --region us-east-1 \ --query 'Functions[?contains(FunctionName, ``)].FunctionName' --output text # When the template uses Lambda EventSourceMapping (orphan ESM = AlreadyExists + rollback): aws lambda list-event-source-mappings --region us-east-1 \ --query 'EventSourceMappings[?contains(FunctionArn, ``)].[UUID,FunctionArn]' --output text # When the template uses VPC + Lambda VpcConfig (hyperplane ENIs outlive the function): aws ec2 describe-network-interfaces --region us-east-1 \ --filters "Name=description,Values=AWS Lambda VPC ENI-*" \ --query 'NetworkInterfaces[].[NetworkInterfaceId,Status]' --output text ``` **Any orphan found → abort** with the orphan list and cleanup commands; do NOT deploy on top of orphans. Under the S3 prefix every key outside `deployments/` counts (`state.json`, `lock.json`, `rollback-journal.json`, a legacy region-less `state.json`); `deployments/**` is event-log history a clean destroy RETAINS unless `--purge-events`. **Except a `lock.json` whose `expiresAt` is in the future: a LIVE peer on the same fixture** — wait, then re-scan. Expired: a killed run's orphan. The LAST call before step 5 re-runs the S3 listing (a batch's earlier scan is stale, #4080) and, since a peer BETWEEN commands or not yet deployed holds no lock, checks this host: a `pgrep -f verify.sh` PID whose cwd (`lsof -a -p -d cwd`) ends in `tests/integration/` is a live peer — wait for it to exit. **A fixture whose assertion runs through the journaled-orphan settle, `cdkd rollback` or `cdkd destroy` reads EVERY stack's record under its state prefix, and fails closed on one it cannot read**, such as a peer session's newer-schema record. Give such a run its own bucket: `cdkd bootstrap --state-bucket --no-assets`, pass `STATE_BUCKET=`, then empty its object versions and delete it after step 6. A fixture already deploying under its own `--state-prefix` (`remove-protection-journaled-orphan`) is immune. 5. **Run the test(s)** **Dispatch**: a `verify.sh` in `tests/integration//` owns its own deploy + verify + destroy cycle; the standard flow below is for plain smoke tests. **Check for `verify.sh` first**: rc=127 from `bash verify.sh` means no such file (said on STDERR); such a fixture takes the standard flow below. - `cd tests/integration//`; `npm install` if no `node_modules`. - **If `verify.sh` exists**: `AWS_REGION=us-east-1 STATE_BUCKET= bash verify.sh` — steps 6/7 STILL run after. Propagate its exit code; never swallow a failure. - **Otherwise** (standard flow): - `node ../../../dist/cli.js synth --region us-east-1` - **Multi-stack apps**: if synth lists more than one stack, pass `--all` to deploy and destroy (otherwise they fail with `Multiple stacks found`). - `node ../../../dist/cli.js deploy [--all] [] --region us-east-1 --state-bucket --verbose` - `node ../../../dist/cli.js destroy [--all] --region us-east-1 --state-bucket --force` **Never run it unwatched, and do not reach for `timeout`** — it is not in stock macOS (its absence is exit 127 in 0s, which reads as instant completion). Shell watchdog, firing made visible. A SET runs this whole block per fixture, in a background subshell `cd`'d into it with its own `T`, stdout to a per-fixture file that gets `echo "$T $VPID $LOG"` right after the spawn; stop one with `kill -9 -- -` (the gate refuses `pkill -f`): ```bash LOG=$(mktemp) # assign HERE: a separate block is a separate shell # Budget: 2x the last PASS's duration, floor 1500s. A FAIL row times the # failure, not a pass: walk the ledger's history back to a numeric PASS. L=../../../docs/_generated/integ-last-run.tsv; T="" P='$1==t && $3=="PASS" && $4~/^[0-9]+$/{print $4; exit}' LAST=$(awk -F'\t' -v t="$T" "$P" "$L") [ -n "$LAST" ] || LAST=$(git log -n 300 --format=%h -- "$L" | while read -r c; do git show "$c:docs/_generated/integ-last-run.tsv" | awk -F'\t' -v t="$T" "$P"; done | head -n 1) case "$LAST" in ''|*[!0-9]*) LAST=750;; esac POLLS=$(( 10#$LAST * 2 / 5 )); [ "$POLLS" -lt 300 ] && POLLS=300 # Own process group (`perl`: zsh refuses `set -m` without a terminal), so a # FIRE also reaches verify's `node` child. `>>`: cleanup output after a FIRE # must not overwrite WATCHDOG_FIRED. Before a re-run, `ps -g ` is rc=1. perl -e 'setpgrp(0,0); exec @ARGV or die' bash verify.sh >> "$LOG" 2>&1 & VPID=$! # 5s polls that end on their own: NEVER kill the watchdog (orphaned `sleep`, # false FIRE). A FIRE sends TERM so verify's trap cleans up, then `kill -9` # after 30 min (~ noecho-parameter-masking's two 15-min cleanup loops). ( i=0; while [ $i -lt $POLLS ]; do sleep 5; kill -0 $VPID 2>/dev/null || exit 0; i=$((i+1)); done kill -0 $VPID 2>/dev/null || exit 0 echo "WATCHDOG_FIRED" >> "$LOG"; kill -TERM -- -$VPID i=0; while [ $i -lt 360 ] && kill -0 $VPID 2>/dev/null; do sleep 5; i=$((i+1)); done kill -9 -- -$VPID 2>/dev/null ) & WPID=$! wait "$VPID"; RC=$? wait "$WPID" # at most 5s more grep -c WATCHDOG_FIRED "$LOG" || echo "watchdog did not fire" echo "verify.sh rc=$RC" # the verdict steps 6-11 read ``` The `grep` and `rc` lines are load-bearing: a grep of 1 is a FIRE whatever the rc (143: cleanup ran; 137: cleanup outlived the grace, so step 7 sweeps). With no FIRE, rc=137 is a manual stop and other non-zero is a FAIL: read the LOG. **Steps 6-11 are LATER calls that read this output** — a marker or a `PASS` ledger row chained into this same call is written before any verdict exists. 6. **Verify cleanup** - Step 4's S3 listing, re-run per stack — empty = no leftover state (a bare `aws s3 ls s3:///cdkd/` is never empty: retained `deployments/`). - **The state bucket is VERSIONED**, so that listing shows nothing while every prior version stays readable (`aws s3 rm` writes a delete marker). For a fixture that WRITES a secret into state (redaction / scrub / drift ones do, deliberately), "the object is gone" is not "the content is gone" — the difference is a disclosure. Check versions: ```bash # Per state/lock key the fixture touched. Non-empty = content still readable. aws s3api list-object-versions --bucket --prefix "cdkd///state.json" \ --query "([Versions, DeleteMarkers][])[?Key=='cdkd///state.json'].VersionId" \ --output text ``` If a fixture seeded a secret, grep the surviving versions for it. - Verify the AWS resources are gone, per stack name from synth output, for the types the test actually created — the per-service listing commands are in `/cleanup` step 4; use the same ones with this run's stack prefix. - FSx tests take an extra check: destroy keeps CFn parity, so `DeleteFileSystem` takes a chargeable FINAL BACKUP by default, and it usually carries NO tags, so a name scan reports clean over a live billing backup. Attribute by the persisted file-system id (`aws fsx describe-backups --region us-east-1 --query 'Backups[?FileSystem.FileSystemId==\`{fs-id}\`].[BackupId,Lifecycle]'`); if the run's fs ids are unknown, list ALL backups and flag any unattributed entry for manual review. 7. **Auto-cleanup orphans (mandatory when destroy didn't fully succeed)** — trigger when the destroy step reported errors, OR step 6 found leftover state or any resource matching the stack prefix. **Not while a PEER runs the fixture** (this run failed on its lock, or a `lock.json` under the prefix has a future `expiresAt`): the scan finds the peer's LIVE fixture, whose lock lapses between commands. Delete nothing until no live lock remains AND the prefix is unchanged for 10 minutes, then re-run step 6; what it still finds is an orphan, cleaned here (#3813): - VPC-attached Lambda failures (commonest), **in delete order**: (1) hyperplane ENIs (`describe-network-interfaces --filters "Name=vpc-id,Values="` → `delete-network-interface`; re-poll `in-use` until `available`), (2) SecurityGroups, (3) Subnets, (4) VPC. - S3 state orphans: `aws s3 rm s3:///cdkd// --recursive` (or `cdkd state orphan ''`, which also handles the lock key). - Other types: infer delete order from CFn dependency rules (children before parents). Always pass `--region`. Re-run step 6 after cleanup. **Never** end the run with orphans present. One that resists deletion after this pass is surfaced with its ID, region and what was tried. 8. **Report results**: pass/fail per test, resource counts, timing. Always state "destroy completed: 0 errors, 0 orphans" or itemize what remained. 9. **Set the `integ-destroy` markgate marker (only on full clean success)** — destroy finished with **0 errors**, step 6 found **0 leftovers**, and step 7 was skipped or re-checked clean. `mise trust` is UNCONDITIONAL and part of the block: an untrusted `.mise.toml` makes `markgate set` die naming no cause, discarding a real-AWS run that cannot be cheaply repeated. ```bash mise trust mise exec -- markgate set integ-destroy || { echo "markgate set integ-destroy FAILED — the marker was NOT recorded." >&2 exit 1 } # ABSENCE of the line is the signal, not the rc (`grep` exits 0 on `no marker`). mise exec -- markgate status | grep integ-destroy \ || echo 'NO integ-destroy LINE — markgate status itself failed' >&2 ``` **Read BOTH the exit code and the status line** — they fail in different directions. `set` exits **2** when `origin/main` is unresolvable in this worktree or the branch has no delta against the merge base; the remedy is `git fetch origin`, never re-running the integ. An untrusted `.mise.toml` is the other direction: `mise` writes to stderr and the rc can still read as success, so only `markgate status` says whether a marker exists. Run from the PR's own worktree on the PR branch, and if any success condition failed, do NOT set the marker. **The agent sets the marker, and commits and pushes step 11's ledger row, itself** — the maintainer explicitly authorizes it. When the auto-mode classifier refuses one, retry it ONCE stating that authorization; only if that is refused too, ask the maintainer (`AskUserQuestion`) to TYPE the authorization in their own words (a picked option has not always cleared a refusal), then run it yourself. Never hand either command to the maintainer; refused even then, the marker stays unset — stop and report. **Also set `integ-schema-migration`, and ONLY for a test named `schema-v-to-v-migration`**, under the same conditions. That test is the only proof a schema bump auto-migrates (deploy under vN, swap binary, read works, the next write persists vN+1, destroy clean), and `integ-schema-migration-gate.sh` blocks `gh pr merge` on a PR bumping the version constant in `src/types/state.ts` until it has run. Never set without that test's clean run. **The test-name condition is IN the block, not only in the sentence above it.** Step 9's block is unconditional, so running both after any clean run flips this marker too — the substitution the gate refuses, as one binary against its own schema proves no round trip. `mise trust` for step 9's reason. ```bash mise trust case "" in schema-v*-to-v*-migration) mise exec -- markgate set integ-schema-migration || { echo "markgate set integ-schema-migration FAILED — the marker was NOT recorded." >&2 exit 1 } mise exec -- markgate status | grep integ-schema-migration \ || echo 'NO integ-schema-migration LINE — markgate status itself failed' >&2 ;; *) echo "not a schema-migration test — integ-schema-migration NOT set" ;; esac ``` 10. **Post-run Docker sweep (mandatory for every `local-*` test)**, on top of step 6's AWS checks. A local run leaves containers and networks behind the way a deploy leaves AWS resources behind, and the run is not clean until all three listings come back empty: ```bash # All three MUST return empty, else show the orphan IDs and clean them up. # `-a`, not bare `docker ps`: a print-and-exit task container is already # `Exited` when this runs, so a running-only sweep reports clean over a # real orphan. docker ps -a --filter name=cdkd-local- --format '{{.ID}}' docker network ls --filter name=cdkd-local-task- --format '{{.ID}}' docker network ls --filter name=cdkd-local-svc- --format '{{.ID}}' ``` Subnet-overlap gotcha: `cdkd local start-service` uses the FIXED subnet `169.254.171.0/24`, so a `local-start-*` test can fail with `Pool overlaps` even when all three are empty — a foreign leftover network (e.g. cdk-local's `cdkl-svc-*`) may own the subnet. Diagnose with `docker network inspect $(docker network ls -q) --format '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}} {{len .Containers}}'` and remove the holder ONLY at 0 attached containers. A purely local run never touches AWS, so it cannot satisfy step 9's destroy conditions and does not set the `integ-destroy` marker; `local-invoke-from-state` is the exception — it exercises a real deploy + destroy as well, so it both sweeps clean here AND qualifies for step 9. 11. **Record the run in the integ ledger (MANDATORY — every run, pass OR fail)**: `docs/_generated/integ-last-run.tsv` is a COMMITTED update-type ledger (one row per test) feeding `/pick-integ`. Write it on EVERY invocation, right after step 9 (or right after a failure). Columns (TAB): `test last_run_iso result duration_s flow note`. `result` is `PASS` only at the same bar as step 9's marker (destroy 0 errors / 0 orphans; verify.sh exit 0), else `FAIL`. `last_run_iso` is UTC; `flow` is `verify.sh` or `standard`. **Use an ABSOLUTE path into the feature worktree for `LEDGER`** — the session's Bash cwd can silently reset to the MAIN worktree, and a relative write then dirties the main tree on `main`. Verify with `pwd`. ```bash LEDGER="/path/to/repo/.claude/worktrees//docs/_generated/integ-last-run.tsv" # The file already exists; if it does not, copy its header from git history # first — `>>` alone creates it headerless and the normalizer preserves # whatever header it finds (none). TEST=""; TS="$(date -u +%Y-%m-%dT%H:%M:%SZ)" RESULT="PASS"; DUR=""; FLOW="verify.sh"; NOTE="rc ok, orph clean" # One printf per row from NAMED variables: zsh never word-splits `set -- $row`. printf '%s\t%s\t%s\t%s\t%s\t%s\n' "$TEST" "$TS" "$RESULT" "$DUR" "$FLOW" "$NOTE" >> "$LEDGER" vp run integ-ledger-normalize ``` Commit the ledger update with the branch's changes. The one-row-per-test invariant is CI-enforced. When two lanes recorded the SAME test, the rebase conflicts and keep-both leaves two rows: re-run `vp run integ-ledger-normalize` after any rebase touching this file **and commit the rewrite before pushing**. Confirm with `git status --porcelain -- docs/_generated/`, never the normalizer's own output. ## Choosing the fixture Which fixture to run is a coverage judgement, not a marker lookup. - **A cross-cutting deploy/destroy change → run a BROAD fixture.** A test is "broad" iff its name is one of: ```text bench-cdk-sample lambda microservices drift-revert drift-revert-vpc multi-stack-deps multi-resource remove-protection export ``` **Only five carry a `verify.sh`; from an agent session the other four cannot be run at all** (step 5's dispatch note). Runnable from a session: **`lambda`** (the cheap default), `drift-revert`, `drift-revert-vpc`, `remove-protection`, `export`. Re-derive the split with `ls tests/integration//verify.sh`; nothing compares the copies of this list, which `/pick-integ` and `/verify-pr` also carry. **A narrow feature fixture is NOT a substitute**: a 2-stack feature fixture destroys cleanly without ever reaching the broad VPC / Lambda / multi-resource / Custom-Resource paths a cross-cutting change can break. - **A local-execution change → run a `local-*` fixture** and complete step 10's Docker sweep. A cross-cutting `src/local/` change wants at least `local-invoke` + `local-start-api`. - **A state schema version bump → run the matching `schema-v-to-v-migration` fixture** (step 9 says what it proves). ## Important - **Run `/review-pr` (and apply its fixes) BEFORE this skill when both are planned for the same PR** — the marker is digest-bound to its src scope, so a post-integ review fix stales it and forces a full real-AWS re-run. - `--region us-east-1` unless the fixture names another (step 4); always destroy after deploy; if deploy fails, still attempt destroy to clean up partial state — unless it failed on a peer's lock (step 7). - **A run blocked BEFORE its assertions is not a test failure — say which it was.** (A peer's lock, step 4's gc rule.) Record it as `FAIL` (the bar is exit-code-based) with a ledger note naming the blocker and any hand-removed resources, WAIT for the blocker to clear, then clean up what the aborted run leaked (step 7 says when). Never `cdkd force-unlock` a lock you did not take. - **Never report success on a successful deploy alone** — destroy must complete and the orphan check must pass. - **Do NOT restart Docker to fix a hung docker-dependent run (`local-*`, or an ECR asset push) — on Docker Desktop the restart IS the likelier cause**: a quit-and-reopen can leave the self-respawning backend up while the app serving the daemon's registry proxy never finishes launching. Host networking, container networking and the daemon's own pull path fail INDEPENDENTLY, so name which is down first: `curl` the registry from the HOST, `curl` it from inside an already-cached container (401 from both means networking is fine), then `docker pull hello-world`. **A pull that hangs while both curls return 401 is NOT yet the daemon: retry it with `DOCKER_CONFIG` pointing at a `mktemp -d` scratch dir whose `config.json` is `{"auths":{"cdkd-verify.invalid":{}}}`** (never `{}`: with no auth, docker falls back to `osxkeychain` and a login's token outlives the dir, go-to-k/cdkd#3651). If that pulls, the `credsStore` helper is hung and waiting will not clear it — run with that override, then `rm -rf` the dir (an ECR login writes its token there in plaintext). Only a pull that ALSO hangs there is the daemon path alone: WAIT, it recovers on its own. Do not pipe the waiting probe through `tail`, which buffers away the progress lines. Never escalate to a factory reset or deleting Docker data (it destroys local images and volumes) — ask the maintainer. Clean up your own probes: `kill`ing a `docker pull` wrapper leaves the `com.docker.cli` child. - **A fixture that discards the CLI's stderr cannot report its own failure.** `RESULT=$(${CDKD} ... 2>/dev/null | tail -1)` under `set -euo pipefail` prints the arm header and exits 1 with NO error text. The shape is banned ([../../rules/abort-capture.md](../../rules/abort-capture.md)); a failing fenced invoke prints `[verify] command exited N` plus the stderr tail, so a log ending at a bare arm header is an unfenced fixture — re-run that command with stderr attached BEFORE concluding anything. - **A FAIL that passes on re-run is not a flake** until its path is traced (#4480).