--- name: break description: Try to break a change the way the real world will - network faults, restarts, two things at once, users doing things out of order - on a disposable local stack, and report what broke with a reproduction. --- # /break `/verify` asks "does the change do what the plan said?". This skill asks the opposite: "how does this change fail once it meets real networks, real users and the other services around it?". Run it after `/verify` passes, on the same local stack. It is not a unit test hunt. Empty strings, unicode and wrong types belong in unit tests, unless a real producer sends them: a model filling in tool arguments is one, so an empty or odd-but-valid value it writes is in scope here. The bugs this skill looks for only exist when the whole system is running: - **Network and infrastructure:** a dependency that is slow, refuses connections or hangs; a reply lost after the side effect happened; a process killed or restarted mid-job. - **Users:** double-clicking, two tabs, going back, refreshing mid-flow, leaving and coming back after something changed. - **Services interacting:** two workers, or a background job running at the same moment as a request, events arriving twice or out of order, old and new versions running together, state changed by someone else between two steps. ## Hard rules - Never fix what you break. Report only. - Only the disposable local stack. Never inject faults into a shared, staging or production system. - Use local stubs for model and paid-provider calls. Real paid calls happen only if the charter lists them with their cost and the user approves. - A stub accepts whatever you send, so it cannot tell you the real provider would refuse the request. When the change alters what is sent to an outside service (a schema, a tool definition, a new field), check the request against the provider's real rules wherever that is free: a token-counting or validation endpoint, a dry-run mode, the provider's published schema. That is not a paid call. - A finding is something you **made happen on the running system**. Something you can only argue from reading code is a *suspicion*: list it separately, with what it would take to demonstrate. - Every finding reproduces twice before it is reported. - A failure in your own harness is not a finding. If the fault you injected also broke the thing you use to observe, say so and fix the harness, not the verdict. ## Before you start Reuse `/verify`'s setup: `.verify/setup.json`, the environment scripts and the run marker. Boot with `env.sh up`, seed with `env.sh seed`, and arm `env.sh down` on every exit path, exactly as `/verify` half two does. If there is no setup contract, stop and ask for `/verify-setup`. Create a run directory: `.verify/breaks//`. Everything you write goes there. Find the plan the same way `/verify` does. Then read the diff and the code around it. Unlike `/verify`, this skill must read code: the plan says what should happen, and only the code shows where it can fail (what calls the changed code, what it trusts, where it gives up). Code is for **aiming**. Judging still comes from the running system. ### Read the app profile `/verify-setup` writes `.verify/profile.json`: the app's services, background actors, status transitions, outside services, behaviour-changing config and the chain in front of production. Check it first: ```bash VERIFY_PIPELINE="${VERIFY_PIPELINE:-$CLAUDE_PLUGIN_ROOT/pipeline}" (cd "$VERIFY_PIPELINE" && npx --no-install tsx src/cli.ts profile-check --repo "$(pwd -P)") ``` If the file is missing or the check fails (the code has moved on), rebuild it before anything else by following section 7 of the `/verify-setup` skill; it takes a few minutes and asks at most one question. If the rebuild fails, carry on without a profile and say so in the report. Use only the slice that matters for this change: the actors and transitions that touch the changed files or the tables they write, the outside services the change calls, and the edge chain. Open each cited line you rely on and drop any entry the code no longer supports. Entries are leads, not proof: a missing entry never means nothing else does this. ### Reaching the state under test The repo's seed usually creates tenants, not the mid-flight state you want to attack. Getting there is most of the work, so do it cheaply: - Use the product's own test doubles first: local sandbox, fake providers, stub servers. - Point model and provider clients at a local stub you control when the attack is about how the system handles their answers (invalid, truncated, refused, slow, hanging). It costs nothing and lets you choose the bad answer. - When an upstream stage is expensive (a real agent run, a paid pipeline), you may seed its output with SQL, writing the same rows the real code writes. Say so in the report. - Give each attack its own project or tenant, so one attack's leftovers do not change the next one's outcome. If the setup contract is missing something the stack needs (an env var, an image that will not pull), work around it and list it under **Setup gaps** in the report. ## 1. Map the seams A seam is anywhere the change hands work or state to something else. Write `seams.md` with one short entry per seam: - **Actors** that touch the changed state: API handlers, workers, the reaper or lease expiry, sweepers, schedulers, boot-time migrations, webhooks, the browser, external services. Include background actors the diff did not touch but that read or write the same rows. They are where the collisions come from. - **State** each actor reads or writes, named as tables, keys, files or messages. - **Moments between steps**: for each changed write, find its transaction boundary and every actor that can read or change the same state before the next write. After a commit and before a reply, between two transactions, between a claim and a completion, between a freeze and a publish. Write down the exact interleaving you would force ("write A commits, then another actor changes the same row, then write B runs") and what each actor should do. These interleavings are your best attacks. ### Other producers A fresh stack holds only what the happy path wrote. Real systems hold what everyone produced. Find the data the change consumes from the diff, not only from the plan: for every function, check or query the change edits, list what already flowed through it on the base commit. A change aimed at a new kind of item often edits code the older kinds still pass through, and the plan will not mention them. A path through code the change edited is never out of scope, whatever the plan is about. For each piece of data the change consumes (rows, messages, files, headers, payloads, config, URLs, anything it parses): 1. **List the producers.** Older versions of this code, other services and background jobs, people (manual fixes, imports, admin tools), outside systems (webhooks, SDKs, browsers, build tools), and failures (a write cut short, a retry, a duplicate delivery, a crash between two writes). 2. **List the shapes each producer could leave.** Ask for an exhaustive list with a source for each item (a tool and version, a doc, a named failure), never for "some sample data": samples cluster on the common case and miss the shape that breaks things. Include at least one shape that contradicts the plan's own examples. 3. **Produce each shape the cheapest honest way.** Run the real producer when you can (the base commit, the provider's CLI, a real build tool, a recorded browser session). Otherwise write the shape directly (SQL, a crafted request) and say so. For your own infrastructure (the proxy in front, the deployed config), test the path the profile names first. For inputs that vary by customer (their build tools, SDK versions, proxies), try the few most common variants and name the ones you skipped. Skip this for data the change does not consume. Run the attacks on top of this data, and do restart and deploy attacks after it is loaded. Put the producers you covered, and the ones you could not, in the charter. ### Other consumers For each piece of data the change produces or changes the meaning of (a row, a stored payload, a message, a status), find everything else that reads it: other endpoints, tools, notifiers, the UI, reports, and older versions still running during a deploy. When the data ends up in a prompt, the model is a consumer too: check what the prompt tells it, not only what is stored. Trace them in the code now, for this change; do not rely on a list from the profile. After each attack, read the data through every consumer and check they agree on what it means (whether it exists, what it holds, what state it is in), even if they show it differently. If the map is empty (no second actor, no gap in time, no outside call touches the changed state), say so and stop. There is nothing for this skill to attack. A change to copy, styling or docs usually ends here. ## 2. Write the charter, then stop ### Which attacks apply Choose from what the change does. The categories are generic; turn each into concrete cases for this app from the profile and the diff. | If the change | Try | | --- | --- | | does two writes in a row | a crash between them | | has a background job, retry or scheduler | a poison item, a crash mid-job, two workers | | lets a person act on something a job is using | the person acts mid-job | | calls an outside service | the service down, slow or hanging | | reads config or a feature flag | the shipped defaults: flags unset, optional keys empty, exactly what the compose file and example env pass through | | ships an operator script or deploy step | run it end to end on the stack, as the operator would | | stores data others read | check all consumers agree | | runs at startup | a restart on data left by older versions | | adds a rule that decides (a threshold, filter, detector, classifier) | both directions: it fires when it should not, and it stays silent when it should fire | | runs a model or an agent, or acts on what one returns | the model and agent attacks below | When the plan states a goal ("no customer is charged twice") as well as a mechanism ("a lock around the charge call"), test the goal. Inputs the mechanism did not foresee are where it fails. ### Starting states Run each attack from the starting states that apply, not only the clean default: - retries: first attempt, and the last attempt before the job gives up - concurrency: one run, and another run of the same thing already in progress - data age: fresh, and data left behind by older versions or crashed runs - config: filled in, and default or missing Name the concrete state in the charter ("fix job at attempt 2 of 3"), and how you reached it. ### Check a guard before building on it When the profile, the plan or your own reading says a write has no guard, open the code and confirm it before planning an attack around it. Reading gets these wrong. ### Hypotheses Pick hypotheses from the attack list below. Each one reads: > If **\** happens during **\**, then after the system settles, > **\** still holds. Rank them by how bad the worst outcome would be for a real user. Run the two or three highest-risk interleavings from section 1 first, then the rest as time and budget allow. Write `charter.md` with the hypotheses, the settled-state rules (section 3), the stack you will use, and a rough time estimate. List every attack from the list below that you considered and skipped, each with a one-line reason ("no second actor writes this row"). Scale the effort to the change: a single-flow change gets two or three attacks, a change to background jobs gets the full list. Show it to the user and stop until they say go. ### The attack list 1. **Poison item.** One input that always fails, however many times it is retried. Count retries and spend, and check whether the work around it still finishes. 2. **Lose the reply after the side effect.** Let the write commit, then kill the process or drop the response before the caller hears back. Re-run, retry or redeliver. Count results: exactly one. 3. **Real network faults, not clean errors.** Point a dependency at a closed port, make it hang (`docker pause`), make it slow. A clean HTTP 500 is the easy case; hangs and refused connections are where the bugs are. 4. **Two at once.** Two actors on the same state at the same moment: two workers, a background job and a request, two tabs, a double-click, two users on one record. 5. **Restart or deploy mid-flight.** Restart the services the change touches while work is in progress. Boot-time migrations and sweepers run again. Run the old and new versions side by side where a deploy would. 6. **Change state between two steps.** Between the moments from section 1, let another actor act: a person clicks a button, a sweeper runs, retention deletes, an external service changes its mind. 7. **Crash at every commit point.** For each database write in the changed flow (from section 1), kill the process right after that write commits, restart, and check the rules. One commit point per attempt. 8. **Prod parity.** Compare the env vars and default hosts the change needs against what the deployed config provides. Put the profile's edge chain in front (a local proxy standing in for each hop) where the change handles requests. ### Model and agent attacks A model or agent returns well-formed things the surrounding code can mishandle. Nothing crashes, so the attacks above miss them. Use these when the change runs a model or an agent, or acts on what one returns. Point the client at a local stub to choose the answer; use a cheap real call only where the charter lists it with its cost. 1. **Every way a run can end.** Finished, hit a limit (turns, tokens, time, money), cut off mid-answer, refused, empty, errored. Produce each one. Only "finished" may count as a result; the others must each land somewhere sensible and be told apart. When the answer is streamed, also produce the reason arriving in pieces: repeated, out of order, or followed by a later piece that leaves it empty. 2. **Checks on answers, both ways.** A wrong answer that passes the format check: it claims evidence it does not give, cites something that does not exist, contradicts itself, or answers a different question. And a right answer the check rejects: build it from real material (the exact text a real tool prints, real names, real quoting), not a tidy example. Follow each rejected answer to where it ends up. 3. **Tool trouble.** A tool errors, returns nothing, returns something huge, is slow, returns plausible but wrong data, is unavailable, or returns text that contains instructions. The run should fail honestly, not reach a confident answer anyway. The other direction too: the arguments the model sends a tool can be empty, missing, extra, or name something that does not exist. The tool should refuse clearly or fall back, not fail the whole run. 4. **Loop control.** The agent repeats a step, never stops, stops early, or calls its final tool more than once. Check the spend cap is enforced before the next call. 5. **What the model is shown.** Two checks. Complete: list what the model needs to make the decision it is asked for, and confirm each item reaches the prompt at realistic sizes (a real-sized repository, history or record, not a fixture), not cut off or left out. True: for each input shape from "Other producers" (odd, empty, huge, cut short to fit), the prompt does not mislead. Where a cheap real call is approved, check what the model concludes. 6. **Retries around a model.** A retry can answer differently, mix with a partial result from the first attempt, or multiply with the client library's own retries. 7. **Same input, several runs.** Run one decision at least three times. Every guarantee must hold on every run, not on the lucky one. 8. **One bad answer, one bad item.** One malformed or rejected output fails its own item, never the batch it arrived in. ### Optional: randomise the timing Forcing the exact moment (below) is the main tool. Random fault times are a cheap extra sweep for moments you did not think of, but they rarely land in a gap that is only a few milliseconds wide, and they never create a starting state such as a last retry. If you run a sweep: 1. Name the fault and the window it may land in, from one observable point to another ("SIGKILL the worker, any time between the job being claimed and it completing"). 2. Run at least 20 attempts. Each one resets to the same starting state, picks a random time inside the window, records that time, injects the fault, lets the system settle, and checks every guarantee. 3. Save a short timeline per attempt: the fault time and the product's own log lines and state changes around it. The timeline of a failing attempt is its explanation. Report the fault times of the failing attempts. Re-running at those times will usually, not always, repeat the failure, because thread scheduling is not controlled; say so. Use a forced window (below) when the random attempts point at a narrow gap you then want to hit every time. ### Widening a race window Real races are narrow. Two concurrent requests or a `docker pause` rarely land in the gap by luck. Force the interleaving you wrote down in section 1, as long as the state you force is one production can reach: - Hold a row lock from a second `psql` session (`BEGIN; SELECT ... FOR UPDATE;`) so an actor blocks at the exact point you want, then release it. - Move a lease or schedule time into the past with SQL so the reaper or scheduler acts now instead of in ten minutes. - `docker pause` a container at the moment you care about and `docker unpause` it after the other actor has run. - Run two requests together with `&` and `wait`, or with a barrier. Check that the first write actually committed before you release the competing actor. A lock can block the write you meant to race and give you a different interleaving than you think. If a SQL edit or pause creates an ordering the application cannot reach, discard the result. Moving a lease, retry or schedule time into the past stands in for time passing and is always fair. Setting a status column to a value the code would never write is not. Record in the finding how you widened the window and what makes the same interleaving happen in production (a slow query, a long pause, a timeout). A forced state production cannot reach is not a finding. ## 3. Guarantees Name each guarantee (`diagnosis_kept`, `charged_once`, `no_stuck_jobs`) and give it the query or request that checks it. Check every guarantee on every attempt, not only the one the attack was aimed at, and report each as a count: "held on 17 of 20 attempts". Count only attempts that actually exercised the guarantee, and list guarantees no attempt exercised separately. A guarantee that held on every attempt is evidence; a failure rate tells the reader how often a user would hit it. `all_consumers_agree` is always one of them when the change stores data others read: every consumer from section 1 agrees on what the data means. Check two things: what each actor saw at each step (the response to the user, the error the worker logged), and the state left behind after a stated recovery deadline. A final row can look fine while a user got a 500 on the way. Write the rules in data terms **before** attacking, from the plan and from common sense about the product, never from what the code happens to do. For each workflow, write down its permitted end states first. - Each item reaches one of its permitted end states, or is still active with a live owner and deadline. `needs_human` is a failure only where the plan says the system recovers on its own. - For each side effect, name what the product promises (at most once, at least once, exactly once) and what makes two of them "the same". Count duplicates against that promise, and keep attempts, committed effects and paid calls apart. - Good data survives. A later failure never overwrites an earlier success. - Spend is accounted: one ledger row per real paid call, and no loop of paid calls. - Only a finished answer is a result. A model or agent run that stopped for any other reason never writes one, and a rejected answer reaches an end state within a fixed number of attempts. - The user was told the truth: no 500 for a normal action, no success shown for a failure, no spinner that never ends. - The views agree once polling or caches have caught up: what the UI shows, what the API returns and what the database holds. Capture anything wrong shown before then. Add rules specific to the change. Each rule names the query or request that checks it. ## 4. Attack Go through the charter in order. For each hypothesis: 1. Record the steady state (the rule queries) before touching anything. 2. Inject the fault at the moment named. 3. Remove the fault and wait for the system to settle, with a deadline. Give the reaper, retries and schedulers the time they would take (shorten it with SQL if needed, and say so). Wait with a polling loop and a deadline, not a single sleep. A missed deadline is inconclusive until you show recovery cannot happen. 4. Run every guarantee's check again and save the raw output under `attacks//`. 5. Write down what happened: for each guarantee, how many attempts it held on, or could not run and why. Keep an action log for each attack (`attacks//log.txt`): every action you took and what came back, in order, with timestamps. Reproductions come from this log. Weave the run marker into everything you create so the evidence is unambiguous. Do not tweak an attack until it breaks something. If a hypothesis holds, it holds; move on. ## 5. Confirm each finding Before a break becomes a finding: - **Reproduce it a second time** from a clean state, with the same forced interleaving and a saved trace of the order things happened in. A failure seen on several randomised attempts already counts, with their timelines. - **Cut it down** to the fewest steps that still cause it. - **Check it against base.** Run the same attack on the base commit when the stack can boot there. Label the finding `new` (the change introduced it), `not fixed by the change` (the old behaviour survives in a case the plan promised to cover), `pre-existing` (unrelated to the plan), or `base not checked` with the reason. - **Name who gets hurt and how.** Severity comes from that: lost or wrong data and money loops are high; a confusing message with a working retry is low. ## 6. Report Write `report.md` in the run directory, in plain language: - **Findings**, worst first. Each has: one-line headline, who gets hurt, a runnable reproduction script saved under `repro/`, the steps in words, how any race window was widened and why production can reach it, the settled-state rule it broke with the raw before and after, `new` or `pre-existing`, and severity. - **Suspicions**: things you believe from the code but could not make happen, with what it would take. - **Guarantees**: a table of every guarantee against every attack, as "held N of M". - **Held**: every attack that did not break anything, one line each. This is what makes an empty findings list believable. - **Not attempted**: each skipped hypothesis and the reason. - **Spend**: paid calls made and roughly what they cost. - **Setup gaps**: anything the setup contract was missing. - **Data coverage**: the producers and shapes you produced, the consumers you checked, and the ones you could not. If you cannot write files (for example, you are running as a subagent), return the report as your final answer instead. Tear the stack down and print the report path. Never file issues or fix anything. Offer to draft an issue for each finding and let the user choose.