--- name: roam-measurement-design description: "Design or interpret a comparison of Roam performance, detector accuracy, retrieval or workflow value. Use to choose a fair measurement and bound its resulting claim, not for ordinary regression tests, product copy, release approval or a new benchmark platform." --- # Roam measurement design Turn a proposed improvement into a comparison that can change a decision. An exploratory result can justify the next check without establishing general benefit. A good design can be retained; more trials, labels or machinery are not automatically better evidence. ## Choose the question before the machinery Use the maintained method, not a copied checklist. Read the sections applicable to the selected claim: - Comparisons: **Comparing analytical approaches** in `docs/concepts/verification-evidence.md` owns controls and acceptance rules. - Harness/model runs: that guide's **Benchmark accounting** owns run populations. - Detector/graph accuracy: `docs/concepts/detector-evidence.md` and **Limits an agent must respect** in `docs/understanding-roam.md` own applicability and the requirements for independent/generalized claims, beyond local exploration. - Performance/workflow value: **Evidence and acceptance** in `internal/WORK-GUIDE.md`; use the bets/accounting in `internal/ROAM-NORTH-STAR.md` when choosing a value experiment. Public numbers also need the orientation's **Numbers** requirements. Read known `internal/` paths directly before declaring them unavailable; default file listings can omit these gitignored authorities. Missing private strategy is a named limit, not a veto of an adequate bounded technical comparison. Frozen sources may stand in for repository reads in a trial. Name the decision, intended population and binding resource. Distinguish an exploratory local comparison, a confirmatory study and a retrospective analysis. Fix prospective acceptance and stop/narrowing criteria before assignment; describe retrospective choices honestly instead of pretending they were preregistered. Use an owner-agreed material margin when needed, not an invented percentage, sample size or composite score. Study design is not approval to run paid trials, contact users, access other repos or change production. Compare the proposed intervention with a capable incumbent and a simpler alternative, using the comparison guide's controls; qualify deliberate changes separately. Separate a warm-query microbenchmark from indexing, startup, integration, fallback and maintenance costs at the repetition rate that matters. Verify output and accepted-outcome equivalence; removing work can be faster without being an improvement. Internal code repair needs no customer or efficacy study. ## Match the observation to the claim Match the population to the question: emitted findings for finding precision, independently sampled eligible sites for bounded recall, accepted task outcomes for workflow value. Apply the source's sampling, adjudication and applicability requirements; developer-selected examples remain exploratory, not general accuracy. Keep failed attempts, false refusals and missing observations visible under the applicable accounting method, with each metric's known denominator. A zero is not a missing value. Separate resources instead of converting tokens, seconds and owner attention into an unexplained score. Apply the comparison guide's development/confirmation separation and healthy-case controls. Choose the smallest study that can resolve the stated decision; preserve the cheaper method when added complexity does not earn its cost. ## Produce a bounded decision record Before running, record the bounded protocol and unresolved choices needed to execute; reuse an adequate design. Afterwards preserve exact artifacts, identities, dates, denominators and metric definitions. State uncertainty appropriate to the design: neither a small identical success count nor a best cell proves parity or general benefit. Do not invent intervals from insufficient records. Write what the result establishes, what it leaves unknown, and whether to retain, narrow, investigate or stop. A descriptive internal result can be useful without a publishable efficacy claim. A new public numerical claim also needs the project's reproducibility/publication requirements; preserve historical records rather than silently updating their numbers. Product wording and release approval are separate jobs, not consequences of a favorable comparison. Use existing measurement/accounting tools after verifying their actual inputs and coverage. Do not build a runner or run an expensive study merely to finish a design request. Keep protocols and dated observations private under `internal/` unless public documentation is explicitly in scope.