--- name: calc-sample-size description: Use when planning how many patients or cases a study needs before data collection (power analysis, IRB justification). Walks a decision tree to the right test and returns reproducible R/Python code and IRB-ready justification text. Analyzing collected data is /analyze-stats. metadata: triggers: "sample size, power analysis, power calculation, how many patients, how many subjects, IRB sample size" --- # Calc-Sample-Size Skill ## Decision Tree Walk the user through this tree one question at a time; do not assume answers. Reader studies, segmentation and model-comparison designs sit outside the tree: see Tests 14–17. ``` What is your primary outcome? | +-- Binary (yes/no, positive/negative) | | | +-- Paired data (same subjects, two methods)? | | +-- YES --> [5] McNemar test | | +-- NO --> How many groups? | | +-- 2 groups, superiority --> [4] Two-proportion comparison (chi-square) | | +-- 2 groups, non-inferiority --> [10] Non-inferiority / equivalence | | +-- Multivariable model --> single-predictor hypothesis test? --> [9] Logistic regression | | --> clinical prediction / AI model for use? | | +-- developing the model --> [12] Prediction-model development (Riley) | | +-- externally validating --> [13] External-validation (Riley) | | +-- Continuous (measurement, score) | | | +-- How many groups? | +-- 2 groups --> [6] Independent t-test | +-- 3+ groups --> [8] One-way ANOVA | +-- Time-to-event (survival, recurrence) | | | +-- Two groups, unadjusted --> [7] Log-rank test | +-- Multivariable / adjusted HR --> [7] Log-rank (Schoenfeld) + [11] Cox EPV | +-- Agreement (inter-rater, reproducibility) | | | +-- Continuous measurements --> [2] ICC | +-- Categorical ratings --> [3] Kappa | +-- Diagnostic accuracy (Se, Sp, AUC precision) | +--> [1] Diagnostic accuracy (precision-based) ``` ## Tests 1–11 Once the test is chosen, read `${CLAUDE_SKILL_DIR}/references/formulas.md` § Test N — the parameter table with defaults, the effect-size interpretation, the formula, R/Python code and the methodological reference. | # | Test | Use when | |---|------|----------| | 1 | Diagnostic accuracy — Se/Sp precision | desired 95% CI half-width for sensitivity or specificity | | 2 | ICC agreement (Walter 1998 test; Bonett 2002 CI width) | inter-/intra-rater agreement on continuous measurements (tumor size, angle) | | 3 | Kappa agreement (Donner & Eliasziw 1992; needs the trait prevalence) | agreement on categorical ratings (BI-RADS category, lesion present/absent) | | 4 | Two-proportion comparison (chi-square) | two independent groups (AI vs conventional detection rate) | | 5 | McNemar (paired proportions) | paired binary outcomes (two readers on the same cases, before/after) | | 6 | Independent t-test | means in two independent groups (lesion size, malignant vs benign) | | 7 | Survival / log-rank (Schoenfeld events, then patients) | time-to-event between two groups | | 8 | One-way ANOVA | means across 3+ independent groups | | 9 | Logistic regression (Peduzzi EPV + Hsieh 1998, continuous or binary predictor) | multivariable binary outcome, single-predictor hypothesis test | | 10 | Non-inferiority / equivalence | new method not worse than standard by more than a pre-specified margin, or equivalent within it | | 11 | Cox regression EPV | multivariable Cox model — enough events for stable estimates | - **Test 9**: Peduzzi EPV ≥ 10 is a minimum baseline for a single-predictor hypothesis test only. For a clinical prediction / medical-AI model intended for use, EPV-10 is outdated and reviewer-vulnerable — use Test 12 (development) / Test 13 (validation). Always report both Peduzzi and Hsieh and recommend the larger N. - **Test 10**: NI alpha is one-sided (typically 0.025); orient the difference so > 0 favours the new method. Equivalence is TOST, powered jointly. The margin must be clinically justified (see `formulas.md` § Margin selection). - **Test 11**: EPV ensures model stability, not power for a specific HR. If an HR is available, also run Test 7 (Schoenfeld) and recommend the larger N. ## Tests 12–17 (specialised designs) Each has its own reference file (parameters, method, reporting); read it once the test is chosen. - **Test 12 — Prediction-model development (Riley).** Developing a clinical prediction / classification model (including a medical-AI model evaluated as one) for use. EPV-10 does not apply: N is the largest satisfying all Riley criteria — three for a binary or time-to-event outcome, four for a continuous one (R `pmsampsize`). - **Test 13 — External validation (Riley).** Validating an existing prediction/AI model: size for the CI width of the C-statistic, calibration slope, O:E and (if claimed) net benefit (R `pmvalsampsize`); ≥ 100 events and ≥ 100 non-events is only a floor. For Tests 12–13 read `${CLAUDE_SKILL_DIR}/references/prediction_model_sample_size.md`. - **Test 14 — MRMC reader study (Obuchowski–Rockette).** Readers with vs without the AI, or AI non-inferior to readers. Test 1 under-sizes it because readers are a random effect: size readers `J` × cases from pilot/literature variance components with `RJafroc` / `MRMCaov` / `iMRMC` (do not hand-roll the OR algebra) and report the `J × N` power grid. Read `${CLAUDE_SKILL_DIR}/references/mrmc_reader_study_sample_size.md`. - **Test 15 — Segmentation-metric precision (Dice / HD95 / NSD).** The outcome is a per-case score, not a proportion: `n ≈ (1.96·SD/δ)²` from the pilot SD of per-case Dice, sized on the worst structure; CI by patient-level bootstrap (BCa). This is precision; a comparison is Test 16. Read `${CLAUDE_SKILL_DIR}/references/segmentation_metric_sample_size.md`. - **Test 16 — Between-model comparison.** Model A beats B, C, …: power the paired per-case *difference*, `n = ((z₁₋α/₂ + z₁₋β)·SD_Δ/Δ)²` (a CI sized to just exclude zero has ~50% power), not each model's precision; for > 2 models pre-specify one primary contrast or pay the family-wise correction; a ranking claim needs multiple seeds. Read `${CLAUDE_SKILL_DIR}/references/multi_model_comparison_sample_size.md`. - **Test 17 — Segmentation usability.** Clinicians can use it (acceptability rate, catastrophic-failure bound, edit time): acceptability is a proportion sized per structure class; `DE ≈ 1 + (m−1)ρ` holds only when each case has its own readers — the same readers on every case (crossed) add a reader term more cases cannot shrink; bounding failures at ≤ 1% needs ~300 clean cases (rule of three). Read `${CLAUDE_SKILL_DIR}/references/segmentation_acceptability_sample_size.md`. ## Out of Scope Do not compute adaptive trials (group-sequential, sample size re-estimation), cluster-randomized trials (design effect, ICC-based inflation), Bayesian sample size determination, crossover designs, or multi-endpoint correction (mention Bonferroni if asked, but do not compute corrected sample sizes). Say the design is beyond this skill and point to G*Power (free, https://www.psychologie.hhu.de/gpower), PASS, or a biostatistician. ## Workflow ### Phase 1: Understand the Study 1. Ask the user to describe the study briefly (design, primary outcome, groups). 2. Walk the decision tree to a test, and confirm it with the user before proceeding. ### Phase 2: Collect Parameters 1. Present the selected test's parameter table. 2. For each parameter without a user-provided value, explain it and offer the default. 3. Estimate effect sizes from prior literature (ask for the references) or pilot data; use Cohen's conventions only as a last resort, noting that convention-based estimates are less precise. ### Phase 2b: Retrospective Studies When the dataset already exists, formal power analysis is often impractical. Offer: - **Fixed extract**: read `${CLAUDE_SKILL_DIR}/references/observational_cohort.md` and report event budget / confidence-interval precision instead of forcing a prospective recruitment-style power calculation. - **Experience-based justification** (acceptable for IRB and many journals): - *Institution volume*: `total exams in period × prevalence × (1 − exclusion rate) = expected N`. Ask for the annual exam volume for the modality, study period, prevalence and exclusion rate. This gives a realistic upper bound for N. - *Prior studies*: report the N of 3–5 comparable published studies and cite them; the user's N should be in the same range or larger. - IRB templates for both: `${CLAUDE_SKILL_DIR}/references/justification_examples.md` § Retrospective. Use a formal calculation (Phase 3) even for a retrospective study when a subset is enrolled prospectively, the primary analysis tests a hypothesis (not just estimation), the journal's Instructions for Authors require a power analysis, or the IRB requires it. ### Phase 3: Calculate and Report 1. Generate R code (primary) and Python code (alternative) from the reference formula. 2. Run the R code via Bash; the reported N is the number it prints, not a hand calculation. 3. Present the result in the Output Format below. Cite methodological sources only from `formulas.md` or the test's reference file; any other reference needs a DOI/PMID confirmed via `/search-lit`, otherwise mark it `[UNVERIFIED - NEEDS MANUAL CHECK]`. Mark an effect size, clinical definition or threshold you could not confirm `[VERIFY]`. 4. In a project, save the IRB text as `protocol/sample_size_justification.md` and the scripts as `protocol/sample_size_calc.R` / `.py`: `/write-protocol` and `/write-paper` embed that text verbatim, so the numbers are never retyped. ### Phase 4: Sensitivity Analysis (Optional) If a parameter is uncertain or the effect-size estimate is vague, flag it and offer a table of N across plausible values (e.g., varying effect size, or power from 0.80 to 0.90). ## Output Format Always structure the final output as follows: ```markdown ## Sample Size Calculation Report ### Study Design [1-2 sentence summary of the design and test selected] ### Parameters | Parameter | Value | Source | |-----------|-------|--------| | ... | ... | user / literature / convention | ### Result - **Required sample size**: N = [value] - **With [X]% attrition adjustment**: N_adj = [value] ### R Code (Reproducible) ```r # [complete, self-contained R script] # Dependencies: [list packages] # Run: Rscript sample_size_calc.R ``` ### Python Code (Alternative) ```python # [complete, self-contained Python script] # Dependencies: [list packages] # Run: python sample_size_calc.py ``` ### IRB Justification Text > A sample of [N] participants is required to detect [effect description] with [power]% power > at a [one/two]-sided significance level of [alpha], assuming [key assumptions]. > Accounting for an estimated [X]% attrition rate, we plan to enroll [N_adj] participants. > This calculation is based on [formula/method reference]. ### Effect Size Interpretation [Cohen's benchmark classification + clinical meaning in the context of this study] ``` The IRB text must state N; name the test and its formula source; give every assumed parameter (effect size, alpha, power); state the attrition adjustment and final enrollment target; cite the methodological reference (e.g., "Schoenfeld, 1981"); and use formal, third-person language. Read `${CLAUDE_SKILL_DIR}/references/justification_examples.md` for per-design exemplars when writing it.