# Aggregate Cohort Evaluation ## Scope This workflow documents cohorts using pre-aggregated counts and summaries. It does not ingest records, classify people, estimate a patient-specific risk, or recommend care. ## Protocol Before Results Pre-specify: - objective and target population; - study design and setting; - index date/time zero; - eligibility and sampling; - exposure, comparator, outcomes, covariates, and time windows; - causal estimand if making a causal claim; - confounding strategy; - missing-data strategy; - subgroup and interaction analyses; - multiplicity control; - sensitivity and negative-control analyses; - disclosure policy. For routinely collected data, document code sets, phenotypes, database versions, linkage quality, data provenance, and validation. ## Participant Flow Report aggregate counts for: 1. source population; 2. eligibility assessed; 3. excluded by reason; 4. included; 5. analysis populations; 6. missing outcome or follow-up; 7. subgroup availability. Apply suppression before releasing the flow. Do not reconstruct suppressed values through totals. ## Table 1 Use summaries appropriate to distributions and measurement: - categorical: count, denominator, percentage, missing; - continuous: mean and standard deviation or median and quartiles; - time-dependent or repeated measures: define the summary window; - assay measurements: units, platform, detection limits, batch, and transformation. Baseline significance tests do not measure meaningful imbalance and are not generated by the bundled table helper. If comparison is needed, pre-specify descriptive standardized differences or another justified measure and interpret it in context. ## Effect Estimation Match measure to question: - prevalence/risk: risk difference and risk ratio; - rates: rate difference and rate ratio; - odds: odds ratio, with care when outcomes are common; - time to event: estimand-aligned survival measures; - repeated outcomes: model and covariance assumptions; - diagnostic accuracy: sensitivity/specificity and predictive values at prespecified thresholds. Report absolute and relative effects with uncertainty when both are relevant. A p-value is not an effect size and “not significant” is not evidence of no difference. ## Confounding and Bias Address: - confounding by indication; - selection and collider bias; - immortal-time and time-varying treatment bias; - informative observation/censoring; - measurement error and misclassification; - missing data; - outcome ascertainment; - site and calendar-time effects; - data-driven subgroup or cut-point selection; - unmeasured confounding. State which variables were selected before analysis and why. Do not select confounders solely by univariable p-values. Distinguish prediction from causal inference. ## Subgroups and Fairness Subgroup work must document: - rationale and prespecification; - representation and missingness; - sample sizes and event counts; - effect estimates with intervals; - interaction tests when effect heterogeneity is the question; - multiplicity; - measurement validity across groups; - intersectional and site effects where feasible; - whether categories are self-reported, assigned, or derived; - risk of reinforcing structural inequities. Do not rank groups or declare fairness from one metric. Small groups may require pooling, secure analysis, or non-release rather than unstable public estimates. ## Biomarker Cohorts Record: - biomarker category using FDA-NIH BEST terminology; - biological and analytical rationale; - specimen collection and handling; - assay platform, version, units, and quality controls; - prespecified threshold and source; - analytical validation; - blinding to outcomes; - missing/failed assays; - internal and external validation; - distinction among prognostic, predictive, and treatment-effect interaction claims. Never derive a threshold on the evaluation cohort and present it as validated without independent confirmation. ## Disclosure Controls The table generator implements: - a configurable minimum cell size; - primary suppression for small nonzero cells; - complementary suppression when one cell could be recovered from a row; - group-level suppression when denominators are too small; - bounded groups and rows; - omission of raw values and identifiers. The default threshold is a conservative operational setting, not a universal rule. It does not address all differencing, linkage, longitudinal, geographic, genomic, or rare-combination risks. Follow an approved disclosure policy and privacy review. ## Interpretation Template Use: > In this aggregate [design] evaluation, [effect/summary] was estimated as [value and interval] for [defined outcome and horizon]. The analysis is [prespecified/exploratory] and is limited by [bias, missingness, precision, transportability]. It does not establish causality, clinical utility, or an action for any person. ## Reporting - STROBE for observational design. - RECORD for routinely collected data. - REMARK for tumor prognostic-marker studies. - TRIPOD+AI for prediction-model development/evaluation. - Appropriate causal-inference and target-trial reporting when making causal claims. See `study_reporting.md` and `privacy_and_disclosure.md`.