--- name: cross-national description: Use when comparing an exposure-outcome association across countries with parallel national surveys (KNHANES, NHANES, CHNS). Harmonizes variables, runs parallel weighted analyses and builds comparison tables for 2-country (KR+US) or 3-country (KR+US+CN) designs. model: opus metadata: triggers: "cross-national, 한미 비교, Korea US comparison, KNHANES NHANES, 양국 비교, binational, cross-country, 비교연구, 3국 비교, CHNS, 한미중" --- # Cross-National Comparison Study Skill ## Inputs 1. **Research question**: exposure → outcome association to compare across countries 2. **Korean data path**: KNHANES CSV file 3. **US data path**: NHANES CSV directory (multiple tables to merge) 4. **Harmonization table** (optional): CSV mapping variables across surveys - Default: `/replicate-study`'s `references/harmonization_knhanes_nhanes.csv` ## Reference Files - `/write-paper`'s `references/paper_types/cross_national.md` — writing template - `/analyze-stats`'s `references/analysis_guides/survey_weighted.md` — survey-weighted analysis guide - `references/chns_coding.md` — CHNS files, merge keys, coding and warnings; read it when the design includes China (3-country design) - `references/additional_variables.md` — asthma, sleep, physical activity, diet, treatment and non-HDL-cholesterol coding, plus composite-score (LE8) warnings; read it when the study uses any of them ## Workflow ### Phase 1: Study Definition 1. Confirm research question: Exposure → Outcome 2. Define variable coding for both countries: - Exposure: PHQ-9, BMI category, smoking, etc. - Outcome: diabetes, hypertension, mortality, etc. - Covariates: age, sex, education, income, smoking, alcohol, obesity, CVD 3. Check harmonization table for variable availability 4. Output: study protocol summary for user approval ### Phase 2: Data Preparation Never guess a variable name, dataset column name, or variable coding. If a mapping is uncertain, output `[VERIFY: variable_name]` and ask the user to confirm it against the data dictionary. **KNHANES (single CSV)**: 1. Load CSV and keep every row — the age ≥20 (or per-protocol) restriction is applied to the design in step 3 2. Derive variables using KNHANES coding: | Variable | Raw Var | Coding | |----------|---------|--------| | Smoking | BS3_1 | 1,2=Current; 3=Former; 8=Never | | Alcohol | BD1_11 | 2-6=Frequent (current drinker); 1=Occasional (past-year abstainer); 8=Never | | Obesity | HE_obe | 1-3=Normal; 4-6=Obesity (BMI≥25, Asian cutoff) | | Depression | BP_PHQ_1~9 | Sum ≥10 = depression | | Diabetes | HE_glu, HE_HbA1c, DE1_dg | FPG≥126 or HbA1c≥6.5 or DE1_dg=1 | | CVD | DI4_dg, DI5_dg, DI6_dg | Any = 1 → CVD yes | | Education | edu | 1-3=Non-college; 4=College | | Income | incm | Quartile (1=lowest … 4=highest): 1-3=Bottom 75%; 4=Top quartile | 3. Set survey design on the full file, then restrict to the analytic domain: des <- svydesign(id=~psu, strata=~kstrata, weights=~wt_itvex, nest=TRUE, data=df); des_ad <- subset(des, age >= 20). Never filter rows before svydesign() — dropping them changes the standard errors (see `survey_weighted.md`, subpopulation analysis) **NHANES (multiple CSVs)**: 1. Load and merge tables by SEQN (DEMO_J, DPQ_J, GHB_J, GLU_J, BMX_J, SMQ_J, ALQ_J, DIQ_J, MCQ_J, BPQ_J, BPXO_J) 2. Derive variables using NHANES coding. **CRITICAL**: NHANES data downloaded via R `nhanesA` package uses TEXT LABELS, not numeric codes. | Variable | Raw Var | Coding (text labels) | |----------|---------|----------------------| | Sex | RIAGENDR | "Male" / "Female" (NOT 1/2) | | Smoking | SMQ020 + SMQ040 | 100 cigs (SMQ020 "Yes" / "No") + now smoke (SMQ040 "Every day" / "Some days" / "Not at all") | | Alcohol | ALQ121 + ALQ111 | Frequent (current drinker): any ALQ121 frequency except "Never in the last year"; Occasional (past-year abstainer): "Never in the last year"; Never (lifetime non-drinker): ALQ111 == "No" (ALQ121 will be NA) | | Obesity | BMXBMI (BMX_J, kg/m²) | ≥30 (WHO cutoff, NOT Asian) | | PHQ-9 | DPQ010~DPQ090 | "Not at all"→0, "Several days"→1, "More than half the days"→2, "Nearly every day"→3; sum ≥10 = depression | | Diabetes | LBXGLU (GLU_J, fasting subsample, mg/dL), LBXGH (GHB_J, %), DIQ010 | LBXGLU≥126 \| LBXGH≥6.5 \| DIQ010=="Yes" (DIQ010: "Yes" / "No" / "Borderline"). CRITICAL: fasting glucose is GLU_J LBXGLU, not BIOPRO_J LBXSGL — CDC says the serum LBXSGL should not be used to determine undiagnosed diabetes; an analysis that uses LBXGLU needs the fasting-subsample weight (step 3). LBXGLU is NA for non-fasters, and in R `NA \| TRUE` is TRUE but `NA \| FALSE` is NA, so this composite goes missing only for non-fasters who would be non-diabetic; dropping those NAs under WTMEC2YR inflates prevalence. Either analyse the FPG composite only in the fasting domain with the fasting weight, or use the non-fasting variant `LBXGH>=6.5 \| DIQ010=="Yes"` on WTMEC2YR, and say which in `variable_mapping.csv` | | CVD | MCQ160B/C/D/E | MCQ160B=="Yes" (CHF) \| MCQ160C=="Yes" (CHD) \| MCQ160D=="Yes" (angina) \| MCQ160E=="Yes" (MI); labels "Yes" / "No" / "Don't know" | | HTN | BPXOSY2+3, BPXODI2+3, BPQ020 | mean(BPXOSY2, BPXOSY3)≥140 \| mean(BPXODI2, BPXODI3)≥90 \| BPQ020=="Yes" (BPXOSY3 is the 3rd reading, not an average; the mean of the 2nd and 3rd matches KNHANES) | | Education | DMDEDUC2 | 5 text levels | 3. Set survey design on the full file, then `subset()` the design object to the analytic domain (e.g. RIDAGEYR >= 20): svydesign(id=~SDMVPSU, strata=~SDMVSTRA, weights=~WTMEC2YR, nest=TRUE). The weight follows the files: the single-cycle `_J` tables above take WTMEC2YR; WTMECPRP goes only with the pre-pandemic `P_` files (P_DEMO, P_BMX, ...). A variable from the fasting subsample (GLU_J LBXGLU, TRIGLY_J LBXTR/LBDLDL) takes the fasting weight instead: WTSAF2YR (`_J`) or WTSAFPRP (`P_`); a variable that combines a fasting-subsample component with full-sample ones (diabetes above) is defined only in that fasting domain, so it cannot enter a WTMEC2YR model. Pooling cycles follows the NCHS rules: divide each cycle's weight by the number of cycles pooled (1999–2002 has its own 4-year weights), and to combine 2015–2016 with 2017–March 2020 use 2/5.2 × WTMEC2YR and 3.2/5.2 × WTMECPRP. **CHNS (3-country design)**: read `references/chns_coding.md` before preparing China data. ### Phase 3: Parallel Analysis For EACH country independently: 1. **Table 1**: Baseline characteristics by exposure (weighted counts + percentages) 2. **Main analysis**: Sequential logistic regression models - Model 1 (unadjusted) - Model 2 (age + sex) - Model 3 (fully adjusted: + education, income, smoking, alcohol, obesity, CVD) 3. **Subgroup analyses**: By sex, age group, education, income, alcohol, smoking, CVD, obesity. Whether the association differs between subgroups is tested with an exposure × subgroup interaction term fitted on the full design (`svyglm` on the whole sample), not by comparing the subgroups' P values. 4. **Dose-response** (if applicable): RCS with 3 knots ### Phase 4: Cross-National Comparison Table Generate a side-by-side comparison: | Analysis | Korea wOR (95% CI) | US wOR (95% CI) | Ratio of wORs (95% CI); P | |----------|-------------------|-----------------|---------------------------| | Overall (fully adjusted) | ... | ... | ... | | Male | ... | ... | ... | | Female | ... | ... | ... | | ... | ... | ... | ... | Compare the countries with the ratio of their odds ratios, not with whether the directions agree: two estimates in the same direction can differ, and opposite directions can be compatible. With log odds ratios b₁, b₂ and their design-based standard errors SE₁, SE₂ from the two independent surveys, the ratio is exp(b₁ − b₂) with 95% CI exp(b₁ − b₂ ± 1.96·√(SE₁² + SE₂²)) and z = (b₁ − b₂)/√(SE₁² + SE₂²) (Altman & Bland, *BMJ* 2003;326:219). Every number comes from executed code output (`analysis_korea.R`, `analysis_us.R`) — never an invented p-value, effect size, confidence interval, or sample size. ### Phase 5: Output Files ``` {working_dir}/ ├── cross_national_report.md — Study summary + comparison tables ├── variable_mapping.csv — Variable mapping with match status ├── analysis_korea.R — KNHANES analysis (self-contained) ├── analysis_us.R — NHANES analysis (self-contained) ├── results/ │ ├── table1_korea.csv │ ├── table1_us.csv │ ├── main_results_comparison.csv │ └── subgroup_comparison.csv └── manuscript_draft/ — Optional: Methods + Results draft ├── methods_draft.md └── results_draft.md ``` Take every citation in the report or draft from `/search-lit` (confirmed DOI/PMID); mark any other `[UNVERIFIED - NEEDS MANUAL CHECK]`, and never generate references from memory. ## Critical Rules 1. **NEVER pool data across countries**. Each country analyzed with its own survey design. 2. **Country-specific BMI cutoffs**: Korea ≥25 (Asian), US ≥30 (WHO). 3. **Country-specific income**: KNHANES quartile, NHANES PIR → harmonize to binary. 4. **Weighted analysis mandatory**: Both KNHANES and NHANES are complex surveys. CHNS has no survey weights — analyse it unweighted (see `references/chns_coding.md`). 5. **Document all harmonization decisions**: What matches, what needed recoding, what differs. 6. **Same analytic approach**: Identical model specifications for both countries for fair comparison.