--- name: analysis-workflow description: "Organize multi-step scientific analyses into reproducible, self-contained modules. Use for workflows such as QC→PCA→DEG→GSEA that produce scripts, inputs, figures, tables, and methods. Creates a stable module layout, records exact inputs/parameters/package and database versions in each module README, keeps large data as references instead of copies, and verifies outputs before completion." license: Apache-2.0 wisp: schema_version: 1 domains: [bioinformatics] research_stages: [observation, analysis, validation] roles: [analyst, validator] evidence_types: [project-data, omics, computational] outputs: [analysis-module] side_effects: code_execution --- # Reproducible Analysis Modules Use this skill for a scientific workflow with two or more analysis stages or when a stage produces scripts plus result files. It defines project organization and methods capture; load `figure-style` as well whenever a stage creates or revises a plot. ## 1. Plan module boundaries Before writing outputs, list the modules and the dependency edges between them. Use stable ASCII names. Conventional acronyms such as `QC`, `PCA`, `DEG`, and `GSEA` may stay uppercase; otherwise prefer a short kebab-case name. Respect a compatible layout that already exists. Do not reorganize unrelated user files merely to impose this convention. ## 2. Default module layout Create only directories the module actually needs: ```text / ├── scripts/ ├── input/ ├── output/ │ ├── figures/ │ └── tables/ └── README.md ``` - `scripts/` contains the executable source for this module. - `input/` contains small module-specific inputs or a manifest/reference to the canonical data. Do not duplicate a large dataset by default. - `output/figures/` contains rendered figures from this module only. - `output/tables/` contains machine-readable results from this module only. - `README.md` is the module's reproducibility record and methods source. Shared immutable/raw data may live in project-level `data/`. A downstream module references an upstream output by a project-relative path; it does not silently copy or rename that output. ## 3. Make outputs attributable Every output must have one producing script or recorded command. Use deterministic filenames that identify the analysis and content. Keep temporary files outside the final output directories or name them clearly as temporary. Before completing a module, verify: 1. every declared output exists and is non-empty; 2. every table can be parsed in its declared format; 3. every figure was rendered and visually inspected using `figure-style`; 4. README input and output paths resolve from the project root; 5. reported thresholds and parameters match the actual script. ## 4. Update README.md at module completion Create or update these sections: ```markdown # ## Purpose ## Inputs - `` — source, upstream module, checksum or version when available ## Methods ## Software and data sources - R/Python package: exact version - External API/database: release or access date - Wisp/model/runtime metadata: exact recorded value when available ## Commands and scripts - `` — how it was executed ## Outputs - `` — meaning and format ## Limitations ``` Write methods from executed code and recorded parameters, not from a generic template. Do not claim a package, database, model, OS, or version that was not actually used or observed. ## 5. Capture exact versions without dumping the world Record direct dependencies used by the module: - R: `packageVersion("")` for named packages and `sessionInfo()` for the runtime context. - Python: `importlib.metadata.version("")`; use the project lock file when it is the authoritative environment record. - External databases/APIs: release identifier when available, otherwise access date plus endpoint/source. - Wisp version and model profile: use runtime/session metadata only when it is available. Write `unavailable` rather than guessing. Do not paste an entire global `pip freeze` into every module. If a complete environment export is useful, save it once as a separate artifact and link it from the README. ## 6. Finish the workflow After all modules pass their checks, summarize the dependency chain and link the module READMEs. Treat those READMEs as the first-version source of truth. Generate a root `METHODS.md` only when the user asks for it or a deterministic project tool can derive it from the module records; do not maintain a second hand-edited copy that can drift.