--- name: sc-markers description: Load when ranking cluster-level marker genes from a clustered single-cell AnnData via Scanpy Wilcoxon / t-test / logreg or COSG specificity. Skip when comparing condition-vs-control with replicates (use sc-de); assigning cell-type labels (use sc-cell-annotation). tags: - singlecell - scrna - markers - cluster-markers - annotation - differential-expression - cosg --- # sc-markers ## When to use The user already has clustering / cell-type labels in `obs` (typically `leiden`, `louvain`, or `cell_type`) and wants ranked marker genes per group as evidence for downstream annotation or interpretation. Four methods: `wilcoxon` (default rank-sum), `t-test` (Welch), `logreg` (multinomial logistic regression — discriminative ranking), `cosg` (fast cosine-specificity scoring without p-values). This is for **cluster markers**, not condition contrasts — for treatment-vs-control DE with replicates use `sc-de`. ## Use from a step ```python markers = load_skill("sc-markers") adata = read_input("results/02_cluster/intermediate/adata_clustered.h5ad") table = markers.find_markers(adata, groupby="leiden") write_output(table, "tables/markers_all.csv") write_output(markers.top_markers(table), "tables/markers_top.csv") write_output(markers.run_info(table), "tables/marker_run.json") ``` A complete step on log-normalized PBMC expression is in `examples/example_step.py`. ## API ### `find_markers(adata, *, groupby: str, method: str='wilcoxon', n_genes: int | None=None, min_in_group_fraction: float=0.25, min_fold_change: float=0.25, max_out_group_fraction: float=0.5, mu: float=1.0) -> pd.DataFrame` Rank marker genes for each group against the rest using adata.X. X should contain normalized expression. Computation uses a copy, so the input's matrices, obs and uns are unchanged. Scanpy methods apply the expression-fraction and fold-change filters; when filtering fails or removes every row, the unfiltered ranking is returned and run_info records filter_fallback and its reason. COSG scores specificity without p-values. :param groupby: The obs column defining clusters or cell types. :param method: wilcoxon (default), t-test, logreg or cosg. :param n_genes: Genes ranked per group. None means all genes for Scanpy methods and 50 for COSG, matching the CLI defaults. :param min_in_group_fraction: Minimum expressing fraction within a group. Default 0.25; unused by COSG. :param min_fold_change: Fold-change filter passed to Scanpy. Default 0.25; unused by COSG. :param max_out_group_fraction: Maximum expressing fraction outside a group. Default 0.5; unused by COSG. :param mu: COSG specificity penalty. Default 1.0; unused by Scanpy methods. :returns: A DataFrame with group, names, scores and method-dependent effect, p-value and fraction columns. COSG pvals and pvals_adj are NaN. The table's attrs hold the run record; use run_info before CSV export. :raises ValueError: Unknown method or missing grouping column. ### `run_info(table: pd.DataFrame, *, keep: bool=True) -> dict` Return the method, grouping and filter-fallback record attached to a marker table. :param keep: False removes the record from table.attrs; default True keeps it. :returns: A copy of the run record, or an empty dict for an unrecorded table. CSV files do not preserve attrs; write this record separately if needed. ### `top_markers(table: pd.DataFrame, *, n_top: int=10) -> pd.DataFrame` Return the top n_top rows per group, using adjusted p-value then effect. Tables without finite adjusted p-values, including COSG, use scores. Otherwise logfoldchanges is the effect when available, falling back to scores. ### `cluster_summary(table: pd.DataFrame) -> pd.DataFrame` Return n_markers, top_gene, top_effect, median_effect and effect_metric per group. The top gene uses the same ordering as top_markers. top_effect is the group's maximum effect, not necessarily that gene's effect. Use a row from top_markers when reporting a gene together with its effect size. Effect statistics use all returned markers in the group. ### `marker_dotplot_figure(adata, table: pd.DataFrame, *, groupby: str, n_top: int=5)` Return a matplotlib Figure of marker expression in adata.X by group. Dot size shows the expressing fraction; color shows mean expression. The top n_top markers per group are selected with top_markers. ## Methods and parameters The four methods and their interpretation are described in `references/methodology.md`. `find_markers` uses `X` with `use_raw=False`. For `pbmc3k_processed` or `pbmc68k_reduced`, take `adata.raw.to_adata()` to use their log-normalized matrix; their X is scaled. - `groupby` explicitly names the labels to compare. - `method="wilcoxon"` preserves the CLI default; `t-test` uses Welch's test, `logreg` gives a discriminative ranking, and `cosg` gives specificity scores. - `n_genes=None` ranks all genes for Scanpy methods and 50 per group for COSG. - The CLI's post-filter defaults are `min_in_group_fraction=0.25`, `min_fold_change=0.25` and `max_out_group_fraction=0.5`. - `mu=1.0` preserves the COSG specificity penalty; the other methods ignore it. - `top_markers(..., n_top=10)` preserves the CLI's compact-summary size. ## Gotchas - In `tables/cluster_summary.csv`, `top_effect` is the group's maximum, not necessarily the effect of `top_gene`. Use the same row from `top_markers` for a gene and its effect size (`_api.py:132`, `cluster_summary`). - `run_info(table)["filter_fallback"]` is true when post-filtering failed or removed every row. The function then returns the unfiltered ranking, as the CLI did previously; read `fallback_reason` before interpreting the table. - COSG's `pvals` and `pvals_adj` columns in `tables/markers_all.csv` contain NaN. `top_markers` ranks those rows by `scores`; p-value filters must handle missing values. - `_api.py:21` (`find_markers`) uses X without checking that it is log-normalized. Supply the intended expression matrix explicitly. - `sc_markers.py:134` (`_resolve_groupby`) can choose a recognized label column for the CLI. The function API requires `groupby` explicitly. - Binary `logreg` can return a Scanpy table without `group`, causing `_api.py:91` to raise `KeyError`. Use `wilcoxon`, `t-test` or COSG for two groups; this existing Scanpy-wrapper limitation is unchanged. - CSV export drops DataFrame attrs. Save `run_info(table)` separately when keeping the fallback record matters; the CLI stores it at `result.json["data"]["run_info"]`. ## Inputs and outputs `find_markers` reads X and the selected obs column, computes on a copy and returns a DataFrame. The input AnnData is unchanged, including its uns entries. `top_markers` and `cluster_summary` return tables; `marker_dotplot_figure` returns a matplotlib Figure from X. These functions do not write files. The CLI writes `processed.h5ad`, `report.md`, `result.json`, `tables/markers_all.csv`, `tables/markers_top.csv`, `tables/cluster_summary.csv`, and figure-data copies of those tables. Conditional figures and R outputs are listed in `references/output_contract.md`. ## Key CLI ```bash # Demo (built-in PBMC3K with leiden labels) python skills/singlecell/scrna/sc-markers/sc_markers.py --demo --output /tmp/sc_markers_demo # Default Wilcoxon on leiden clusters python skills/singlecell/scrna/sc-markers/sc_markers.py \ --input clustered.h5ad --output results/ --groupby leiden # COSG fast specificity ranking on a labelled AnnData python skills/singlecell/scrna/sc-markers/sc_markers.py \ --input annotated.h5ad --output results/ \ --groupby cell_type --method cosg --mu 1.0 # Strict marker filtering (high fold-change, low out-group fraction) python skills/singlecell/scrna/sc-markers/sc_markers.py \ --input clustered.h5ad --output results/ \ --min-fold-change 1.0 --max-out-group-fraction 0.2 ``` ## See also - `references/parameters.md` — every CLI flag and per-method tuning hint - `references/methodology.md` — Wilcoxon vs t-test vs logreg vs COSG; when each wins - `references/output_contract.md` — `markers_all.csv` column schema; figures' figure_data CSVs - Adjacent skills: `sc-clustering` (upstream — produces the `leiden` / `louvain` column), `sc-cell-annotation` (downstream — uses these markers as evidence for label assignment), `sc-de` (parallel — replicate-aware condition contrasts, NOT cluster markers) ## Dependencies Python packages this skill's script needs. They are not installed for you — check before a long run. `anndata`, `matplotlib`, `numpy`, `pandas`, `scanpy`, `scikit-learn`, `scipy`, `seaborn`