# Cancer Immunotherapy Data Science Challenge ## Code for Paper **Title:** *A community machine learning challenge to predict the effects of gene perturbations on T cell differentiation for cancer immunotherapy* **biorxiv link:** [10.64898](https://www.biorxiv.org/content/10.64898/2026.05.21.726863v1.abstract). **Abstract:** Perturbations of genes with functional importance in T cells could be used to change the distribution of CD8 T cell states to enhance anti-tumor functions for cancer immunotherapies. We launched a world-wide computational challenge to predict the effects of gene perturbations and to devise objective functions for prioritizing gene perturbations that lead to desired T-cell state distributions. We supported the challenge by generating a single-cell Perturb-seq dataset profiling the effect of knocking out 73 individual expert-defined genes in T cells transferred into a mouse melanoma model. We compared the top algorithms developed by participants, and found that performance was primarily determined by the prior data used for gene feature representation, with perturbational data derived features, proving most effective. Experimental validation of the top 61 genes nominated by the algorithms revealed that perturbation of Ndufv2 and Dimt1 reached the defined objective and biased differentiation toward desired states. **Raw and processed data:** Available at [GSE327731](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE327731). --- ## Contents of this Repository ### Data - **`data_screen1`**: Perturb-seq from our first experiment targeting 73 genes (screen 1), and how it is split in our benchmark - **`data_screen2`**: Perturb-seq from our second experiment targeting 59 genes (screen 2), and how it is split in our benchmark - **`data_online`**: Public perturb-seq from [PMID 37968405](https://pubmed.ncbi.nlm.nih.gov/37968405/) (online screen), and how it is split in our benchmark; includes other online data sources used by the top methods ### Implementation of Winning Methods - **`methods_run_screen1`**: Benchmarking top methods on screen 1 - **`methods_run_screen2`**: Benchmarking top methods on screen 2 - **`methods_run_online`**: Benchmarking top methods on the online screen - **`methods_run_screen1-screen2`**: Training top methods on screen 1 and applying them to predict on screen 2 ### Analysis - **`notebooks`**: Data analysis and code for reproducing figures - **`feature_analysis`**: Analysis of all features - **`cell_cycle`**: Cell cycle analysis - **`annotation`**: Annotating the online screen based on our cell states **Notes:** - Some of the data files are stored on this Google Drive [folder](https://drive.google.com/drive/folders/1ShaZ0Ql0JVKoncvO_khMzYnP0KbMpKkf?usp=sharing) due to their size. To view and use these files, you can download them; the folders on Google Drive correspond to the folders here. ## Environment setup To run any of the scripts, follow this setup to create a Conda environment named **`sca`** (Python 3.10, CUDA 11.6 toolchain via conda, PyTorch and scientific stack via conda + pip). **Create the environment from the exported spec** (from the repository root, which typically takes 10-60 minute): ```bash conda env create -f environment.yml conda activate sca ``` **Or clone from your existing `sca` env** on another machine: ```bash conda env export -n sca --no-builds | grep -v '^prefix:' > environment.yml ``` **Notes:** - `environment.yml` is a full export (conda packages + pinned pip packages). Recreating on a different OS or CUDA driver may require adjusting PyTorch/CUDA-related lines; in that case install a matching [PyTorch build](https://pytorch.org/get-started/locally/) first, then install remaining pip dependencies. - After edits to the env, refresh the file with: `conda env export -n sca --no-builds | grep -v '^prefix:' > environment.yml` **Key stack (high level):** Python 3.10 · NumPy / SciPy / pandas · PyTorch (pip: `torch` 2.5.x with CUDA 12.1 wheels; conda also provides CUDA 11.6 dev libraries) · scanpy / scvi-tools · Jupyter (ipykernel) · transformers, and related ML dependencies—see `environment.yml` for the full list. --- ## Citation and Beyond If you find this useful, we appreciate your citation using: ``` TBD ``` **Cell Perturbation Prediction Challenge (CPPC)** This project constitutes the initial effort of the Cell Perturbation Prediction Challenge (CPPC). Stay tuned for future challenges and consider joining the CPPC community by participating! **_About CPPC_:** We established the Cell Perturbation Prediction Challenge (CPPC) as an ongoing series of yearly competitions designed to generate biomedically relevant perturbation datasets and provide an objective platform for iterative machine learning model testing and improvement. More information about this initiative can be found at [TBD]. **_CPPC Community_:** We acknowledge the following participants, grouped by teams (affiliations are as they were at the conclusion of the challenges), who won prizes in the cancer immunotherapy challenges: - Marios Gavrielatos and Konstantinos Kyriakidis, Greece - Yuzhou Gu, Anzo Teh, Yanjun Han, and Brandon Wang, U.S. (MIT) - Peter Novotný, Poland - Brody Langille, Jordan Trajkovski, and Elizabeth Hudson - Marc Glettig - Ai Vu Hong, researcher at Genethon, France - Saket Kunwar, independent researcher, Nepal - John Gardner, freelance data scientist - Basak Eraslan, postdoctoral researcher holding a joint position at the Regev Lab in Genentech and Kundaje Lab at Stanford University - Haoyue Dai, Kun Zhang, Ignavier Ng, Yujia Zheng, Xinshuai Dong, and Yewen Fan from Carnegie Mellon University; Petar Stojanov, postdoctoral fellow at the Eric and Wendy Schmidt Center; Gongxu Luo, Mohamed bin Zayed University of Artificial Intelligence; and Biwei Huang, University of California, San Diego - Liu Xindi, freelance programmer - Johnson Zhou, Camille Sayoc, and Yi-Cheng Peng, Master’s students at the Faculty of Engineering and IT at the University of Melbourne, Victoria, Australia - Dariusz Brzeziński and Wojciech Kotlowski from Poznań University of Technology in Poland - Salil Bhate, postdoctoral fellow at the Eric and Wendy Schmidt Center - Irene Bonafonte Pardàs, Artur Szalata, and Benjamin Schubert from Helmholtz Center Munich, and Miriam Lyzotte from Mila - Quebec AI Institute