--- name: active-swe-eval description: Prepare and run the complete local Docker evaluation for Active-SWE. Claude Code is the fixed executor for Recorded, Potential, and Judge. --- # Active-SWE Evaluation With Claude Code Use this skill to prepare an Active-SWE project checkout and evaluate one or more models. Claude Code is both the outer orchestrator and the fixed executor inside every task container. ## Locate the project First look for an existing project root containing both `pyproject.toml` and the `active_swe/` package. Reuse it without pulling, resetting, or replacing local files. If no checkout exists, clone into a new relative workspace and enter it: ```bash git clone --depth 1 https://github.com/XLearning-SCU/Active-SWE.git Active-SWE cd Active-SWE ``` Do not clone over a non-empty path. If Git or repository access is unavailable, stop and ask the user for an existing checkout. After locating or cloning the project, open `.claude/skills/active-swe-eval/SKILL.md` from that checkout and use the local copy as the authority for all remaining steps; do not continue from an older remote or cached copy. Then read these local files in order: 1. `.claude/skills/active-swe-eval/ENVIRONMENT.md`: Docker, Claude Code, dataset, images, model setup, and network boundaries. 2. `.claude/skills/active-swe-eval/EVALUATION.md`: Recorded, Potential, Judge, and the six paper metrics. ## Public Sources ```text Source code: https://github.com/XLearning-SCU/Active-SWE Dataset: https://huggingface.co/datasets/XLearning-SCU/Active-SWE ``` When this skill is installed separately from the repository, its optional `bin/bootstrap_active_swe.py` performs the same checkout check and clone. Runtime tools belong under `./.tools`, and run artifacts belong under `./runs`; neither is part of a source package. ## Required User Configuration Check `config/evaluation.json`. If absent, copy `config/evaluation.example.json` plus `config/.env.example`, then request the missing values. One task file contains a non-empty `evaluation_models` list, exactly one `judge_model`, plus `input`, `output`, and host-side `execution` settings. Keep multiple evaluation tasks as separate JSON files and select one with `--config`. Never overwrite an existing file silently. Put only credential variable names in task files. Check `config/.env`; if a required variable is absent from both that file and the host environment, ask the user to populate it without sending the value in chat. Keep `config/.env` at permission `0600`. If the user already identifies a task JSON and dotenv file, use them directly with `--config` and `--env-file`; do not ask for the same settings again. ```text id: stable, unique run/output identifier input: project-relative data path and row limit output: project-relative output root execution: host-side image-pull concurrency model: Anthropic-compatible model name base_url: fixed model API destination credential: API-key environment variable or local key file; otherwise prompt concurrency: optional per-model task concurrency; default 4 ``` Run evaluation models sequentially. Recorded and Potential use the current evaluation model; every evaluation model uses the same selected Judge. Defaults are per-model concurrency 4, maximum 300 turns, and 5,400 seconds per task. When the user requests a subset, keep the shared config and pass one `--only-model ID` per selected model. Do not edit away unselected entries. Keep API-key values in project-local `config/.env`, host environment variables, user-owned local key files, or hidden terminal input. The controller loads `config/.env` automatically, while existing host variables take precedence. Hidden input is only for a user launching the controller directly in an interactive terminal. Never put key values directly in chat, `models.json`, a skill command line, dataset, output, log, or generated result file. ## Workflow 1. Locate or bootstrap the project checkout. 2. Check standard local Docker and Claude Code. Install project-local tools under `./.tools` when needed. 3. Copy or check the tracked task example, update `config/evaluation.json`, then verify that `config/.env` or the host environment supplies each referenced credential variable without reading it back to the conversation. 4. Download the requested public dataset configuration to `data/Active-SWE.parquet`, convert it to `data/Active-SWE.jsonl`, copy the exact run input below `///inputs/`, validate it, and pull exactly the referenced images. Review the image preflight report before starting stages. The task's `input.limit` controls the default scope; pass `--limit N` only for an intentional one-run override. 5. Run each evaluation model through Recorded and Potential, use the selected fixed judge for Judge, then compute that evaluation model's metrics. 6. Report artifact locations, stage outcome counts, and computed metrics relative to the run root, without exposing host paths, credentials, or endpoint details. Treat controller `ok` as pipeline completion, not proof that every sample generated all expected artifacts. Use the controller unless the user explicitly requests stage-level debugging: ```bash python -m active_swe.run_evaluation \ --config config/evaluation.json ``` ## Isolation Recorded and Potential containers always use Docker `--network none`. Model traffic reaches only the configured fixed destination through the bundled byte-transparent API tunnel. WebSearch, WebFetch, general network clients, and Git history or network operations fail closed. The controller manages the tunnel and must not print or persist its destination or payloads. Judge is outside the Recorded/Potential isolation boundary. It retains the existing host-network behavior and evaluates the image worktree without the generation command policy. Never use Judge networking as a fallback for Recorded or Potential.