--- name: active-swe-eval description: Prepare and orchestrate the complete local Docker evaluation for Active-SWE. Claude Code remains the fixed executor for Recorded, Potential, and Judge. --- # Active-SWE Evaluation With Codex Use this skill when Codex organizes Active-SWE on the host. Claude Code remains the fixed executor inside every task container. ## Locate the project First look for an existing project root containing both `pyproject.toml` and the `active_swe/` package. Reuse it without pulling, resetting, or replacing local files. If no checkout exists, clone into a new relative workspace and enter it: ```bash git clone --depth 1 https://github.com/XLearning-SCU/Active-SWE.git Active-SWE cd Active-SWE ``` Do not clone over a non-empty path. If Git or repository access is unavailable, stop and ask the user for an existing checkout. After locating or cloning the project, open `.codex/skills/active-swe-eval/SKILL.md` from that checkout and use the local copy as the authority for all remaining steps. Then read its local `ENVIRONMENT.md` and `EVALUATION.md`; do not continue from an older remote or cached copy of the skill. Public sources are: ```text Source code: https://github.com/XLearning-SCU/Active-SWE Dataset: https://huggingface.co/datasets/XLearning-SCU/Active-SWE ``` Check `config/evaluation.json`. If absent, copy `config/evaluation.example.json` and `config/.env.example`, then ask for missing values without silently overwriting files. One task JSON contains `evaluation_models`, exactly one `judge_model`, plus `input`, `output`, and host-side `execution` settings; keep multiple tasks as separate JSON files. A credential may use `api_key_env` loaded from project-local `config/.env` or the host environment; a permission-restricted `api_key_file` remains available. If a referenced variable is missing, ask the user to populate `config/.env` without sending the value in chat. Hidden input is only for a user launching the controller directly in an interactive terminal. Models run sequentially; per-model concurrency defaults to 4, maximum turns to 300, and timeout to 5,400 seconds per task. Never put an API-key value in chat, `models.json`, or a command line. When the user selects only part of the configured models, pass one `--only-model ID` per selection; keep the full config intact. When the user identifies an existing task JSON and dotenv file, reuse them with `--config` and `--env-file`. Use the public controller from the `active_swe` package: ```bash python -m active_swe.run_evaluation \ --config config/evaluation.json ``` Check standard local Docker and Claude Code, download the requested public dataset configuration to `data/Active-SWE.parquet`, convert it to `data/Active-SWE.jsonl`, copy the exact run input below `///inputs/`, pull and validate referenced images, review the preflight image report, then run Recorded and Potential with each evaluation model, Judge with the same selected fixed judge, then metrics. The task's `input.limit` controls the default scope; pass `--limit N` only for an intentional one-run override. The materialized input is also the exact source used to choose preflight images. Project-local tools belong under `./.tools`; run artifacts belong under `./runs`. Recorded and Potential use Docker `--network none`; only their bundled fixed-destination API tunnel may carry model traffic. Judge has a separate network setting and retains host-network behavior. Report artifacts using paths, stage outcome counts, and metrics relative to the run root, without exposing credentials, endpoint details, or host paths. Treat controller `ok` as pipeline completion, not proof that every sample produced all artifacts.