--- name: skill-optimization-study description: Re-runnable measurement loop for agent-driven JetBrains MPS work over mps_mcp_* tools — baseline headless worker runs on fixed scenarios, server call log + transcripts, hotspot ranking, remedy classification (docs / server tool / offline script / online script / template), optional A/B. Use when tool descriptions or mps-* skills changed, before a release, or when agents seem slow or retry-prone on MPS tasks. --- # Skill optimisation study (MPS MCP) A repeatable procedure for finding where agents waste turns and tokens when driving MPS through `mps_mcp_*`, and for choosing the cheapest fix. This skill is self-sufficient: the 2026-09 study documents (`plugins/mcp-tools/docs/skill-script-automation-study.md` and its runbook) are history and evidence, not required reading. Scripts and scenario prompts live in `plugins/mcp-tools/study/`; if that directory is removed, move `scripts/` here and the prompts into `assets/`. ## Conventions used below Agent shells reset the working directory per call, so every command uses absolute paths through two variables — set them at the start of each Bash call (or export them in a wrapper script): ``` STUDY=/Users/vaclav/work/MPS/myMPS-fix/plugins/mcp-tools/study # adjust to the checkout RUNS=$HOME/MPSProjects/mcp-study/runs # evidence dir, outside the repo ``` `$RUNS/inventory.json` is a load-bearing name: `run_worker.sh` records its sha in every run's meta. ## Roles - **Observer** (this session, Opus-class): orchestrates, never performs the MPS task, never tells workers they are measured, evaluates results read-only, writes the report. Owns the whole lifecycle: **synthesizes** empty projects (`scripts/new_study_project.py`), **opens** them via CLI (`mps-project-management`), **closes** them with `mps_mcp_close_project`, and **starts, restarts and shuts MPS down** with `scripts/mps_control.sh`. Before each swap, tell the user the absolute path about to close (if any) and the absolute path about to open — do not ask them to perform the swap and do not wait for approval. - **Workers**: headless CLI processes (`claude -p` or `junie --task`), one per (scenario, model, run), launched by `study/scripts/run_worker.sh`. Evidence = their transcript + the server call log. - **Human**: answers the gate questions, approves pushes, and dismisses MPS dialogs when a call returns `MODAL_BLOCKED` or a close/exit hangs on a confirmation. Does **not** open, close or create projects and does **not** restart MPS — those are observer actions now. ## Gate questions to ask before starting (use them verbatim) Before question 2, run `python3 $STUDY/scripts/list_worker_models.py` and present `models` as a multi-select. The orchestrator model is first and marked; list it first with "(Recommended)" (the picker cannot preselect). Extra ids the user types are allowed. Ask question 2a after question 2, as one single-select question per selected model (batch them, at most 4 per AskUserQuestion call). The recommended level is the one the user-level settings give that model today: its `settingsEffort.perModel` entry for the full model id, else `settingsEffort.default`. List it first with "(Recommended)". Fill the other options from `effortLevels`, at most 4 in all: on Claude leave `max` to "Other", unless `max` is the recommended level, in which case drop `low` instead. With no settings level (or on Junie) there is no recommendation; the CLI's built-in default is unknown, so the user picks. Always pin it, because an unpinned worker takes the observer's last `/effort` for that model, which changes between rounds without a trace (round 19). Ask question 3 only when the detected harness is Claude; Junie non-interactive has no `bypassPermissions` equivalent (`--brave` is interactive-only). 1. Instrumentation: server call log first (needs a plugin rebuild; the observer restarts MPS itself with `mps_control.sh restart`) or transcript-only? 2. Worker models (run list_worker_models.py; default: the orchestrator model from that list). 2a. Effort level for `` (one per selected model; options from `effortLevels`; recommended: the level the user-level settings give that model today, if any). Every run of that model gets `EFFORT=`. 3. Permission mode for workers (default: `bypassPermissions` on the developer's machine). 4. Scope of the first pass before gate 1 (default: S1 + S3 on the selected models). Gate 1 (after the pilot): matrix size. Gate 2 (after the report): which remedies; A/B yes/no. ## Procedure (tick as you go; details in the references) 1. **Preflight** — MPS running with the MCP server enabled. The port belongs to the IDE selector (64343 on the 261 from-sources MPS, 64344 on 262); `mps_control.sh` and `run_worker.sh` detect it from the live launcher themselves (`mps_control.sh url --json` shows what they find). **Do not export `MPS_MCP_URL` during preflight**: MPS may not be up yet, and an exported unconfirmed value pins the whole round to the wrong port. Set it by hand only to override detection (e.g. several MPS processes), or after `mps_control.sh wait` has reported `confirmed: true`. SMOKE targets the harness project, never a developer checkout. Check the toolchain: `claude --version` (≥ 2.1; must accept `--output-format stream-json --strict-mcp-config`) when the detected harness is Claude, or `junie --version` when it is Junie, `python3 -c 'import sys; assert sys.version_info >= (3, 9)'`, `jq --version`. Preflight is self-healing, and every step of it is yours: - MPS not running → `mps_control.sh start` (or, with no capture on file, the IDEA `MPS` run configuration), then `mps_control.sh wait`. - Welcome screen → **synthesize** a *harness project* and open it: `python3 $STUDY/scripts/new_study_project.py --dir ~/MPSProjects/mcp-study/proj/harness` writes the three descriptor files (`migration.xml` derived from this MPS, so no Migration Assistant), then open it via CLI (`mps-project-management`). Announce the path first. Do not retry MCP until that open has landed — Welcome-screen calls are rejected before dispatch. - `mps_mcp_list_open_projects(projectPath=)` must then list it. A synthesized project is empty by construction (`mps_mcp_get_project_structure` returns no modules); nothing in the study depends on a hand-maintained one, though an existing empty project may be substituted if synthesis fails. **Every other project must be closed first** — including the developer's own checkout, the common case when MPS was started from the IDEA run configuration. Disjoint module names are not enough: with `confirmOpenNewProject2 = -1` (the default) the second open raises the modal New Window / This Window prompt and blocks the round (lesson 30). Announce the path you close; it is an observer action, not one to ask for. - Is the call log on? Ask the process, not the log: `ps -ww -p $(pgrep -f '[j]etbrains\.mps\.Launcher') -o args= | tr ' ' '\n' | grep calllog` must print the option, and the file must grow after a tool call. If it is off and gate question 1 said call-log, step 2 turns it on; if gate 1 said transcript-only, expect 0-line `*-server.jsonl` slices and skip the call-log checks below. Record the tool inventory: `MPS_MCP_URL=$($STUDY/scripts/mps_control.sh url) python3 $STUDY/scripts/tools_inventory.py --out $RUNS/inventory.json` (it does not detect the port itself). Run the harness's own unit tests once (`cd $STUDY/scripts && python3 -m unittest discover -s tests -p 'test_*.py'`) — a broken script is cheaper to find here than in the evidence. Then run the contamination guard yourself: `python3 $STUDY/scripts/check_user_agents.py` (exit 0 clean, 3 contaminated). It rejects MPS-related Markdown definitions below `~/.claude/agents` / `~/.junie/agents` — a filename matching `*mps*` or a body containing `mps_mcp`, both case-insensitively — **and** any `mps-*` folder in `~/.claude/skills` or `~/.junie/skills`, which would shadow the per-project catalog and silently replace the thing being measured (lessons 26, 32). `run_worker.sh` runs the same guard before any run side effect. The guard never modifies anything: move an offending user skill out of the skills directory for the round and restore it at wrap-up. Built-in `Explore` and `Task` agents are outside this pin and remain enabled. 2. **Instrument** — the plugin logs one JSON line per dispatched call when MPS runs with `-Dmps.mcp.calllog=` (`McpCallLogListener`, off by default). Turn it on without a human and without touching a tracked file: `mps_control.sh capture` **while MPS is still alive**, then `mps_control.sh calllog $RUNS/server-calllog.jsonl` (writes the option into the capture), then `shutdown` (the close of the last project carries the exit), `start `, `wait`, and one SMOKE run as the readiness gate. `capture` only preserves VM options the live process already carries, which is why `calllog` exists — adding the option to the `MPS` run configuration works too but is study-only and must be reverted at wrap-up (lesson 13), so prefer the capture route. Confirm the relaunched process actually carries it (`ps -ww -p -o args= | tr ' ' '\n' | grep calllog`) and that the file grows. 3. **Template** — the empty fixture is **synthesized, not snapshotted** (`scripts/new_study_project.py`): three descriptor files, no doc surface possible, and a `migration.xml` derived from the MPS that will open it. Module-bearing fixtures (`statechart`, `recipes*`) are still tarballs, snapshotted **without any agent doc surface**: exclude `.git`, `workspace.xml`, and also `.agents/`, `.claude/`, `AGENTS.md`, `CLAUDE.md`. A tarball is a point-in-time copy, so a catalog inside it is what every later round measures no matter how far the bundled skills have moved (lesson 20). Instead, `run_worker.sh` installs the **live** catalog into each run's project right before launching the worker — `scripts/install_skills.py` purges every `mps-*` folder plus both guides and calls `mps_mcp_initialize_project_for_agents`, then records `skillsSha256` in the meta. Verify every tarball: `tar -tzf .tar.gz | grep -E '(^|/)(\.claude|\.agents|AGENTS\.md|CLAUDE\.md)'` must be empty (a synthesized project has nothing to verify). Do NOT put `.mcp.json` in the template; `run_worker.sh` generates the worker's MCP config per run into `$RUNS/-mcp/` from the detected URL and passes it with `--strict-mcp-config`. 4. **Smoke** — `SMOKE` is a harness check, not a scenario: a read-only prompt that lists open projects and stops, so it runs against the harness project itself (no template copy, no evaluation, `pass` stays empty). It is also the **readiness gate after every MPS start or restart** — a live process is not readiness. If the harness project is not open, announce its path, synthesize it if needed and open it via CLI; do not close it afterwards unless the next run needs a different project. `RUNS=$RUNS EFFORT= MAX_TURNS=6 PROJECT_SYNTHESIZED=1 $STUDY/scripts/run_worker.sh SMOKE $MODEL ` — bump `` on every re-run (the harness refuses an existing run id); drop `PROJECT_SYNTHESIZED=1` if the harness project was not synthesized. The transcript must contain `tool_use`, `tool_result`, per-message `usage`; exactly one MCP server; and, when the call log is on, a `SMOKE-…-server.jsonl` slice of ≥ 1 line. 5. **Scenarios** — `study/scenarios/S1..S10/{worker_prompt.md,done_criteria.md}`. Which cells a changed skill actually forces is `study/scenarios.md` (brief list, then the directory). Look that up before picking the matrix; the gate-4 default (S1 + S3) is a pilot default, not that lookup. Add a scenario for whatever skill/tool the directory does not cover. **S10 (project lifecycle) runs last in a round**: it is the only scenario whose worker closes and opens projects, and a mistake in it can leave a modal dialog that blocks every later `mps_mcp_*` call. Prompts are developer-voice, fixed names, explicit "done", NO reporting requirements. Fixtures: `empty-project` (synthesized per run, not a tarball), `statechart` (Projectxx5), `recipes` (a passing S1) — regenerated per `study/fixtures/README.md`, not stored in git. 6. **Runs** — ONE scratch project open at a time (see lessons: shared module repository leaks across projects; S10 honours this by being sequential — it closes one project before opening the next). Restart MPS before a cell whose fixture language an earlier cell already loaded in the current process — any two of S3/S5/S6/S7/S9 on `recipes*`, S2/S8 on `statechart` (`shutdown` → `start` harness → `wait` → SMOKE; launch with `ISOLATION=per-shared-fixture-restart`, recorded per run beside `mpsPid`; lessons 40, 42). A read-only cell (S9) may precede a language-changing one in the same process; synthesized cells (S1, S10) need no restart. Per run: **synthesize** the empty project or copy the fixture tarball (`PROJECT_SYNTHESIZED=1` when synthesized) → announce the scratch path (and any path you will close first) → close a previous scratch with `mps_mcp_close_project` if one is still open → open the new copy via CLI (`mps-project-management`) → confirm with `list_open_projects` → launch detached (`run_worker.sh` first rejects MPS-related user agents, then installs the live skills; either guard failure aborts with exit 3) → poll the PID in bounded loops → evaluate with an Opus subagent using the `done_criteria.md` (read-only `mps_mcp_*`, always with `projectPath`) → record pass/evidence in `.meta.json` → announce the path and close with `mps_mcp_close_project` (`force=false`; on `MODAL_BLOCKED` ask the user only to dismiss the dialog). Sequential, never two workers against one MPS. Pass the model's gate-2a level as `EFFORT` on every launch. Check that every meta's `effort` is that level and that every meta's `skillsSha256` and `guidesSha256` are each the same value before comparing runs; a differing one means the catalog or the installed `AGENTS.md` / `CLAUDE.md` moved mid-round. Open/close details: `references/harness.md`. 7. **Analyse** — `python3 $STUDY/scripts/analyze_runs.py $RUNS [--out DIR]` (default `$RUNS/analysis`) → `metrics.csv`, `tools.json`, `chains.json`, `errors.json`, `hotspots.md`; `pass` is filled from each run's meta after evaluation. Then `python3 $STUDY/scripts/families.py $RUNS > $RUNS/analysis/families.tsv` (per-run family counters, `references/analysis.md`). Filter chains containing `mps_mcp`, group into families, have an Opus reviewer inspect 3 instances per family with `study/scripts/show_steps.py` and assign determinism {1.0, 0.5, 0}. Rank by avoidable turns (fixed context ≈ 150 K cache-read tokens per turn dominates) as well as by the study formula. 8. **Classify** each hotspot D → S → P-off → P-on → T (first fit). Write `HOTSPOT_REPORT.md`: baseline table, ranked hotspots with `run:step` evidence, hypotheses, defects, remedies with owner/contract/saving/risk. Keep a separate `docs-defects.md` from day one. 9. **Treat** — parallel Opus implementers with DISJOINT file sets (skill docs / skill scripts + packaging + drift test / server batch). Never two agents in the same Kotlin toolset; the observer registers new tests in `McpToolsIntegrationTestSuite` and runs the suite between batches. A server-side change needs the plugin rebuilt and MPS restarted: `mps_control.sh restart` + SMOKE, no human step. 10. **A/B** (optional) — restart MPS onto the treated plugin (`mps_control.sh restart`, then SMOKE), then the same runs against the treated tools; success = ≥ 30 % fewer tool calls and ≥ 25 % fewer context tokens on treated scenarios, no drop in pass rate; delete remedies that do not pay. 11. **Wrap up** — fold conclusions into the study doc; revert the VM option; delete fixture tarballs (keep the SMOKE scenario — step 4 needs it); announce and close any remaining scratch with `mps_mcp_close_project`, **closing the last one with `shutdownWithLastProject=true`** if MPS should go down — the shutdown rides on that last close, because a Welcome-screen MPS cannot be stopped over MCP; delete any `-target` directory an S10 run left; clean `~/MPSProjects/mcp-study/`, `~/.claude.json` project entries, and `~/.claude/projects/-…-mcp-study-proj-*/` memory dirs; Junie workers may leave `~/.junie/sessions` — do not auto-delete them; keep the call-log listener. The capture file lives in `$TMPDIR`, outside that cleanup, so a later round can still relaunch. ## References - `references/harness.md` — run_worker.sh, analyze_runs.py, show_steps.py, tools_inventory.py, new_study_project.py, mps_control.sh usage; clean-environment rule; the observer's project and MPS lifecycle protocols (create / open / close / shutdown + relaunch); per-run procedure card. - `references/scenarios.md` — how to run and add a scenario, fixtures, orchestrator project swap, done-criteria style. Which scenario covers which skill: `plugins/mcp-tools/study/scenarios.md`. - `references/analysis.md` — metrics, chain scoring, rubric, report template, thresholds. - `references/lessons.md` — what went wrong the first time and the rule that came out of it.