--- name: loopx-performance-diagnosis description: Diagnose an owned slow command or runtime with language-appropriate profilers, independent baseline measurements and local-private evidence. Use for demonstrated CPU, wall-time, allocation or IO regressions; profiling does not qualify a budget or grant execution authority. --- # LoopX performance diagnosis Read the current Goal/Todo contract and repository optimization evidence rules. State the real user operation, observed failure and owning acceptance. Preserve the exact source, runtime/interpreter, backend/history, inputs and concurrency. Use an owned disposable target for writes. Do not attach to a production process, elevate privileges, weaken OS policy, or start paid/remote workloads without the existing authorization. Raw stacks, paths, arguments and profiles stay ignored and local-private; public summaries contain only generalized evidence. Read `loopx/capabilities/performance_diagnosis/README.md` for tool selection and blind spots. On installed copies use `loopx capability show performance-diagnosis` to locate its canonical documentation. Python waits/startup: Pyinstrument; Python threads/native on Linux: py-spy; allocations: Memray; Node/TS: V8 CPU/heap. Go/JVM/kernel costs need their native tools. Research official sources when the installed runtime or missing evidence makes the documented choice uncertain. Do not choose a tool from popularity or claim one profiler covers every layer. 1. Run the uninstrumented target and record ordinary elapsed time. Use repeated controlled samples for comparisons; profiling time is not a baseline or p95. 2. Verify the optional tool version locally. Preserve the selected interpreter; install optional tools in an isolated environment when authorized. Do not silently install them into product dependencies or fall back on permissions. 3. Write the exact target argv to an ignored JSON array. Run `loopx performance-diagnosis plan --tool TOOL --command-json FILE --output-directory FRESH_IGNORED_DIRECTORY --format json`. Read the whole plan. A plan has not executed or verified tool readiness. 4. Execute `profile_argv` through the Host executor as an argv array without shell interpolation. Capture only the owned process; record failures and coverage gaps. Profile a separate Node worker instead of inferring its CPU from Python transport waits. Do not collect locals or automatically profile unrelated children. 5. Require actual successful target exit and nonempty artifact. Read Speedscope or V8 CPU using `loopx performance-diagnosis inspect --profile-json FILE --format json`; Memray/heap use their own reporters. Keep threads independent and self/inclusive time distinct. Never sum inclusive rows to obtain latency. 6. Turn the hotspot into a falsifiable hypothesis. Use a controlled intervention to distinguish caller/transport, shared semantics and backend cost. Reuse the owning contract; do not skip freshness, authority checks or decision inputs. 7. Repeat the original uninstrumented workload and semantic checks after a fix. Preserve passed, failed and untested results and update the existing checkpoint. Readback of a stack does not prove root cause, improvement or provider admission. Do not add a new receipt/approval requirement or automatically mutate scheduling, Goal state or telemetry. Stop invoking the workflow to disable it; remove optional tools/captures from their owned local environment when no longer needed.