Local runtime evidence for coding agents investigating performance, memory, execution, concurrency, and reliability.
Language / 语言: English | 简体中文
Flameox connects profilers, benchmark tools, and trace processors to a local evidence record. It preserves their native artifacts and provenance, then exposes bounded evidence to the agent. The agent states what it wants to test; Flameox captures the measurements and preserves the experiment record for review. ## Quick start Install the local runtime and connect a supported MCP client through the guided setup: ```console npx flameox@latest setup ``` Restart the client, open the project you intend to inspect, and ask it to: > Initialize Flameox in this project and list the available profiling capabilities. The setup command installs a versioned local runtime and changes only approved client configuration. Project initialization is separate and creates `.diagnostics/` only after the client calls the initialization workflow for its fixed project root. For source development: ```console uv sync --extra dev uv run flameox init . uv run flameox status ``` Python 3.12 or newer and the committed `uv.lock` are required. ## Investigation path ```text symptom → capture or import → bounded evidence → hypothesis → discriminating experiment → supported, refuted, or inconclusive finding ``` Evidence sources include pyperf, py-spy, pytest-reportlog, coverage.py, Memray, Perfetto, torch.profiler, Nsight Systems, Nsight Compute, ROCprofiler, Compute Sanitizer, NVBench, and typed inference-provider exports. Availability depends on the host, permissions, installed extras, and selected adapter. Flameox reports missing evidence instead of silently substituting a weaker source. A profile helps explore a problem; it does not establish a performance or correctness conclusion. That requires a representative workload, a declared metric and estimand, compatible run identity, preserved samples, and an appropriate semantic oracle. ## Named workloads Commands live in `flameox.toml` as argument arrays. Parameters are declared scalars; there is no shell expansion. ```toml schema_version = 1 [workloads.scan] argv = ["python", "bench.py", "--implementation", "{implementation}"] cwd = "." timeout_seconds = 60 [workloads.scan.parameters] implementation = ["baseline", "candidate"] [workloads.scan.oracle] strength = "cross_treatment_equivalence" argv = ["python", "validate.py", "--implementation", "{implementation}"] [experiments.scan_comparison] workload = "scan" design = "randomized_complete_blocks" blocks = 10 treatment_factor = "implementation" combination_policy = "cartesian" primary_metric = "pyperf.workload" polarity = "lower_is_better" estimand = "median_paired_log_ratio" practical_threshold = 0.05 confidence_level = 0.95 random_seed = 1984 [experiments.scan_comparison.factors] implementation = ["baseline", "candidate"] ``` The MCP `configure_workload` tool validates and writes the canonical definition without executing it. A manually authored valid definition is active immediately; there is no approval copy or secondary workload registry. ```console uv run flameox workload show scan --json uv run flameox capture plan pyperf --workload scan \ --parameters '{"implementation":"baseline"}' --json uv run flameox capture run pyperf --workload scan \ --parameters '{"implementation":"baseline"}' --json ``` Planning resolves every executable once. The resulting binding contains the exact invocation path, canonical target, trust decision, and file identity. Execution revalidates that binding instead of searching `PATH` again. Plans are short-lived, single-use capabilities whose complete intent is retained in the workspace SQLite control plane. ## Experiments and analysis ```console uv run flameox investigations create \ '{"question":"Does the candidate remove reverse-scan overhead?"}' --json uv run flameox hypotheses record @hypothesis.json --json uv run flameox experiment plan scan_comparison \ --investigation