--- name: review-experiment-results description: >- Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Classifies each metric's outcome and produces plain-language readout by default, surfacing underlying stats (p-value, CI, sample size) on request. Use when asked to explain, review, or interpret experiment results; "did treatment X win"; "is this experiment significant"; "what happened to metric Y in this experiment"; or "should we ship this experiment". Do NOT use for changing experiments (manage-experiments), choosing metrics (choose-metric), or diagnosing bias (SRM root cause, peeking, Simpson's paradox). Trigger phrases: experiment results, readout, winner, statistical significance, guardrail metric, ship this experiment. metadata: author: Harness version: 2.1.2 mcp-server: harness-mcp license: Apache-2.0 compatibility: Requires the Harness MCP server or the Harness CLI --- # Review Experiment Results Explain an FME experiment's outcome in plain language: whether there's a winner, why or why not, how guardrail metrics were affected, and any data-quality caveats - with underlying statistics available on request. Related: [manage-experiments](../manage-experiments/SKILL.md) (experiment changes), [choose-metric](../choose-metric/SKILL.md) (metric selection). ## Tools Works through the Harness MCP server or the Harness CLI; names are from [tool-map.md](../../references/fme/tool-map.md). | Operation | MCP | CLI | |-----------|-----|-----| | **List experiments** | `harness_list` · `fme_experiment` · `filters: { parent_type, name?, match_type: "contains", status?: ["ACTIVE", "PAUSED", "COMPLETED", "ARCHIVED"], offset: 0, limit: 100 }` · `compact: false` | `harness list experiment --parent-type FEATURE_FLAG [--search ] --status ACTIVE`, then repeat with `--status PAUSED`, `--status COMPLETED`, `--status ARCHIVED` (CLI `--status` is single-value; API defaults to `ACTIVE` if omitted) | | **Get experiment** | `harness_get` · `fme_experiment` · `params: { experiment_id }` | `harness get experiment ` | | **Get settings** | `harness_get` · `fme_experiment_settings` · `params: { experiment_id }` | `harness get experiment:settings ` | | **List results** | `harness_list` · `fme_experiment_result` · `filters: { experiment_id, comparisons? }` | `harness list experiment:results --raw` | | **Get/list metrics** | `harness_get` · `fme_metric` · `params.metric_id` (one call per id), or `harness_list` · `fme_metric` · `filters: { ids }` · `compact: false` | `harness get metric ` once per id, or `harness list metric --id ` once per id - the CLI's `--id` flag is repeatable but only the first value reaches the API (spec bug), so repeating `--id id1 --id id2` on one call silently drops `id2` | ## Instructions ### Phase 1: Establish scope Follow [scope-establishment.md](../../references/scope-establishment.md). ### Step 1: Resolve the experiment If the user gives an exact experiment ID, **Get experiment** directly - skip list. Otherwise **List experiments** for parent type FEATURE_FLAG (or AI_CONFIG if user specified), with name substring match if provided (MCP: `match_type: "contains"`). Default parent type: FEATURE_FLAG. If user didn't specify status, list ACTIVE, PAUSED, COMPLETED, **and ARCHIVED** - a finished or archived experiment is exactly the kind a results question is usually about, so never default to active-only here. Fully paginate all selected statuses per [pagination](../../references/fme/tool-map.md#pagination) before deciding uniqueness or absence; CLI needs a separate paginated query per status. More than one match or no name given → ask which experiment. If a scan is incomplete, report that instead of silently selecting a first-page match. ### Step 2: Fetch experiment definition **Get experiment**. Key fields: `status`, `environment`, `startAt`/`endAt`, `baselineTreatment`, `comparisonTreatments[]`, `keyMetrics[]`, `supportingMetrics[]`. `keyMetrics`/`supportingMetrics` are the only metric lists on the experiment. `GUARDRAIL` and `ALERT` metrics are workspace-wide categories applying to every experiment, so they never appear here - read role from each result's `category` (Step 4). Empty `keyMetrics` → no winner criterion; ask which metric(s) should drive verdict. More than one `comparisonTreatments` and user didn't name one → report every treatment (one verdict each, see Output Format). If question presumes a single treatment ("did it win?") → ask which treatment or confirm per-treatment readout. ### Step 3: Fetch experiment settings **Get settings**. Always returns applied settings (404 only if experiment doesn't exist). `source: "EXPERIMENT_OVERRIDE"` means experiment has its own override; other values mean it inherits org defaults. Fields to carry forward: `significanceThreshold` (not always 0.05), `multipleComparisonCorrection` (`NONE` or `GROUPWISE_HOCHBERG`; applies only to KEY/SUPPORTING, never GUARDRAIL/ALERT), `minimumSampleSize` (per treatment), `statisticalTestType` (`FIXED_HORIZON`/`SEQUENTIAL`), `reviewPeriod` (ISO-8601 duration), `varianceReduction.method` (`NONE`/`CUPED`). ### Step 4: Fetch evaluated results **List results** for the experiment, optionally narrowed to specific comparison treatments. List-only (no single-result get, no per-environment filter). Returns latest calculation run only, as one row per (metric, comparison treatment) pair, plus top-level `calculatedAt` (`null` if never calculated). If narrowing to specific treatments, filter the returned rows on each item's `comparison` (see [tool-map.md](../../references/fme/tool-map.md)). Per-row fields: `category` (`KEY`, `SUPPORTING`, `GUARDRAIL`, or `ALERT`), `comparison` (comparison treatment this row is for), `metricId.id` / `metricId.name` (`name` can be `null`; resolve in Step 5), `metricResultState` (stats engine verdict; `null` if none yet), `pvalue`, `impactLower`/`impactUpper`, `value`, `errorMargin`, `baselineMean`/`comparisonMean` (each with `Lower`/`Upper`), `baselineSampleSize`/`comparisonSampleSize`, `varianceReduction`, `positive` (metric's configured `isPositive`, not observed outcome; non-null even when numeric fields are `null`). No SRM field in response, so can't confirm observed treatment allocation matched configured treatment allocation - Step 8 says so whenever it reports a winner. Report `metricResultState`/`pvalue` as given; never recompute significance or claim to correct for peeking. **Reconcile before scoring.** The expected row set is the Cartesian product of every configured metric (`keyMetrics` ∪ `supportingMetrics` from Step 2) × every comparison treatment in scope (all of them, or the one(s) picked in Step 2); it must be non-empty, since Step 2 already stops on empty `keyMetrics`. Match returned rows against it: - Two or more rows for the same (`metricId.id`, `comparison`) pair → duplicate/conflicting data; don't pick one arbitrarily - treat that combination as `needs_more_data` and note the conflict in Step 8. - An expected (metric, comparison) combination has no row at all → treat as `needs_more_data`, same as a row with `metricResultState: null`; never treat a missing combination as "nothing to report." - A row's `comparison` isn't in the experiment's current `comparisonTreatments` → drop it (stale data from a removed treatment). - `calculatedAt` is `null`, or the whole result list is empty → no calculation has run; every KEY result is `needs_more_data`, so Step 7's verdict is `NEED_MORE_DATA` - never report `WINNER` from an empty or uncalculated result set. GUARDRAIL/ALERT rows aren't part of the expected set (they're workspace-wide, not listed on the experiment) but must still flow through to Step 7's `DATA_QUALITY_CONCERN` modifier when present with no data - don't drop them while reconciling KEY/SUPPORTING coverage. ### Step 5: Resolve metric names and descriptions Resolve every distinct metric ID from returned rows **and** configured metrics missing from those rows, so gaps can be named. MCP may **List metrics** by IDs, fully paginating until every requested ID is resolved or the complete inventory establishes a missing ID; alternatively get each ID directly. CLI must **Get metric** separately for each ID (or one fully paginated single-ID list per call). An unresolved ID stays unverified, not absent from a first-page miss. Include returned GUARDRAIL/ALERT IDs too. Descriptions help judge severity: a 2% dip on "leading indicator, noisy" reads differently from same dip on "primary revenue guardrail". ### Step 6: Classify each result Map each `metricResultState` to `desired`, `undesired`, `inconclusive`, or `needs_more_data` using [state-classification.md](./references/state-classification.md). For `needs_more_data` results, compare `baselineSampleSize`/`comparisonSampleSize` to `minimumSampleSize` so readout can say how far along collection is. ### Step 7: Determine the verdict One verdict per selected comparison. **Completeness gate runs first:** empty/uncalculated results, missing or duplicate configured metric/comparison pairs, or mismatched expected categories force `NEED_MORE_DATA` + `DATA_QUALITY_CONCERN` for the affected comparison. Show observed regressions/guardrails separately; don't hide them. Only a non-empty reconciled set with every configured KEY and SUPPORTING pair present exactly once enters the table below. A missing workspace-wide guardrail inventory remains unverified, not evidence of no breaches. Then evaluate top to bottom; first match wins: | Verdict | Condition | |---|---| | `MIXED` | At least one KEY result is `desired` and at least one result of any category is `undesired` | | `WINNER` | Completeness gate passed; at least one configured KEY result exists and every configured KEY result is `desired` | | `REGRESSION` | At least one KEY result is `undesired` | | `NEED_MORE_DATA` | At least one KEY result is `needs_more_data` (including missing or conflicting combinations caught by Step 4's reconciliation) | | `NO_WINNER` | Otherwise (KEY results are `inconclusive`, or mix of `desired` and `inconclusive`) | Modifiers (attach to any verdict): - `GUARDRAIL_BREACH` - any `category: GUARDRAIL` result is `undesired`. An `undesired` SUPPORTING or ALERT result affects the verdict but is not a guardrail breach. By row order, `WINNER` never carries this modifier. - `DATA_QUALITY_CONCERN` - more than half of KEY results are in a `NO_DATA_*` state; any result is `NO_DATA_SERVER_ERROR`/`FAILED_METRIC`; or any GUARDRAIL/ALERT result is in a `NO_DATA_*` state at all. A guardrail that never received data is a monitoring gap, not just a data point to omit. ### Step 8: Explain the results Default to plain language. Link [concepts.md](../../references/fme/concepts.md#experiments-and-metrics) for SRM and peeking caveats. Lead with verdict, then why. Describe trade-offs, not business decision. `DATA_QUALITY_CONCERN`: say so plainly, and call out any duplicate/conflicting rows or missing metric×treatment combinations Step 4 found. `WINNER` or `MIXED`: add one line: SRM isn't exposed through API; check experiment's results page in Harness UI before acting. Fixed-horizon + before review period: say readout is preliminary - if the verdict is `WINNER`, report it as `WINNER (preliminary)` in the Output Format and don't call it actionable until the review period ends or the test type is `SEQUENTIAL`. Inconclusive but not significance-tested: say why (e.g. `ACROSS` metric). `calculatedAt` non-null: include timestamp. `pvalue` null (e.g. `WAITING_NORMALITY`): don't quote `value`/impact as confident. `GUARDRAIL_BREACH` + `MIXED`: lead with breach. Guardrail or supporting metrics with data (any category present in Step 4's reconciled rows) must appear in the Metric Impact table - never omit a row just because it isn't KEY. If `multipleComparisonCorrection` is `NONE` and the experiment has more than one key metric, add one line flagging the false-positive risk. Stats detail: add p-value, CI, sample sizes, `significanceThreshold`, `multipleComparisonCorrection`, settings `source` when question uses statistical terms or user asks. No stats detail: close with one-line note that stats available on request. ## Output Format For a single comparison treatment: ``` ## Experiment Readout - Experiment: () - Comparing: vs - Calculated: ## Verdict **** [+ `(preliminary)` if fixed-horizon and before the review period ends] [+ modifiers if any] <2-4 sentence explanation> ## Metric Impact | Metric | Role | Result | |---|---|---| | | Key / Supporting / Guardrail / Alert | | ## Notes ``` For more than one comparison treatment, keep one shared experiment header and Notes section, but repeat verdict + table block per treatment under its own `### ` subheading (independent verdict per treatment). Add p-value/CI/sample-size columns to Metric Impact only when Step 8's stats condition is met. ## Examples - "Explain the results of checkout-redesign experiment" → resolve by name, Steps 2-8, plain-language output - "Did treatment B win in exp_8f2a1c?" → resolve by ID, verdict for treatment B only - "Is the signup-flow experiment statistically significant?" → statistical wording triggers stats detail - "Should we ship the onboarding experiment?" → verdict + trade-offs; no ship/no-ship recommendation ## Performance Notes - Resolve concrete experiment id before Step 4; never fetch results by name - Fetch settings (Step 3) and metric names (Step 5) once per experiment, not per treatment - Narrow Step 4 to specific comparison treatments once user picks a treatment instead of fetching every comparison ## Troubleshooting | Issue | Resolution | |-------|------------| | An experiment operation fails | Report error and stop | | Step 2 404s but Steps 3/4 succeed | Without Step 2 there are no treatment or key-metric definitions; tell user full readout isn't possible for this experiment | | Experiment still ACTIVE or PAUSED | If verdict is NEED_MORE_DATA or NO_WINNER, ask whether user wants preliminary readout (caveated) or would rather wait; use `endAt` (and `reviewPeriod`) to say how much longer it's configured to run | | Unrecognized or null `metricResultState` | Treat as `needs_more_data`; don't guess a direction; see [state-classification.md](./references/state-classification.md) | | Duplicate or conflicting rows for the same metric + comparison | Treat that combination as `needs_more_data`; note the conflict; never pick one row arbitrarily | | Expected metric×treatment combination missing from results | Treat as `needs_more_data`; never report `WINNER` from partial coverage | | Results are empty and experiment's `rule` is `default` or another label with no traffic | Results only count impressions whose label matches the experiment's `rule`; offer to update the experiment's `rule` to `"default rule"` via [manage-experiments](../manage-experiments/SKILL.md) | | User asks about SRM | Results API exposes no SRM value; say so and point to experiment's results page in Harness UI; don't diagnose causes | | User wants to change experiment | Route to [manage-experiments](../manage-experiments/SKILL.md) | | User asks which metric to use | Route to [choose-metric](../choose-metric/SKILL.md) |