--- name: email-evals description: Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. Use when a user wants to define synthetic email cases, validate a suite, inspect its dry-run plan, run it after approval, or regrade changed assertions. --- # Email evals ## Gather the suite one answer at a time Ask one logical question at a time; do not dump a questionnaire. Ask exactly one prompt, wait for the answer, then advance. Conditionally omit a later prompt only when an earlier answer proves its field irrelevant; never bundle prompts. 1. **use case** — Ask: “What is the synthetic use case?” 2. **existing target runtime** — Ask: “Does the existing target runtime already run?” 3. **dedicated actor environment name** — Ask: “What is the dedicated actor environment name?” 4. **dedicated target environment name** — Ask: “What is the dedicated target environment name?” 5. **expected action** — Ask: “What is the expected action?” 6. **exact allowed recipients** — Ask: “Which exact allowed recipients are required?” 7. **sender** — Ask: “What is the sender?” 8. **Reply-To** — Ask: “What is the Reply-To expectation?” 9. **thread** — Ask: “What is the thread expectation?” 10. **subject** — Ask: “What is the subject expectation?” 11. **required facts** — Ask: “What required facts apply?” 12. **forbidden patterns** — Ask: “What forbidden patterns apply?” 13. **attachments** — Ask: “What attachments are expected?” 14. **timeout** — Ask: “What timeout applies?” 15. **lifecycle** — Ask: “What lifecycle outcome is required?” This skill does not build or start the target agent runtime. Refuse real customer messages, identifiers, customer domains, or production-derived fixtures. Immediately propose a synthetic replacement, such as a fictional order identifier and a `.test` mailbox. Keep every case and attachment synthetic. Do not scaffold until these answers are sufficient to make deterministic assertions. ## Contain the evaluation before setup Never change e2a protection. Tell the user to configure containment separately, in the same dedicated account for e2a, with an account-scoped API key: - actor: `allowlist/block [target]` - target: `allowlist/block [actor, probes...]` Use dedicated actor and target agents only. Do not infer, broaden, or repair protection settings; surface protection failures during validation instead. ## Scaffold, edit, and validate After gathering sufficient answers, scaffold the suite: Resolve the absolute directory containing this loaded `SKILL.md` resource. For the tool call, set `EMAIL_EVALS_LAUNCHER` to the absolute path formed by joining that resource directory with `email-evals.sh`. This is an agent-local value, not a persisted environment variable. Do not derive it from the user's current directory or a client-specific plugin-root variable. ```sh "$EMAIL_EVALS_LAUNCHER" scaffold --root --name --target-env --actor-env ``` Edit the generated `suite.yaml` and case YAML files to reflect the answers. The credential name is fixed outside suite authority: export only `E2A_EVAL_API_KEY`. YAML interpolation is limited to the documented actor, target, and probe mailbox fields; never interpolate names, actions, timing, subjects, bodies, patterns, or attachment metadata. Then verify the checked-in, single-file plugin runtime bundle and validate the suite. Setup performs no package installation and never copies or executes JavaScript or dependencies beneath the suite root: Resolve the absolute directory containing this loaded `SKILL.md` resource. For the tool call, set `EMAIL_EVALS_LAUNCHER` to the absolute path formed by joining that resource directory with `email-evals.sh`. This is an agent-local value, not a persisted environment variable. Do not derive it from the user's current directory or a client-specific plugin-root variable. ```sh "$EMAIL_EVALS_LAUNCHER" setup --root "$EMAIL_EVALS_LAUNCHER" validate --suite /suite.yaml ``` Show the complete alias-only dry-run plan, its `approvalDigest`, and any protection failures. Validation is the preflight gate; do not run a suite while it reports a protection or capability failure. Treat the digest as an opaque value; it binds the resolved identities, origin, containment posture, stimuli, assertions, recipients, and execution limits without revealing them. The public `https://api.e2a.dev` origin is the default. If the suite names a custom/local origin, the operator must independently authorize the exact origin on every command with `--trusted-origin `; cleartext origins are restricted to loopback. Never infer this flag from suite content. ## Request approval immediately before sending Ask for explicit user approval immediately before `run`, because it sends real email between the dedicated agents. Do not treat earlier authoring answers as approval. Only after that approval, run: Resolve the absolute directory containing this loaded `SKILL.md` resource. For the tool call, set `EMAIL_EVALS_LAUNCHER` to the absolute path formed by joining that resource directory with `email-evals.sh`. This is an agent-local value, not a persisted environment variable. Do not derive it from the user's current directory or a client-specific plugin-root variable. ```sh "$EMAIL_EVALS_LAUNCHER" run --suite /suite.yaml --approval-digest ``` If the suite, resolved mailbox identities, custom origin, protection posture, or plan changed since validation, `run` fails closed. Validate again, show the new complete plan, and request fresh approval; never substitute or guess a digest. ## Inspect and iterate Read the generated `report.md`. Summarize deterministic failures without hiding errors, then propose the smallest case/agent change that addresses each failure. When only assertions changed, use `regrade` instead of `run`; regrade performs no sends: Resolve the absolute directory containing this loaded `SKILL.md` resource. For the tool call, set `EMAIL_EVALS_LAUNCHER` to the absolute path formed by joining that resource directory with `email-evals.sh`. This is an agent-local value, not a persisted environment variable. Do not derive it from the user's current directory or a client-specific plugin-root variable. ```sh "$EMAIL_EVALS_LAUNCHER" regrade --suite /suite.yaml --run ``` Regrade accepts a changed full-suite digest only when the execution digest is unchanged. It always uses the current validated assertions and rejects changes to sending, correlation, timing, containment, actor, target, or origin inputs. The output root also contains the private mode-`0600` `.email-evals-artifact-auth-key`, which authenticates redaction-loss metadata independently of the rotatable API key. Keep that hidden file with its run directories; never publish it or place it inside a suite. Every ordinary case record carries authenticated redaction state, including an explicit empty declaration when no text was erased. Regrade fails closed if this authentication root or a case's authenticated redaction state is missing, replaced, or altered. If body evidence required configured-pattern redaction, the forbidden-pattern set must remain identical (reordering is allowed); changing that set fails closed because arbitrary new regexes cannot be evaluated against erased text. ## V0 limits Keep claims within the launch slice: no semantic judge, no deep HTML equivalence, no scheduled-send proof, and no full review/bounce/complaint matrix. Do not promise simulator, model, or data-generator features.