--- name: interego-evaluate-retain description: Evaluate an Interego intervention, workflow, skill, or agent configuration and retain a versioned change only to the level supported by evidence. Use for before-and-after comparisons, benchmark reviews, retained lessons, and decisions to adopt, revise, or retire improvements. --- Use the connected Interego tools; hosts may prefix their names. Recover the current evaluation and artifact heads before using older reports. 1. Specify the question, comparison, representative tasks, scoring rubric, resource limits, and adoption criterion. Use the same conditions across arms; record differences that cannot be controlled. Separate validation of connectivity or persistence from evaluation of performance. 2. Collect actual outputs, scores, failures, elapsed time, cost where available, and assistance. Use unseen or delayed tasks for transfer and retention. Record missing metrics as unknown, not zero, and avoid attributing a gain to Interego when the comparison tests only a lesson or delivery method. 3. Inspect provenance and integrity for the evidence that matters. Distinguish self-reports, relay-signed authorship, independent review, and replayable proof. A verified signature or successful write does not attest truth, capability, safety, or causal benefit. 4. Decide to adopt, revise, retire, or gather more evidence. Preserve the intervention's hypothesis status until the actual criterion is met. Keep domain modal status and review requirements from the live Foxxi/Interego contract; do not invent a promotion rule or bypass one. 5. Persist a concise evaluation with artifact version, baseline and comparison references, results, limitations, decision, and next action. Use `remember` for an ordinary report. For a mutable project or retained lesson chain, resolve `get_current_head`, apply the live supersession contract with `if_match`, and handle a conflict by rereading and reconciling. 6. Read back the durable result. For cross-session acceptance, test retrieval and reuse in an independent authorized session when available. Do not label a second read in the same session as fresh-session continuation. If that test is unavailable, report the exact narrower result. Keep effective artifacts linked to their version and evidence so future work can find them. Do not erase negative results, promise improvement from packaging alone, or silently change sharing to obtain independent access.