--- name: evolve-chief-of-staff description: Review normalized Chief-of-Staff execution evidence, maintain a procedural evolution wiki, and propose or evaluate narrow skill improvements. Use when explicitly analyzing CoS performance or evolving its skills; do not use for ordinary operational reviews or as a path for storing personal facts. --- # Evolve Chief of Staff Improve the Chief-of-Staff skill through evidence-linked procedural learning. Keep raw runs, accumulated lessons, and deployed skills distinct. ## Required Inputs Work only from the materials available for the requested evolution pass: - normalized CoS run records and delayed outcome annotations; - the existing CoS evolution wiki; - the current operational skill version; - prior proposal and evaluation history; - retained validation fixtures or independent comparison episodes when available. Do not treat Codex memory summaries as raw traces. Do not infer missing source content or reconstruct hidden reasoning. Read [references/evolution-records.md](references/evolution-records.md) when creating or changing wiki entries, skill proposals, or evaluation records. Read [references/wiki-memory-adapter.md](references/wiki-memory-adapter.md) before consolidating operational evidence into the evolution wiki. ## Maintain the Wiki Review successful, failed, and ambiguous evidence. Update procedural patterns, counterexamples, scope conditions, open questions, and evidence references before proposing a skill change. Admit a generalization through either: - an explicit standing instruction from the user; or - sufficiently independent outcomes supporting a scoped induced lesson. Keep particular people, projects, messages, meetings, papers, and deadlines in the vault or their authoritative source. The wiki may record how to retrieve, verify, prioritize, or route such objects, but not their individual contents. An individual example normally remains evidence. Do not turn it into a universal rule without justified scope. ## Propose a Skill Change When the wiki supports a change: 1. Select one coherent behavioral improvement. 2. Identify the supporting patterns, evidence, and known counterexamples. 3. Make the smallest change likely to improve the targeted behavior. 4. Preserve existing authorization, privacy, provenance, and semantic-memory boundaries. 5. Explain the expected effect and likely regression risks. Do not update the deployed operational skill merely because a proposal was generated. ## Evaluate and Gate Evaluate the candidate against independent evidence when available. Early passes may use human review and subsequent real episodes; label them exploratory rather than pretending they prove general improvement. Convert representative successes, failures, and boundary cases into protected fixtures only when doing so materially improves future comparison. Never test mutating behavior against live systems. Use frozen, redacted, or synthetic sources. Record whether the proposal was accepted, rejected, or rolled back and why. Preserve wiki learning even when the candidate skill is rejected. Adoption requires explicit human approval until the user establishes a different reviewed policy.