--- name: aeon-doctor description: Static config-correctness linter for this instance - catches the silent-failure class (unquoted schedules, duplicate keys, unconfigured skills, mode typos, broken requires/MCP refs) that no run-based health skill can see. Notifies only on problems. metadata: title: Aeon Doctor category: evolution var: "" # ""=lint the whole instance config | =lint one skill's entry + SKILL.md tags: - meta - health mode: read-only --- > **${var}** — scope. **Empty (default)** = lint the entire instance config (`aeon.yml` + every `skills/*/SKILL.md` + `.mcp.json`). A **skill slug** (e.g. `digest`) = lint just that one skill's `aeon.yml` entry and its `SKILL.md`. Today is ${today}. You are this instance's **config doctor**. Every other health skill (`heartbeat`, `skill-health`) reads *run outcomes* — did a skill fire, did it pass. You read the **config itself**, before anything runs, for the class of bug where a skill is silently misconfigured and **never fires at all** — no error, no failed run, nothing in the Actions tab to notice. That class is invisible to run-based observability *by construction*, and it is the single most common reason an Aeon instance quietly stops doing what its operator thinks it does. You do **not** fix anything — a diagnostic that inspects config must never mutate it. You surface precise, actionable findings; the operator (or `skill-repair`) applies the fix. ## Preamble (always) 1. Read `memory/MEMORY.md` for context and scan the last ~3 days of `memory/logs/` — **drop any finding you already reported** so you don't re-nag a known-but-unfixed issue every run. (A finding is "the same" if it's the same check on the same skill.) 2. Resolve scope from `${var}`: empty → all skills; a slug → restrict every check to that skill (skip fleet-wide-only checks like duplicate-key detection unless they touch the target). 3. Every check below is a **pure local file read** - `grep`, `sort`, and `node` (inline or `node scripts/*.js`). This skill is `read-only`, whose tool allowlist (`scripts/skill_mode.sh`) has `grep`/`sort`/`head`/`tail`/`wc`/`cat`/`jq`/`node` but not `awk`/`sed`/`comm`/`bash`, so the checks use those. No network, no secrets, no GitHub API. If a referenced script is missing, skip that check and note it; **never let one check's failure stop the others**. ## Steps — run every check, collect findings Each finding = **{check, skill, severity, one-line what's-wrong, exact fix}**. Severity: - **critical** — an `enabled: true` skill that will **never fire** or will run with the **wrong privilege**. Live breakage. - **warn** — a latent trap: the same defect on a *disabled* skill, or a correctness issue that degrades silently rather than killing the run. ### 1 · Unquoted `schedule:` on the old regex scheduler (critical / warn) Applies only when `scripts/parse-aeon-config.sh` is **missing**. With it, `scheduler.yml` reads `aeon.yml` through yq and a bare `schedule: 0 12 * * *` fires like a quoted one, so the command below prints nothing. Without it, `scheduler.yml` matches schedules with the bash regex `schedule: *"([^"]+)"`: an unquoted value doesn't match, is read as empty, and the skill is **skipped every tick, forever** - the file is still valid YAML, so nothing else notices. ```bash [ -f scripts/parse-aeon-config.sh ] || grep -nE '^\s+[a-z0-9-]+:\s*\{[^}]*schedule:' aeon.yml | grep -vE 'schedule: *"' ``` Each printed line is an entry whose `schedule:` isn't double-quoted. **critical** if that entry is `enabled: true`; **warn** if disabled (it'll be dead the moment it's enabled). Fix: add the quotes (`schedule: "0 12 * * *"`), or pull upstream to get the yq scheduler. ### 2 · Duplicate skill keys — silent shadow (critical) A repeated skill name under the `skills:` map silently disables the first copy (last-wins YAML). ```bash node scripts/validate-config.js # authoritative — dup keys + checkout ordering node -e 'const k=require("fs").readFileSync("aeon.yml","utf8").match(/^ [a-z0-9-]+:/gm)||[];const c={};for(const x of k)c[x]=(c[x]||0)+1;for(const x in c)if(c[x]>1)console.log(x.trim().slice(0,-1))' # names appearing more than once ``` Any printed name (or a dup-key error from the validator) is a finding. **critical** if either copy is enabled. Fix: remove the shadow copy. ### 3 · On disk but unconfigured — invisible skills (warn) A skill with a `SKILL.md` but no `aeon.yml` entry defaults to disabled, so "not configured" and "deliberately off" look identical. ```bash node -e 'const fs=require("fs");const k=new Set((fs.readFileSync("aeon.yml","utf8").match(/^ [a-z0-9-]+:/gm)||[]).map(x=>x.trim().slice(0,-1)));for(const d of fs.readdirSync("skills").sort())if(fs.existsSync("skills/"+d+"/SKILL.md")&&!k.has(d))console.log(d)' ``` Each printed name exists on disk but has no config entry. **warn** — list them so the operator can decide (enable, or accept it's intentionally uninstalled). ### 4 · Enabled skill with no `SKILL.md` — broken entry (critical) The inverse: an `aeon.yml` entry pointing at a skill dir that doesn't exist. For every `enabled: true` key, confirm `skills//SKILL.md` is present. Missing → **critical** (the run fails or no-ops). ### 5 · `requires:` entry the allowlist silently drops (warn) Both list forms parse — inline (`requires: [KEY?]`) and block (`- KEY` on its own line), top-level or nested under `metadata:`. What still bites is the *value*: `scripts/skill_requires.sh` injects only names matching `^[A-Z][A-Z0-9_]{2,}$` (a trailing `?` = "works better with" is allowed). An entry that fails the filter — lowercase, fewer than 3 chars, a leading digit, or stray punctuation — is silently dropped, so the skill declares a credential it never receives and fails or degrades with a confusing auth error. ```bash node -e 'const fs=require("fs");for(const d of fs.readdirSync("skills")){const p="skills/"+d+"/SKILL.md";if(!fs.existsSync(p))continue;const fm=(fs.readFileSync(p,"utf8").split(/^---$/m)[1]||"").split("\n");fm.forEach((l,i)=>{const m=l.match(/^\s*requires:\s*(.*)$/);if(!m)return;let it=[];if(m[1].includes("["))it=m[1].replace(/.*\[/,"").replace(/\].*/,"").split(",");else for(let j=i+1;j` — health loop can't key it (warn) `CLAUDE.md` mandates each skill append its daily-log entry under a **`### `** heading — "the health loop parses this shape", and `skill-health` / `heartbeat` key skills by **slug**. A skill that logs under `## ` (wrong level *and* wrong identifier) still runs, but its narrative log is harder for the health view to attribute and for the cross-skill dedup ("read the last 3 days of logs") to match — a silent degrade, never an error. ```bash node -e 'const fs=require("fs");const e=x=>x.replace(/[.*+?^${}()|[\]\\]/g,"\\$&");for(const s of fs.readdirSync("skills")){const p="skills/"+s+"/SKILL.md";if(!fs.existsSync(p))continue;const t=fs.readFileSync(p,"utf8");if(!/memory\/logs\/\$\{today\}/.test(t))continue;if(new RegExp("###\\s+"+e(s)+"\\b").test(t))continue;const n=((t.match(/^name:\s*(.*)$/m)||[])[1]||s).trim();const h=t.match(new RegExp("^##\\s+("+e(s)+"|"+e(n)+")\\b","m"));if(h)console.log(s+" logs under \""+h[0]+"\" - should be \"### "+s+"\"")}' ``` Each hit → **warn**. Fix: change the Log-section heading (the instruction line *and* the example block) to `### `, and demote any sub-sections inside the block to `####`. ## Report - **No findings → send nothing and exit.** A clean config is the common case; silence is correct and keeps this channel trustworthy. - **Findings → one consolidated `./notify`**, most-severe first. Write the body to a scratch file and send with `-f` (never a long argv): ```bash ./notify -f \ --title "aeon-doctor: config issue(s)" \ --severity \ --mute-key "aeon-doctor" ``` Group by severity. For each finding give: the skill, one line on what breaks (and that it's **silent** — the operator won't see it in the Actions tab), and the exact one-command fix. If a `critical` exists, lead with it — an enabled skill that never fires is the whole reason this skill exists. - Do **not** open a PR or edit any file. Point mechanical fixes at `skill-repair` / the `./aeon` CLI; leave the fix to the operator. ## Constraints - Read-only by contract — inspect config, never mutate it. No `Write` / `Edit` / `git` / `gh`. - Every finding must cite the **exact** file + line and a **copy-pasteable** fix. A config finding with no fix is noise. - Don't invent problems: only report what a check actually matched. If every check is clean, say nothing. - Fully local — no network, no secrets, no GitHub API. Run-outcome health is `skill-health`'s job; live attention is `heartbeat`'s. Stay in your lane: the **static config**. ## Log This skill is `read-only`, so the workflow's read-only guard writes its `### aeon-doctor` log entry from your captured output; a self-written entry would be a duplicate. Don't append to `memory/logs/` yourself - put this record in your **final output** as bullets: checks run, findings by severity (or `clean`), and whether a notification was sent. End-states: `AEON_DOCTOR_CLEAN`, `AEON_DOCTOR_FINDINGS`, `AEON_DOCTOR_ERROR`.