# Current delivery goals Re-scoped on 2026-09-07 after the independent goal audit. The product helps a user develop and revise their intent inside an existing task. It must follow ordinary feedback and retain the current direction, not require repeated manual bookkeeping. 1. **Connect the loop.** After one explicit `/intent start`, let ordinary goals and material feedback maintain small sourced records. Execution corrections keep the goal; real intent changes supersede it. Preserve unknowns and disagreements. Resume/compaction reads current records with their identifiers and sources. 2. **Keep the implementation small.** Reuse the existing host, Hooks, MCP tools and record store. No new tool, storage format, semantic Hook, transcript parser, autonomous Harness, GUI, or general instruction-following framework. 3. **Test the remaining uncertainty once.** Target regressions at this connection and real retained failures. The supplementary live check remains two synthetic cases and at most six user turns. The authorized ten old probes may use at most three policy variants for removal/calibration. Then run the one approved fresh, independent 80-case confirmation. No rolling series of replacement holdouts. 4. **Publish useful development work.** Push reviewed source, failed and passing diagnostic evidence, and an honest status to GitHub now. Update the profile's Research and applied systems entry. A development push is not a release or an efficacy claim. Tag a new prerelease only after the unchanged applicable gates. 5. **Leave no local installation.** Test in isolated homes, remove task-created installations and state, then delete no-longer-needed dependencies and caches. Retain source and sanitized evidence; do not touch unrelated installations. ## Acceptance and stopping rules - Natural feedback, record changes, provenance, supersession and recovery must be observed in a real host run, not inferred from a green unit test. Report separately any recovery check that still exposes conversation history to the model. - Unstarted/private/off tasks must not acquire an automatic persistence path. Hooks remain read-only for ordinary prompts, byte-bounded and fail-open. - Do not enlarge the scope for isolated word-count imperfections shared with the baseline, or add another evaluation abstraction to explain a failed metric. - The original efficacy thresholds remain unchanged. Automated synthetic rework scores are not measured human hours or proof of market demand. - If ordinary use still needs manual bookkeeping, fix that bounded defect before a final study. If the final clean confirmation shows no reliable benefit, publish the failure and stop expansion; do not tune against it and call it independent. Windows, Linux/macOS packaging and the existing DeepSeek adapter stay in scope. New host expansion, fresh marketing assets and unrelated framework work are deferred.