# v0.3.0-beta.1 release decision As of 2026-09-07: **ITERATE / NOT RELEASED.** Reviewed development source and honest diagnostic evidence may be pushed to main. Do not tag or announce v0.3 as a release until every applicable gate passes. Existing v0.2 public artifacts are separate. The independent goal audit found that manual state tools were implemented but ordinary feedback was not connected to them. The [re-scoped delivery goals](current-goals.md) prioritize that connection, provenance and recovery before more efficacy testing. Progress is not a passing product evaluation. ## Public development integration Development commit `c970c28` and the revised GitHub profile entry are public. Its [first 18-job CI run](https://github.com/rrrrrredy/intent-loop/actions/runs/34104135229) failed overall: the nine Codex source-test steps passed, but six Unix generated-file checks found a missing executable Git mode; root evidence checks could not fetch a detached historical ablation commit; and the fresh DeepSeek Corepack home selected an unpinned newer pnpm. These are integration failures, not failed model conversations. The follow-up fixes the generated server's Git mode, pins and verifies pnpm 11.7.0 inside the actual isolated home, and preserves original experiment commits on the separate `evidence/history` branch without merging their runtime changes into main. The corrected local DeepSeek pack/add/compose/help/remove lifecycle passed without an API key and removed its temporary home. The next exact-commit CI run is still required. Full-history clones fetch the evidence branch; shallow/single-branch clones must also fetch `evidence/history` before verifying historical evidence. ## Current verified implementation and development evidence - Repaired candidate `b025222` repeated the 17 authorized controls with five follow-ups and 17 successful native task deletions, no primary retries and no first-turn actions. The 1,496-byte policy retains the observed cheap-draft/comparison/feedback routes; this does not verify costly-branch improvements. A 131-word response to a 130-word request and an unsupplied game mechanism remain visible in the raw outputs. Core source/Git/installed SHA-256 is `7048d10c8f1f501200967dc00c0bb796b90ff0f275652d768ac4cdbc2dec1363`. - Candidate `2b9101f9fc99ab0874f05a5d1e80b03b070de6a0`: privacy fixes cover conditional controls, private false receipts and metadata, interrupted erasure, unrelated recovery records, and concurrent recreation. The adversarial agent reran its fixed 14-case set with no remaining blocker in that scope. - Complete Codex source suites: 120 / 120 on Node 20.19.1; 120 / 120 on Node 22.19.0, including the current v7 local validation. These prove implementation behavior, not product value. - Fresh installed core: 17 first turns, five follow-ups, and 17 task deletions completed with no first-turn actions. Three requested comparisons include mix/reject/free-reply invitations. Raw outputs retain remaining prose/rule-following imperfections; this is not an all-content-pass score. - Selected post-v6 development evidence is preserved as 18 trials / 180 conversations and is explicitly excluded from efficacy claims. The earlier seven-case run finished seven first turns, four follow-ups and seven native task deletions without first-turn actions. The invalid long-interrupted batch is excluded from timing and quality metrics. - Independent v7 author artifacts are sealed and preserved. Candidate `16c4b19` completed 160 attempts but only 155 usable conversations / 75 complete pairs: four explicit capacity errors and one fixed-timeout failure, with no primary retry. All 160 native tasks were deleted. Its 75-pair blind diagnostic does not meet the predeclared 80-pair design and also retains two raw failed gates (inference denial and premature actions). See [preserved v7 failure evidence](../evidence/failed-holdout-v7/README.md). - Both exact lockfile snapshots returned HTTP 200 / zero advisories from the authorized npm official bulk endpoint on 2026-09-07. Release CI must still obtain a fresh advisory result. - Root checks passed for evidence hashes, 8 / 8 DeepSeek adapter tests, legal inventory, and the 18-file DeepSeek package. X and Xiaohongshu drafts still satisfy their local length checks, with four image assets present. ## Remaining gates - Diagnose the retained v7 failures, fix reproducible product defects, and repair any prospective evaluation defects without rewriting original grades. A repaired policy requires a new independent sealed confirmation; v7 is now exposed development material. The user authorized one new independent 80-case OpenAI gpt-5.6-sol paired confirmation and blind grading after clarification on 2026-09-07. Regression, component removal and prospective grader calibration must precede its seal. - All original joint efficacy thresholds must pass on a clean candidate bound to the complete installed plugin tree. Do not change thresholds or present selected pairs, retries or a tuned corpus as the original complete study. - User-perspective review has completed six synthetic cases / 22 actual model turns using normally reviewed Hooks; finish auditing its source-bound evidence, cleanup and accepted usability findings. Preserve the earlier zero-model local-component snapshot separately. - Final exact-candidate source/package/secret checks, real Codex State lifecycle, current DeepSeek host lifecycle, and generated/SBOM/notice consistency. - Push the verified candidate and obtain all 18 main CI jobs, including real Linux/macOS runners; then an annotated exact tag, all 18 tag CI jobs, verified assets and attestations, prerelease, and fresh public installs. - Refresh the existing Research and applied systems profile entry for the new public version. It currently links v0.2.0-beta.5. - Finish deleting task-created isolated installations, state, caches, and dependencies after the work is complete, retaining source and sanitized evidence. Everyday Codex core/State installations and the intent-loop marketplace entry have already been removed and verified absent. ## Evidence limits V5 failed final-match gain, clear latency, and inference denial. V6 also returned STOP: clear paired latency +6.89%, wrong interventions 24.14%, and inference denial 62.5%. At least eleven v6 final requirements were invisible to the user-facing model, independently invalidating its efficacy result. Neither run authorizes publication. V7 is an incomplete primary with a diagnostic subset, not a passing release study. Any later passing study remains a synthetic, automated evaluation on one model/host configuration. Real users may behave differently, models and graders can drift, and DeepSeek compatibility does not inherit a Codex efficacy result. No new platform adapter beyond the existing requested DeepSeek/Linux/macOS scope is authorized by these local checks.