# Local Acceptance ## 0.6.3 core-alignment batch (2026-09-16) Source candidate, uncommitted, not frozen, not installed, not published. macOS, Node v25.1.0, pnpm 11.22.0. Baseline `63326f22d40407099baa70c8947c37029749588e` (0.6.2); the four documents handed over with the batch were preserved untouched. ### Reproduced defects and their regressions | Defect | Regression that exercises the old reading | Fixed behaviour | | --- | --- | --- | | F062-01 question marker swallowed the mixed request | `legacyQuestionReadingIsInformational` (the 0.6.2 rule, retained in `src/domain/semantics.ts`) returns true for the three recorded inputs, while the current reading keeps one information range and at least two execution obligations | `tests/domain/v063-core-alignment.test.ts` | | F062-02 session directory was a `resolved` request target | the recorded capture now yields `targetSource.kind = environment_default` and `targetCaptureStatus = clarification_required` instead of `resolved` | `tests/domain/v063-core-alignment.test.ts` (K2) | | F062-03 prepare rendered a recipe the gate refuses | the same item/action pair now returns `incompatible` with `action_not_compatible_with_item` and no `evidence_input_contract` | `tests/domain/v063-core-alignment.test.ts` (K3) | | K4 records an earlier version closed as answered were inherited | an answered mixed record is marked `needs_review` before terminal filtering and blocks the certificate and Goal completion | `tests/domain/v063-core-alignment.test.ts` (K4) | ### Hold-out bookkeeping, stated exactly Fifty set files exist in this batch — thirty-five hold-out sets, fourteen reviewer-probe regression files and one self-review file — and they are not interchangeable: | Set | File | Status | | --- | --- | --- | | hold-out round 1 | `tests/domain/v063-holdout.test.ts` | **Now regression coverage.** It produced three findings; repairing two of them changed the source, and the third was an incorrect case in the set itself. Applying the plan's rule, the finding-affected cases stay as regressions and the set no longer counts as untuned hold-out evidence. | | independent review | `tests/domain/v063-review-regressions.test.ts` | **Regression coverage.** The reviewer's own nine failing probes, kept with the contract expectations the reviewer stated. | | second review | `tests/domain/v063-review2-regressions.test.ts` | **Regression coverage.** The second review's five failing probes and their positive controls — an embedded interrogative read as a clause question, a `check if …` condition, inheritance that fills rather than overwrites, a standing prohibition in preparation. | | third review | `tests/domain/v063-review3-regressions.test.ts` | **Regression coverage.** The third review's four failing probes with the controls that keep the repair from over-reaching — a purpose clause behind the action, a preface that moved the verb offset, a tautological preparation, and the work vocabulary. | | fourth review | `tests/domain/v063-review4-regressions.test.ts` | **Regression coverage.** The fourth review's four failing probes with the controls that keep the structural rule from over-reaching — a question with no subordinate span before it, a real conditional order, and an obligation that DID name the branch. | | hold-out round 2 | `tests/domain/v063-holdout-round2.test.ts` | **NOW REGRESSION COVERAGE.** Its nine findings drove source repairs in the first repair round, so it can no longer be hold-out evidence. | | hold-out round 3 | `tests/domain/v063-holdout-round3.test.ts` | **NOW REGRESSION COVERAGE.** Its findings drove the third repair round. | | hold-out round 4 | `tests/domain/v063-holdout-round4.test.ts` | **NOW REGRESSION COVERAGE.** Its findings drove the fourth repair round. | | hold-out round 5 | `tests/domain/v063-holdout-round5.test.ts` | **NOW REGRESSION COVERAGE.** It found the classifier defect (a purpose clause was masked before segmentation) and that repair changed the source. | | hold-out round 6 | `tests/domain/v063-holdout-round6.test.ts` | **NOW REGRESSION COVERAGE.** The fifth review then returned three source defects (F1-F3), and this set's own oracle was revised three times, so it cannot be untuned evidence. | | hold-out round 7 | `tests/domain/v063-holdout-round7.test.ts` | **NOW REGRESSION COVERAGE.** It found the `then` conflict (the English sequencing preface was also read as a comparative subordinate boundary, so `Then check whether the build passed.` was an acceptance order while the Chinese spelling was an information request) and that repair changed the source. Two of its own expectations were also wrong and are recorded in its header. | | fifth review | `tests/domain/v063-review5-regressions.test.ts` | **Regression coverage.** The fifth review's three defects and four failing assertions, with the probes that already passed. | | hold-out round 8 | `tests/domain/v063-holdout-round8.test.ts` | **NOW REGRESSION COVERAGE.** It found no defect of its own and its expectations were never revised, but the sixth review then returned three counterexamples against the same invariant and those repairs changed the source. | | sixth review | `tests/domain/v063-review6-regressions.test.ts` | **Regression coverage.** The sixth review's three defects and their controls. | | hold-out round 9 | `tests/domain/v063-holdout-round9.test.ts` | **NOW REGRESSION COVERAGE.** It found nothing of its own, but the seventh review then returned three counterexamples against the two-sided invariant and those repairs changed the source. | | seventh review | `tests/domain/v063-review7-regressions.test.ts` | **Regression coverage.** The seventh review's three defects, with both sides of the invariant: an explanation creates no authority, and a following order still does. | | hold-out round 10 | `tests/domain/v063-holdout-round10.test.ts` | **NOW REGRESSION COVERAGE.** It found nothing of its own, but the eighth review then found the explanation scope was still pattern-based and tightened the "and then" control, and that repair changed one of this set's shapes. | | eighth review | `tests/domain/v063-review8-regressions.test.ts` | **Regression coverage.** The eighth review's three counterexamples (a finite complement, a `whether` complement, a longer object), the tightened `and then` reading, the separate-instruction positive controls and the closed-complement controls. | | hold-out round 11 | `tests/domain/v063-holdout-round11.test.ts` | **NOW REGRESSION COVERAGE.** It found nothing of its own, but the ninth review then required that an explanation's scope never authorize an action it mentions, and that repair changed this set's own reading. | | ninth review | `tests/domain/v063-review9-regressions.test.ts` | **Regression coverage.** The ninth review's two counterexamples, the same heads followed by a real instruction, the multi-action plan refusal, and the two scope controls (an unrecognised instruction form keeps its path, a pure reported question keeps its closable lane). | | hold-out round 12 | `tests/domain/v063-holdout-round12.test.ts` | **NOW REGRESSION COVERAGE.** The tenth review then showed a bare question head must carry its non-execution qualification into every child, and that repair changed the source. | | tenth review | `tests/domain/v063-review10-regressions.test.ts` | **Regression coverage.** The tenth review's three question-scope counterexamples, the refusal decided against the obligation's own target, and the investigation-imperative contrast. | | hold-out rounds 13-17 | `tests/domain/v063-holdout-round13.test.ts` … `-round17.test.ts` | **REGRESSION COVERAGE.** Each of these sets found one or more source defects in the question-scope family while the family was being closed: a temporal interrogative read as a condition, the Chinese temporal and subject-prefixed heads, the interrogative vocabulary (`谁`, `怎样`, `何时`…), a question word that doubles as a relative pronoun, a verb-fronted interrogative, and the modal that can stand between the action and the interrogative. Every defect was repaired in the source; each set's own oracle corrections are recorded in its header. | | hold-out round 18 | `tests/domain/v063-holdout-round18.test.ts` | **NOW REGRESSION COVERAGE.** The eleventh review then showed an investigation imperative governs an OPEN complement, and that repair changed the source. | | eleventh review | `tests/domain/v063-review11-regressions.test.ts` | **Regression coverage.** The eleventh review's three investigation-complement counterexamples, the fact-stating contrast, the separate-instruction positive and the gate/preparation agreement. | | hold-out rounds 19-24 | `tests/domain/v063-holdout-round19.test.ts` … `-round24.test.ts` | **REGRESSION COVERAGE.** Closing the investigation boundary exposed six more defects of the same family, each repaired in the source: the `if`-complement taken by the condition splitter, the modal-bearing subordinators (能否/可否/能不能), the postposed Chinese interrogative, the yes/no interrogatives (是否/是不是), the Chinese A-不-A class (要不要/该不该/需不需要/可不可以/对不对), a question about a single action read as an instruction, and an OBJECT list (plugins AND skins) mistaken for an action list. Each set's own oracle corrections are recorded in its header. | | hold-out round 25 | `tests/domain/v063-holdout-round25.test.ts` | **NOW REGRESSION COVERAGE.** The twelfth review then showed a DECLARATIVE investigation complement carries its own actor, and that repair changed the source. | | twelfth review | `tests/domain/v063-review12-regressions.test.ts` | **Regression coverage.** The twelfth review's two actor-complement counterexamples, the state-question contrast, the separate-instruction positive and the gate/preparation agreement. | | hold-out rounds 26-28 | `tests/domain/v063-holdout-round26.test.ts` … `-round28.test.ts` | **REGRESSION COVERAGE.** Closing the actor boundary exposed three more defects of the same family, each repaired in the source: an investigation of a DECLARATIVE `that` clause was authorized, a coordination-free `that` verification was read as an instruction, and the Chinese subject rule needed an action predicate (the rotation verb was also missing from the work vocabulary) so that a state question about an OBJECT keeps its coordinated order. Each set's own oracle corrections are recorded in its header. | | hold-out round 29 | `tests/domain/v063-holdout-round29.test.ts` | **NOW REGRESSION COVERAGE.** The thirteenth review then inverted the complement rule (governing by default, closing only on positive proof of a state question), and that repair changed the source. | | thirteenth review | `tests/domain/v063-review13-regressions.test.ts` | **Regression coverage.** The thirteenth review's five counterexamples — an unrecognised predicate and a subject position after the subordinator — plus the proven-state contrasts and the separate instruction. | | self-review (fifth round) | `tests/domain/v063-repair5-regressions.test.ts` | **Regression coverage.** The three counterexamples found by adjacency self-review while the round-5 set was being written (the English sentence end, the coordinated ordering fragment, the two-repository clause), with the abbreviation, version, single-repository and preface controls. | | hold-out round 30 | `tests/domain/v063-holdout-round30.test.ts` | **NOW REGRESSION COVERAGE.** It required no source change of its own, but the fourteenth review then showed that closure still accepted a state WORD as proof, and that repair changed the source. | | fourteenth review | `tests/domain/v063-review14-regressions.test.ts` | **Regression coverage.** The fourteenth review's four counterexamples — a state word as an action's object or modifier, in both languages — plus the siblings of the same root cause (attributive, relative clause, causative, a Chinese predicate with an action object), the proven-state contrasts, the answerable lane and the plain order. | | hold-out rounds 31-34 | `tests/domain/v063-holdout-round31.test.ts` … `-round34.test.ts` | **REGRESSION COVERAGE.** Each of these sets found one or more source defects while the predicate proof was being closed: the English infinitive test firing on the `to` inside `up-to-date`, a lazy Chinese modifier strip that stopped at a possessive, a state-noun proof that could not name a mirror or a lease, a postposed interrogative whose questioned span carried a verb outside every vocabulary, a yes/no question with an explicit subject that authorized the order beside it, and a governed clause that the message classifier discarded as session talk. Every defect was repaired in the source; each set's own oracle corrections are recorded in its header. | | hold-out round 35 | `tests/domain/v063-holdout-round35.test.ts` | **The current independent set**, written after the fifteenth repair round with wording this batch has never used: the stative and adjective shapes, the subject question beside the imperative question, the classifier and the reader agreeing that a governed clause is work, the pure question that adds no obligation, and inheritance beside them. It required NO source change, so it is the untuned set. | Two hold-out round 1 expectations WERE changed during that round, so the earlier "no expectation was relaxed" claim was wrong and is corrected here: the conditional-request case was rewritten to assert the contract's conditional reading instead of an executable one, and one legacy case was replaced because the sentence it used is genuinely a pure information request rather than a mixed one. Both are corrections of an incorrect oracle, recorded rather than hidden; the case that a conjunction was read as a repository name was a real source defect and was repaired in `src/domain/capture.ts`. Hold-out round 2 initially produced nine failing assertions. Every one was repaired in the source, not in the expectations: - 呢 and 吧 were treated as interrogatives, so "安装这个主题呢。" closed as an answer. Only 吗/? ask on their own; 呢 asks only when the clause carries its own interrogative content, and 吧 never does. - The state vocabulary was too narrow: "是否有新版本" and "Check for a new version" did not read as state questions. - 并且 split as 并 + 且, leaving a stray fragment, and 以及/而后 were not boundaries at all. - 是不是 was read as a negation, turning a question into a prohibition. - A request preface (先/然后/请) hid an investigation opener behind an action verb, so "先检查是否有新版本" read as an order. - A targeted inheritance repair (below) returned to the informational-only clause head, so a mixed clause starting with an order kept only a partial reading. ### Targeted repair round after the independent review The review returned the batch with nine failing probes. The root causes and repairs: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `安装新主题吧。` and `Check whether an update exists and install the package.` closed after a zero-tool final answer | 吧/呢 counted as interrogatives; an unpunctuated English conjunction was not a boundary | `endsOnInterrogative` (only 吗/? ask alone; 呢 needs question content; 吧 never), and a `CONJUNCT_BOUNDARY` that also covers 并且/以及/而后 and bare 并/且 followed by a distinct clause | | A prohibition naming `/repo-b` became the inherited target of a later commit | inheritance accepted any same-unit git item that mentioned a path | candidates must be pending, non-legacy requirements whose disposition is `executable_now`, with `targetCaptureStatus === 'resolved'` and no wait or condition | | `提交分支 release。` bound `release` as the repository | the bare-object reader stepped over the 分支 label but not its VALUE | `labeledTokenRange` records what each field already claims, and a claimed value is never re-read as the repository | | Three references to one repository reported ambiguity | uniqueness was counted per item | candidates are deduplicated by canonical repository identity | | A passed `needsReview` record of the current unit was outside the blocking set, and another unit's answered record was inside it | the blocking set used the pending-only closure plus every answered record | the set is selected by the record's own unit scope (current unit, required descendants, and unit-less legacy records), independent of terminal status | | An English unpunctuated mixed legacy record escaped the upgrade check | the check reused the current fragment splitter, which returned one fragment | the check tests the run's informational fragments against the whole-clause reading, so a historically mixed record is caught whatever rule produced it | | prepare reported `compatible` while execution denied with `mutation_host_lock_unavailable` | prepare did not pass the projection-level snapshot facts | `enabled`, `integrity` and `hostStatus` are passed to the shared judgement | After those repairs the reviewer's 11 probes pass, the round 2 hold-out set passes, and the pre-existing suite is unchanged. ### Second repair round after the second independent review The batch was returned again with five failing probes covering three defects; all three were repaired in the source. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Create a file /tmp/test-status.txt recording whether the tests passed.` closed as `informational / answered` | an `whether`/`if` **inside the action's object** was read as the clause's own question (`INVESTIGATION_THEN_QUESTION` matched the investigation word mid-clause, and the embedded-interrogative rule fired on the whole clause) | the English investigation form has to open the clause, and an embedded interrogative is a question only when it is NOT behind the clause's action (`englishInterrogativeIsMatrix`); a leading interrogation verb followed by if/whether/when/… is handled as an interrogation instead of a condition (`interrogativeTakesIfObject`) | | `提交仓库 /repo-b 分支 main。` then `提交分支 release。` produced `branch: main` | inheritance copied the source's target over the item's, discarding a field the follow-up named | inheritance FILLS unset fields only; the item's own captured selection is kept, and its environment-default identity placeholder is excluded from what it "owns" | | prepare reported `compatible` while execution denied with `mutation_conflicting_prohibition` | `prepare` never passed the standing-prohibition input the shared judgement already supported | `prepare` computes the same prohibition conflict the gate does and passes it in, so the verdict is `blocked` | The second repair round also corrected two things the review did not name: `context_guard_prepare` was feeding the CALLER's proposed target into the compatibility judgement, which made a mere proposal read as `incompatible`; it now judges against the obligation's own target and reports the divergence as a proposal. And `推送分支 release` was recording `refspec: release`, because a bare label value was accepted as a refspec; a refspec must now be spelled like one (a `src:dst` pair or a ref path), while a branch named inside a transfer order is still that order's refspec. The reviewer's five probes are kept in `tests/domain/v063-review2-regressions.test.ts`, and round 3 is the replacement independent set. ### Third repair round after the third independent review Four failing probes covering three defects; all three were repaired in the source. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Create a file /tmp/status.txt to show what changed.` and `Create a script /tmp/check.sh to check the status.` closed as `informational / answered` | `REPORTED_QUESTION` and `INVESTIGATION_OF_STATE` still matched ANYWHERE in the clause, so a purpose clause behind the action ("… to show what changed") was read as the clause's question; only the `whether` branch had been fixed | every English question branch now has to open the clause (`opensClause`), and `INFO_OPENING` is gated the same way | | `Please check if the package is installed.` became `pending / conditional_wait` | `interrogativeTakesIfObject` compared the verb's absolute character offset with the matched head's LENGTH, so a preface pushed it out of range and the condition splitter took over | the verb's offset is located inside the match itself, so a preface cannot move it | | prepare reported `compatible` for a target the gate denied | the previous round's fix passed the obligation's target as BOTH sides of the comparison, making it a tautology | preparation compares the SUPPLIED target exactly as the gate will, and uses the gate's own authorizing predicate, so an under-specified target is reported `blocked / target_not_authorizing` rather than compatible | The third repair round also extended the work vocabulary (`draft`, `emit`, `produce`, `log`, `起草`, `拟定`), without which a purpose clause attached to a verb outside the list fell through to `unresolved` and its question word still matched. The reviewer's probes are kept in `tests/domain/v063-review3-regressions.test.ts`, and round 4 is the replacement independent set. ### Fourth repair round after the fourth independent review Four failing probes covering three defects; all three were repaired in the source. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Archive /tmp/logs to show what changed.` and `Compress /tmp/logs to check the status.` still closed as `informational / answered` | every earlier fix located the clause head by looking for a KNOWN action verb, so a purpose clause behind a verb outside the vocabulary kept its question word | the gate is structural now: a question word behind a subordinate span — `to `, a relative pronoun (`which`/`who`/`that`/`than`), a participle, or a prepositional opener (`after`/`before`/`about`/`for`/…), in English, or 为了/用来/以便/从而/进而/用于 in Chinese — belongs to that span whatever the main verb is (`headOpensClause`, `ENGLISH_SUBORDINATE_BOUNDARY`, `CJK_SUBORDINATE_BOUNDARY`). An unknown main verb now yields `unresolved`, never `answered` | | preparation reported a target verdict for a caller-supplied target the gate had already denied | the previous round's fix fed the CALLER's target in as the obligation's own selection as well, so the authorizing predicate compared a value with itself | preparation uses the gate's own `requestedTargetAuthorizesMutation` against the OBLIGATION's selection and reports `blocked / target_not_authorizing` when the obligation never named that field; an obligation that did name it stays authorizable from the caller's target | | `Create /tmp/check.sh to determine if the service is running.` became `pending / conditional_wait` | `if` inside the purpose span was read as a condition on the main clause, because the condition splitter does not know about subordinate spans | the condition marker counts only when it is not inside a subordinate span opened before it (`conditionMarkerIsClauseLevel`), so the purpose reading survives while `Install the package if available.` stays conditional | The reviewer's four probes are kept in `tests/domain/v063-review4-regressions.test.ts`, and round 5 is the replacement independent set. ### Fifth repair round, found by adjacency self-review Round 5 found one source defect of its own (the interaction classifier masked question terms anywhere in the message, so `打包日志以便确认哪些请求失败。` was dropped before segmentation; subordinate spans are now masked before the question-term test and restored for everything else). Writing round 6 against the fixed source then exposed three MORE defects by adjacency — the English sentence boundary, the coordinated ordering fragment, and the two-repository clause. None of the three was returned by a reviewer; all three are in the fail-OPEN direction, so they are recorded with the same weight. | Defect | Root cause | Repair | | --- | --- | --- | | `Install the package. What changed?` was read as ONE information range: the order disappeared and the record closed as answered | the clause splitter treated `。`/`!`/`?`/`!`/`?` as sentence ends but never the ASCII `.`, so an English two-sentence message stayed one run and its interrogative ENDING decided the whole run. The Chinese spelling of the same message split correctly, which is how the asymmetry survived four rounds | an ASCII full stop is a sentence end when a new clause follows it — whitespace, then a capital or non-Latin letter, a quote or a bracket, or a lower-case word that opens a question or its own instruction (`sentencePeriodEnd`). A decimal, a version number (`0.6.3`), a file name (`README.md`) and an abbreviation whose continuation is an ordinary word (`See e.g. the log`) keep the run whole | | `What changed, and update the README?` answered the update away | the coordinated fragment ended the sentence, so `endsOnInterrogative` claimed it even though its own head is an instruction | a fragment that opens with a coordinating conjunction and then carries its own action head is an instruction (`coordinatedFragmentOrdersWork`); a coordinated clause that is still an investigation (`然后检查是否有新版本`) stays a question | | `提交仓库 /repo-b 与 /repo-c。` silently resolved to `/repo-b` | the labelled repository reader returns the FIRST value and never looked for a second candidate | a second repository-looking token joined by 与/和/及/或/、/,/and/or records `requested_target_repository_ambiguous` instead of a guess (`namesSeveralRepositories`); a file extension and a claimed field value (`分支 main`, `remote origin`) are not candidates | The three counterexamples are kept in `tests/domain/v063-repair5-regressions.test.ts`. ### Sixth repair round after the fifth independent review The fifth review returned four failing assertions covering three defects. All three were repaired in the source, and the reviewer's probes are kept. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Install the package. please report what changed?` was ONE information range: the install was answered away | the sentence boundary was decided from the NEXT word (a word list of question words, conjunctions and verbs at offset 0), so a lower-case request preface left the run whole and the trailing `?` decided it | `sentencePeriodEnd` no longer reads the next word's class at all: a period followed by whitespace and further text ends the sentence, and only a period with no space after it or one that belongs to an abbreviation (`e.g.`, `i.e.`, `cf.`, `etc.`, honorifics, dotted initialisms) keeps the run whole — a closed class instead of an open word list | | `What changed, and archive the logs?` produced NO item | the whole-message interaction classifier returned `conversational` as soon as a question term appeared, before any decomposition, so the coordinated order never reached capture | a conversational verdict now has to show that every fragment either asks or says nothing (`ordersWorkBesideQuestion`); a fragment whose head is a Latin word that is not a question, a descriptive opener or a Chinese statement is work the capture layer sees, and the clause-level rule (`coordinatedFragmentOrdersWork`) treats the same head as an instruction | | `提交仓库 /repo-b、/repo-c。` and `提交仓库 /repo-b 和仓库 /repo-c。` still resolved to `/repo-b` | the alternative-repository rule required whitespace before the coordinator and matched the token immediately after it, so a glued enumeration mark and a repeated field label both hid the second candidate | candidates are enumerated by structure (`repositoryCandidates`): every repository-looking token plus the current-repository deixis, then a pair joined by a coordinator, an optional repeated field label, or a coordinator glued inside one token (Han characters are legal in paths, so the token is split instead of excluded) is an alternative | Writing round 7 after those repairs then found one more source defect of the same family: `then` was both a request preface and a "comparative" subordinate boundary, so `Then check whether the build passed.` stayed an acceptance order while the Chinese `然后检查是否有新版本。` was an information request. The boundary list now holds `than` only, and the English sequencing words join `REQUEST_PREFACE`, so a preface cannot change what a clause asks. Round 8 is the replacement independent set and required no further source change. ### Seventh repair round after the sixth independent review The sixth review returned three counterexamples, all of the same invariant — an execution residue must survive whatever question, punctuation or abbreviation stands around it, and a name the root gave must not be re-judged by its extension. All three were repaired in the source. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `What changed and archive the logs?` produced NO item, while the comma spelling kept the archive | the session-talk classifier split fragments on punctuation only, so the coordinator was invisible to it and the semantic layer never ran | the classifier now decomposes with the SEMANTIC layer's own `splitTextFragments` (coordinators and list separators), and the clause layer's rule was generalized from "a coordinated fragment" to any fragment that asks nothing and carries its own action head (`fragmentOrdersWorkOnItsOwn`), with question CONTENT — not a bare `?` — deciding whether a fragment asks | | `Install the package etc. What changed?` was one information range | the abbreviation exception applied unconditionally, so `etc.` swallowed the sentence end and the trailing question mark decided the whole message | an abbreviation keeps the run whole only when what follows continues the sentence, and a question that follows it in either case opens a new one (`CONTINUES_SENTENCE`, `QUESTION_OPENER`) | | `提交仓库 /repo-b.js 与 /repo-c.js。` silently resolved to the first | the candidate list filtered out anything with a file extension while the first-object reader accepted it, so the two judgements contradicted each other | extensions are no longer identity evidence: a repository may be called `/repo-a.js`. A file argument in another clause is excluded by the join rule instead ("提交仓库 /repo-a,运行 /tmp/script.sh" is not a coordinator list) | Writing the Chinese half of the first finding also exposed a fourth gap in the same place: `什么变了并归档日志?` stayed one information range, because a bare 并 split only before a KNOWN action head. A coordinator that separates an ASKING clause from a clause that does not ask is now a boundary whatever verb the second clause uses, so the unknown Chinese verb survives too. Round 9 is the replacement independent set and required no further source change. ### Eighth repair round after the seventh independent review The seventh review returned three counterexamples and stated the requirement on both sides: an execution obligation must not be closed by an answer, and a request for an explanation must not become execution authority. All three were repaired in the source. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Explain how to install foo and restart service api.` produced a restart obligation that the production authorizer returned `authorized` for | the coordinator sits inside the question's own object ("how to install X AND restart Y"), but every decomposition rule treated the second verb phrase as an independent instruction | a question whose object IS a coordinated action list governs the whole list (`governsActionList`, checked before partitioning and inside the clause splitter). It requires a question head — a reported question, an investigation or a question word — plus an interrogative object (`how to `, 怎么/如何/…), so a question that merely stands beside an order still keeps the order, and `Explain the deploy. Then restart service api.` still authorizes the restart in English and Chinese | | `Install the package etc. please tell me what changed?` was again one information range | the abbreviation tie-breaker only recognised a question that STARTS with a question word, so the politeness preface hid it | the tie-breaker no longer looks for words at all: when a fragment is informational only because it ends interrogatively, a yes/no question (final 吗/呢, or an English auxiliary) covers the clause, while a trailing wh-question leaves any action head before it as a residue (`questionCoversWholeClause`, `executionResidueBeforeQuestion`) | | `提交仓库 /repo-a 分支 main。提交仓库 /repo-a 分支 release。提交。` made the last clause inherit `main` | candidates were deduplicated by repository alone and the first source's whole field set was copied, so clause order decided which branch a later clause authorized | uniqueness is judged PER FIELD: candidates are grouped by canonical repository, and a field is inherited only when every candidate of the group that names it agrees. A conflicting branch (or remote or refspec) is left unset, the clause keeps the fields that were agreed and reports `requested_target_field_ambiguous`, and authorization is denied rather than guessed | The self-review sweep that followed the repairs found one more instance of the same invariant, this time in the SESSION layer: `Archive the logs etc. 请说明一下哪些请求失败了?` and `什么变了并归档日志?` produced NO item at all, because the classifier's fragment test called the whole fragment a question before the reader ever ran. The classifier now applies the same residue rule as the reader (question content only speaks for the text before it), keeps each sentence's own closing mark, and both cases are regressions. The repairs also completed two closed classes rather than patching samples: the reported-question head now includes the Chinese reporting verbs (解释/说明/描述/讲解/ 说说/告诉我), and the identity comparison treats two spellings of one repository as one value. ### Ninth repair round after the eighth independent review The eighth review returned three counterexamples that all escaped the SAME approximation, and named the reason: the explanation scope was a PATTERN — it required `to ` and a fixed character window — so a finite complement ("how I CAN install …"), a `whether` complement and a longer object each let the scope escape and the coordinated action became authority again. It also tightened one control: `and then` can belong to the operation order BEING EXPLAINED, so it must not be read as an actual instruction. The repair replaces the pattern with the sentence: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Explain how I can install foo and restart service api.` | the scope needed an infinitive (`to `) | an explanation head governs its SENTENCE; the coordination belongs to the explanation exactly when the complement is still OPEN at the coordinator — an infinitive or a modal in English, a manner interrogative in Chinese (`OPEN_COMPLEMENT_BEFORE_COORDINATOR`). No phrase list, no window | | `Explain whether I should install foo and restart service api.` | the same | a `whether` complement with a modal is open, so it governs | | `Explain how to install the optional development package with its recommended configuration and restart service api.` | the object exceeded the character window | the window is gone: the test looks only at whether the complement was open where the coordinator appears | | `Explain how to deploy and then restart …` must not authorize | `and then` was read as a sequence instruction | while the complement is open, a coordinator — `and then` included — describes the explained operation order, so it creates no authority. The positive control is an explicitly SEPARATE instruction (its own sentence), in both languages | Two closed-complement controls are kept so the repair cannot over-reach: `Tell me what changed and install the package.` and `Explain the incident, rotate every credential and redeploy.` open a NEW predicate and keep their orders, and an explanation of a quoted command (`Explain \`git rebase\`.`) stays undecidable rather than answerable. Round 11 is the replacement independent set and required no further source change. ### Tenth repair round after the ninth independent review The ninth review showed the explanation scope was still keyed on WORDS: the open-complement test looked for `to` or a modal, so a finite complement (`Explain how you install foo and restart service api.`) escaped it and the restart became authority. It named the rule that closes the family: **only proof that an action has LEFT the explanation's scope may grant execution authority, and a protection pattern that did not match is never that proof.** | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Explain how you install foo and restart service api.` | the scope decision required a modal or an infinitive in the complement | the decision is structural and needs no words: a sentence an explanation heads is decided as ONE scope before any partition, and it is `unresolved` whenever a coordinated part carries an action of its own (`isExplanationScope`). A pure reported question with no action residue keeps its closable lane, and an explanation of a quoted command stays undecidable | | `Explain why we install foo and restart service api.` | the same | the same rule; the complement's form no longer matters at all | | the restart must not authorize even though the clause is undecided | the gate authorized any pending requirement whose action and target matched, so a disposition of `unresolved` was not itself a bar | the mutation gate and `evaluateCompatibility` refuse an explanation's scope as a LAST RESORT, so every earlier refusal keeps reporting the reason it always did. The rule is scoped: an `unresolved` clause that is NOT an explanation's scope (an unrecognised instruction form such as `应用包 foo 版本 0.6.3 配置档 default。`) keeps the path it always had, and the action PLAN of an explanation is refused action by action | Round 12 is the replacement independent set and required no further source change. ### Eleventh repair round after the tenth independent review The tenth review showed the protection still keyed on an EXPLANATION head: a bare question (`如何…?`, `How do I …?`, `Can you explain …?`) was decomposed into an information range plus a child that had LOST its parent question scope and was marked `executable_now`, so the gate's last-resort refusal never saw it. It also noted that some earlier negatives only looked safe because `api?` did not match the target `api` — a text artifact, not semantic protection. The requirement was explicit: **propagate the original question scope's non-execution qualification to the children and consume it at the authorization entry; do not re-guess scope from the split text.** | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `如何安装 foo 并重启 api 服务?` / `How do I install foo and restart service api safely?` / `Can you explain how to install foo and restart service api safely?` | only an explanation head was in the governing set, and the clause was partitioned before any qualification could be attached | a clause a QUESTION heads is decided ONCE from that head, before any partition or split, and is `unresolved` whenever it also carries an action (`questionHeadsClause`, `isQuestionScopeNeedingReview`). No child is created, so nothing can lose the qualification, and the gate and preparation consume the same predicate | | a target that matches must not change the verdict | the refusal depended on the child's own target matching | the rule is read from the item's own captured target and action plan, action by action, so a perfect target cannot hide it | | the boundary with the earlier reviews | an investigation imperative (`Check whether …`, `检查是否…`) coordinates two IMPERATIVES, so its second order must keep its authority | the governing set holds question words, auxiliaries, subject-prefixed and verb-fronted interrogatives, and explanation heads — never investigation imperatives. Order-headed mixed messages keep their partition and their execution | Closing the family exposed and fixed four more defects of the same shape, each recorded as its own round: a temporal interrogative read as a condition (`When should I install … ?`), a Chinese question whose subject pronoun stands before the interrogative (`你们如何安装 … 并重启 …`), missing members of the interrogative vocabulary (`谁`, `怎样`, `何时`, `多少`…), a question word that is also a relative pronoun (`Who owns … ?`, which `englishInterrogativeIsMatrix` treated as a subordinate boundary and therefore read as an instruction), and the modal that can stand between the action and the interrogative (`需要安装多少依赖并重启 …?`). Round 18 is the replacement independent set and required no further source change. ### Twelfth repair round after the eleventh independent review The eleventh review showed the last gap of the family: an INVESTIGATION imperative was excluded from the governing set wholesale, so `Check whether it is safe to install foo and restart service api.` split at the coordinator and both actions became authority. The imperative head does not authorize its embedded actions: the root asked to CHECK whether installing and restarting is safe, should happen, or is needed. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Check whether it is safe to install foo and restart service api.` / `Check whether we should install foo and restart service api.` / `检查是否需要安装 foo 并重启 api 服务。` | the governing set held question words, auxiliaries and explanation heads but never an investigation whose complement is OPEN | the complement is now read STRUCTURALLY: the subordinator (`whether`/`if`, 是否/有没有/能否/…) must open the complement AND a modal or an infinitive must stand before the coordinator (`to install`, `we should install`, 需要安装). A complement that merely states a fact (`whether an update exists`) keeps the second order it always had | | the refusal must not depend on the target text | as before, the rule is read from the item's own captured target and action plan, action by action | unchanged, and now exercised on both sides of the boundary | Closing that boundary exposed six more defects of the same family, each repaired and recorded as its own round: an `if`-complement still taken by the condition splitter when a second instruction was coordinated (the head test and the clause test had been conflated), the Chinese modal-bearing subordinators (能否/可否/能不能), the POSTPOSED interrogative (`检查一下[安装 foo 并重启 api 服务]是否安全`), the yes/no interrogatives missing from the question vocabulary (是否/是不是), the Chinese A-不-A class read by the negation path as prohibitions (要不要/该不该/需不需要/ 可不可以/对不对), a question about a single action still read as an instruction, and an ordinary OBJECT list (`检查一下本地插件和皮肤是否有更新`, where 和 joins two nouns) mistaken for a coordination of actions. Round 25 is the replacement independent set and required no further source change. ### Thirteenth repair round after the twelfth independent review The twelfth review closed the last gap of the family: an investigation's complement can have its OWN actor — a deployment script, a migration script, the operations staff — so a DECLARATIVE complement coordinates its own predicates and none of them is root authority. Requiring a modal or an infinitive in the complement was still a pattern, and a plain declarative complement escaped it. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Check whether the deployment scripts install foo and restart service api.` | the complement test only recognised a modal or an infinitive | a declarative complement is now read by the actor signal, per language: Chinese puts the subject BEFORE the subordinator and the action directly after it (确认[运维人员]是否[轮换密钥]…), English puts the subject after the subordinator and the action takes an object (whether [the deployment scripts] [install foo]). The Chinese side additionally requires an ACTION predicate in the complement, so a state question about an object (确认[缓存]是否[有效]并安装依赖) keeps its coordinated order | | `检查运维人员是否安装 foo 并重启 api 服务。` | the same | the same rule; the clause is one undecided obligation and every action in it is refused | | the refusal must not depend on the target text | unchanged: the rule is read from the item's own captured target and action plan, action by action | unchanged | Closing that boundary exposed three more defects of the same family, each repaired and recorded as its own round: an investigation of a DECLARATIVE `that` clause (`Verify that the operator rotates the credentials and redeploys the service`) was authorized, a coordination-free `that` verification was read as an instruction instead of an answerable check, and the Chinese subject rule needed the action predicate (which also required the rotation verb in the work vocabulary). The cross-end ledger records the tightened DSH reading for the English mixed-conjunction case (one undecided obligation) with revision 3. Round 29 is the replacement independent set and required no further source change. ### Fourteenth repair round after the thirteenth independent review The thirteenth review showed the rule still needed evidence the surface cannot supply: the complement closed only when a KNOWN action predicate was recognised (`archive`, `归档` and `bootstrap` were not), and the Chinese subject only counted in one position (before the subordinator, while `是否有人安装…` and `是否由运维人员安装…` put it after). Two lines of attack, one root cause. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Check whether the operators archive the logs and restart service api.` / `检查运维人员是否归档日志并重启 api 服务。` / `… bootstrap the environment …` | closure required a recognised action predicate in the complement | the rule is INVERTED: an investigation's complement GOVERNS by default, and closing it requires positive proof that the question is about a STATE — the subordinator opens the complement, no actor of its own stands on either side of it, no modal or infinitive governs an action there, and the complement names a state (`complementIsProvenStateQuestion`). An unrecognised predicate is no longer evidence of anything | | `确认是否有人安装 foo 并重启 api 服务。` / `检查是否由运维人员安装 foo 并重启 api 服务。` | the subject had to stand before the subordinator | the actor test is position-independent: it looks on BOTH sides of the subordinator, and an actor introduced by 由 counts. A proven state question (`是否有新版本`, `缓存是否有效`, `whether the lock file is current`) keeps its coordinated order, and a coordination-free verification stays answerable | Round 30 is the replacement independent set and required no further source change. ### Fifteenth repair round after the fourteenth independent review The fourteenth review returned the batch with one P1 and four counterexamples: the previous round still asked the surface for evidence it does not carry. Closure had been granted when a state WORD appeared anywhere in the complement and no word from an actor list stood in it, so a state word that is an action's OBJECT or MODIFIER was accepted as proof and the coordinated order was authorized. | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Check whether the technicians install available updates and restart service api.` / `Check whether the technicians archive completed jobs and restart service api.` / `检查程序是否安装更新并重启 api 服务。` / `确认小王是否更新依赖并重启 api 服务。` | closure was proven by the PRESENCE of a state word anywhere in the complement plus the ABSENCE of a word from an actor list — both are absence-of-pattern tests, and `available`/`completed`/`更新`/`依赖` were doing duty as an action's modifier or object | the proof is now the PREDICATE: the complement itself must be a state predication. Chinese places the subject before the subordinator, so the predicate follows it directly and needs no copula (`是否正常`, `是否已经完成`), or a stative verb (有/是/为/存在…) governs a state noun (`是否有新版本`, `是否有更新`). English places the subject after `whether`, so the predicate is the complement's LAST word and must be a state adjective reached through a copula (`is valid`, `is current`), with no action verb before it. The actor list is gone entirely — any actor, in any position, is irrelevant once the predicate is proven | | the refusal must not depend on the target text | unchanged: the rule is read from the item's own captured target and action plan, action by action | unchanged, and now exercised on both sides of the proof | Closing that proof exposed six more defects of the same family, each repaired and recorded as its own round, and each pinned by the set that found it: - the English infinitive test read the `to` inside `up-to-date` as a modal, so a copula state question never closed; - the Chinese modifier strip was lazy and stopped at a possessive, so `有可用的镜像` left `的镜像` and the coordinated order lost its authority; - the state-noun proof could not name ordinary state objects (`镜像`, `租约`), and the Chinese predicate class could not name `为空`/`满`; - the POSTPOSED interrogative was recognized only when the action stood on the far side of the coordinator, so `核对一下重启 api 服务并归档日志是否安全。` was read as an order — and `classifyPositive` now decides the postposed form as undecided directly instead of through a vocabulary-dependent residue test; - a yes/no question that states its own subject and carries no imperative at all (`审计人是否加盖有效的印章并重启 api 服务。`) governed nothing, so the order coordinated beside it was authorized; the yes/no subject form is now an investigation head, and a separator no longer opens a clause inside a governed investigation; - and the message classifier discarded such a governed clause as session talk (`审计员是不是替换凭据并轮换密钥。` produced NO obligation at all, so the order vanished instead of staying visible and undecided); the classifier now consumes the reader's own predicates, exactly as the gate and preparation do. Round 35 is the replacement independent set and required no further source change. ### Sixteenth round: the contracted automatic authorization is withdrawn This round is a **user-approved contract change**, not a repair: the requirement that a coordinated second action inside one governed clause be auto-authorized is replaced by the narrowed contract in [CONTRACT_REVISION_0_6_3.md](CONTRACT_REVISION_0_6_3.md). The fourteen independent reviews that preceded it each showed that the closure it depended on could only be guessed (a state word, an actor list, a work verb, a position window); the new reading keeps the whole clause as ONE undecided obligation and refuses execution until the root writes the action as its own instruction. | Superseded (old contract) | Now (narrowed contract) | Keeper | | --- | --- | --- | | `Check whether an update exists and install the package.` keeps its install as an order | one undecided obligation; the install stays visible in its action plan and authorizes nothing | — | | `检查是否有新版本并且安装这个主题。` / `检查是否有更新并安装新主题。` keep the second order | same | — | | a "proven state question" (`检查服务是否正常并记录变更。`, `确认缓存是否有效并安装依赖。`, `Check whether the cache is valid and install the package.`, `Verify whether the lock file is current and install the package.`) closes its complement and keeps the coordinated order | same-clause coordination is undecided; there is no complement-closure proof any more | — | | an explanation whose complement is closed opens a new predicate (`说明…的作用,然后更新 README。`) | the explanation head governs its SENTENCE, so the same-sentence order is undecided | a separate sentence or clause still authorizes | | a trailing question mark keeps the order in front of it (`What changed, and update the README?`) | punctuation alone never authorizes: one undecided obligation | — | | Kept and re-pinned | Where | | --- | --- | | pure questions are answerable (`检查是否有新版本。`, `Is there any update for the plugin?`, an object-list question) | `tests/domain/v063-narrowed-contract.test.ts` | | a plain explicit instruction with a legal unique target is executable (`重启 api 服务。`) | same | | an instruction that clearly left the governed scope is executable (its own sentence, or its own clause: `检查是否存在更新;安装新主题;`) | same | | an undecided obligation is captured, cannot be closed by an ordinary answer, and cannot take a certificate | same | | the qualification is established once, inherited by partition children, and CONSUMED by the gate and preparation (the gate follows the stored field, not the item text) | same | | target ambiguity, K2 provenance, K3 prepare/execute agreement and K4 legacy fidelity | the existing suites, unchanged | The four counterexamples of the fourteenth review (`Check whether the technicians archive records that are complete and restart service api.`, `Check whether the technicians ensure the cache is valid and restart service api.`, `负责生产环境维护的运维人员是否安装 foo 并重启 api 服务。`, `负责生产环境维护的运维人员是不是安装 foo 并重启 api 服务。`) are pinned in that suite together with the group transformations the review asked for: subject substitution and lengthening, known and unknown verbs, object modifiers, relative clauses and nested complements, Chinese and English, punctuation and conjunction variants, the reading before and after the split and after a replay, and the main action plus every `actionPlan` action. Superseded expectations in the earlier hold-out and review files are relabelled **合同调整** in place rather than deleted; the two expectations that were my own oracle errors are labelled as such. Total independent sets in this batch: fifty set files (thirty-five hold-outs, fourteen reviewer-probe regression files and one self-review file). ### Seventeenth round: qualification became a positive finding The concentrated review of the narrowed-contract batch returned three P1s about how the qualification is ESTABLISHED (delivery and consumption were already accepted): | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `I wonder whether the technicians restart service api.` was `granted` and authorized the restart | `qualificationOfClause` granted whenever no governed pattern matched — a default, not a finding | `granted` now requires POSITIVE evidence: no question content anywhere in the clause's own span and a recognised action. Every other reading is `restricted` (`unproven_scope`), so the reviewer's exact input is denied with `mutation_item_not_executable` | | `Explain this instruction: "Install foo. Restart service api."` re-qualified the second sentence inside the quote and authorized the restart | the qualification was computed AFTER the sentence split, so the quoted parent scope was already lost: the quoted `.` ended the clause | the parent span is established first: a quote (or code span) owns its own punctuation, so a mark inside it never opens a clause of the parent, and a protected clause is indivisible | | `安装 foo 并重启 api 服务是否可行。` split into two `granted` items and authorized the restart | a postposed interrogative was only recognised with a trailing `?`, so the coordinator split the clause | the protection now follows the clause's OWN question content (`clauseAsksOwnQuestion`), with or without the mark; the whole clause is one restricted obligation | Group transformations of the same invariant are pinned in `tests/domain/v063-narrowed-contract.test.ts`: several uncertainty heads (`I wonder`, `I am not sure`, `想问一下`), the quote styles and a longer quote, the postposed form with and without `?`, and the pure-quote control. Two further consequences are recorded rather than hidden: an order whose coordinated object carries an interrogative (`创建文件 /tmp/x 并记录测试是否通过。`) is now protected and undecided (relabelled 合同调整 in place), and an imperative whose verb is outside the action vocabulary is refused rather than guessed, so the purpose-clause control in the invariant suite uses a recognised verb. ### Eighteenth round: a directive, not a mention; the parent scope first A second concentrated review of the narrowed contract returned two more P1s about ESTABLISHMENT, plus the quote-style gap: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `The technicians restart service api every night.` and `日志显示运维人员重启 api 服务。` authorized a restart | `granted` was established by `namesWork`, which only proves the text MENTIONS an action, not that the root ordered it | `granted` now requires a DIRECTIVE (`opensWithDirective`): an imperative in the root's voice — the action opens the clause after any request preface, or the clause opens with a closed-class actor/locative/time/object phrase (`由你…`, `在仓库…`, `按 P0—P4…`, `明天推送`) — and the clause is not a report. A statement is `restricted`, and the disposition is forced to agree with the stored qualification, so a restricted clause can no longer be `executable_now` | | `Explain this instruction: 'Install foo. Restart service api.'` and `解释这条指令:「安装 foo。重启 api 服务。」` re-qualified the restart inside the quote | only straight/curly double quotes were protected | every quote style is protected (`"…"`, `“…”`, a word-boundary `'…'`, `「…」`, `『…』`), and the wait/resume heuristics read the quote-blanked text, so a quoted echo of a confirmation no longer reserves anything | | (follow-on) the bare `应用包 foo 版本 0.6.3 配置档 default。` form lost its grant under the directive rule | an unrecognised operation word cannot be a positive directive finding | the sanctioned route is the root's explicit restatement (`把应用包 foo 版本 0.6.3 配置档 default 明确为 apply`), which is granted and keeps the K3/K4 fixtures meaningful; the retired boundary is recorded rather than hidden | Group transformations are pinned in `tests/domain/v063-narrowed-contract.test.ts`: statements that mention work (four), every quote style (four), directives that keep their authority (including a fronted actor or locative), and the explicit-restatement route. ### Nineteenth round: authorization intent, not wording The third concentrated review closed the last two gaps in the positive findings: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `把重启 api 服务记为待讨论事项。`, `Record restart service api as a hypothetical example.`, `把重启 api 服务明确为禁止操作。` authorized the restart | the restatement route granted on its own wording ("明确为/记为/record … as"), but recording, classifying or forbidding an action is not ordering it | a restatement is granted only when its RESTATED CONTENT is itself an instruction (`restatedContentIsInstruction`: the content opens with an action the reader can identify, or resolves to an operation). A category, a topic or a ban is restricted | | `重启 api 服务是一个危险操作。` was executable | an action at the clause head was taken for an imperative, but here the action is the clause's SUBJECT | `opensWithDirective` now rejects a descriptive matrix predicate (`是/属于/意味着/导致…`, `is/are/means/causes…`), read only from the clause the action head belongs to, so a relative clause or a following clause cannot make an order look descriptive | Controls kept: `把更新插件明确为 apply package demo@2.0.0 profile web` and `把应用包 foo 版本 0.6.3 配置档 default 明确为 apply` remain granted (they restate an INSTRUCTION), `重启 api 服务。` remains executable, and a directive beside a description keeps its own authority while the description adds no action. All are pinned in `tests/domain/v063-narrowed-contract.test.ts`. ### Twentieth round: the restatement is judged, and its authority is bound The fourth concentrated review showed the restatement route still took a shortcut ("the content mentions an action") and handed the whole item one reusable grant: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `把重启 api 服务记为需要讨论的重启操作。`, `把重启 api 服务明确为解释重启流程。`, `Record restart service api as a description of how technicians restart service api.` authorized the restart | the content test accepted any action word anywhere in the restated span | the restated span (`restatedContentOf`) must now pass the SAME judgement a clause passes (`qualificationOfClause`) or be a canonical OPERATION spec whose HEAD is the operation (`restatedContentIsOperation`, which also rejects content that asks or describes). Prose that mentions an operation stays restricted | | `把重启 api 服务明确为检查日志。` authorized the restart even though the restated directive is `检查日志` | the qualification was a blanket `granted` for the item, whose action was still read from the pre-restatement text | capture now takes `semanticAction` and `actionPlan` from the RESTATED span (the target identity still comes from the whole clause, because a restatement clarifies the obligation it names). The item's identity is the inspection, so a restart mutation is denied — the action named before the restatement is never authorized | Pinned in `tests/domain/v063-narrowed-contract.test.ts` with the four reviewer inputs, the identity-binding assertion and the sanctioned-restatement controls. ### Twenty-first round: the restatement binds its target too The fifth concentrated review found one binding defect in that repair: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `把重启 api 服务明确为重启 worker 服务。` and `Rebind restart service api as restart service worker.` still captured `service_id=api` and authorized the OLD target | capture took the ACTION from the restated span but the TARGET from the whole clause, so the pre-restatement target won | capture now binds action, `actionPlan` and TARGET together to the restated span: a field the new span names wins, and only the fields it OMITS are inherited from the clarified obligation — and only when that older selection is unambiguous. The old `api` target is denied, the new `worker` target is what the item authorizes, and `把应用包 foo 版本 0.6.3 配置档 default 明确为 apply` still inherits package/version/profile (pinned in `tests/domain/v063-narrowed-contract.test.ts`) | ### Twenty-second round: per-span uniqueness, per-field merge The sixth concentrated review found the merge rule was still conditional and the uniqueness check was missing on both spans: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `把重启 api 服务或 worker 服务明确为 restart。` authorized `api` and `把重启 api 服务明确为重启 worker 服务或 cache 服务。` authorized `worker`, both `resolved` | neither span was checked for a UNIQUE selection: the extractor silently took the first candidate | `identityCandidates` counts the distinct candidates a span names for the action's identity field; two or more on EITHER span is reported as `requested_target_field_ambiguous` with `clarification_required`, so the gate refuses every candidate | | `把应用包 foo 版本 0.6.3 配置档 default 明确为 apply package foo version 0.6.4。` lost `profile=default` | inheritance ran only when the restated span named NOTHING | the merge is now PER FIELD: the restated span's own fields win, every field it OMITS is inherited from the clarified obligation while that older selection is unique. The main action and the `actionPlan` share the rule | ### Twenty-third round: the enumeration shares the extractor's grammar The seventh concentrated review showed the candidate enumeration and the target extractor used different syntax: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Rebind restart service api or worker as restart.` authorized `api`, and `Rebind restart service api as restart service worker or cache.` authorized `worker`, both `resolved` | `identityCandidates` matched "name before `service`", while the extractor accepts `service `; it read `restart` as a candidate, missed `api or worker`, and the shared-label list went unchecked | the candidate set is now enumerated with the EXTRACTOR's own parsers (`labeledTokens` with a coordinated continuation list, `actionObjectTokens`, and the noun-suffix scan), each normalized by the same helper, and ambiguity is judged within one surface form so mixing forms cannot invent a pair. `buildActionPlan` calls the same check per entry (`restatedSpanAmbiguous`), so the main action and the plan share one rule | ### Twenty-fourth round: uniqueness covers every identity field The eighth concentrated review showed the audit stopped at the object name: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `把应用包 foo 版本 0.6.3 配置档 default 明确为 apply package foo version 0.6.4 or 0.6.5 profile default。` authorized `version=0.6.4`, and `… profile web or prod。` authorized `profile=web` | only service and package names were enumerated; version and profile were still read as single values | the enumeration now follows the action's whole identity contract (`identityFieldLabels`: service, package, version, profile, registry, repository, branch, remote, refspec), each field read with the extractor's labels and normalization, so a coordinated list after any label is ambiguous. The merge reads the CLARIFIED span rather than the whole clause, because the whole clause necessarily names both the old and the new value | Boundary recorded: the audit applies to the restatement merge, where two spans must be reconciled; a plain clause keeps the extractor's single reading (the fixture families that depend on it are unchanged). ### Twenty-fifth round: the uniqueness result is consumed by every entry The ninth concentrated review showed the ordinary capture bypassed the audit: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `Restart service api or worker.` authorized `api` and `提交仓库 /repo-a 分支 main 或 release。` authorized `main`, both `resolved` | the plain branch took the extractor's single value; the audit ran only for a restatement | the ordinary capture, the restatement merge and every `actionPlan` entry now consume the same uniqueness result. The enumeration and the extractor share ONE label/verb definition (`IDENTITY_LABELS`) with boundaries, and a label's value must be separated from it, so a path (`/repo-a`), a compound token (`synthetic-plugin`) or a sub-path (`repo/sub`) is not read as a second candidate; earlier reason codes (an ambiguous repository, a missing field) are preserved | The 33 fixture rejections from the first attempt were traced to exactly those parse false positives and fixed there, not by skipping the ordinary entry; the positive controls (`重启 api 服务。`, `提交仓库 /repo-a 分支 main。`, the per-field inheritance and the install fixture form) all still resolve. ### Twenty-sixth round: identity values keep their case The tenth concentrated review found the enumerators lowercased every candidate: | Reviewer finding | Root cause | Repair | | --- | --- | --- | | `提交仓库 /repo-a 分支 main 或 Main。` authorized `main` and `Push repository /repo-a remote origin refspec main:main or main:Main.` authorized `main:main`, both `resolved` | `labeledTokens` and `fieldCandidates` applied a blanket `toLowerCase()`, so two distinct git identities collapsed into one candidate | candidates keep their value; each field is normalized by its OWN rule (`fieldNormalizer`): packages through the package-spec parser, registries through their canonical form, and every git identity — repository, branch, remote, refspec — case-sensitively. A repeated EXACT value is still one candidate (`main 或 main` stays `resolved`), and the positive controls are unchanged | ### Twenty-seventh round: package-spec tuples and the shared label set (closing checklist) The closing review fixed the last two items on its checklist: | Blocker | Root cause | Repair | | --- | --- | --- | | `foo@1.0.0 or foo@2.0.0` collapsed into one package name and authorized `1.0.0` (install, restated apply and publish), and `foo@1.0.0 version 2.0.0` was not refused as a conflict | the package candidate was normalized to its NAME only, and a version inside a spec never reached the version field | a package candidate is the whole SPEC (`id@version`), and the spec's version joins the labelled versions as ONE candidate set for the version field, so a same-name/different-version list and a spec/label conflict are both ambiguous, while an exactly repeated version stays unique | | `推送仓库 /repo-a 远端 origin 引用规范 main:main 或 main:release。` authorized `main:main` | the audit's label set lacked the Chinese `引用规范` that the extractor already accepted | the audit and the extractor now share the whole field-label set (repository, branch, remote, refspec incl. `引用规范`, package, artifact, version, profile, registry, service, plus the guarded verb forms), so no label can exist on one side only | The closing review's other fixed checks — qualification propagation, prohibition/wait/human actions, ordinary instructions, explicit follow-up authorization, the upgrade and certificate suites — reported no new blockers, and the repository-wide suite is green. ### T06 evidence: a real 0.6.2 projection, upgraded by the production entry The earlier T06 evidence derived a log with the CURRENT reader and then assigned the old shape onto the result by hand, and its "history is preserved" case compared two snapshots with no operation between them. That is withdrawn; what replaces it: - `scripts/record_legacy_upgrade_fixture.mjs` EXECUTES the committed 0.6.2 build (`deriveProjection` from the baseline `dist/`, extracted with `git archive 63326f2 dist`) and stores the records that build emitted, with the baseline commit and a SHA-256 for every module file. Re-running it is deterministic; the recorded file is `tests/fixtures/upgrade/legacy-0.6.2.json`. The recorder covers BOTH real protocols, because the earlier release behaves differently on each: a session with no boundary notice leaves the mixed record `pending`, and a session that announced the v5 boundary applies its delivery pass and writes the record CLOSED — `status: answered` with a delivered `answeredBy` (turn, response sequence and response digest). The v5 notice is written out in the recorder so it never imports the current source; the sixth review's own v5 comparison is what showed the first recording had omitted it. - `tests/domain/v063-legacy-upgrade.test.ts` loads those records as the persisted state of an earlier session and runs `applyUpgradeEligibility` — the exported entry `deriveProjection` itself calls — BETWEEN the two snapshots, so the diff shows exactly what the upgrade changed: the eligibility fact and nothing else. It asserts the recorded ID, revision, `normalizedText`, `textSha256`, spans, status, action, target and `answeredBy` are untouched (including the CLOSED v5 records, whose delivered answer does not rescue them), that no certificate, checkpoint or effect appears, that the flag blocks the certificate and the Goal with `legacy_record_needs_review`, that a second upgrade keeps the original reason and revision, and that the pure-question control is never flagged. The unit attribution the v5 recording carries is part of the restored state, so the record is in the eligibility scope exactly as it is in a live session. - What this file does NOT claim: it drives the eligibility entry over the persisted records directly. Native host recovery — restoring a real profile and replaying a real session through the runtime — remains unmeasured below. - `tests/domain/v063-upgrade-chain.test.ts` keeps only the durable-log cases and labels every constructed record as predicate-level. ### Evidence added in the repair round - `tests/domain/v063-review-regressions.test.ts` — the reviewer's nine failing probes plus two positive controls, with the reviewer's own contract expectations. - `tests/domain/v063-holdout-round2.test.ts` — 28 assertions across 19 cases in wording families absent from every earlier set: 呢/吧 as particles, 并且/以及/ 而后/then chains, an enumerated repository, a subdirectory that must not collapse onto its parent, inheritance across an unrelated clause, prepare/action agreement on stale revision, non-pending status and a cross-family action, and eligibility scope across unit boundaries. - `tests/domain/v063-legacy-upgrade.test.ts` — the T06 evidence: records emitted by executing the committed 0.6.2 build, upgraded by the production entry with the before/after comparison bracketing that call. The construction and its limits are described in the T06 section above. - `scripts/record_legacy_upgrade_fixture.mjs` + `tests/fixtures/upgrade/legacy-0.6.2.json` — the recorder and the recording it produces (baseline commit, every module hash, the records as the earlier release left them). - `tests/domain/v063-upgrade-chain.test.ts` — the durable-log half of T06: the v5 boundary, root message, completed turn and delivery, idempotent re-derivation, the v4 (pre-v5) scope, and the predicate-level shapes a recorder cannot emit, each labelled as such. - `tests/tools/v063-host-lifecycle.test.ts` — T07 over ONE durable log: the registered tools materialized through the real `ToolRuntime`, their call/result round-trips APPENDED to the log in the persisted wire shape, a restart-style re-derivation, a `compaction/summary`, a second re-derivation, the Recovery packet and digest from the post-compaction projection, the production Stop decision, and the same refusals after the replay. - `tests/domain/v063-review2-regressions.test.ts` — the second review's five failing probes plus positive controls (a conditional order, genuine questions, a prohibition on another repository, a prohibition of another action, an unset field still inheriting). - `tests/domain/v063-holdout-round3.test.ts` — the round-3 independent set (28 assertions), now marked in its own header as regression coverage. - `tests/domain/v063-review3-regressions.test.ts` — the third review's four failing probes plus the positive controls that keep the repairs from over-reaching (a genuine question, an order with a preface, a complete selector). - `tests/domain/v063-holdout-round4.test.ts` — the round-4 set (19 cases), now marked in its own header as regression coverage. - `tests/domain/v063-review4-regressions.test.ts` — the fourth review's four failing probes plus the positive controls that keep the structural rule from over-reaching (a question with no subordinate span before it, a real conditional order, an obligation that DID name the branch). - `tests/domain/v063-holdout-round5.test.ts` — the round-5 set, now marked in its own header as regression coverage; it found the classifier defect. - `tests/domain/v063-repair5-regressions.test.ts` — the fifth repair round's three self-review counterexamples (the English sentence end, the coordinated ordering fragment, the two-repository clause) with the abbreviation, version, single-repository and preface controls. - `tests/domain/v063-holdout-round6.test.ts` — the round-6 set (16 cases), now regression coverage. - `tests/domain/v063-review5-regressions.test.ts` — the fifth review's four failing assertions and three defects plus the probes that already passed. - `tests/domain/v063-holdout-round7.test.ts` — the round-7 set, now regression coverage; it found the English `then` preface/boundary conflict. - `tests/domain/v063-holdout-round8.test.ts` — the round-8 set (16 cases), now regression coverage. - `tests/domain/v063-review6-regressions.test.ts` — the sixth review's three defects plus the controls that keep the repairs from over-reaching (a genuine question, an abbreviation that continues its sentence, a claimed field value, a file argument in another clause, two spellings of one repository). - `tests/domain/v063-holdout-round9.test.ts` — the round-9 set (19 cases), now regression coverage. - `tests/domain/v063-review7-regressions.test.ts` — the seventh review's three defects, with the explanation-versus-order pair in both languages and the per-field inheritance controls. - `tests/domain/v063-holdout-round10.test.ts` — the round-10 set (17 cases), now regression coverage. - `tests/domain/v063-review8-regressions.test.ts` — the eighth review's three counterexamples plus the tightened `and then` reading, the separate-instruction positive controls and the closed-complement controls. - `tests/domain/v063-holdout-round11.test.ts` — the round-11 set (18 cases), now regression coverage. - `tests/domain/v063-review9-regressions.test.ts` — the ninth review's two counterexamples plus the separate-instruction positives and the two scope controls. - `tests/domain/v063-holdout-round12.test.ts` — the round-12 set (16 cases), now regression coverage. - `tests/domain/v063-review10-regressions.test.ts` — the tenth review's question scopes, the target-independent refusal and the imperative contrast. - `tests/domain/v063-holdout-round13.test.ts` … `-round17.test.ts` — the five sets that closed the question-scope family, each recording the defect it found and any oracle correction of its own. - `tests/domain/v063-holdout-round18.test.ts` — the round-18 set (11 cases), now regression coverage. - `tests/domain/v063-review11-regressions.test.ts` — the eleventh review's investigation-complement counterexamples and the contrasts that keep the rule honest. - `tests/domain/v063-holdout-round19.test.ts` … `-round24.test.ts` — the six sets that closed the investigation boundary, each recording the defect it found and any oracle correction of its own. - `tests/domain/v063-holdout-round25.test.ts` — the round-25 set (12 cases), now regression coverage. - `tests/domain/v063-review12-regressions.test.ts` — the twelfth review's actor-complement counterexamples and their contrasts. - `tests/domain/v063-holdout-round26.test.ts` … `-round28.test.ts` — the three sets that closed the actor boundary, each recording the defect it found and any oracle correction of its own. - `tests/domain/v063-holdout-round29.test.ts` — the round-29 set (12 cases), now regression coverage. - `tests/domain/v063-review13-regressions.test.ts` — the thirteenth review's five counterexamples, the proven-state contrasts and the separate instruction. - `tests/domain/v063-holdout-round30.test.ts` — the round-30 set (17 cases), now regression coverage. - `tests/domain/v063-review14-regressions.test.ts` — the fourteenth review's four counterexamples, the siblings of the same root cause, the proven-state contrasts, the answerable lane and the plain order. - `tests/domain/v063-holdout-round31.test.ts` … `-round34.test.ts` — the sets that closed the predicate proof, each marked in its own header as regression coverage with the defects it found and its own oracle corrections. - `tests/domain/v063-holdout-round35.test.ts` — the current independent set (20 cases), the only set in this batch that required no source change. - `tests/fixtures/cross-end/codex-0.13.9.shape.json` — a NEW Codex recording made in this batch by `scripts/record_cross_end_oracle.py --mode shape`, which executed the installed module's `_reply_only_request_shape` plus `clause_metadata`/`verification_contract` on twelve shared inputs and bound the result to the module SHA-256 `4b895aad…`. T08 reads it, so the comparison is measured rather than inferred: on this batch Codex's reply-only judge returns false for all twelve, which is why the mixed-request and pure-question families are recorded `not-aligned` — DSH closes a range Codex would not. ### What remains unmeasured after this round - Native Codex host turns: the shape recording executes entry points, not a session. No native Codex acceptance is claimed. - Cross-repository follow-up reference and prepare/execute comparison on the Codex side: no equivalent entry point exists; recorded as not measured and not applicable respectively. - Candidate CI, the frozen tgz, macOS/Windows exact-artifact acceptance, real-model behaviour, plugin installation, DSH restart, tag, npm publication and GitHub Release. ### Known boundaries found while testing, and deliberately not repaired - A publish target's registry identity is the canonical base the repository's own canonicalizer records (`https://registry.npmjs.org/`). A caller that supplies the root's UNNORMALIZED spelling therefore does not match the obligation. Both lanes agree — so T05 holds — and preparation reports the recorded identity, so the workflow proceeds from the target it returns. Loosening the comparison would change an identity predicate the gate shares, which this batch does not do. - An English investigation behind a preface USED TO keep the pre-existing `ACCEPTANCE_LEAD` lane; the fifth review's follow-up and round 7 showed that made the English reading differ from the Chinese one, so it is repaired: `Then check whether the disk is full.` is an information request, like `然后检查是否有新版本。`. - An investigation's complement GOVERNS by default. It closes only on positive proof that the complement PREDICATES a state — a Chinese state adjective, a stative verb governing a state noun, or an English state adjective reached through a copula, with no modal, infinitive or action verb in the way. A state WORD that is an action's object or modifier (`install available updates`, `安装更新`, `更新依赖`) proves nothing, an unrecognised predicate proves nothing, an actor anywhere proves nothing, and the actor list the earlier rounds leaned on is gone: proof is positive, never the absence of a protection pattern. - The proof vocabulary is itself a positive list, so a state the lists cannot name (`有可用的镜像` before the round that added it) is refused rather than authorized. The asymmetry is deliberate: an unprovable complement leaves its coordinated part inside the question, where it is never authority. - A clause a QUESTION heads is one scope: it is never partitioned into an executable child, and when it carries an action it is `unresolved`. An investigation IMPERATIVE is different — `Check whether … and install …` coordinates two orders — so its second order keeps its own authority — UNLESS its complement is open (`Check whether it is safe TO INSTALL … AND restart …`, `检查是否需要安装…并重启…`), in which case the actions are what the root asked to have checked and none of them authorizes. - Only an EXPLICITLY executable reading authorizes. An explanation whose sentence mentions an action is refused action by action, whatever its target says, and a question is refused for the same reason; authority requires an explicitly separate instruction. An `unresolved` clause that is not an explanation's scope keeps the behaviour the gate always had, so an unrecognised instruction form is still judged on its action and target rather than refused on its disposition alone. - Inside the sentence an explanation heads, a coordinator belongs to the explanation while its complement is open, so `Explain how to install X and then restart Y` creates no authority even though a reader might take it as an instruction. The safe direction is deliberate: an instruction must be its own sentence (`Explain the deploy. Then restart service api.`). - A period ends a sentence whenever whitespace and further text follow it. The only non-boundaries are a period with no space after it (a decimal, a version number, a file name) and an abbreviation from the closed English class (`e.g.`, `i.e.`, `cf.`, `etc.`, honorifics, dotted initialisms). An ordinary lower-case word after a period therefore STARTS a new sentence, which is the fail-closed direction. - A repository candidate must be spelled like one: a path or a Latin name that is not a value another field already claims. An extension is NOT evidence of identity — a repository may be called `/repo-a.js` — so a clause that offers several candidates records the ambiguity instead of choosing, and two spellings of the SAME path (`/repo-a` and `/repo-a/`) are one candidate. A file argument in another clause is not a candidate because the join rule requires a coordinator and nothing else between them. - A period that belongs to an abbreviation keeps the run whole only when what follows continues the sentence; a question after it opens a new sentence. (An earlier revision of this record claimed the 0.6.2 build never wrote `answered` records for these shapes. That was wrong — it followed from a recording made without the v5 boundary notice — and the recorded v5 fixture now carries the closed records instead.) ### Gates run | Gate | Command | Result | | --- | --- | --- | | typecheck | `pnpm run typecheck` | exit 0 | | lint | `pnpm run lint` | exit 0; 27 warnings, 0 errors — the same count as the 0.6.2 baseline, so the batch adds no new warning | | unit/deterministic tests | `pnpm test` | 127 files, 2140 passed, 1 skipped | | release packer | `pnpm run test:release-pack` | 2/2 passed | | download stats | `pnpm run test:stats` | 10/10 passed | | build | `pnpm run build` | exit 0; 6 files, 1105.16 kB; two consecutive builds are byte-identical | | pack inventory | `pnpm run pack:check` | name `dsh-completion-guard`, version `0.6.3`, 35 files (the contract revision note is packaged) | | documentation audit | `python3 scripts/audit_repository_documentation.py .` | 27 markdown files, 0 errors, 0 warnings | | documentation audit test | `python3 -m pytest tests -q` | 84 passed | | whitespace | `git diff --check` | clean | | reviewer probes (supplementary evidence, not the acceptance) | the fourteen counterexample files under `/tmp/dsh-063-review`, `/private/tmp/dsh-063-independent-review`, `/private/tmp/dsh-063-review3`, `/private/tmp/dsh-063-review4`, `/tmp/dsh-063-review-20260917`, `/tmp/dsh-063-review-20260917-next`, `/tmp/dsh-063-review-20260917-r8`, `/tmp/dsh-063-review-20260917-r9`, `/tmp/dsh-063-review-20260917-r10`, `/tmp/dsh-063-review-20260917-r11`, `/tmp/dsh-063-review-20260917-r12`, `/tmp/dsh-063-review-20260917-r13`, `/tmp/dsh-063-review-20260917-r14` and `/tmp/dsh-063-review-20260917-r15`, run from `tests/domain/` | all fourteen files passed against the final source; every copy was removed afterwards | | recorded 0.6.2 fixture | `node scripts/record_legacy_upgrade_fixture.mjs --module-dir --commit 63326f2… --output tests/fixtures/upgrade/legacy-0.6.2.json` | exit 0; re-running reproduces the committed fixture byte for byte | `NODE_PATH` was unset for the test runs. The DSH harness injects a `NODE_PATH` that points at the runtime store, and `tests/domain/host-target-preflight.test.ts` resolves a visible parent package through it and reports `target_profile_module_shadow` for eight of its 46 cases. With `NODE_PATH` unset — the environment every CI lane and the 0.6.2 recheck used — the file passes 46/46. This is an environment artifact of the developing session, not a repository defect, and it is recorded here rather than worked around in the source. ### T01–T08 coverage - T01/T02 — `tests/domain/v063-core-alignment.test.ts` (paraphrase, punctuation, word order, negation, condition, quoted command, unknown tail) and `tests/domain/v063-holdout.test.ts`. - T03 — delivery closes only its own information range, including the zero-tool final and the partial-completion counterexamples. - T04 — target provenance, unique inheritance, two-candidate ambiguity, explicit-current-repository resolution, and a model-supplied target. - T05 — one compatibility judgement for prepare and the mutation gate, with the same-snapshot equality and the revision/condition change counterexamples. - T06 — the pre-terminal eligibility check, its idempotence across a reload, and the certificate/Goal block. - T07 — `tests/tools/v063-host-materialization.test.ts` drives the registered tools through the real `ToolRuntime` output contract and reads the production Stop decision; the producer-reference and boundary negatives stay refused. - T08 — `tests/domain/v063-cross-end-projection.test.ts` reads the recorded Codex facts and the case-level ledger `tests/fixtures/cross-end/core_alignment_0_6_3.json`. The mixed-request and unnamed-repository families are `not-applicable`/`not-aligned`; nothing is claimed as aligned without a measured Codex counterpart. Not executed, and not implied by the above: candidate CI, the frozen tgz, macOS/Windows exact-artifact acceptance, real-model behaviour, plugin installation, DSH restart, tag, npm publication and GitHub Release. ## 0.6.2 lifecycle and documentation consolidation (2026-09-16) The coordinator added a third Codex recording (`--mode lifecycle`) using product-owned prompt ingestion and disposable synthetic ledgers. Its four cases measure silent pending, missing-proof correction, user wait and one opaque mixed-result command. DSH receives the same recorded prompt/command bytes; its checkpoints remain incomplete and its boundary decisions preserve pending/wait. See [the result contract](CROSS_END_RESULT_CONTRACT.md). The earlier in-memory-only probes remain historical evidence, including their unexecuted ledger-dependent branch. That limitation is resolved by the later lifecycle recording, not by rewriting the old result. T06 now has an owned held-handle OS fixture behind `native_acceptance.py --t06`; the source probe passed on macOS. Windows and exact-artifact results are not yet claimed. Duplicate repair notes and the superseded 0.6.1 plan were removed. Historical references point to their exact Git snapshot; current semantic and release limitations remain in maintained documents. ## 0.6.2 release-preparation recheck (2026-09-16) The coordinating Codex run on macOS (Node 25.1.0, Python 3.12.2) independently passed the full local Vitest suite: 70 files, 1092 tests passed, one skipped. The previously reported eight `host-target-preflight` failures did not reproduce; that file separately passed all 46 tests. This establishes the result in this execution environment, not the root cause of the earlier DSH failures or Node 22/24 and Windows portability. Typecheck passed; lint reported 27 warnings and no errors. Release-pack tests passed 2/2; stats tests passed 10/10; build and pack dry-run passed. Rebuilt `dist` matched the staged runtime bytes. Documentation audit reported zero errors/warnings, 81 Python tests passed, and both Codex recording checks were current. The recorder tests are now included in candidate CI's static job. The uncommitted candidate is not frozen, installed, or published. D062-04 was partial at that checkpoint; subsequent disposable-ledger probes above close its bounded lifecycle gap. T06/exact-artifact native acceptance and candidate CI remain pending. See the [0.6.2 release plan](RELEASE_PLAN_0_6_2.md) for the remaining ordered gates. Each section names its evidence boundary. Deterministic checks, isolated DSH_HOME composition, native-platform lifecycle runs, model sessions, CI, and public release readback are separate claims; none substitutes for another. ### Third repair round after targeted review (2026-09-14) Two findings from the targeted review of `fcc3813`: an adoption could be ratified by a certificate minted AFTER it (the revision comparison was satisfiable by a later log entry), and a mistyped `/context-guard release adopt` payload marked the persisted release state damaged, so one typo blocked every later publication. Both are fixed: the adoption freezes the closure certificate's identity at its own watermark, and only unreadable persisted records damage the release state. Regression cases live in `tests/domain/v060-release-migration.test.ts` (31 cases). ### Second repair round after follow-up review (2026-09-14) The follow-up review of the repair commit found five remaining wiring defects (pre-v5 delivery retro-closure, the unobserved registry and unreachable ref, the closure revision that made publishing depend on publishing, read-only status poisoning the release state, and a proof that was checked when signed but not when replayed). All five are fixed and the reviewer's counterexamples remain as regression cases. Added deterministic evidence for this round: - `tests/tools/v060-release-chain.test.ts` — the production release chain: a real root publish instruction, a real certified preparation closure, a real contract adoption, a real tgz and resolution through the registered evidence tool, and the registered action tool consulting the runtime's own gate (reserve, execute once, stay in flight, reconcile through the registered recovery tool, refuse the replay). Also the revoked-but-in-flight restart recovery and a mismatching readback. - `tests/domain/v060-proof-production-chain.test.ts` — a proof-bound certificate signed, persisted and replayed through the real tool, plus the missing-proof, tampered-proof and Goal-consumption negatives. ### Repair round after concentrated review (2026-09-14) The `0dce898` candidate failed a concentrated review: fifteen counterexamples, all reproducing real defects. The findings were repaired as eight families (F01–F08) and the reviewer's suite is now a permanent regression gate (`tests/domain/review-counterexamples.test.ts`, 15 cases, all passing). One reviewer expectation was deliberately changed and is marked in place. Added deterministic evidence for this round: - `tests/domain/v060-proof-production-chain.test.ts` — the C09 proof entry through real DSH tool registration: a real read fact, the real binder, the real checkpoint tool, a persisted certificate and a full log replay. - `tests/domain/v060-release-migration.test.ts` — rewritten (26 cases) around a REAL certified closure, with every candidate identity refusal, the reconciliation of an unconfirmed attempt by a trusted readback, damaged-state scoping, revocation, and the migration rule sets. - The v2 fixture grew to 27 cases: S09 now carries a real read fact and asserts the binder's outcome, the S11 probes are evaluated at their own point in the log with real closure certificates, and the coverage table records each operation's honest attribution. The two scope facts decided on 2026-09-14 (candidate-only v2 plus open parity; `npm_publish` as the only protectable release surface, with the missing routes attributed as an approved scope reduction) are recorded in [SEMANTIC_COMPATIBILITY.md](SEMANTIC_COMPATIBILITY.md). ## 0.6.0 publication facts (verified 2026-09-15) The "not established" list of the 2026-09-14 source-candidate section below was written before publication and is superseded by the following readback, recorded the same day it was performed. The historical section text is kept as written. - The annotated tag `v0.6.0` resolves to commit `cc5cbc6d408664172d9383de7c83c55ec6dfd602`, and the GitHub Release `v0.6.0` targets the same commit. - The npm registry carries `dsh-completion-guard@0.6.0`; its registry `dist.integrity` is `sha512-xIV5wAmDhdXGr7xlpOAceaVJUYgkxmqO2/7zuIat95T+QH2k0e09oYRGokUpH18NNNC2p6UVqnjO8fcpp1V58Q==` and `dist.shasum` is `4573474855d2f36c011cb64cc7139ec2bf68d613`. - The byte-level comparison of the published artifact against the release receipt's frozen tgz belongs to the release archive and is not re-derived here. ## 0.6.2 source candidate (2026-09-16) Development baseline: `d11009d8f755ecee7d288cff18250c7b372cfd2a`. The capability, process-layer and recovery fixes are covered by `v062-capability-and-layers.test.ts`; the seven review findings were repaired, including frozen-outcome preservation and checkpoint source/conflict output. The later lifecycle recording and source probe are described above. Earlier DSH runs reported eight `host-target-preflight` failures. The Codex full run passed those cases; the environment-dependent cause is unconfirmed. Do not carry the earlier speculative realpath/module-shadow diagnosis forward as an established production defect. Current source tests do not establish candidate CI, native Web/Headless or a frozen tgz. The ordered remaining gates are in the release plan. Historical release bytes and their native results below do not certify this candidate. ## 0.6.1 source candidate (2026-09-15) This section records a **source and deterministic** claim only. It is not a release, not an artifact acceptance, and not a native-platform result; each of those is a separate fact that, when established, gets its own section and its own frozen-artifact identity. Established at this point: - The five W060-01–05 repairs from the Windows 0.6.0 session review ([historical plan](https://github.com/GreenLv/dsh-completion-guard/blob/d11009d8f755ecee7d288cff18250c7b372cfd2a/docs/WINDOWS_0_6_0_REPAIR_PLAN.md)) are implemented and covered by dedicated positive/negative suites: `tests/domain/v061-attachment-interpretation.test.ts`, `tests/domain/v061-conservative-interpretation.test.ts`, `tests/domain/v061-discovery-pagination.test.ts`, `tests/domain/v061-evidence-contract.test.ts`, and `tests/domain/v061-ordinary-observation.test.ts`, alongside the full pre-existing regression matrix. The suites include the review rounds' counterexamples: an answer that states the images were not viewed closes no asset; a replayed log without interpretation records keeps assets `pending`; an OLD asset closes through the turn that explicitly re-interpreted it, and the closing turn must be the interpreting turn; a receipt contradicting the call, item, revision, or asset identity is an `interpretation_receipt_mismatch` integrity violation that records nothing, and so is a transplanted result whose call turn and receipt turn disagree, while a valid-identity receipt with a MISSING turn association simply records nothing and keeps the log valid; a real order containing a relative wh-clause ("Create a file where logs are stored") stays executable; unknown requests stay unresolved regardless of wording, under the final interpretation contract: a grammatical declarative can express a task requirement, so no surface shape proves a clause is closable information — the review rounds' holdouts all stay unresolved and unanswered ("Please sanitize inputs that are untrusted", "请处理被遮挡的面板", "Have these inputs sanitized", "处理被遮挡的面板", "处理没有标签的输入", "Sanitize inputs I have received", the long-attributive "处理被异常宽大…遮挡的面板", "I need you to sanitize these inputs", "Our requirement is to sanitize all inputs", "避免面板被遮挡", "没有标签的输入也要处理", "Sanitize all inputs", and declarative contexts like "设置面板被遮挡。请修复登录页。" whose undecidable context stays pending) — while clauses with POSITIVE information grounds (questions, quoted actions, past/aspect reports as the clause's entire predicate such as "我刚才已经推送过了") keep closing through their own turn's answer, a bare whole-message acknowledgment ("当然。") is session talk that is never captured, and an unresolved clause — including explanation/investigation openers and every holdout above — is closable only in parts: the implemented structured-interpretation pathway (`context_guard_interpret`) requires a span partition (information_spans/unknown_spans) validated against the full input spans, and only the information sub-item closes with that turn's answer while every unknown/undeclared sub-span stays pending ("Explain the issue, sanitize all inputs" keeps the sanitize demand open); a verbatim concrete clarification supersedes an unresolved clause; a successful command whose text merely contains quoted action text flags nothing; and a git effect that already holds (push with the remote at the local head, pull/fetch already up to date) is refused with `effect_already_applied` before any command runs. The test host lifecycle now disposes every started loop host, which removes the intermittent FileHandle garbage-collection failure of the full-suite run. - Compatibility boundaries hold: the asset obligation's contract text is byte-identical to 0.6.0 (replayed contract digests do not move, recorded certificates still verify); asset closure requires a per-asset `context_guard_interpret` record plus the turn's trusted delivery, so logs written by 0.6.0 replay with their assets `pending` — the same status 0.6.0 produced; new fields, the interpretation tool, and reason codes are additive; legacy sessions keep their historical reading; and no digest domain was reused. - The full local deterministic matrix (this section's top lists it in [historical notes](https://github.com/GreenLv/dsh-completion-guard/blob/d11009d8f755ecee7d288cff18250c7b372cfd2a/docs/NEXT_VERSION_REPAIR_NOTES.md) and the repository instructions) and byte-identical repeated builds of the candidate dist. The dist differs from the 0.6.0 base until the 0.6.1 candidate is committed. Not established here, and deliberately not claimed: - No 0.6.1 tgz has been packed with `scripts/release-pack.mjs`, so there is no frozen artifact digest for 0.6.1. - No native macOS or Windows Web/Headless run has been executed against a 0.6.1 artifact; the v5-session, attachment, discovery, and unattributed-observation behaviours above have deterministic coverage only. - No tag, npm publication, GitHub Release, or consumer installation exists for 0.6.1 in this state. ## 0.6.0 source candidate (2026-09-14) This section records a **source and deterministic** claim only. It is not a release, not an artifact acceptance, and not a native-platform result. Established at this point: - The full local deterministic matrix on the frozen candidate: `typecheck`, `lint`, `vitest`, `test:release-pack`, `test:stats`, `build`, `pack:check`, the repository documentation audit and its unit test, `git diff --check`, and a byte-clean `git diff --exit-code -- dist` after the build. - The v2 conformance candidate (25 cases, S01–S12) and its independence, proof-v2, release/migration, strict-policy, unit-closure, and release-gate suites, all through production entry points. - `pack:check` reports the 0.6.0 payload inventory (28 files) with `dist/`, both changelogs, `docs/`, and `manifests/` included. Not established here, and deliberately not claimed: - No tgz has been packed with `scripts/release-pack.mjs`, so there is no frozen artifact digest and no repeated-pack byte identity for 0.6.0. - No native macOS or Windows Web/Headless run has been executed against a 0.6.0 artifact. The 0.5.3 native annexes are bound to the 0.5.3 bytes and do not transfer. - No tag, npm publication, GitHub Release, or consumer installation exists for 0.6.0. - Cross-language parity for the v2 input family is not established, because the upstream has not landed a canonical v2 fixture. The next stages are P5 (one canonical pack, two platforms, same tgz) and P6 (publication and consumption), each requiring its own authorization. ## 0.5.3 published release (2026-09-14) [Version 0.5.3](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.5.3) is published on GitHub and [npm](https://www.npmjs.com/package/dsh-completion-guard/v/0.5.3). The annotated tag and npm `gitHead` identify commit `a7baccdfbc538aa071941ebb36aa41fa9bf6e709`. The frozen 308178-byte, 28-file tgz has SHA-256 `972269489ba65092bf13daf2993c162b6f509d5f49040412376e6ae21d957798`. The local deterministic checks and [exact-candidate CI](https://github.com/GreenLv/dsh-completion-guard/actions/runs/34813886909) passed, including Ubuntu/macOS/Windows on Node.js 22 and 24. Separate macOS and Windows native runs on DSH `0.1.5-rc.2` each passed all 28 required host-bound gates on the same CI-frozen package, with no reused gates and successful cleanup. The real host tool path covers prepare without an optional reason code and Git missing-input replies. Both annexes retain the `real_model_request` skip; these runs do not establish daily-profile adoption, rc.1 native acceptance, or general answer-delivery semantics. Anonymous readback verified the annotated tag target, npm version and latest tag, embedded commit, registry integrity and downloaded package bytes. The GitHub Release title and bilingual body match the reviewed candidate. All seven attachments match the accepted files: package, checksum, artifact manifest, and the separate macOS/Windows annexes and transfer receipts. Windows result hashes were also matched to the original remote files before publication. This section and the updated main-branch installation instructions are post-release documentation; the tag and npm package retain their original bytes. The patch fixes prepare serialization and Git input guidance. The answer-delivery design remains future work in [the repair notes](https://github.com/GreenLv/dsh-completion-guard/blob/d11009d8f755ecee7d288cff18250c7b372cfd2a/docs/NEXT_VERSION_REPAIR_NOTES.md). ## 0.5.2 published release (2026-09-11) [Version 0.5.2](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.5.2) is published on GitHub and [npm](https://www.npmjs.com/package/dsh-completion-guard/v/0.5.2). The annotated tag and npm `gitHead` identify commit `e8b21b3d0bb5b3864c856cc9ea6bd19d5ab066ef`. The frozen 303412-byte, 27-file tgz has SHA-256 `d975f17dc3bfed999cbe6c0e638b226c3dcf8651f288ea7ab83faa5ba7903d31`. The local deterministic matrix and [exact-candidate portability CI](https://github.com/GreenLv/dsh-completion-guard/actions/runs/34577054241) passed. Separate macOS and Windows runs on DSH `0.1.5-rc.2` each passed all 28 required host-bound gates for these same bytes, with cleanup passed. Both annexes retain the `real_model_request` skip; no real-model task, daily-profile adoption, or rc.1 native run is claimed for this artifact. Anonymous public readback verified the annotated tag target, npm version and `latest` tag, embedded `gitHead`, registry integrity, and downloaded tgz bytes. The GitHub Release title is `DSH Completion Guard 0.5.2`. All seven attachments matched the accepted files: the tgz, checksum, artifact manifest, and separate native annexes and transfer receipts for macOS and Windows. The public package publishes `0.1.5-rc.2 || 0.1.5-rc.1` consistently in the top-level DSH engine field, nested plugin engine field, and all seven DSH peer dependencies. This section and the updated main-branch installation guidance are post-release documentation. The published tag and npm package retain their original bytes. ## 0.5.1 published release (2026-09-11) [Version 0.5.1](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.5.1) is published on GitHub and [npm](https://www.npmjs.com/package/dsh-completion-guard/v/0.5.1). The annotated tag and npm `gitHead` identify commit `aefdeaf2737ef1c99f1170085140c93a515e1125`. The frozen 303460-byte, 27-file tgz has SHA-256 `5e00ddf2f9772b1ae2ca3665d7c4ce858b814dc26a4116388082db2d625da28d`. The local deterministic matrix and [exact-candidate portability CI](https://github.com/GreenLv/dsh-completion-guard/actions/runs/34553623984) passed. Separate macOS and Windows runs on DSH `0.1.5-rc.2` each passed all 28 required host-bound gates for these same bytes, with cleanup passed. Both annexes retain the `real_model_request` skip; no real-model task is claimed. These isolated runs do not establish daily-profile adoption or native acceptance on rc.1. Anonymous public readback verified the annotated tag target, npm version and `latest` tag, embedded `gitHead`, registry integrity and downloaded tgz bytes. All seven GitHub Release attachments matched the accepted files: the tgz, checksum, artifact manifest, and separate native annexes and transfer receipts for macOS and Windows. The packaged registry provenance remains `registry-derived-pending-native-audit`; the later native results are carried by the exact-artifact annexes, not by a changed host-lock digest. This section and the updated main-branch installation guidance are post-release documentation. The published tag and npm package retain their original bytes. ## 0.5.1 validation scope The release supports exact DSH `0.1.5-rc.1` and `0.1.5-rc.2` core graphs. Development tests retain the rc.1 baseline; rc.2 graph tests additionally reject missing, mixed, forged and unregistered package identities. Dependency-free Headless preflight selects the matching registered cohort and checks its launcher version. Candidate results are recorded outside packaged documents after the source and package bytes are frozen. Require the local deterministic matrix, exact-candidate portability CI, and separate macOS and Windows native annexes for the same frozen tgz. An earlier rc.1 artifact's result does not establish acceptance of an artifact that adds rc.2 support. Real-model behaviour, daily profile adoption and publication remain separate gates. The shipped graphs retain their registry-derived provenance in the host-lock digest. Read a matching native annex for platform acceptance; the provenance field alone establishes neither a native pass nor a failed native run. ## Host-bound acceptance for the 0.4.3 line This line separates core and optional-provider acceptance while retaining completion feedback and root-confirmed rebinding. Local regression tests exercise compound capture, source authority, one-to-many replacement, stale/partial replay, target matching, large pages, snapshot changes, small recovery budgets and qualified pending boundaries. Local tests and the existence of a host driver do not establish native DSH acceptance. The versioned entrypoint is `scripts/native_acceptance.py`. Its default `portable_artifact` profile checks exact source/tgz identity, isolated installation, installed-file parity, second-install no-op and JavaScript syntax. `--gate-profile host_bound --runtime-root ` additionally creates isolated Web and Headless profiles, injects and reads back their host locks, then runs `scripts/native_host_probe.mjs` inside the real DSH composition. The probe uses the host AgentRegistry, ToolRuntime and durable Session services for nonempty test certification, generic refusal, rebind confirmation, history queries, compact/resume and a qualified pending boundary. Web checks require an owned-host restart, a different host process, persisted-session recovery and listener cleanup; they do not depend on a market restart API. The installed launcher shim is checked separately. Supplied daily target paths use the [read-only installation preflight](HOST_LOCK_UPGRADE.md#check-a-headless-profile-before-installation). A dependency-free `0.1.5-rc.1` or `0.1.5-rc.2` Headless target may have no private map or lockfile; the preflight verifies its installation-owned bundles and labels that state separately. The isolated lifecycle still installs Guard before checking its private graph, lock and second-install no-op. A successful target preflight alone is not a native gate result or proof of live adoption. The host driver deliberately makes no model request. It first checks that the normal Headless task driver stops with `MISSING_CREDENTIAL` in the isolated environment, then disables that task driver for the separate real-service probe. A required package-update probe installs an inert local fixture at version 1, then clarifies and confirms a generic requirement, applies version 2 through the real producer/action tools, independently reads it back and requires a nonempty certificate. Capability skips appear in the returned annex. Never treat a synthetic test or an empty probe case set as a passed native run. After exact-candidate portability CI passes and one clean commit's tgz has been frozen, a native owner can run the following with Python 3.11+ (the same artifact is used on both platforms): ```text python scripts/native_acceptance.py --gate-profile host_bound --repo-root --runtime-root --web-cohort --headless-cohort --web-market-version --artifact --artifact-sha256 --source-commit --output ``` Use `--target-web-profile` and `--target-headless-profile` together to compare actual consumer targets read-only before isolated installation. Both core identities are explicit; the Web market version is also explicit, and Headless installs no market. A target mismatch fails instead of selecting a historical fixture. The driver creates a new temporary DSH_HOME, explicit temporary HOME/USERPROFILE and credential-free environment; it does not select or enable a daily profile. It binds the annex to the artifact, source commit, host lock and probe bytes, and checks owned processes, ports and temporary-file cleanup. Do not run it from a dirty source tree, substitute a rebuilt package, or copy raw temporary logs into the repository. The historical 0.4.2 candidate at `6b92b3b7a5eb642686df9f2a1b4d66d54455f503` passed [candidate CI](https://github.com/GreenLv/dsh-completion-guard/actions/runs/34128172642). Its frozen tgz SHA-256 was `9d3e0a0bb948b137f27303a98ad38d3ecc8a901c5c579ce7e3b3c664530c3b4a`. Native macOS and Windows each passed all 28 required host-bound gates, including the normal Headless credential boundary and package-update probe, with cleanup passed. Both annexes retained the `real_model_request` capability skip. These are historical candidate results, not a release or evidence for later package bytes. The core annex uses `native-acceptance/v2`, gate profile `host_bound_core`, and the DSH-specific `capability_skips`, `host_driver_sha256`, `host_lock_digests`, `host_lock_policy` and `market_interface` fields. A generic closed v2 validator correctly rejects those extensions. The historical 0.4.2 `host_bound` annex retains its original `dsh-host-bound/v1` contract; the historical v1/v2 profiles and the current v3 profile are not interchangeable. A consumer that provides the explicit `dsh-host-bound/v3` contract profile must validate the original annex with independently supplied expected commit, artifact SHA-256, and probe SHA-256; do not remove fields to make it pass. `host_driver_sha256` identifies `native_host_probe.mjs`, not the Python wrapper. Windows checkout CRLF bytes can produce a different probe hash from macOS LF bytes; bind each platform to its actual reviewed file bytes. Failed probe annexes may additionally contain bounded `host_probe_failures` diagnostics. The profile consumer is separate tooling and is not installed by this npm package. Every subsequent frozen package needs its own CI and same-byte native annexes, recorded outside its packaged documentation. Publication and public readback are separate gates. A disabled daily installation remains disabled until the user separately requests an upgrade and enablement. ## Historical 0.4.2 release The [0.4.2 release](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.4.2) binds commit `df19d84db35369afb5b3a8041b6e5051db367155` and tgz SHA-256 `87deac307922fdff3391f1d6e0fa1b537c5b57e3a99e1d0943b0dc44b4e90068`. Its published annexes and public readback are specific to those bytes and the old market-1.41 fixture. They do not establish compatibility with a market-1.44 daily profile or the new core-lock policy. Direct market HTTP lifecycle checks exercise the installed market's actual interface. They do not substitute for the independent loaded-instance binding required by Guard's optional restart adapter. An unavailable adapter must leave an explicit restart requirement pending while core work remains protected. ## v0.4.0 release gates (2026-09-02, passed) Version 0.4.0 targets DSH `0.1.2-alpha.3` with dshmarket `1.39.0` and Cordis `4.0.2`. The repository does not place a candidate's own commit, checksum, or public status inside that candidate's packaged documentation: those facts are generated after the package bytes are frozen and are attached to the GitHub Release. | Gate | Required evidence and location | | --- | --- | | Source and CI | one clean release commit; repository matrix plus Ubuntu, macOS, and Windows Node.js 22/24 CI | | Exact artifact | one deterministic 26-file tgz; full `gitHead`, SHA-256, npm shasum, integrity, size, and file count in `release-artifact.json` and `SHA256SUMS.txt` | | Native macOS and Windows | the same tgz on both platforms; separate machine-readable annexes covering isolated Web and Headless lifecycle, host-lock, restart, daily-profile preservation, and cleanup | | Tag, npm, and GitHub Release | annotated tag, npm manifest and downloaded package, Release target and assets, and anonymous public readback all resolve to the same commit and bytes | | Daily-profile upgrade | separate from release acceptance; never implied by a passing artifact run | The release commit is `d5cd0ca17833d05bfdf41457ae203864bde8056b`. Candidate CI run 33626069583, main CI run 33633570436, and tag CI run 33633866279 each passed Ubuntu, macOS, and Windows on Node.js 22 and 24. The annotated `v0.4.0` tag has tag-object SHA `6130d53bdede2399b01f05e75a91bf3fadecb34e` and peels to the release commit. The frozen `dsh-completion-guard-0.4.0.tgz` contains 26 files, is 196166 bytes, and has SHA-256 `71ce205dedeffe72566ad399e001f0b337d057cede297e1d2249acec67cde1f2`. The package manifest carries the release commit as `gitHead`. These exact bytes passed isolated Web and Headless acceptance on native macOS and Windows. Both platforms checked installed-file parity, strict Web reinstall no-op, the alpha.3 host lock, injection and dump readback, a real Web restart with wrong-origin rejection and same-origin acceptance, changed boot ID and listener PID, intentional Headless `MISSING_CREDENTIAL`, preservation of the daily profiles, and cleanup. The attached macOS annex records 11 required gates; the Windows R5 annex records 24. The macOS Web run used Chokidar's polling mode after the host hit its file-watcher limit; the annex keeps that environment detail visible. npm published this exact tgz at `2026-09-02T13:18:57.245Z`. Anonymous registry readback returned `latest=0.4.0`, the release commit as `gitHead`, shasum `0980d44fa21b1f6bc225c96a4c52dcf1c6b3834a`, and the expected integrity. A fresh registry download was byte-identical to the frozen file. The non-draft, non-prerelease [`v0.4.0` GitHub Release](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.4.0) was published at `2026-09-02T13:21:09Z`. Its five assets are the frozen tgz, `SHA256SUMS.txt`, `release-artifact.json`, and the macOS and Windows native annexes. Anonymous Release readback returned the expected target commit and asset digests; a fresh Release tgz download was byte-identical to the frozen and npm files. This is a post-candidate evidence record and is not part of the accepted tgz. Changing the package bytes or release commit would require new artifact and native-platform acceptance. No daily Web or Headless profile was upgraded, and no real model request was made; those remain separate actions. ### Superseded pre-release artifacts The first 26-file candidate, SHA-256 `33757ce11633ac167fa9c5a6fdb7a8938578fcb871edd9afdb2f89dcb51b4316`, was built from implementation baseline `ffc6fe9e1246a815f0bb630943c59d14b6505716`. Native macOS passed. Native Windows reached the restart endpoint but did not prove a changed boot ID, so its Headless phase correctly remained unrun. That candidate was superseded when packaged documentation changed. The next 26-file candidate, SHA-256 `f3489e0f1edbb81e0ebdbd33e43a14e46913496cb3e1c799115143a70d2ab4c9`, was built from `ae20bffd06af45d4a0c60519f5ab7a7af6076ec8`. Its exact bytes passed isolated native macOS and Windows Web and Headless acceptance; the Windows R2 record contained 25 passing gates. Publication was still rejected because the package itself repeated the superseded candidate's blocked status and said another package still had to be built. That documentation defect is why a later release candidate is required; the native results remain valid only for `f3489e0f...`. ## v0.3.2 release gates (2026-09-01, passed) The release candidate is commit `22cde6106cdca511265bb4103375a263a0762b9c`. Its frozen `dsh-completion-guard-0.3.2.tgz` contains 26 files, is 181157 bytes, and has SHA-256 `feb7fc29799820e08dfe6d2bdb94823e745df9b5aa7c34d46262e5df30dabac4`. The package manifest carries the same full commit as `gitHead`. Candidate CI run 33461879125 passed Ubuntu, macOS, and Windows on Node.js 22 and 24. Those exact package bytes passed the required isolated Web and Headless lifecycle on native Windows and macOS: artifact and commit parity, clean install, strict Web reinstall no-op, installed-file parity, complete 34-row host checks, plugin load and configuration readback, real Web restart with 403 rejection and 202 acceptance, changed boot ID, HTTP recovery, intentional Headless `MISSING_CREDENTIAL`, and scoped cleanup. The platform-specific host digests differ because platform identity is part of the lock; both match their own complete checked package graph. The optional credentialed model-session gate was not run and is not claimed. `inspect` and `verify-dump` answer different questions. `inspect` reports whether the active runtime and profile package graph is a supported complete set, so it can correctly return `supported` before injection. The pre-injection fail-closed check comes from reading the composed config and from `verify-dump`, which rejects missing or mismatched injected host-lock data. The annotated `v0.3.2` tag has tag-object SHA `753fd9625c58c03de2660464c7f7c86462a42eb1` and peels to the release commit above. Tag CI run 33474979019 passed the same six Ubuntu, macOS, and Windows Node.js 22/24 combinations. npm published this exact tgz at `2026-09-01T05:54:28.309Z`; anonymous registry readback returned `latest=0.3.2`, the expected `gitHead`, shasum and integrity, and a fresh download was byte-identical to the frozen file. The non-draft, non-prerelease [`v0.3.2` GitHub Release](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.3.2) was published at `2026-09-01T05:56:24Z`. Its 181157-byte tgz and 97-byte `SHA256SUMS.txt` have GitHub asset digests `sha256:feb7fc29799820e08dfe6d2bdb94823e745df9b5aa7c34d46262e5df30dabac4` and `sha256:82f839d3ccf5f429b10ae18cc54350379097f0f59b01a462ed222a869886a451`. Anonymous downloads of both assets were byte-identical to the frozen files. This section is a post-candidate evidence record and is not part of the accepted tgz. Changing the package bytes or release commit would still require new artifact and native-platform acceptance. Consumer pin and managed runtime installation remain separate actions. ### Earlier source and host-cohort gates Version 0.3.2 lets the Guard recognize either of two complete DSH package sets. It never combines packages from different sets, and a missing or mismatched package leaves the whole host unavailable. The initial host-cohort checks below were run on macOS from a working tree based on `844b62848c1e2685e0574b660dc4546b6bf6dbac`, which was `origin/main` when implementation began. They cover source behavior, not a release artifact. A later native Windows release-pack run exposed an archive-extraction portability defect: the packer passed absolute archive paths to tar. Commit `a3e77de6d8260f16f0723491495cacab57b9f62d` now runs tar from the temporary package directory with relative archive and output paths. GitHub Actions run `33413968461` passed that exact fix on Ubuntu, macOS, and Windows with Node.js 22 and 24. This is CI evidence for the packaging fix, not native alpha.2 host acceptance or final documentation-inclusive artifact evidence. Package sets checked: - The `dsh-0.1.1-rc.2` set has the same 34 package rows used by 0.3.1 and already checked on macOS and Windows. - The `dsh-0.1.2-alpha.2` set contains 34 exact package name, version, and integrity rows originally extracted from the active macOS runtime and Web profile with dshmarket `1.38.1`, then matched in full on native Windows. Tests confirm that the TypeScript and JSON copies match. - A source comparison found no change in the session events, Goal calls, tool definitions, or terminal results that the Guard uses. Internal DSH changes outside those inputs are not treated as compatibility evidence. Verified source gates (macOS, Node.js 25.1.0, pnpm 11.x): - Type checking, lint, build, and `git diff --check` pass. The 20-file suite passes 359 tests and skips one Windows-only test on macOS. - `pnpm run pack:check` lists the expected 26 package files. `pnpm run test:release-pack` passes exact-commit binding, repeatable package output, dirty-tree rejection, and the portable relative-path extraction flow; it does not establish a final release artifact from the current documentation-inclusive tree. - Read-only checking against the daily macOS Web profile reports `dsh-0.1.2-alpha.2` as supported with all 34 rows, Web control, and Goal support. The installed 0.2.1 plugin still rejects a generator-version mismatch as expected. Native Windows source evidence (2026-09-01, PowerShell 5.1): - Fresh checkout `a3e77de6d8260f16f0723491495cacab57b9f62d` passed install, build, release-pack (2/2), and `git diff --check`. - The active DSH `0.1.2-alpha.2` / dshmarket `1.38.1` graph matched all 34 candidate rows; missing, extra, and duplicate counts were zero. This authorizes the Windows cohort registration, but it is source/host evidence rather than final package evidence. The frozen artifact above closes the package and native-platform lifecycle gates. Release mutation and public readback remain separate. ## v0.3.1 release-repair gates Version 0.3.1 preserves the v0.3 runtime and digest behavior while repairing the public provenance path. `scripts/release-pack.mjs` requires a clean Git root, resolves the full 40-character HEAD, stages the exact npm file set, injects that HEAD as `gitHead` only in the staged package manifest, and packs the staged package twice. It fails unless both tgz outputs are byte-identical and retain the same file count, then emits the single frozen tgz, `SHA256SUMS.txt`, and `release-artifact.json`. Publishing must use that tgz without repacking. The registry manifest, downloaded registry tgz, annotated tag, GitHub Release target, and checksum must all read back to the same commit and bytes. The focused Node test covers exact-HEAD injection, repeated-pack byte identity, checksum and artifact-record output, and dirty-tree rejection. Acceptance of the final tgz requires the full isolated Web/Headless install, no-op, package parity, host-lock, real restart, HTTP recovery, and cleanup lifecycle on native macOS and Windows before publication. CI, native-platform evidence, registry publication, tag identity, and GitHub Release remain separate gates. ### v0.3.1 completed public release (2026-08-31) The release commit is `00ed5c6456e15f0859c1ef7731157d07a3903af9`. Candidate CI run 33349603269, main CI run 33351918318, and tag CI run 33352017747 each passed the six Ubuntu, macOS, and Windows combinations for Node.js 22 and 24. Public `main` readback returned the same commit. Annotated tag `v0.3.1` has tag-object SHA `88e6842b289a0bf0df68fcc25d09e5358d7f457d` and peels to that release commit. The frozen `dsh-completion-guard-0.3.1.tgz` contains 26 files and is 172923 bytes. Its identities are: - SHA-256 `df3c0cae29fdfa0014d5cfdb6ade72c42386555779f9f4d8b59e37a2557c5d7e`; - npm shasum `b09f3844de9d957b5e15aa432833baf978483c55`; - npm integrity `sha512-TgCFFzIoj4tpDKIGr3QrgVjUq+nPySEFWS3f9RJR9VuNqBiWUvOBZrzS5QSd4+jXp3ug7uq91fPgdh44yVROcw==`; - staged manifest `gitHead` `00ed5c6456e15f0859c1ef7731157d07a3903af9`. That exact tgz passed isolated Web and Headless installation, strict second Web-install no-op, 26-of-26 package-byte parity, commit-blob parity, host-lock inspection/injection/dump verification, package import, intentional Headless `MISSING_CREDENTIAL`, fixed-port restart, HTTP recovery, and scoped cleanup on native macOS. Native Windows repeated the same exact-artifact lifecycle under an isolated supported DSH `0.1.1-rc.2` runtime after the daily runtime had advanced to an unsupported alpha cohort; the fail-closed version mismatch and the isolated rerun are both recorded in the Windows annex. The reported annex is 45531 bytes with SHA-256 `5b46cac12973c287dd72445f7d0b118d91a9329d77fae1ad56e756f24a99365c`. The raw Windows annex remains on the Windows host, so this repository records its immutable identity and bounded result rather than claiming an independent macOS read of that file. The exact tgz was published as [`dsh-completion-guard@0.3.1`](https://www.npmjs.com/package/dsh-completion-guard/v/0.3.1). Anonymous registry readback returned `latest=0.3.1`, the exact `gitHead`, shasum, integrity, 26-file count, and tarball URL. A fresh public-registry download had the expected SHA-256 and was byte-identical to the frozen tgz. The non-draft, non-prerelease [`v0.3.1` GitHub Release](https://github.com/GreenLv/dsh-completion-guard/releases/tag/v0.3.1) was published at `2026-08-31T03:04:09Z` and is the repository's latest public Release. Its only assets are the 172923-byte frozen tgz and the 97-byte `SHA256SUMS.txt`. Anonymous downloads of both assets were byte-identical to the frozen local files; the GitHub asset digests are respectively `sha256:df3c0cae29fdfa0014d5cfdb6ade72c42386555779f9f4d8b59e37a2557c5d7e` and `sha256:699764cfe8887f2a3abbaa35028abea18d6105794ea85983dff768d87239152f`. The release is therefore closed across source commit, annotated tag, CI, native-platform acceptance, npm metadata and bytes, GitHub metadata, and both public assets. These facts do not make the separate 0.3.0 npm publication a completed release and do not move or reuse its immutable identity. ## v0.3.0 incomplete publication (2026-08-31) Release preparation froze one canonical pre-release package from commit `a33b69326eb46fbefc56affc55e2a486695f545c`. The 26-file, 170158-byte tgz had SHA-256 `72d848e313a0e35e06fd1f493215cc0338b86a79a8001a4f07156e782157fe08`. The same bytes passed isolated Web and Headless installation, strict second no-op, package-file parity, host-lock inspect/inject/dump/verify, plugin load, real dshmarket restart, HTTP recovery, process cleanup, and the intentional Headless `MISSING_CREDENTIAL` boundary on native macOS and Windows. The Windows source run passed all 352 tests without skips; macOS passed 351 with the Windows-only command-shim test capability-skipped. GitHub Actions run 33320743166 passed that exact commit on Ubuntu, macOS, and Windows with Node.js 22 and 24. The four mirrored conformance files remained byte-identical to their upstream pin, all 37 portable semantic cases ran without skips, and all 29 digest-v3 vectors re-derived successfully. A separate credentialed model-session gate used that exact package. A real test result produced an accepted evidence binding with no rejected binding, and a persisted typed `user_wait` boundary was accepted; independent post-hook observation read back the same Goal as disarmed. An intentionally over-broad prompt captured additional non-certifiable clauses, so its checkpoint remained incomplete and no completion certificate was issued. This proves the bounded evidence, boundary, and disarm paths without claiming that arbitrary model instructions are semantically certifiable. The final documentation-inclusive source was commit `12e8411537b7f843aed267bc150a9403ddbb04c9`. Its frozen 26-file, 171419-byte tgz had SHA-256 `416b3539d38c13ea0e01b2154f342d911d4c14b81cc90d71f99e8b9bd6d6de45` and passed the same isolated lifecycle and same-byte package readback on native macOS and Windows. Main CI run 33347875843 and tag CI run 33347976931 passed Ubuntu, macOS, and Windows with Node.js 22 and 24. Annotated tag `v0.3.0` peels to that commit. The exact tgz was published as `dsh-completion-guard@0.3.0`; registry shasum `4c881b83b6046833229d5f54a062bfa2eea5be6f` and integrity `sha512-WSxbxD5N/79SJ/6UW0/XonCUlmckUhMHtj4LhUpa9u3XGC+MFZ/GUX935Y4Aqp8HJmxnj2gdHmAYmk8pcccykg==` identify the validated bytes. However, npm did not populate registry `gitHead` when publishing the prebuilt tgz. That fails this release's frozen public-identity gate. The version and tag remain immutable historical facts, but no GitHub Release is created for `v0.3.0`, and 0.3.0 is not claimed as a completed release. Version 0.3.1 repairs the packaging path instead of moving the tag, reusing the npm version, or weakening the gate after publication. ## Windows exact-source readback (2026-08-30) Direct readback of the isolated Windows TEMP acceptance evidence verifies source commit `b75868e9e73d29f50530ddaba15cfaef82e03ece` (`origin/main` at checkout time), tested in a fresh Windows 11 checkout with Windows PowerShell 5.1, Python 3.12.10, Node.js 24.18.0, pnpm 11.22.0, and effective `core.autocrlf=true`. The source and conformance matrix is verified: - `pnpm test` -> 0; 7 test files, 170 tests passed. - `pnpm run typecheck` -> 0. - `pnpm run lint` -> 0. - `pnpm run build` -> 0. - `pnpm run pack:check` -> 0. - `git diff --check` -> 0. - The raw SHA-256 values of all four mirrored conformance fixture files matched their entries in `tests/fixtures/conformance/UPSTREAM_PIN.json`. The exact-source artifact and load chain is also verified from the TEMP clone/log evidence. It used a tarball built from that checkout and a fresh isolated `DSH_HOME`: - All six tracked `dist/` files matched across the HEAD blobs, local build, tarball, and isolated installation. - `dsh --profile web --dump-config` read back the `context-guard` bundle, and a Node import smoke loaded the installed package. - `dsh --profile web --dump-config`, Web startup logging, process stop, and cleanup completed. An HTTP 200 appeared only in the first-run stdout; that response was not persisted and the readback did not rerun the GET, so HTTP 200 itself is not independently confirmed. With effective `core.autocrlf=true`, the build made all six tracked `dist/` paths appear as `M` even though `git hash-object` matched the corresponding HEAD blob for 6/6 files and `git diff` contained no content changes. This was an EOL status phantom, not an artifact mismatch; `dist/** text eol=lf` now pins the generated files to LF while preserving the existing conformance fixture LF rule. Evidence boundary: source checks and the exact-source install/load/startup and cleanup chain are verified. A real model-session smoke was not run, and the HTTP response lacks independently persisted/read-back evidence. This does not establish v0.3 runtime behavior, npm publication, a tag, or a GitHub Release. ## v0.3.0 source candidate gates (2026-08-30) Source commit `4f079499509822425c80e0b5ab98d1ebc58da9d5` on `codex/v0.3.0-sequence-2` passed the deterministic source matrix on macOS and native Windows. The commit remains an unpublished source candidate; it is not a tag, npm artifact, GitHub Release, or public-package readback. The macOS source matrix used Node.js 25.1.0 and pnpm 11.22.0: - `pnpm install --frozen-lockfile` -> 0. - `pnpm run typecheck` -> 0. - `pnpm test` -> 0; 19 test files, 351 tests passed and the one Windows-only command-shim test was capability-skipped. - `pnpm run lint` -> 0 with no warnings. - `pnpm run build` -> 0; a second build produced the same generated file names and SHA-256 values. - `pnpm run pack:check` -> 0; package identity is `dsh-completion-guard@0.3.0`, and the action, Git command, and supported-host manifests plus the host-lock CLI are included. - The portable runner executed all 37 mirrored semantic cases without skips. - The DSH digest runner re-derived all 29 digest-v3 vectors; the four mirror file SHA-256 values remain identical to `UPSTREAM_PIN.json`. The native Windows matrix used a fresh Windows 11 checkout, Windows PowerShell 5.1, Python 3.12.10, Node.js 24.18.0, pnpm 11.22.0, Git 2.53.0.windows.2, and effective `core.autocrlf=true`: - The focused host-lock/evidence suite passed 33/33 tests with no skips. It executed a `.cmd` shim from a path containing spaces and parentheses, bound both the shim and canonical `SystemRoot\\System32\\cmd.exe` identities, ignored later `PATH`/`ComSpec` substitution, rejected expansion characters, detected a fake `git.cmd` identity swap, and completed the real Git commit/push/fetch/pull round trip within its Windows timeout. - `pnpm test` -> 0; all 19 test files and all 352 tests passed with no skips. - Typecheck, lint, build, package dry-run, documentation audit, documentation unit tests, `git diff --check`, generated-`dist` parity, and final clean-tree readback passed. - All 37 portable semantic cases, all 29 digest-v3 vectors, the four pinned mirror hashes, and cross-repository fixture byte equality passed. Each platform built and recorded its own local exact-source tarball. The macOS artifact SHA-256 is `d613d88edbc44ccc020ad48dff6d180e79c04ed5bcabb73d7fae09d449500890`; the Windows artifact SHA-256 is `397975f720f0c6d734e7faf7279fab3feee1f92433ff457122755d9296adb19f`. These raw archive hashes are platform-local build provenance, not a requirement that independently packed gzip/tar containers be byte-identical. Both were built from a clean checkout of the exact commit, reported the same package identity and 26-file package list, retained build-to-`dist` parity, and were installed from the artifact that was hashed on that platform. A future release must instead freeze one canonical tarball, bind it to the release commit, tag, npm `gitHead`, and registry integrity, and verify that same artifact on every required native platform. Both platform-local exact-source artifacts were installed into fresh isolated Web and Headless profiles without modifying the user profile. For both profiles, the packaged host-lock CLI completed `inspect -> inject -> inject (idempotence) -> dsh --dump-config -> verify-dump` against the active runtime/profile graphs. The Web tuple contained 34 exact package rows with Web control available; the Headless tuple contained 33 rows with Web control unavailable by profile while the other applicable capabilities remained supported. On both macOS and Windows, an isolated Web profile loaded real `dshmarket@1.36.0`, returned HTTP 200, accepted the correct restart request with HTTP 202, replaced the process, and returned a different boot ID after restart. The Windows run also verified that a wrong restart origin was rejected. Each replacement process returned HTTP 200 before teardown; the temporary listener, helper, and profile processes were then read back as stopped. Each isolated Headless profile loaded the plugin and advanced to the expected `MISSING_CREDENTIAL` boundary with `DEEPSEEK_API_KEY` removed from the child environment. This establishes deterministic source behavior plus native macOS and Windows package composition, host-lock readback, load, and Web restart lifecycle for commit `4f079499509822425c80e0b5ab98d1ebc58da9d5`. It is not a credentialed model-session Goal/checkpoint/boundary round. CI, a canonical release artifact, npm/GitHub publication, tag identity, and a real model-session smoke remain independent pending release gates. ### npm statistics integration boundary (2026-08-30) The validation branch later integrated the bilingual npm download chart in `131b2db7f7ee555e6e4794395a8aa5118275fa80` and hardened its collector and publication workflow in `0ebda88312bf226869543c65de70159fe0abdaca`. The collector's 8 focused tests pass on macOS and cover leap-year date chunks, missing and unordered days, duplicate/out-of-range/negative rows, range/point reconciliation, scoped package URL encoding, HTTP and JSON failures, and preservation of the previous output set when collection fails. The fixed-date 2026-08-29 render reconciled 780 requests for `dsh-context-guard` and 140 for `dsh-completion-guard` into a 920-request project line. Both 960 x 540 SVGs parsed as XML, retained `` and `<desc>` accessibility text, rendered without invalid numeric tokens, and were visually inspected in English and Simplified Chinese. Publication is fail-closed to the repository default branch before checkout. The collection job has read-only contents permission and does not retain Git credentials; only the downstream publication job receives contents write permission. The four official Actions used by the workflow are pinned to immutable commits. No `stats` branch or chart asset was published during this candidate integration. Evidence boundary: the runtime source paths, `dist/`, CLI, manifests, and lock file remain unchanged from the native-platform candidate above. The chart integration does change packaged README and package metadata, and this acceptance document is itself shipped in the package. The earlier macOS and Windows tarball hashes therefore cannot be inherited by the final candidate. One canonical tarball must be built from the final exact commit and the same bytes installed and read back on macOS and native Windows before release. ## macOS v0.2.0 acceptance (2026-08-28) Verified source commit: `c107cd8ead97988f6a71cab8182edb16b23b086b`, on `main`, with a clean worktree and `main == origin/main` at the release preflight. Node v25.1.0 and pnpm 11.22.0 were used on macOS. Repository gates passed before publication: - `pnpm install --frozen-lockfile` -> 0. - `pnpm test` -> 0; 5 test files, 124 tests passed. - `pnpm run typecheck` -> 0. - `pnpm run lint` -> 0 errors / 0 warnings. - `pnpm run build` -> 0. - `pnpm run pack:check` -> 0. - `validateManifest()` returned no errors. The semantic checks covered informational receipt filtering, deterministic-check classification, `2>&1` support, and rejection of compound `&&` syntax. Live Web acceptance after installing `dsh-context-guard@0.2.0` into the managed `web` profile and restarting DSH: - The live profile reads installed version `0.2.0`; `dsh --profile web --dump-config` includes the `context-guard` bundle. - A real Web session in this repository executed one foreground Bash `pnpm test` call. The persisted result reported 5 test files and 124 tests passed. - The next checkpoint bound the successful evidence (`E0001`) to the two current session items. The checkpoint returned `status: certified`, `contract_revision: 4`, with `rejected_bindings: []` and `open_items: []`. Publication and consumer readback: - `npm view dsh-context-guard dist-tags.latest` -> `0.2.0`. - npm `gitHead` for `0.2.0` -> `c107cd8ead97988f6a71cab8182edb16b23b086b`. - The official `dsh-context-guard-0.2.0.tgz` unpacks to version `0.2.0`; all six packaged `dist/` files are byte-identical to the local build by md5. The local-only `dist/.DS_Store` is not part of the npm package and is excluded from this comparison. - Annotated tag `v0.2.0` was pushed. Remote readback returned tag object `dc1b13e00b2e1adadacb2c8de0a533ec8f51ac22` peeling to commit `c107cd8ead97988f6a71cab8182edb16b23b086b`. - codex-sync update discovery reported exactly `dsh-context-guard 0.1.2 -> 0.2.0`; the dry run planned one update, apply succeeded, the installed profile package reads `0.2.0`, the bundle contains `context-guard`, and the follow-up dry run returned `No DSH plugin changes needed`. Evidence boundaries: - The live 0.2.0 checkpoint certifies the real `pnpm test` workflow and its bound session items. It does not certify every release requirement in the recovered handoff contract. - The rejected compound-command case with an actionable `rejected_bindings[].hint`, and informational receipt filtering, are covered by the 0.2.0 semantic/regression checks and earlier real-session evidence; they were not re-created as additional live Web calls in this acceptance run. - Windows native acceptance remains the documented 0.1.x bounded PowerShell result; no Windows 0.2.0 run is claimed here. ## macOS v0.2.1 acceptance (2026-08-28) Verified source commit: `ba8f05dc6a922b8a12f4dc211d9a888d2ece526a` (= annotated tag `v0.2.1`, tag object `5bccac9dd77549a929137befd808dc1701d77093`). Repository gates and package verification ran on a clean worktree at that commit; the npm publication ran from `806aa863234a064b5c42391fa1288998abfc846e` as recorded below. Node v25.1.0, pnpm 11.22.0, npm 11.17.0 on macOS 26.6.2 (arm64). The pre-publish artifact `dsh-context-guard-0.2.1.tgz` had SHA-256 `ea2b6a0bc1db82150af9d88494b6c9d2a1e0243d33ed21df48d3c23d7c52a895`. Repository gates passed at `ba8f05d` (each exit code recorded at run time; the untracked `docs/PROBLEM_REPORT_v0.2.1.md` was moved out of the worktree for the pack steps and restored afterwards): - `pnpm install --frozen-lockfile` -> 0. - `pnpm run typecheck` -> 0. - `pnpm test` -> 0; 6 test files, 138 tests passed (`tests/domain/conversation.test.ts` is the new v0.2.1 conversational-capture matrix with 7 tests). - `pnpm run lint` -> 0; 0 warnings / 0 errors (26 files, 96 rules). - `pnpm run build` -> 0; 6 files, 151.33 kB total. - `pnpm run pack:check` -> 0; `dsh-context-guard-0.2.1.tgz`, 19 files, and the untracked problem report is not part of the package. Installed-state verification: the tgz above was added to the real `web` profile (`dsh plugin --profile web add file:...`) and to the `headless` profile; both installed packages read `0.2.1`, `dsh --profile web --dump-config` still composes the `context-guard` bundle, and DSH web was restarted after the install (new process PID 97136 started 09:55, listening on 127.0.0.1:3080, after the 09:53 install) so the live instance loads 0.2.1. ### Live multi-turn Web session (one guarded session, all steps) A brand-new guarded session in the real Web UI (workspace `dsh-context-guard`, DeepSeek-V4-Flash-Vision-Exp, workspace-write), driven through a dedicated Chrome instance over the Chrome DevTools Protocol — the same browser-automation lane this DSH install uses for its own browser skill (`chrome-devtools-mcp`). User messages were submitted with CDP-trusted pointer events on the composer's send control; every acceptance observable below was double-checked against the server-side session log (`~/.dsh/sessions/…/session-6dcd1274-d4bb-430f-a07c-bcddda8a3bca/session.jsonl.zstd`). The profile's managed patch layer already sets `context-guard` to `activation: always`, so the session was guarded from its first event. 1. Baseline: `/context-guard status` -> `{"enabled":true,"epoch":0, "contract_revision":0,"pending":0,"passed":0,"evidence":0, "integrity":"valid"}`. 2. Change 1, clarifying message: `这个收尾具体要做什么` was sent; the model answered (and explored the repository, producing 32 evidence entries). The next `/context-guard status` returned `pending:0`, `contract_revision:0` — the clarifying question was not captured as a contract item. 3. Change 1, progression phrase: `继续` was sent; the model continued its analysis. `/context-guard status` again returned `pending:0` (evidence 37) — the progression phrase was not captured. 4. Evidence production and certified checkpoint: the single-clause task "在本仓库前台单一命令运行 pnpm test 且不得使用管道分号或重定向,然后用本次运行产生 的证据调用 context_guard_checkpoint 绑定全部开放项完成认证" was captured as `R001` (checkpoint with empty bindings returned `incomplete`, `open_items:["R001"]`); the model executed a clean foreground `pnpm test` (single command, no pipes or redirects; 6 files / 138 tests passed), which produced durable evidence `E0038` (bash, scope, outcome success, shell + deterministic-check), and the binding `{"item_id":"R001", "evidence_ids":["E0038"]}` returned `{"status":"certified", "contract_revision":1,"open_items":[],"rejected_bindings":[]}`. Both raw tool results are persisted in the session log. 5. Pre-clear goal-gate denial (fail-closed evidence): a verification message was itself captured as `R002` (revision 1 -> 2); the model's empty-binding checkpoint returned `incomplete, open_items:["R002"]`, and `update_goal(action=complete)` was denied with "Context Guard requires a current completion certificate before Goal completion." — the gate denies before checking whether a goal exists (`get_goal` returned `{"goal":null}`). 6. Clear remediation: `/context-guard clear` returned "Context Guard contract cleared: 1 requirement/acceptance item(s) superseded; 0 pending remain (prohibitions retained)." A follow-up verification (phrased as a question, which the conversational filter does not capture) re-ran both tools: the empty-binding checkpoint returned `{"status":"certified", "contract_revision":8,"open_items":[],"rejected_bindings":[]}`, and `update_goal(action=complete)` was no longer blocked by the guard — it reached the goal tool's own argument validation (`goal_id` placeholder rejected) because this session has no active goal. The final `/context-guard status` read `{"enabled":true,"epoch":0, "contract_revision":8,"pending":0,"passed":1,"evidence":43, "integrity":"valid"}` — the guard stayed enabled across the clear. Headless cross-checks (fresh guarded session per run, `--patch` overlay with `activation: always`, real model turns; server logs under `~/.dsh/sessions/`): - Compound-command rejection: `pnpm test 2>&1 | tee …; echo …` produced scope-only evidence and the binding attempt was rejected with `incomplete` and per-item hints ("needs a scope run effect: a whitelisted executable … without pipes, `;` or `&&`"). No binding rule was changed for it. - Recovery-injection dedup: across three headless sessions, six `context_guard_checkpoint` rejections (nonexistent evidence `E9999`, both missing-item and evidence-mismatch reasons) produced exactly one "Open task requirements (recovered after compaction or resume):" plugin notice per session — no duplicate recovery packet was ever injected live. The step-separated identical-rejection scenario is covered deterministically by `tests/runtime.test.ts` ("recovery injection dedup (v0.2.1)"). Evidence boundaries: - The live Web session covers the work order's steps 1-5 (status baseline, clarifying message, progression phrase, supported-task evidence with a certified checkpoint, and clear -> empty-binding certified -> goal-gate release with the guard still enabled). The `update_goal` leg ends at the goal tool's own argument validation because the acceptance session had no active goal; the gate's release is nonetheless demonstrated, and the denial path was captured live pre-clear. - Evidence remains session-scoped by design; cross-session evidence IDs are rejected and no cross-session import exists. Windows native behavior is not re-claimed here; the 0.2.0 Windows records stand. ### Publication and consumer readback The npm publication was authorized and executed from `806aa863234a064b5c42391fa1288998abfc846e` ("docs: update READMEs to v0.2.1"), a post-release README follow-up commit that was already on local `main` (not pushed at the time) when the release was approved. Per the maintainer's explicit choice, the registry `gitHead` therefore points at `806aa86…` rather than the tag commit `ba8f05d…`; the annotated tag `v0.2.1` itself was not moved and still peels to `ba8f05d…`. - `npm publish` -> `+ dsh-context-guard@0.2.1` (19 files, package size 70.5 kB, shasum `2a32ce46fd00fd87d144011653bd99fe2c5a9bd5`). The account's npm token had expired and the publish required npm's EOTP web authentication, completed in the browser. - `npm view dsh-context-guard dist-tags.latest` -> `0.2.1`; versions sequence `0.1.2, 0.2.0, 0.2.1`. - `npm view dsh-context-guard@0.2.1 gitHead` -> `806aa863234a064b5c42391fa1288998abfc846e`. - The official `dsh-context-guard-0.2.1.tgz` unpacks to version `0.2.1`; all six packaged `dist/` files are byte-identical (md5) to the `ba8f05d` gate-run build, and both README files are byte-identical to the `806aa86` content. No `docs/PROBLEM_REPORT_v0.2.1.md` is part of the package. ## 0.1.1 release candidate (2026-08-27) The candidate was built on macOS from branch `codex/0.1.1-bash-success-evidence`. The tested runtime domain bundle is `dist/domain-DEXOzCqH.js` with SHA-256 `e8ad974b8263a2a25c7ef08ca3839097d3312c10d450b8520163c623e54881f9`. The exact pre-commit source-tree and tgz digests are recorded in the generated candidate manifest rather than embedded here: this document is itself part of the npm archive, so embedding that archive's digest would change the digest. Verified for this candidate: - `pnpm install --frozen-lockfile`, typecheck, lint, all 106 tests, build, and package-content validation completed successfully on macOS. - The candidate tgz composed into fresh isolated `web` and `headless` profiles. Both installed packages read back as `0.1.1`; a second identical add reported `Already up to date`, and hashes of the profile manifests, lockfile, composed patch files, and installed package manifest did not change. - The 0.1.1 domain runtime replayed a completed real DSH Web log captured with `activation: always` and confirmed durability. Its foreground `bash` `pnpm typecheck` result contained ordinary stderr output and no `[exit code: 0]` marker; replay derived successful scope evidence and issued a certificate for all four matching open items with no rejected bindings. - A clean Windows checkout of commit `16ea9e7a5d088ce6b7e09617f15acd771f57ff40` passed the frozen install, typecheck, lint, all 106 tests, build, and package-content validation with DSH `0.1.1-rc.2`. The candidate package was `dsh-context-guard@0.1.1` with SHA-256 `6c0aed77a9fc7f43f6d8bc11e4355d465d1146c81fe10276ad3e523f0abd8a83`. Fresh isolated Web and headless profiles both read back version `0.1.1`; a second identical installation was a strict no-op with unchanged lockfile hashes. - A separate minimal Windows real-model session loaded that candidate with `activation: always`. One supported foreground `pwsh Set-Content -LiteralPath ... -Encoding ascii -NoNewline` call created `D:\dsh-pwsh-acceptance\pwsh-acceptance.txt`, and an independent `read` returned `guard-test`. The persisted create evidence and read evidence were bound to the single current contract; the checkpoint returned `certified` at contract revision 1 with no open items or rejected bindings. Evidence boundary: - The macOS Web result is a candidate-runtime replay of a genuine closed Web log, combined with isolated Web composition. It is not a claim that the unpublished candidate was installed into the user's live Web profile. - The Windows certificate covers only the documented PowerShell subset and the exact candidate identity above. The first workspace-confined write attempt was rejected because the target was outside the checkout; the same single command succeeded only after an explicitly approved elevated retry. The rejected call was not available for certification. This run does not expand the Windows Bash claim or certify compound PowerShell commands. - An earlier Windows session completed its model turn without a certificate, and a later diagnostic session produced only scope evidence for compound PowerShell calls. Neither is counted as acceptance. The successful result is the fresh revision-1 session whose supported write and independent read produced artifact-matching evidence. - These pre-release native results do not by themselves establish CI, an npm publication, a tag, or a GitHub Release; those publication identities require separate readback. ## Deterministic checks ```sh pnpm install pnpm run typecheck # tsc --noEmit pnpm test # vitest pnpm run lint # oxlint pnpm run build # tsdown pnpm pack --dry-run --json ``` ## Isolated profile composition (macOS, no real ~/.dsh) ```sh export DSH_HOME=/tmp/dsh-context-guard-smoke DSH=/path/to/dsh "$DSH" plugin --profile web add "file:/abs/path/to/dsh-context-guard" "$DSH" --profile web --dump-config | grep context-guard "$DSH" plugin --profile headless add "file:/abs/path/to/dsh-context-guard" "$DSH" --profile headless --dump-config | grep context-guard ``` Verified result: both profiles contain `id: context-guard, name: dsh-context-guard` in the composed tree. ## Native load A real headless boot loads the plugin and advances to the model call; it stops at `MISSING_CREDENTIAL` when no provider key is configured. That is the expected boundary for a load check and is unrelated to plugin correctness. ## Verified - Web UI command rendering and the `/context-guard on|off|status|diagnose` subcommands produce the expected `command/run`/`command/done` (round-5 isolated profile); `/context-guard diagnose` was re-verified in round-8. - A complete macOS headless task in an isolated profile created one artifact, read it back as durable evidence, rejected an incomplete first binding, then certified the complete requirement and acceptance set before `turn/end`. ## v0.2.0 Windows acceptance status The GitHub tag `v0.2.0` resolves to release commit `c107cd8`. The GitHub Release is published at https://github.com/GreenLv/dsh-context-guard/releases/tag/v0.2.0. Windows 0.2.0 native runtime acceptance completed in the separate live DSH Web session. The loaded package version was `0.2.0`. The supported PowerShell command `Set-Content -LiteralPath 'D:\dsh-pwsh-acceptance\pwsh-acceptance-020.txt' -Value 'guard-test' -Encoding ascii -NoNewline` created the artifact successfully after the documented permission-boundary retry. An independent `read` returned exactly `guard-test` with no trailing newline. The durable evidence IDs from that session were `E0009` for the successful PowerShell write and `E0011` for the independent read. The session's checkpoint attempt was `incomplete`, and a follow-up session also rejected `E0009` and `E0011` because those IDs are not present in its current session projection. Context Guard currently derives evidence from the active DSH event log only; it has no cross-session evidence import or lookup contract. The runtime actions and their outputs are therefore recorded as completed runtime acceptance, but the v0.2.0 checkpoint remains `incomplete` in this later session by design. To obtain a certificate, the commands and checkpoint must be performed in one session. The Windows scope is limited to the documented PowerShell subset and does not expand to Windows Bash or compound PowerShell syntax. ## Native platform acceptance - macOS: an isolated real-model headless task used a supported POSIX shell write and an independent read, then persisted a certified checkpoint before the completed turn. - Windows: the v0.1.1 candidate had an isolated real-model task using `pwsh Set-Content -LiteralPath` and an independent read, with a certified checkpoint. The v0.2.0 runtime actions are recorded in the dedicated status section above, but their cross-session checkpoint repair remains incomplete. Both final-SHA runs used clean public checkouts and isolated `DSH_HOME` directories. They establish native behavior for the bounded v0.1 command subset; they do not claim support for shell or PowerShell syntax outside the subset documented in `COMPATIBILITY.md`. ## Public-package profile readback On macOS, the published `dsh-context-guard@0.1.0` npm package was installed into a real DSH Web profile through the pinned `codex-sync` plugin reconciler. A second dry run was a strict no-op; direct package and bundle readback reported version `0.1.0`; `dsh --profile web --dump-config` included the `context-guard` bundle; and the restarted Web command directory exposed `/context-guard`, whose `status` subcommand returned a valid projection summary. This verifies public-package consumption and real-profile loading on macOS. It does not replace the isolated model-task evidence above and does not claim a second Windows run from the public npm package. ## macOS v0.1.2 acceptance (2026-08-27) Verified commit: `dd402fbaa5d37dd246056d8ecd66430e5f75f412` (annotated tag `v0.1.2`, clean checkout, `main` fast-forwarded to the tag). macOS, Node v25.1.0, pnpm 11.22.0. Gate run (exit codes and summary): - `pnpm install --frozen-lockfile` → 0. - `pnpm test` → 0; 5 test files, 107 tests passed. - `pnpm run typecheck` → 0. - `pnpm run lint` → 0 errors / 0 warnings (23 files, 96 rules). - `pnpm run build` → 0; 6 files, 123.60 kB total (`dist/index.js`, `dist/domain/index.js`, `dist/domain-CJulh_RZ.js`, `dist/index.d.ts`, `dist/domain/index.d.ts`, `dist/index-CA7Z-W_A.d.ts`). - `pnpm run pack:check` → 0. Artifact check: `shell exited: code` is present in `dist/`. The literal string `persistent bash shell was reset` cannot appear in any build of commit dd402fb, because the implementation parameterizes the reset line as `/^The persistent (?:bash|pwsh) shell was reset;/` to cover both persistent renderers (`src/domain/evidence.ts` `PERSISTENT_RESET_LINE`). Equivalent checks pass: the regex is present in `dist/domain-CJulh_RZ.js` (line 1059), it matches the rendered Bash and pwsh prose lines (run through `RegExp.test`), and the built `dist` module replays all eight 0.1.2 cases (`[shell exited: code 1]`, `[shell killed by signal: SIGTERM]`, `[shell exited]`, timeout intro — each with and without the reset prose; plus `[shell exited: code 0]` → success and reset-prose-only → success) with the expected outcomes. Publication readback: `npm view dsh-context-guard dist-tags.latest` = `0.1.2`; registry `gitHead` = `dd402fbaa5d37dd246056d8ecd66430e5f75f412`; the official tarball `dsh-context-guard-0.1.2.tgz` unpacks to version `0.1.2` with both artifacts above, and all six `dist/` files are byte-identical (md5) to the local build — the published artifact matches the gate-run build. Installed state (macOS, real `~/.dsh`, `codex-sync` managed): pin raised `0.1.1` → `0.1.2` in `config/dsh/plugins.toml`; `--check-updates` reported the pin current; the pre-apply dry run planned exactly one UPDATE; `--apply` ran the single `dsh plugin --profile web add dsh-context-guard@0.1.2` pass; `~/.dsh/profiles/web/node_modules/dsh-context-guard/package.json` reads `0.1.2`, the installed `dist/` carries both vocabulary markers, the package stays in `dsh.profile.bundles`, and a follow-up dry run is a no-op (`No DSH plugin changes needed`). DSH was restarted (process started 23:41:05, after the 23:39:17 install) so the live instance loads 0.1.2. Smoke (in the live guarded session, 0.1.2 loaded): a clean foreground `pnpm test` in this repository completed with 5 files / 107 tests passed and produced no terminal marker; the guard-derived evidence (`E0065`, bash, repo scope) carries outcome `success`, and the checkpoint binding `{R042, [E0065]}` was accepted with no rejected bindings — the clean foreground bash result both scores `success` and certifies its matching item, confirming the 0.1.1 clean-success contract still holds under 0.1.2 (the 0.1.1 → 0.1.2 core regression point). One earlier binding was properly rejected and is recorded, not hidden: the first smoke attempt wrapped the command in `2>&1 | tail -6; echo ...`, which the guard classifies as non-deterministic with no supported operations; the binding `{R041, [E0058]}` failed the run contract, and no binding rule was changed for it. A subsequent clean run was used for the accepted binding. This section records the live-profile acceptance of v0.1.2 on macOS. It does not claim completion of the session's full 55-item contract, which was not part of this smoke; it does not re-claim any Windows result.