# Hermes R-round References: Transferable Lessons for Raven ## Scope and Raven filter This audit read the 26-line compatibility stub and all 25 Markdown files under `C:\Users\lhan\AppData\Local\hermes\profiles\nana\skills\research\r-round-plan-lifecycle\references\`. It treats those files as an operational incident record, not as a specification to copy. Raven's contract is intentionally smaller: one continuing Raven Task; progressive immutable Checkpoints; Steering Revisions on the same Task; normalized Sources with bounded verbatim excerpts that are re-fetched and literally matched; Claims with explicit dispositions; visible Limitations; and Completion only when the candidate Artifact bytes equal the exact latest post-steer Checkpoint. Raven has no public stage approvals, external ledger files, lease/monitor protocol, or agent topology. Verdicts below therefore mean: - **Keep** — directly encode or test the integrity property in Raven. - **Change** — preserve the underlying lesson but express it through Raven's compact state, Source/Claim checks, Checkpoint fingerprint, or Limitations rather than Nana workflow ceremony. - **Drop** — Nana-specific lifecycle, path, fan-out, or document bureaucracy that does not belong in Raven's public contract. ## Per-reference extraction | Reference | Concrete operational lesson | Class | Raven verdict and evidence | |---|---|---|---| | `r19-plan-revision-vs-execution.md` | Do not confuse revision/correction intent with authorization to execute a different lifecycle; accidental execution should not erase prior state. For Raven, a Steering Revision must update the same Task and invalidate stale verification rather than launch a replacement Task. | **GENERAL** core; quarantine/archive mechanics are Nana-specific. | **Change.** Keep intent/identity continuity and non-destructive recovery, but drop launch gates and plan files. Evidence: `C:\Users\lhan\AppData\Local\hermes\profiles\nana\skills\research\r-round-plan-lifecycle\references\r19-plan-revision-vs-execution.md:3-20,33-42`. | | `verdict-shard-authoring.md` | Evidence must be classified by what it can support: source prestige is not source intent, negative searches are bounded, sample insufficiency caps confidence, and untouched falsifiers remain untouched. The prescribed 12-section page, pool tags, fixed verdict vocabulary, shard destinations, and closing sentence are Nana workflow conventions. | **MIXED**, predominantly **NANA-SPECIFIC bureaucracy**. | **Change.** Retain claim ceilings, bounded negative evidence, non-independence, and honest deferral as Claim dispositions/Limitations; drop the shard template and exact prose. Evidence: `...\verdict-shard-authoring.md:36-49,53-63,75-89`. | | `supplement-pass-and-subagent-scope.md` | A later gap-fill should state whether it changes the conclusion, and work must be checked against its authorized scope rather than trusted from a worker summary. Raven should preserve new evidence on the same Task and expose any conclusion change, but it should not model supplement rounds or subagent file ownership. | **MIXED**. | **Change.** Keep explicit gap/limitation handling and stale/out-of-scope contribution rejection at the main-agent boundary; drop S-suffixed rounds, archive integration, git revert protocol, and topology. Evidence: `...\supplement-pass-and-subagent-scope.md:9-23,89-95,98-169`. | | `bilingual-breadth-challenge-shard.md` | A translation or external framework must not silently redefine the authoritative source concept; inspected originals control, snippets remain leads, and evidence lanes must remain separate. Mechanical checks should not be overstated as project tests. | **GENERAL**. | **Change.** Model translation scope and non-applicability in Claim text/disposition and Source metadata/Limitations; no card taxonomy or temporary JSONL verifier is needed. Evidence: `...\bilingual-breadth-challenge-shard.md:7-22,24-30`. | | `broad-to-top-n-plan-review-pattern.md` | Ranking must not manufacture quota fulfillment; candidates need semantic admission, source-family independence, claim-type ceilings, and bounded retrieval. Much of the review/scoring/fan-out apparatus is specific to top-N R-rounds. | **MIXED**. | **Change.** Keep semantic admission, source-family non-independence, claim ceilings, inspected-primary requirements, and honest partial coverage; drop scorers, sensitivity runs, pilots, and launch blockers. Evidence: `...\broad-to-top-n-plan-review-pattern.md:14-34,36-47,67-79,81-92`. | | `broad-to-top-n-execution-normalization.md` | Accepted evidence must bind to frozen inputs; duplicate URLs/publication families are not independent corroboration; heuristic similarity cannot silently merge claims; named dimensions must not be confused by positional order. | **MIXED** with strong general integrity lessons. | **Change.** Raven should canonicalize Source identity/family and reject duplicate Source IDs or mutated identity, but it needs no candidate pool, scorers, manifests, or pilot releases. Evidence: `...\broad-to-top-n-execution-normalization.md:5-21,23-43,45-66,68-79,93-105`. | | `artifact-only-breadth-shard-verification.md` | Search snippets are leads only; persisted excerpts must occur literally in reopened source bodies, while literal occurrence remains separate from semantic support. Non-contiguous excerpts, punctuation changes, compound claims, stale post-edit verification, schema aliases, dates, and source roles are recurring failure points. | **GENERAL**. | **Keep** the core, **Change** the machinery. Raven directly implements bounded excerpt re-fetch/match, immutable Source identity, Claim dispositions, current/as-of metadata, and fresh verification at checkpoint/complete; it should not import JSONL, stubs, child scripts, or git checks. Evidence: `...\artifact-only-breadth-shard-verification.md:7-18,20-56,58-71,112-160,189-204`. | | `structured-ledger-repair-and-reverify.md` | Repair must be bounded and honest: replace with an exact excerpt, split/narrow the claim, downgrade support, or defer; then reverify persisted repaired bytes and distinguish closure of old findings from newly found defects. | **GENERAL**. | **Keep.** Raven's equivalent is a new Checkpoint with new immutable Source IDs when evidence changes, disposition downgrades for unsupported Claims, complete source rechecks, and no mutation on failed Completion. Drop separate repair workers/verifier files. Evidence: `...\structured-ledger-repair-and-reverify.md:5-19,21-37,39-80,82-110`. | | `claimed-repair-vs-final-byte-verification.md` | Claimed repair, worker messages, and tickets are hypotheses; final persisted bytes are authority. Semantic acceptance and integration cleanliness are distinct, and verification must bind to exact hashes/fields. | **GENERAL**. | **Keep.** This directly supports Raven's exact latest-Checkpoint SHA-256 completion gate and rechecking of current Source/Claim state; no external checker artifact is needed. Evidence: `...\claimed-repair-vs-final-byte-verification.md:5-22,24-48`. | | `deep-dive-publicity-terminology-and-role-boundaries.md` | Preserve source-language meaning and separate responsibility design, interpretation, implementation, and measured effect. Hosting, official wording, and self-report do not establish independent outcomes; adjacent concepts must not be inferred as equivalent. | **GENERAL**. | **Change.** Encode as narrow Claims, source-role/family metadata, qualification/deferment, and Limitations; drop dossier quote caps and Nana acceptance cues. Evidence: `...\deep-dive-publicity-terminology-and-role-boundaries.md:5-19,21-38`. | | `deep-dive-public-policy-benefits-evidence.md` | Current law, implementation, fiscal rules, and outcomes are separate evidence lanes; policy existence is not measured effect. Compound claims require clause-complete persisted evidence, and any later edit invalidates prior verification. | **GENERAL**. | **Keep** the integrity rules, **Change** the format. Raven should use `as_of`, narrow Claims, Source families, exact excerpts, Claim disposition/Limitations, and post-edit Checkpoint/Completion freshness; drop three-file closeout and child-verifier ceremony. Evidence: `...\deep-dive-public-policy-benefits-evidence.md:5-23,25-36,38-51`. | | `frozen-label-and-adjacent-boundary-verification.md` | Verify identity before substance: an internally coherent artifact can answer a narrower or sibling question. Exact excerpt presence and full Claim support are separate booleans; a boundary decision cannot silently delete requested scope. | **GENERAL**. | **Change.** Raven has no frozen direction manifest, but it must bind every Checkpoint to the Task request plus latest Steering Revision and reject completion if no post-steer Checkpoint exists; semantic scope remains the agent's responsibility and should surface as Limitations. Evidence: `...\frozen-label-and-adjacent-boundary-verification.md:5-17,19-40,42-53`. | | `frozen-direction-deep-dive-rebuild.md` | Renaming a title does not repair semantic drift; claims, sources, query framing, and conclusions must all be rebuilt consistently. Final source/excerpt verification must occur after the last edit. | **GENERAL** core; workspace and artifact-bundle protocol are Nana-specific. | **Change.** Raven should require a revised Checkpoint after Steering and exact-byte Completion, while changed evidence gets new Source IDs; drop three-file synchronized rebuilds, field-order checks, and temp verifier scripts. Evidence: `...\frozen-direction-deep-dive-rebuild.md:5-16,18-34`. | | `artifact-only-deep-dive-pilot-execution.md` | Current sources, source-family independence, bounded non-hits, exact excerpts, claim ceilings, and clause-level support must be verified; mechanical/schema success does not equal independent semantic acceptance. Pilot manifests, stubs, caps, helper scripts, and writer/verifier role topology are Nana-specific. | **MIXED**. | **Change.** Preserve Source/Claim integrity and Limitations; drop pilot/fan-out and external artifact protocol. Raven should explicitly remember that literal matching proves quotation fidelity, not entailment. Evidence: `...\artifact-only-deep-dive-pilot-execution.md:7-19,23-45,50-73`. | | `public-party-pla-responsibility-deep-dive.md` | Formal responsibility, interpretation, implementation, and measured outcome must not be collapsed; exact phrases should not be manufactured, one page is one source family, and bounded non-hits are not absence proofs. | **GENERAL**. | **Change.** Map to Claim kind/disposition, Source family/role, exact excerpt checks, and Limitations. Drop topic-specific organizational taxonomy and child-verifier closeout. Evidence: `...\public-party-pla-responsibility-deep-dive.md:5-18,20-30,32-43`. | | `corrected-governance-participation-bundle-verification.md` | A corrected artifact can still carry old-direction claims; inspect residual wording semantically, not with blind zero-token rules. A 100% live/literal census can still hide compound overclaims, so track whole-Claim support separately. | **GENERAL**. | **Keep.** Raven needs tests where all excerpts match but Claim scope exceeds them, and Steering leaves stale final text; affected Claims should be qualified/deferred and Completion blocked or limited. Evidence: `...\corrected-governance-participation-bundle-verification.md:5-23,25-49`. | | `corrected-frozen-direction-independent-verification.md` | Current bytes, exact requested identity, semantic residue, literal excerpts, and whole-Claim support must all be checked independently. Negative/boundary mentions are not contamination, while a renamed old thesis can be. | **GENERAL** core; independent-verifier artifacts and migrated workspace reconciliation are Nana-specific. | **Change.** Raven should bind latest Checkpoint to latest Steering Revision and inspect actual current Artifact/Claims, but not create independent verifier files or cross-workspace hashes. Evidence: `...\corrected-frozen-direction-independent-verification.md:5-26,28-41`. | | `deep-dive-pilot-independent-verification.md` | Never infer acceptance from missing/stale/wrong-path inputs, HTTP success, or literal excerpt success. Verify current authoritative bytes, source identity, clause-level support, locators, current status, and honest coverage; preserve failures rather than forcing expected counts. | **GENERAL** integrity embedded in extensive **NANA-SPECIFIC** verifier bureaucracy. | **Change.** Keep fail-closed restore/codec validation, exact Source identity, complete relevant Source rechecks, semantic support separation, and Limitations; drop path stubs, manifests, verifier JSON, query-family tables, and multi-pass topology. Evidence: `...\deep-dive-pilot-independent-verification.md:7-19,43-65,67-88,111-159,161-195,198-255`. | | `final-reverify-punctuation-and-residual-closure.md` | Repair has ledger-level and bundle-level closure: narrowing a Claim is insufficient if the Artifact still asserts the removed predicate. Exact excerpts must preserve punctuation at every boundary, and segmented excerpts must not masquerade as continuous quotations. | **GENERAL**. | **Keep.** Raven should scan final Artifact citations/Claim trace after disposition changes and reject stale accepted support; normalization must not fold punctuation. Raven currently models one bounded excerpt per Source, so segmented quotations should be represented as separate Sources rather than imported segment bureaucracy. Evidence: `...\final-reverify-punctuation-and-residual-closure.md:5-29,31-52,54-65`. | | `final-segmented-anchor-reverify-and-contract-regression.md` | A named repair passing does not imply the whole contract still passes; rerun all relevant checks on final bytes. Delimiters and HTML boundaries must be represented honestly, with row-level and bundle-level counts kept distinct. | **GENERAL** core. | **Change.** Raven should rerun source/citation/claim/completion invariants after any substantive Checkpoint and use new Source IDs for changed excerpts. Drop query-family matrices and verifier-output ceremony. Evidence: `...\final-segmented-anchor-reverify-and-contract-regression.md:5-14,16-46,48-58`. | | `deep-dive-fanout-identity-and-path-binding.md` | Identity and bytes must come from authoritative frozen inputs, not chat memory or worker summaries; a valid artifact for the wrong target is misdelivery. Verifier results are valid only for the exact input hashes they inspected. | **GENERAL** lesson expressed through **NANA-SPECIFIC** fan-out/path machinery. | **Change.** Raven already needs Task ID binding, replay-safe state, immutable Checkpoints, monotonic revision, and exact final fingerprint. Drop dispatch maps, worker paths, salvage, wave manifests, and topology. Evidence: `...\deep-dive-fanout-identity-and-path-binding.md:5-31,33-71,73-103`. | | `final-independent-major-round-acceptance-audit.md` | Final acceptance must verify current bytes, accepted identity, source/claim boundaries, and risk-shaped evidence rather than trusting prior green verifiers. A complete census and an independent sample answer different questions. | **GENERAL** integrity wrapped in **NANA-SPECIFIC** major-round acceptance ceremony. | **Change.** Raven should deterministically reverify every cited/relevant Source at Checkpoint and Completion and expose Limitations; it should not add independent audit stages, manifests, leakage scans, or approval gates. Evidence: `...\final-independent-major-round-acceptance-audit.md:5-45,47-76`. | | `major-artifact-only-round-closeout.md` | Acceptance must bind passing verification to current bytes; final-byte edits stale old verification; source decoding/segmentation defects must not be mistaken for evidence defects. The archive/status/audit sequence is Nana lifecycle bureaucracy. | **MIXED**. | **Change.** Keep current-byte binding, fresh source verification, exact-byte Completion, and conservative extraction; drop acceptance manifests, independent major audits, archive transitions, logs, lint and commit protocol. Evidence: `...\major-artifact-only-round-closeout.md:5-35,37-58,60-89`. | | `harness-tool-promotion-and-anchor-census.md` | Repeated evidence verification deserves one tested implementation; exact matching must handle encoding, segmented anchors, user-agent differences, genuine throttling, and only presentation normalization. A checker needs positive and negative tests and must be rerun after changes. | **GENERAL**. | **Keep.** This strongly validates Raven's single `SourceVerifier` seam and deterministic adapter tests. Add fixtures for mojibake, punctuation drift, dead hosts, redirect/identity drift, UA-dependent responses, cancellation, and retry-only recovery; do not copy Nana `.harness/` promotion policy. Evidence: `...\harness-tool-promotion-and-anchor-census.md:7-34,36-113`. | | `major-plan-review-closure-loop.md` | Review results apply only to the exact bytes reviewed; concurrent writers can overwrite one another; closing known findings does not preclude new defects; administrative edits after PASS can stale the review. The durable plan-review directory, Q1-Q5 approval, and launch key are Nana-specific gates. | **MIXED**. | **Change.** Keep hash/freshness and non-inherited acceptance as Raven Checkpoint/Completion invariants; drop review passes, stage status, launch approval, and external review artifacts. Evidence: `...\major-plan-review-closure-loop.md:5-16,17-36`. | ## Failure modes that can also occur in Raven Raven is simpler, but simplification removes coordination surfaces—not evidence failure modes. The following incidents remain possible. ### 1. Claimed repair is absent from final Artifact bytes **R-round symptom:** a ticket or worker says the repair landed, while the persisted file still contains the old wording (`claimed-repair-vs-final-byte-verification.md:5-13`). **Raven exposure:** the agent may describe a correction in a Checkpoint summary or Steering response but submit an unchanged Artifact to `complete`. **Detection needed:** - Store the exact latest Checkpoint SHA-256 and applied Steering Revision. - Reject Completion unless candidate Artifact bytes equal that exact latest Checkpoint fingerprint. - Reject Completion when the latest Steering Revision has no subsequent Checkpoint. - Treat summaries/messages as non-authoritative. This is already central to the architecture and should receive acceptance tests for unchanged, one-character-changed, normalized-but-not-byte-equal, and pre-steer candidate Artifacts. ### 2. Repair closes a Claim but stale broader wording survives in the Artifact **R-round symptom:** ledger-level closure is green while dossier/query text still asserts the removed predicate (`final-reverify-punctuation-and-residual-closure.md:5-14,43-52`). **Raven exposure:** a Claim is changed to `qualified` or `deferred`, or a failed Source is removed from usable support, but the latest Artifact still cites or states the original supported conclusion. **Detection needed:** - Validate all `[@source-id]` tokens and generated Claim trace against current dispositions. - Require every material supported/qualified external Claim to have a cited currently usable Source. - When source failure empties a Claim's usable support set, automatically defer it and reject any final Artifact that still presents/cites it as accepted support. - Add tests where the Claim record is repaired but prose/citation remains stale. Raven cannot mechanically prove every uncited paraphrase corresponds to a Claim. That semantic gap should remain an explicit main-agent duty and, when uncertain, a Limitation—not a reason to add an external ledger. ### 3. Literal excerpt matches, but the Claim overreaches it **R-round symptom:** HTTP 200 and exact substring match pass while one or more material Claim clauses are supported only elsewhere on the page or not at all (`artifact-only-breadth-shard-verification.md:20-27`; `deep-dive-pilot-independent-verification.md:121-129`). **Raven exposure:** Source verification proves quotation fidelity, not entailment. The main agent can still attach a broad Claim to a narrow excerpt. **Detection needed:** - Keep SourceVerifier's contract observational: reachable, identity-safe, excerpt-matched. - Make prompt/tool guidance explicit that the main agent must judge clause-level entailment before marking `supported`/`qualified`. - Test deterministic engine invariants that prevent `supported`/`qualified` external Claims with empty, unknown, or failed Source sets. - Consider a compact completion issue/limitation when Claims are marked qualified but their narrow scope is not reflected in the Artifact. Do **not** pretend string matching solves semantic support, and do not add a second model/verifier workflow to v1 merely to simulate certainty. ### 4. Excerpt is semantically right but not verbatim **R-round symptom:** model-added punctuation, quote-mark substitution, stitched paragraphs, missing block-boundary space, or ellipsis makes a non-literal quotation look exact (`artifact-only-breadth-shard-verification.md:24-56,151-158`; `final-reverify-punctuation-and-residual-closure.md:16-39`). **Raven exposure:** bounded Source excerpts can fail literal re-fetch even when humans see equivalent wording. **Detection needed:** - Normalize only HTML/entity and whitespace presentation; never normalize punctuation, quotation marks, digits, or wording. - Require one contiguous excerpt per Source. If evidence is genuinely segmented, use separate Source records/IDs rather than inventing a composite quote format in v1. - Return an actionable mismatch issue and leave Task state intact. - Test CJK punctuation, smart quotes, DOM inline/block boundaries, omitted middle paragraphs, and model-added terminal marks. ### 5. Source identity or independence is wrong **R-round symptom:** mirrors, translations, reposts, or multiple rows from one document are counted as independent sources; hosting domain is mistaken for originating publisher (`broad-to-top-n-execution-normalization.md:23-33`; `deep-dive-pilot-independent-verification.md:111-119`). **Raven exposure:** multiple Source IDs may refer to one underlying family, inflating apparent support; redirects may cross hosts or drift identity. **Detection needed:** - Canonicalize URLs and reject credential-bearing identities. - Preserve optional `family` and role metadata; prompts should require shared-family labeling when known. - Reject cross-host resolution and report resolved URL/identity drift conservatively. - Never infer independence from Source count alone; generated Claim trace should show IDs, while Claim reasoning and Limitations explain shared origin. - Add tests for exact duplicate URLs, canonical equivalents, same-family distinct URLs, redirects, and reused Source ID mutation. Full publication-family resolution cannot be perfect mechanically. Raven should expose the limitation rather than construct a normalization bureaucracy. ### 6. Current-status evidence is stale or uses the wrong date **R-round symptom:** historical translation, issuance/effective date, event date, or old page is represented as current law/status (`artifact-only-breadth-shard-verification.md:141-160,176-185`). **Raven exposure:** a reachable, exact excerpt can still be temporally inapplicable. **Detection needed:** - Preserve `as_of` metadata and inspection/check times. - Re-fetch at Checkpoint and Completion; a match confirms page presence, not supersession absence. - Require the agent to qualify current-status Claims and record bounded supersession uncertainty as a Limitation. - Add tests that retain exact excerpts while forcing explicit `as_of`/qualification behavior in acceptance scenarios. Do not claim that URL reachability proves current legal force. ### 7. Steering changes scope, but the Task silently answers the old or narrower question **R-round symptom:** exact label is renamed while old mechanism/claim set survives, or a sibling boundary silently removes requested scope (`frozen-label-and-adjacent-boundary-verification.md:5-17`; `corrected-frozen-direction-independent-verification.md:18-26`). **Raven exposure:** the latest Checkpoint is post-steer but only cosmetically changed; the artifact remains semantically stale. **Detection needed:** - Bind each Checkpoint descriptor to the latest Steering Revision ordinal. - Inject the compact active-Task request plus latest Steering summary before subsequent work. - Require every substantive final edit to become a new Checkpoint. - Acceptance tests should verify identity continuity, no Task restart, and rejection of pre-steer Completion. - Semantic compliance with the correction remains the main agent's responsibility; if only partial, record a Limitation explicitly. The hash gate detects stale bytes, not inadequate reasoning. Raven should not introduce approval stages to compensate. ### 8. Verification passes, then final bytes change **R-round symptom:** a green verifier applies only to bytes it inspected; any later edit makes it stale (`artifact-only-breadth-shard-verification.md:189-200`; `major-plan-review-closure-loop.md:24-27`). **Raven exposure:** citations, generated Sources/Claim trace, or prose may be altered after source verification. **Detection needed:** - `checkpoint` and `complete` both re-run SourceVerifier on relevant Sources. - Completion compares exact candidate hash to the latest Checkpoint hash after rendering policy is applied consistently. - Ensure the fingerprint owner and rendering order are unambiguous: hash the same canonical Artifact bytes the user sees and supplies to Completion. - Test post-verification one-byte edits, citation rendering changes, and source-check record freshness. ### 9. Restore/replay selects malformed, stale, or wrong-Task state **R-round symptom:** wrong path/worktree or stale mirrored bundle exists and parses, but is not the authoritative repaired state (`deep-dive-fanout-identity-and-path-binding.md:46-69`; `deep-dive-pilot-independent-verification.md:43-55`). **Raven exposure:** compacted/restarted sessions could restore malformed metadata, wrong current Task, or a stale snapshot. **Detection needed:** - Versioned codec with exact allowed keys, size limits, invariants, IDs, hashes, and phase checks. - Skip malformed/unknown snapshots and continue scanning for an older valid one. - Rebuild the full session→task registry, not merely “last Raven result.” - Keep in-memory calls serialized and state revision monotonic. - Test multiple Tasks, stopped/resumed history, malformed newest snapshots, unknown versions, duplicate identities, and replay after steering. This replaces path/manifest bureaucracy with Raven-owned replay validation. ### 10. Source fetch failures are verifier defects, not evidence defects **R-round symptom:** CJK mojibake, user-agent-dependent 403, parser boundary mistakes, or throttling create false failures (`harness-tool-promotion-and-anchor-census.md:36-102`). **Raven exposure:** production SourceVerifier may misclassify valid Sources as mismatched or unavailable. **Detection needed:** - Decode using trustworthy response/HTML charset signals and test CJK fixtures. - Distinguish unavailable, failed, redirected/identity-drifted, excerpt-mismatch, cancelled, and adapter-protocol failure. - Validate exactly one result for each requested Source ID and reject extra/missing results conservatively. - Exercise cancellation and adapters that ignore AbortSignal. - Test deterministic UA/header-sensitive and retry-only cases without silently smoothing recoveries into success. Raven should keep these adapter concerns behind the existing `SourceVerifier` seam, not expose retry policy in the public Task contract. ### 11. Honest partial result is mistaken for full completion, or zero valid support is accepted **R-round symptom:** quota pressure, missing inputs, or green mechanical checks lead to acceptance despite no valid support (`broad-to-top-n-execution-normalization.md:93-105`; `deep-dive-pilot-independent-verification.md:9-19`). **Raven exposure:** a grounding-required Task could complete with only deferred/rejected Claims or all Sources failed. **Detection needed:** - Enforce at least one material supported/qualified external Claim with a currently reachable excerpt-matched Source for grounding-required Tasks. - Keep zero-valid-work Tasks active even when their Limitations are honest. - Allow `completed-with-limits` only after failed support is removed, affected Claims are deferred, and useful independent sections remain. - Test all-sources-failed, one-independent-section-survives, optional grounding, and explicit limitation paths. ### 12. A worker/tool ending is mistaken for business Completion **R-round symptom:** execution, verifier, or archive state is conflated with research acceptance. **Raven exposure:** Harness subagent/workflow/goal completion could be presented as Raven completion. **Detection needed:** - Keep Raven Completion exclusively behind `raven_task({ action: "complete" })` and its exact invariants. - Prompt explicitly that tool, worker, scheduler, and goal termination are never Raven Completion. - Do not store topology in Task state. - Add acceptance tests where a subagent/tool reports success but no valid latest Checkpoint exists. ## What not to import The references repeatedly solve real problems with mechanisms Raven intentionally does not own. Raven should **not** import: - launch phrases, stage approval gates, Q1-Q5 approvals, or `ready/in-flight/verification/closeout/archive` states; - external Markdown/JSONL evidence ledgers, plan indices, review directories, acceptance manifests, audit sidecars, or archive logs; - leases, monitors, controller receipts, shard counts, dispatch maps, worker path ownership, retries as logical workers, or multi-agent topology; - exact page-section templates, pool-tag taxonomies, frozen field order, temporary child-verifier scripts, or “independent verifier” artifacts; - generic top-N scoring, sensitivity analysis, pilot releases, supplement rounds, or quota accounting. Those mechanisms are not necessary to preserve the transferable invariants. Raven can own the invariants locally through its Task engine, immutable Source/Claim identity, bounded re-fetch-and-match adapter, Checkpoints, Steering Revision binding, Limitations, replay codec, and exact final fingerprint. ## Recommended Raven implementation/test emphasis 1. **Make final-byte authority non-negotiable.** Test exact latest post-steer Checkpoint equality exhaustively. 2. **Keep quotation fidelity and semantic entailment separate.** SourceVerifier proves retrieval/identity/excerpt match only; prompt and Claim disposition own semantic judgment. 3. **Make evidence changes immutable.** Changed URL/title/locator/excerpt/role/family/as-of data requires a new Source ID; Claim ID text mutation is rejected. 4. **Reverify after every relevant final state.** Both Checkpoint and Completion reopen applicable Sources; a failure never fabricates success or destroys prior state. 5. **Propagate support loss.** If a Source fails, defer Claims whose usable support becomes empty, remove them from accepted final support, and create a visible Limitation. 6. **Treat extraction as a production-grade seam.** Add deterministic tests for encoding, HTML/entity/whitespace presentation, punctuation drift, redirects, cross-host drift, malformed adapter responses, cancellation, unavailable provider, and dead/mismatched sources. 7. **Test replay as an integrity surface.** Malformed newest metadata, unknown versions, multiple historical Tasks, Steering/Checkpoint ordering, and currentTaskId restoration deserve first-class tests. 8. **Prefer honest incompleteness over ceremony.** A narrow supported Artifact with explicit Limitations is better than a fabricated full answer; zero valid grounding must remain active.