--- name: doctrine-research description: Use when the user asks for deep research done with the doctrine: a multi-source research question that needs a grounded, fact-checked recommendation, dual-engine with red-teaming and looping until confident. --- # Doctrine Research Research-synthesis under the doctrine: two independent engines, one merged claim table, honest gaps, a grounded recommendation. **REQUIRED BACKGROUND:** Read the `doctrine` skill first. There is no code diff here, but the gate is still the full three parts — mapping onto the posture: - A **phase** is one research question; step 5's loop runs inside it. - **Native checks** (doctrine step 3) = the docs-diff form the doctrine names, run against the report: **re-verify every edited claim against the source it cites, then run a link and path check** — every URL fetched and still carrying the cited content, every repo path or filename in the report resolved on disk. A citation that 404s, redirects to a homepage, or no longer says what you quoted is a blocking finding, and **nothing else in this flow looks for one** — a red team argues about claims and a logic critique argues about reasoning; neither opens a link. **Plus every check the project itself documents as a gate** for the directory the report lands in — a prose or markdown linter, a docs build, a frontmatter or link checker — read its task runner rather than assuming a report has no gates, because a report is a file in someone's repo. Where the destination has none, the docs-diff form is the whole of native checks: a complete answer, not a gap to keep hunting. - **Designated review** = the **logic critique** (step 4). That is the named equivalent doctrine step 5 permits in the slot, and **you say so in the Method appendix** — a gate reported as three parts that ran two is a pass you did not earn, and from outside, a two-part gate and a three-part one produce the same clean report. - The **red team** (doctrine step 4) attacks the merged claim table and the drafted report together, never a single source or a lone finding. ## Flow 1. Ask the user what the request left open: the decision this research informs, its success criteria, and the scope bounds (timeframe, region, budget, constraints) — **all three together are the anchor**, plus any Done means items the phase serves (doctrine step 1), written down verbatim (every drift check compares against the anchor, never against intermediate findings). Scope belongs inside it, not beside it: research that answers the right question about the wrong region has drifted, and a critique holding an anchor without the bounds has no way to see that. Then the report destination (suggest a `docs/` subfolder of the current project, e.g. `docs/research/YYYY-MM-DD-.md`). 2. **Round 1 — two engines, blind, concurrent, same refined question:** - Engine 1 (your own model family): the host's deep-research workflow, stock and unmodified, where the hub's host reference names one. Otherwise, a fan-out of parallel web-search seats with per-claim adversarial verification. - Engine 2 (the other model family): the red-team seat the hub's host reference names, given a read-only research task answering the same question independently, **every claim returned tagged `[web]` with a fetchable URL or `[model-knowledge]`**. **Check that it can actually search before you count it as a second sourced engine**, as the host reference says, and check again on return: a `[web]` tag carrying no URL you can fetch is model-knowledge, whatever it says. A return that is only a job handle or an acknowledgement is not an answer, and a handle merged as an answer is an empty engine reported as a full one. Retrieve the result as the host reference says before merging. If it has not returned by the time you would otherwise merge, treat engine 2 as unavailable, say so, and go on. Where that seat is absent, use a fresh-context subagent with web access, and note in the report that engine diversity was reduced. - Neither engine sees the other's output in round 1. Independence is what makes cross-examination meaningful; a second engine that critiques your draft is anchored to your frame and finds only your kind of gaps. 3. **Merge** into one claim table: **agree** (both engines) / **conflict** (keep both sides with sources — conflicts are report content, never silently resolved) / **single-engine** (unverified until a second source or check confirms) / **unknown**. **An unsourced engine 2 does not make an agreement corroborated.** Where engine 2 answered from model knowledge, a claim both engines assert is one search result plus one recollection: file it as **single-engine, model-corroborated** — never as agree — and say plainly in the report that the run had one sourced engine. Read the other way it is *worse* than running one engine, because the report now says corroborated and nobody goes back. What an unsourced engine 2 is still worth is **divergence**: where it contradicts engine 1 it has found you something to go and check, so that goes in **conflict** on those terms and gets a fetch, not a vote. 4. **Logic critique**: a dedicated agent reviews your own merged research for unsupported leaps, circular sourcing (many citations tracing to one origin), survivorship and recency bias, conflated adjacent questions, and drift from the anchor. If the data suggests the real question differs from the anchor, that critique returns a **divergence flag** as a named finding — a dispatched agent has no user to raise anything with, so **you put it to the user on receipt**: re-aim or stay is their call, and never silently pivot the objective. Record the ruling in the dossier's settled section the moment it lands; the next loop's critique is a fresh context that will otherwise re-open the same question with full confidence, and an accepted re-file is a blocking finding against the pass it lands in. 5. **Gap ledger**: every unknown and unverified claim gets a ledger entry: what's missing, what was already tried. Each loop iteration = targeted gap-fill fan-out (ledger items only, not a full re-run) → red-team round (refute AND extend the claim table; verify anything it adds from its cited source before it counts) → logic critique re-run on whatever changed. 6. **Exit gate** (doctrine step 5 — this states the replacement explicitly, as the hub requires, and replaces its exit condition only; nothing else in the posture is overridden): the hub's **prose-deliverable exit, adopted here by name**, where a full pass is native checks + logic critique + red team, with one research-shaped precondition on its closing round: every remaining gap closed or written up as unanswerable with the attempts shown. **A pass is clean when it produces no blocking finding and leaves none outstanding from an earlier pass**, whoever raised it and however long ago. "No *new* findings" is not the test: an unsupported leap the logic critique files in loop 3 and files again in loop 4 is outstanding, not new, and it blocks both loops until the claim is sourced, softened, or moved to the gaps section. **Blocking** is doctrine step 5's test, and its research-shaped instances are: a dead, redirected or misread citation; a chain of citations that widens to four sources and narrows to one origin while the report calls it corroborated; a conclusion the claim table does not carry; drift off the anchor; a gap neither closed nor written up; a failed native check. One more confirming source on a claim already sourced is **non-blocking** and goes on the deferred list you name at delivery; a cleaner phrasing is not a finding, per step 5's three buckets. Never fill a gap with a plausible guess — an honest "unknown" outranks an invented fact. At the round alarm's first firing (doctrine step 5), put the gap ledger as it stands to the user inside the diagnosis that step requires. The ledger is this wrapper's open-findings list, not a substitute for the rest of that diagnosis, and what happens next is the prose-deliverable exit's own rule, not restated here. **Write the loop number and whatever doctrine step 5 puts in that section into the dossier's run-state block every loop**, never only into your context, since a compaction takes in-context state and announces the loss to nobody. 7. **Deliver**: write the report to the path agreed in step 1 and commit it per the project's norms (doctrine step 7 — the destination was agreed at kickoff, so landing it needs no second ask). Report shape: **Recommendation** (BLUF, grounded in the claim table) → **Findings** ranked by confidence, each cited and tagged both-engines / single-engine / contested → **Conflicts** → **Gaps & unknowns** → **Divergence notes** (if any, and how the user ruled) → **Method appendix** (loops run, sources, red-team rounds, **which review filled the designated-review slot — the logic critique — and which native checks ran, plus any that could not run and why**, so a reader can see the gate was three-part rather than taking it on the word "clean"; plus the deferred non-blocking list and the dossier's path). In chat: the recommendation plus anything that needs the user's judgment. ## The dossier One file, built in step 1 **before the first engine is dispatched**, carrying three things: the **anchor** in the user's own words (the decision this informs, its success criteria, the scope bounds, and any Done means items the phase serves, per doctrine step 1), the **claim table** as it currently stands, and the **gap ledger**. Every consumer of all three is a fresh-context agent you dispatch, and **a seat pointed at a file reads its own summary of it** (doctrine step 4) — so the file is where the facts *survive*, and the prompt is where they *arrive*. Both jobs are yours. **It lives beside the report** — the destination agreed in step 1 with a `-worknotes` suffix (`docs/research/YYYY-MM-DD--worknotes.md`) — or, where that destination is not writable or the user does not want working files in the repo, in **the scratchpad**: the session temp directory if there is one, else a temp directory outside every repo. Name the path in the Method appendix either way, and at delivery ask whether it ships beside the report or is dropped. It is amended, never frozen: each merge, each loop's gap-fill results, each red-team finding you verified or refuted with its evidence, each divergence the user ruled on. The loop pauses for a ruling at doctrine step 5's alarms rather than ending there, and can outlive the context that opened it; a claim table held only in conversation goes with the compaction that takes it. **Paste, do not reference.** Into the **logic critique**: the anchor verbatim — its whole drift axis is "drift from the anchor", and a critique handed the merged findings and no anchor grades the findings against themselves, which is the one comparison the logic critique (step 4 above) exists to avoid; it returns clean, and two exit-gate reviewers then pass a report that has drifted off the question — **and the divergence flag named as a return format**: a critique never told the flag exists folds the re-aim signal into ordinary findings, and the report's Divergence notes section then reads "none" over a live one. Into the **red team**: the claim table **and the drafted report** — the red-team sentence above orders the two attacked together, it can fetch neither, and a table without the report's conclusion cannot catch the one blocking class step 6 names, a conclusion the table does not carry. Into each **gap-fill agent**: the anchor verbatim — a gap chased without the question's own timeframe, region and constraints comes back answering a different question — and its ledger entries, including the *what was already tried* column, which is the only thing stopping the fan-out from re-running last loop's dead searches — **and engine 2's tag law, stated in the prompt in the same words engine 2's dispatch sentence defines it in** — the same line goes into the red team's prompt for anything it *extends* the table with, because the merge taxonomy is keyed on sourcing status and a claim arriving untagged can only be dropped or guessed at. Into **both** reviewers: the settled section — every divergence and every finding the user has already ruled closed, one line each, since a fresh reviewer cannot know a question was answered a loop ago and files it again in good faith. **One carve-out: the counters and the round history are yours alone.** Keep them under a trailing `## Run state (orchestrator only)` heading and hand nothing below it to any reviewer. A red team told the next blocking finding fires the round alarm has been told how badly the pass needs to be clean, which is the one fact a reviewer must not hold — and the counters belong nowhere near the claim table in any case, because that table is what the report may assert and a counter asserts nothing. Not for quick fact lookups (plain search) or codebase questions (doctrine-audit). ## Red flags - Both engines "agree" and neither one fetched a page. - Engine 2 returning a job handle, and the handle merged as an answer. - The logic critique dispatched with the merged findings and no anchor. - Four citations, one origin, reported as corroborated. - A gap filled with a plausible number because the loop was otherwise clean. - A finding re-filed by the next loop and waved through as not new. - A report shipped with a URL nobody re-fetched or a repo path nobody resolved. - The gap ledger and the counters living only in the conversation on loop 3.