--- name: agent-wiki-ingest description: Ingest one or more agent trajectories (raw bob/claude traces or normalized JSON) into an agent-wiki end-to-end — convert, summarize, extract guidelines, synthesize skills, optionally compare outcomes, consolidate into clusters, and catalog. Use when you have a batch of traces to turn into a wiki in one pass. --- # Agent Wiki — Ingest (end-to-end orchestrator) ## Overview This is the **one-pass entry point** for turning a batch of raw trajectories into a fully-built wiki. It orchestrates the rest of the `agent-wiki` family in the right order so no pass is skipped — in particular the cross-trajectory **consolidation** pass, which is easy to forget when each skill is invoked by hand. You — the driving agent — run this by **spawning one subagent per (trace × pass)**, not by doing the work inline. That keeps your own context small (you never load every trace's full JSON) and lets independent passes run in parallel. Each subagent acts as the corresponding single-purpose skill (`agent-wiki-summarize`, `-extract-guidelines`, `-synthesize-skill`, `-compare-outcomes`, `-consolidate-guidelines`); this skill only sequences them and passes the per-trace adapter notes. The pipeline: ``` 0. Convert raw bob / claude traces → normalized analysis JSON (skip if already normalized) 1. Bootstrap create wiki scaffold + seed catalog (skip if wiki exists) 1.5 Skip drop traces whose summaries/.md already exists [pre-flight — idempotency] 2. Summarize 1 subagent / new-trace → summaries/.md [PARALLEL] 3. Extract 1 subagent / new-trace → guidelines/*.md (+tags) [SEQUENTIAL] 4. Synthesize 1 subagent / new-trace → skills// --archive-covered [SEQUENTIAL] 4.5 Compare success/failure contrasts → contrastive guidelines [CONDITIONAL] 5. Consolidate 1 subagent over the whole corpus → cluster pages [SINGLE — MANDATORY] 6. Catalog final bookkeeping → indexes, used-by, priority [you run this directly] ``` **Idempotent by default.** Re-running on the same source dir reprocesses nothing: Step 1.5 filters out every trace that already has a summary page, so Steps 2–4 only touch genuinely new traces. The consolidate + catalog tail always runs (it's cheap and self-idempotent). To force a redo of an already- ingested trace, keep it in the list and pass `--rewrite` to its `render-*` calls. **Why this order.** `synthesize-skill` runs *before* `consolidate-guidelines` so skills claim recipe-level territory first (and archive the atomics they cover via `--archive-covered`); consolidation then clusters only the *surviving* atomics. This matches the consolidate skill's own rule — "don't propose clusters that overlap a skill's territory." **Why parallel vs sequential.** Summarize writes one independent file per trace (`summaries/.md`) → safe to parallelize. Extract and synthesize both mutate shared state (`guidelines/_id_index.json`, `skills/_id_index.json`, `_config.yaml`, and the `_archived/` moves) → run them **one trace at a time** to avoid lost-update races. ## Input One of: - a list of trace file paths - a directory of traces (the skill globs it) - already-normalized analysis JSON files …plus a target `--wiki-root` (e.g. `wiki-twobatch-skills`). ### Detecting trace shape (Step 0 dispatch) Read the top-level JSON keys of each input to classify it: | Shape | Signature | Conversion | |---|---|---| | **bob session JSON** | top-level `sessionId` + `messages` | `bob-trace-converter` | | **claude stream-json** | JSONL lines with `{"type":"system"/"assistant"/"result"}` | `normalize_stream_json_transcripts.py` | | **normalized analysis JSON** | top-level `model` + `messages` + `metadata.id` | pass through (no conversion) | ## Step 0 — Convert Write converted output under a stable corpus dir: `trajectories/normalized/