--- name: lecture-to-notes description: "Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file. Handles batch runs." allowed-tools: Read Write Edit Bash Glob Grep Agent --- # Lecture-to-Notes Turn a lecture recording (video or audio-only) into structured notes. Every heavy stage runs locally at 0 Claude tokens; Claude only does the final synthesis. This page is the map; detail lives in `reference/`, one topic per file. | File | What is in it | |---|---| | `reference/pipeline.md` | Per-stage flags, thresholds, JSON schemas, timeouts, observability | | `reference/note-spec.md` | Note quality spec, tier scoring, width table, synthesis prompt requirements | | `reference/segmented-mode.md` | Multi-talk workshop folders → per-segment L2/L3 + Hub + web viewer | | `reference/multi-camera.md` | One long recording + many phone clips/photos → one timeline | | `reference/decisions.md` | Post-mortems, benchmarks, wrong turns, VRAM measurements | ## HARD RULES 1. ==ASK the user what language the speaker(s) used== (English / Mandarin / bilingual code-switching) before transcribing. There is no default and `transcribe_video.py` exits without `--lang`. A wrong guess makes Whisper hallucinate Chinese from accented English and the transcript is unusable. 2. ==Never skip Stage D (VLM) or Stage E (grounding)== for speed or for a deadline. ==The user has not set a deadline; do not invent one.== If a stage really is too slow (>2 h ETA), report the ETA and ask. 3. ==Never auto-correct the transcript.== Flag suspects, let synthesis resolve them. Both auto-correction passes ever built were measured and retired — see `reference/decisions.md#asr-auto-correction`. Do not add an auto-apply mode. 4. ==Do not bypass the collapse auto-retry.== `transcribe_video.py` auto-runs `retranscribe_segment.py --auto` on detected token-collapse. If collapses survive that, escalate (wider beam + `--no-repeat-ngram-size` + a glossary), never skip. 5. ==The VLM does not do OCR.== Stage D asks only for semantic signals. Text comes from Stage B `quick_text`, Stage B2 `clean_text`, or `pdf_text.json`. `ocr.vlm_text` is an always-empty compatibility field. 6. ==Serialize all GPU work.== Whisper and the VLM may not run concurrently on an 8 GB card, and frame extraction must not run alongside transcription. 7. ==On an 8 GB card, never exceed `--batch-size 4` with `--beam-size 10`==, and never combine `--beam-size 15` with sequential mode — that crashes (`0xC0000005`). The measured sweet spot is `--batch-size 3 --beam-size 10`. 8. ==Director batch dispatch==: for a multi-lecture batch, dispatch Steps 1–9 as one subagent and Step 10 synthesis as a separate fresh subagent, spawned only after `slides_grounded.json` exists. A single subagent bounces during the long GPU waits and burns 30+ min of wall time per lecture. 9. ==Order material by REAL CAPTURE TIME — never by filename, never by the printed agenda.== Filenames are labels, not clocks: a camcorder counter restarts across days (a two-day shoot has two `00000`), a recorder's `240526_1119.mp3` sorts into the middle of the video files, and on-site `-1-`/`-2-` labels get stuck on the wrong file. Run `scripts/batch/course_timeline.py` BEFORE segmenting any multi-source course; `manifest.json` clip order and `L1_coarse.md` section order must both be built from it. An agenda is not a clock either — the 2024-05 Conference-Y conference ran ~25 min early on day 1 and ~35 min late on day 2, while its break gaps matched to the minute. 10. ==🚫 PHI red line==: if the recording contains patient-identifiable content (case discussion, ward rounds, named patients), transcribe LOCAL ONLY — drop `--engine groq`. When unsure, ask; default to local. ## Input types and routing ==Start here. `route_inputs.py` is the front door== — it classifies a folder and prints the ordered commands plus the questions a human must answer. It is plan-only: it never runs anything and never writes a file. ```bash python /scripts/route_inputs.py [--recursive] [--out-dir DIR] [--json] ``` | What is in the folder | Slide source | Route | |---|---|---| | Video, no deck | frames from the video | Path A — Steps 5–7 | | Audio/video **+ PDF deck** (==preferred==) | PDF text + page renders | Path B — `build_slides_from_pdf.py` | | Audio/video **+ loose slide images** (≥3) | the images themselves | Path B-images — `build_slides_from_images.py` | | Audio only, no deck | none | Path C — transcript-only note | | N-up handout PDF | cropped tiles | Path B-multi — `crop_multiup_pdf.py` first | | **Multi-talk workshop folder** | per segment | `reference/segmented-mode.md` | | One long recording + many phone clips/photos | per source | `reference/multi-camera.md` | | `.pptx` / `.docx` / `.key` | — | convert to PDF yourself first; there is no conversion step here | ==Multi-source contract==: when two or more independent sources are present, establish the timeline BEFORE anything else — `course_timeline.py ` for a course folder with a manifest (it writes `_seg/real_timeline.json`, maps photos onto the recordings, and with `--reorder-manifest` fixes clip order at the root), or `media_capture_index.py --emit-alignment alignment.json` for a loose material folder. ==Capture timestamps are HYPOTHESES; transcript cross-correlation (`xcorr_media_offsets.py`) is EVIDENCE.== A source whose `reliable` flag is false got its start from mtime or has none — it must not be aligned on. Nothing is ever auto-corrected: a claimed-vs-measured disagreement >5 s is flagged `"conflict": true` for a human to judge. Details in `reference/multi-camera.md`. ## Pipeline One command plus its purpose per step; flags, thresholds and outputs are in `reference/pipeline.md`. ### Step 1 — Ask the language (mandatory, no command) English / Mandarin / bilingual? Accented speakers? Code-switching mid-sentence? Use AskUserQuestion if the user has not said. HARD RULE 1. ### Step 2 — Set up the lecture directory One directory per lecture holds every intermediate; name it `{date}_{speaker}_{topic}`, the shape `finalize_to_vault.py` parses. ### Step 3 — GPU pre-flight ```bash python /scripts/gpu_check.py --out-dir "$OUT_DIR" --min-free-mb 6000 ``` Gate before transcription and again before Stage D. Exit `0` proceed, `1` warn and proceed, `2` blocked — surface it, ==do not retry in a loop==. A card whose *total* VRAM is under the threshold (2–4 GB laptops) is `GPU_TOO_SMALL`, exit `0`: not contention, nothing will free up — proceed with the CPU path (`transcribe_video.py --device cpu --model small`, or `--engine groq`). → `reference/pipeline.md#gpu-check` ### Step 4 — Transcribe ```bash python /scripts/transcribe_video.py "" \ --output-dir "$OUT_DIR" --lang \ --batch-size 3 --beam-size 10 ``` Local faster-whisper by default; `--engine groq` is an optional offload (HARD RULE 9). Default model alias is `breeze25` (needs a local model dir); on a machine without one, pass `--model large-v3`, which faster-whisper downloads. Recordings over ~30 min go through the chunked runner instead. → `reference/pipeline.md#transcription` ### Step 5 — Stage A: frame extraction (Path A only) ```bash python /scripts/extract_slides.py "