--- name: claude-watch description: Watch a tutorial or lecture video (URL or local path) and produce structured study notes. Downloads with yt-dlp, detects scene changes with ffmpeg, pulls a timestamped transcript (captions or Whisper API fallback), and writes a section-by-section markdown notes file with embedded screenshots to ~/claude-watch/library//. argument-hint: " [topic-or-question]" allowed-tools: Bash, Read, Write, AskUserQuestion homepage: https://github.com/devinilabs/claude-watch repository: https://github.com/devinilabs/claude-watch license: MIT user-invocable: true --- # /claude-watch — Claude turns a video into study notes You don't have a video input. This skill gives you one *and* turns each viewing into a saved notes artifact. ## Step 0 — Setup preflight (silent on success) Run on every `/claude-watch` invocation: ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py" --check ``` Exit codes: `0` ready (silent — proceed), `2` missing binaries, `3` missing API key, `4` both. On non-zero, run the installer: ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py" ``` On macOS this auto-`brew install`s ffmpeg + yt-dlp. On Linux/Windows it prints the right commands. It scaffolds `~/.config/claude-watch/.env` (mode 0600) with commented placeholders. If a Whisper key is still missing afterwards, use `AskUserQuestion` to ask whether the user has a Groq key (preferred — cheaper, faster) or an OpenAI key, and write it to `~/.config/claude-watch/.env`. If they don't want to, run with `--no-whisper`; videos without native captions will come back frames-only. ## When to use - User pastes a tutorial / lecture / talk URL and asks to study it - User points at a local screen recording or video and wants notes - User types `/claude-watch [topic]` ## How to invoke **Step 1 — parse input.** Separate the source (URL or path) from any topic the user mentioned. The topic shapes which sections you emphasize in the notes — pass it through to your synthesis, not to the script. **Step 2 — run the watch script.** ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "" ``` Optional flags: - `--start T` / `--end T` — focus on a section (`SS`, `MM:SS`, or `HH:MM:SS`) - `--max-frames N` — lower budget (default 80) - `--resolution W` — bump frame width to 1024 px when on-screen text is tiny - `--scene-threshold X` — sensitivity (default 0.30; raise for fewer cuts, lower for more) - `--max-gap S` — coverage floor in seconds (default 45) - `--whisper groq|openai` — force backend - `--no-whisper` — disable Whisper entirely - `--out-dir DIR` — override library root **Step 3 — read every frame.** The script ends with a structured `=== frames ===` block listing each frame's path and timestamp. `Read` them all in parallel — they render as images in your context. **Step 4 — load the transcript.** The `=== transcript ===` block points to `transcript.json` (or `transcript.window.json` for focused mode). `Read` it — it's a list of `{t_start, t_end, text, speaker_break}`. **Step 5 — write `notes.md` to the library directory.** Use the **strict template** below. Save to `/notes.md`. Then print a 3-line summary to chat: 1. Title and slug 2. Number of sections + key concepts 3. Path to the notes file Do **not** delete the library dir. It is the artifact. ## Notes template (non-negotiable structure) ````markdown #