--- name: youtube-srt description: Download YouTube subtitles (SRT + timestamped TXT) for a channel, playlist or video via yt-dlp, no media download; optionally mine them into a categorized catalog (e.g. AI models by function, open weights, local runnability, hardware tier). Use on "download the srt of this video", "transcripts of channel X for the last N months", "catalog the models AI Search talked about", "update docs/models". --- # YouTube SRT → catalog Two stages. Stage 1 is generic (any channel/video). Stage 2 is the model-catalog pipeline that produced `docs/models/` from [@theAIsearch](https://www.youtube.com/@theAIsearch). ## Prerequisites - `yt-dlp` on PATH (`pipx install yt-dlp`; if pipx complains about uv: `pipx install --backend pip yt-dlp`). - `python3` (stdlib only) for stage 2. No API keys — YouTube captions are free. ## Stage 1 — fetch subtitles ```bash S=packages/video-transcription/.pi/skills/youtube-srt/scripts $S/yt-srt.sh -s now-6months "https://www.youtube.com/@theAIsearch" # channel, date window $S/yt-srt.sh "https://www.youtube.com/watch?v=VIDEO_ID" # single video $S/yt-srt.sh -o ~/somewhere -l "de.*,de" -n 50 "https://youtube.com/playlist?list=..." ``` | Flag | Meaning | Default | |---|---|---| | `-o DIR` | output dir | `~/Documents/Media/youtube/<@handle>` (or `misc`) | | `-s SINCE` | lower upload-date bound (`YYYYMMDD`, `now-6months`, `today-30days`) — stops at first older video | none | | `-l LANGS` | yt-dlp `--sub-langs` | `en.*,en` | | `-n MAX` | max playlist items scanned | 200 | Output per video: `__.<lang>.srt` + `.txt` (`[mm:ss] line`, rolling auto-caption duplicates removed, ~4× smaller — feed THIS to LLMs). `index.tsv` = id, date, duration, url, title. Manual subs preferred, auto-captions fallback; when both `en-orig` and `en` exist only `en-orig` is kept. Idempotent: `.archive.txt` skips fetched IDs; IDs whose subtitles failed (HTTP 429) are removed from the archive and the script prints `WARN: N video(s) without subtitles` → just rerun later. ## Stage 2 — model catalog (`docs/models/`) 1. Fetch: `$S/yt-srt.sh -s <window> "https://www.youtube.com/@theAIsearch"`. 2. Batch: in `<out>/extract/`, `ls ../*.txt | sort | split -l 6 -a 1 - batch-` (~6 videos ≈ 60k tokens per batch; for an incremental update list only the new `.txt` files and pick unused batch letters). 3. Extract: one `general-purpose` subagent per batch (host cap ≈ 2 concurrent), prompt: `Working dir: <out>. Read extract/PROMPT.md and follow it exactly. Batch file extract/batch-X; write extract/X.jsonl.` `PROMPT.md` holds the schema below; copy it from an existing `<out>/extract/PROMPT.md` or recreate from this section. 4. Render: `python3 $S/build-catalog.py <out>/extract docs/models` → `README.md`, one `<category>.md` per function, `local-hardware.md` (tiers T0–T5), `models.jsonl` (merged, one model per line). Hand-curated `picks.md` is never touched by the renderer — refresh it via DocScribe when data changes. ### Extraction schema (one JSON object per model per video) `model, family, vendor, category, functions[], open_weights, license, local_runnable, params, size_gb, hardware, tooling, access, strengths, weaknesses, benchmarks, video_id, video_date, timestamp` `category` ∈ `llm coding multimodal-omni image-gen image-edit video-gen video-edit music tts-voice asr-speech realtime-voice 3d world-model avatar-lipsync agent robotics science-medical embedding-other`. Null when the video doesn't say — never invent numbers. Skip sponsor segments. ### Merge + tier rules (`build-catalog.py`) - Same model across videos = same normalized `model` string; latest non-null fact wins; category = most frequent; every mention kept as a timestamped source link. - Hardware tier (local models only): explicit GB → tier; else presenter wording (per `;`-clause strongest requirement, cheapest clause wins); else checkpoint size as *est.* T0 phone/CPU · T1 ≤8 GB · T2 ≤16 GB · T3 ≤32 GB · T4 ≤128 GB workstation · T5 multi-GPU. ## Pitfalls - HTTP 429 from YouTube on bulk runs: built-in `--sleep-subtitles 2`; rerun to fill gaps. - `--break-on-reject` makes yt-dlp exit non-zero on a date-bounded run — expected, ignored. - Auto-captions mishear names ("clawed" = Claude, "Quen" = Qwen); extraction prompt tells the agent to correct obvious ones; residual oddities stay as heard. - Names that differ between videos ("Qwen 3.6" vs "Qwen3.6 35B-A3B") stay separate rows — normalization is deliberately conservative. - Specs are the presenter's claims, unverified. The catalog says so; keep that caveat. - Keep raw transcripts outside the repo; the catalog cites YouTube URLs with `&t=` only.