--- name: "subscription-videos-metadata" description: "Pull metadata from YouTube subscription videos into a second-brain vault as a real cross-linked knowledge graph — not just a flat archive. Trigger on 'pull my subscription videos', 'fetch latest subscription videos', 'ingest my videos', 'update my video list', 'check new subscription videos', 'show my saved videos', 'fetch transcripts', 'get transcripts for my videos'. Three-phase: fetch pulls verbatim metadata into a configured raw/ folder; transcript pulls each video's caption track via youtube_transcript_api into a sidecar file (cached/throttled, no audio transcription, reports unavailable if captions are off); ingest processes raw/ into wiki/sources/ (one page per video, transcript folded in if fetched), auto-detecting and linking wiki/concepts/ and wiki/entities/ pages via keyword taxonomy — real [[Obsidian links]], not flat tags. Dedup by Video ID at every phase. Not tied to any one project — first run asks where raw/ and wiki output should go." allowed-tools: "Bash(python \"${CLAUDE_PROJECT_DIR}/scripts/youtube_subscriptions.py\":*), Bash(python \"${CLAUDE_PROJECT_DIR}/scripts/fetch_transcript.py\":*), Read, Grep, AskUserQuestion, Artifact" --- # Subscription Videos Metadata → Second-Brain Ingest Pulls video metadata from a real YouTube account's subscriptions (via OAuth) using `scripts/youtube_subscriptions.py`, and turns it into a real knowledge graph in a second-brain vault — not just a searchable list. Replaces manually checking the YouTube subscription feed. **Not hardcoded to any one project or vault.** The script asks where to put things on first run (see "First-time setup" below) and remembers — portable to any machine/checkout, any Brain-Matter-style vault. For Jeffrey's own setup specifically, it's currently configured with `--raw-dir "Brain Matter/raw/youtube-videos"` and `--wiki-root "Brain Matter"`, but nothing in the script assumes that path — treat every path in this document as "wherever it's configured," not a hardcoded constant. **This is the working copy** — the one Claude actually loads each session in this project. The canonical, publicly-released source lives at **https://github.com/SomewhereSimulated/youtube-subscriptions-ingest** (public repo, MIT license, generalized from this project's version 2026-08-11 — first of hopefully many standalone releases from here). That repo's `SKILL.md`/`README.md`/`GUIDE.md`/`scripts/youtube_subscriptions.py` are the source of truth for the *portable* version; this copy is specifically wired for Jeffrey's own Brain Matter vault (frontmatter, `allowed-tools`, and behavior are otherwise identical). **When either copy changes in a way that should apply to both — a bug fix, a new capability, a taxonomy improvement — port the change to the other manually and note it in `decisions/log.md` here.** Don't let them silently drift: Jeffrey-specific config/paths stay local-only, but logic/behavior fixes belong in both places. **Three-phase, mirroring Brain Matter's own `raw/` → `wiki/` convention** (`Brain Matter/CLAUDE.md` — read that file if you haven't; this skill follows its schema, doesn't replace it): - **Fetch** — YouTube → `/`. Verbatim metadata only, no interpretation. Raw source material is immutable per Brain Matter's rule. - **Transcript** (added 2026-08-14) — for each raw video, pulls its caption track via `scripts/fetch_transcript.py` (`youtube_transcript_api`, cached/throttled/retried per video — a straight caption scrape, no audio transcription, no fallback if captions are off) and appends it directly onto the raw `.md` file as a `## Transcript` section, right after the description (plus a plain-text `.transcript.txt` convenience copy — see Archive structure). A video with no captions gets a `## Transcript` section too, reading "_Unavailable_" — recorded as unavailable, not an error. - **Ingest** — `/` → `/wiki/sources/`, `/wiki/concepts/`, `/wiki/entities/`, `/index.md`, `/log.md`. Auto-detects concept/entity matches via keyword taxonomy and writes real `[[Page Name]]` Obsidian links between a video's source page and the concepts/entities it touches — creating stub pages for new ones, appending to existing ones (hand-written or previously auto-created) without disturbing their prose. Carries the raw file's `## Transcript` section (if the `transcript` phase already ran for that video) onto the source page in the same position — directly after `## Description`, before `## Concepts`/`## Entities`. This is deliberately the **lightweight, automated** tier of Brain Matter ingest — real links and auto-created stubs, but no hand-written synthesis prose (that's the full manual process: "read it, talk through takeaways, write real analysis," which doesn't scale to hundreds of videos per fetch). Full manual treatment for one specific video that actually matters is still just a normal conversation — ask for it directly, video by video, following `Brain Matter/CLAUDE.md`'s real ingest workflow instead of this one. ## Archive structure Relative layout under wherever `configure` points — shown here with Jeffrey's actual current configuration (`raw_dir = Brain Matter/raw/youtube-videos`, `wiki_root = Brain Matter`): ``` / e.g. Brain Matter/raw/youtube-videos/ /-.md verbatim — written by fetch; transcript phase APPENDS a "## Transcript" section directly onto this file, right after the description — this is the canonical copy /-.transcript.txt plain-text convenience copy of the same transcript, written alongside it (may not exist: unavailable or transcript phase hasn't run yet) — not the source of truth, just grep-able without frontmatter parsing _fetch-index.tsv internal: fetch dedup (not a real source, don't ingest it) _transcript-index.tsv internal: transcript dedup — only ok/cached/unavailable are terminal; errors retry next run _ingest-index.tsv internal: ingest dedup / e.g. Brain Matter/ wiki/ sources/youtube-videos/ /