--- name: q-multimodal description: Extract visual, video, and audio features from media. Use for pixel features (Pillow), video frames (FFmpeg+Pillow), speech/audio features (openSMILE), music features (librosa), and visual semantic analysis (Gemini API batch or standard). --- # Q-Multimodal Multimodal media analysis: local low-level features (Pillow, openSMILE), mid/high-level visual semantic analysis (Gemini API). Local pipelines are fully generic and CLI-driven. Gemini pipelines are config-driven: copy `scripts/gemini/pipeline_config.py` to your project, customize, and run with `--config `. ## Setup (first time in a project) Do this once when adopting the skill in a new project. The canonical layout is a target only for `scripts/` and `output/` — user assets (input data, media, `.env`, system prompt) can stay wherever they already live; point the scripts at them via absolute paths. - **Identify** ``: read the project's CLAUDE.md if it exists. If `BASE_DIR` isn't defined, ask the user which directory is the project root. - **Locate existing user assets — do not move them.** Search, confirm each location with the user before proceeding: - Input dataset file (xlsx/csv/json/parquet) - Media directory (grouping structure — one subfolder per subject is ideal) - `.env` with `GOOGLE_API_KEY1`-`4` (Gemini only; check project root, home directory, common locations) - System prompt file (Gemini only) - **Default: point at files in place.** Set `pipeline_config.py` fields or CLI `--input` / `--base-dir` arguments to the absolute paths you found. Never move user data without explicit confirmation. - **Materialize** only `scripts/` and `output/` under ``. Copy the pipelines actually being used from `${SKILL_DIR}/scripts/` into `/scripts/`: - **Local pipelines**: `pillow/`, `opensmile/`, `librosa/`, `common.py` - **Gemini pipelines**: `gemini/batch/`, `gemini/standard/`, `gemini/pipeline_config.py` (template → adapt in place or copy to `/scripts/pipeline_config.py`) - `output/` is auto-created by scripts on first run - **Scan input columns** (adapt reader to file format): `python -c "import pandas as pd; print(list(pd.read_('INPUT', nrows=1).columns))"` - **Confirm** `--id-cols` with the user. The file column (`--file-col`) is always retained in output — every row identifies its exact media asset — and `--id-cols` adds further columns (e.g. a post id) carried through checkpoints and merges. - **Confirm** `--features` with the user as well, so the feature categories a pipeline extracts match what the analysis needs. ## References Read the relevant reference file **before** executing a pipeline. These contain all flags, output column definitions, edge cases, and validation rules. **Local pipelines:** - `references/image-visual-features.md` — all feature categories, column definitions, computation notes - `references/video-visual-features.md` — frame extraction, aggregation logic, dual output format - `references/audio-features.md` — openSMILE feature sets, interpretable scores, feature levels - `references/music-features.md` — librosa feature sets, tier-1 music scores, raw tonal/timbre block **Gemini pipelines:** - `references/gemini-batch-workflow.md` — full 6-step batch pipeline, retry workflow, error handling - `references/gemini-standard.md` — standard pipeline details, model config, adapting for new projects - `references/multi-key-management.md` — multi-key quota strategy, retry threshold decision table **Shared:** - `references/checkpoint-format.md` — column order, validation rules, output directory structure ## Dependencies | Pipeline | Python packages | System | |----------|----------------|--------| | Image visual | `Pillow`, `numpy`, `pandas`, `tqdm`, `openpyxl` | — | | Video visual | (same as image) + `scenedetect[opencv]` | `ffmpeg` on PATH (for `--extractor ffmpeg`) | | Audio | `opensmile`, `pandas`, `tqdm`, `openpyxl` | `ffmpeg` + `ffprobe` on PATH (preflight-checked; both ship with any FFmpeg install) | | Music | `librosa`, `numpy`, `scipy`, `soundfile`, `pandas`, `tqdm`, `openpyxl` | `ffmpeg` on PATH (compressed/video formats, via audioread) | | Gemini | `google-genai`, `python-dotenv` (+ above) | `.env` with `GOOGLE_API_KEY1`-`4` | ## Pipelines Script path = `${SKILL_DIR}/scripts/`. Read the pipeline's reference file before running. ### Local Pipelines (generic, CLI-driven) | Script | Input | Output | Reference | |--------|-------|--------|-----------| | `pillow/visual_features.py` | Images | 47 pixel features (color, texture, spatial, quality) | `image-visual-features.md` | | `pillow/video_features.py` | Videos | Frame-level + video-level aggregated features (scene-based extraction by default, FFmpeg fixed-interval optional) | `video-visual-features.md` | | `opensmile/audio_features.py` | Video/audio | 8 interpretable scores + raw openSMILE features + stream/signal diagnostics (`audio_status`, configurable silence threshold) | `audio-features.md` | | `librosa/music_features.py` | Audio/video | 13 music-native scores + raw librosa features | `music-features.md` | `librosa/music_features.py` complements `opensmile/audio_features.py`: openSMILE covers speech/prosody, librosa covers music-native features (tempo, key/mode, harmony, timbre). Shared utilities: `common.py` — `read_input()`, `save_excel()`, `derive_subject()`, `merge_checkpoints()` **Command pattern**: `python