--- name: review-ugc-render description: Mandatory pre-publish review gate for a UGC video render. Transcribes the finished render's AUDIO with Whisper and word-diffs it against the approved spoken script, then gates pinning the final render (video_project_upsert patch.final_render_id) — blocking a render whose generated audio mis-voices a word (e.g. the approved "human-vetted" spoken as "human witted"), says a different number or brand name, flips a negation, drops an approved phrase, or comes back silent. Correct speech written differently ("5mg" said "five milligrams", "30%" said "thirty percent", a spoken URL, "don't" said "do not", "braxleybands" said "braxley bands", a confirmed pronunciation like "AG1" said "A G one") passes. Runnable, gating counterpart to content-goose's review-transcript-integrity atom. Every ugc-video-formats recipe runs this after render and BEFORE pinning the final render. owner: akhil status: superseded version: 2.0.1 created: 2026-07-04 updated: 2026-10-06 superseded_by: check-layer@1.1.1 --- > **Superseded:** the video kit now does this with the check-layer part, version 1.1.1, in the parts folder of this repository. It runs the same speech-against-script rules on every video with speech: numbers, units, negations and brand names must match, and confirmed pronunciations count as the written name. This atom stays, unchanged in behaviour, for skills outside the kit until they move; its scripts still run. # review-ugc-render > The QC gate every UGC video recipe MUST clear before it publishes. Not an > eyeball `/watch` — a deterministic transcript-vs-script diff that exits non-zero > on a defect so the recipe can hard-stop pinning the final render > (`video_project_upsert` `patch.final_render_id`). **In short:** it compares what the render actually SAYS with the script the user approved. Correct speech that is only written differently passes (numbers, units, URLs, contractions, fused brand names, confirmed pronunciations). A wrong number, a flipped negation, a mis-voiced word or brand, a dropped phrase or silence fails. ## Why this exists Seedance generates the audio natively. It sometimes **mis-voices a word** — the approved line `human-vetted` comes back spoken as `human witted`; documented siblings: `Hume`→`Hune`, `Alitu`→`al-too`. The defect lives in the render's audio, so an eyeball `/watch` ("dialogue matches the script") slips it through, and a downstream caption pass then bakes the wrong word in verbatim. Nothing was comparing the **actual spoken audio** against the **script the user approved**. This gate does exactly that, deterministically, and refuses to publish on a miss. ## When to run - **MANDATORY** in every `remix-ugc-*-from-sample` and `create-ugc-*-video-from-refs` recipe, in the QC phase, **after** the master render exists and **before** pinning it as the final render (`video_project_upsert` `patch.final_render_id`). - Re-run after every fix / re-roll until it PASSES. ## Contract Before rendering, persist the exact approved spoken lines (the verbatim utterance, no beat notes) to `working/approved-script.txt`. Then, after render: ```bash python3 /review-ugc-render/scripts/review_render.py \ --video working/final.mp4 \ --script-file working/approved-script.txt \ --json working/review-verdict.json ``` **When the voice-over used confirmed pronunciations, add them (recommended):** `--pronunciations working/brand-rules.json` — the same file `create-vo-elevenlabs` read (its `--rules` brand-rules.json, or the `read_pronunciations.py` output; both carry a `pronunciations: [{term, say_as}]` list). Leave the flag out when there is no such file: a missing file is an ERROR (exit 3). - **exit 0 → PASS** — proceed: pin the final render (`video_project_upsert` `patch.final_render_id`). - **exit 2 → FAIL** — do NOT pin it. Read the report, fix, re-run. - **exit 3 → ERROR** — the check could not run (see below); fix the environment or the input, do not publish blind. Transcription backend (in priority order): the GooseWorks whisper-proxy (CLI credentials or the sandbox token) → `OPENAI_API_KEY` (honors `OPENAI_BASE_URL`) → local `whisper` CLI. `ffmpeg` must be on PATH. ## What counts as the same speech Both the script and the transcript are put in one canonical spoken form before the diff. The rules are **bounded** — each is an exact rewrite, never a fuzzy match. | Written | Heard | Rule | |---|---|---| | `49`, `105`, `2,500`, `1 million` | `forty-nine`, `one hundred and five`, `two thousand five hundred`, `a million` | number words = digits | | `2.5`, `2026`, `249`, `1st`, `2nd` | `two point five`, `twenty twenty six`, `two forty-nine`, `first`, `second` | decimals, years and prices read in pairs (only when spoken as words: written `2 20-minute` is never `220`), ordinals (`second` = `2nd` only when the script writes `2nd`) | | `No. 1`, `No.1`, `#1` | `number one` | number sign, only when written with the dot or `#` (`no 1-star reviews` and `no one` stay negations) | | `5mg`, `30g`, `500ml`, `12oz`, `10 lbs`, `30-day` | `five milligrams`, `thirty grams`, …, `thirty days` | unit **after a quantity** (weights, volumes, %, money, hours/minutes/seconds, days/weeks/months/years, calories, `x` times) | | `30%` | `thirty percent` / `30 per cent` | percent | | `$49`, `$49.99` | `forty nine dollars`, `forty nine dollars and ninety nine cents` | money | | `braxleybands.com`, `www.example.com` | `braxleybands dot com`, `w w w dot example dot com`, `example dot com` | URL; `www.` is optional | | `don't`, `can't`, `it's`, `you're` | `do not`, `cannot` / `can not`, `it is`, `you are` | contractions | | `Braxleybands`, `Gooseworks`, `everyone`, `OneSkin` | `Braxley Bands`, `goose works`, `every one`, `One Skin` | fused/split: exact join of 2–3 words | | `AG1` | `AG1`, `AG one`, `A.G. one` | written forms, no alias needed | | `AG1` | `A G one`, `A G 1` (letters spaced out) | **only** with a confirmed alias | Guards that keep the rules honest: - A unit word **not** after a quantity is left alone — a brand "MG" never becomes "milligrams". - Letters spelled one by one ("A G") are **not** fused into a word. That needs a confirmed pronunciation. - Fusion is exact concatenation. "Braxly Bands" is not "Braxleybands". - A join never swallows a negation: "no table" is not "notable" ("no thing" is "nothing"). - A number word joins a word only when every part has 2+ letters: "every one" is "everyone", but "G one" is not "gone". ## What still FAILS | Report line | Root cause | Fix | |---|---|---| | `[high] said "59" where script has "49" — number differs…` | Wrong, added or dropped number | **Re-roll.** A number is never a benign paraphrase. | | `[high] dropped "5mg" — unit differs…` | A unit after a number was changed, added or dropped ("5mg" said "five") | **Re-roll.** Only a dropped `dollars`/`euros`/`pounds` alone ("$9.99" said "nine ninety-nine") is not HIGH; dropped `cents` is. | | `[high] extra "doesn't" … negation changed` | A `not`/`never`/`no`/`without` was added or lost — the claim flips | Re-roll. | | `[high] said "Hune" where script has "Hume" — brand name not heard as approved` | Brand mis-voiced or dropped (`--brand-term` / confirmed pronunciation) | Re-roll; spell it phonetically in the `SPOKEN LINE` (e.g. `Ali-too`, never a `(pronounced …)` parenthetical). See `create-video-seedance-2-fal` Failure Modes. | | `[high] said "witted" where script has "vetted" — audio likely mis-voices…` | Seedance mis-voiced a similar-looking word | **Re-roll a new seed.** | | `[medium] dropped "…"` / low similarity | Seedance dropped an approved phrase | Re-roll; if only a tail word, a surgical `stitch_replacement.py` window fix may recover it. | | `[low] extra "…"` + low similarity | Extra speech beyond benign filler | Re-roll. A single filler word ("so", "okay") in a normal-length line is LOW and passes. | | `⚠ audio is effectively silent` | Wrong render / audio track lost in post | Re-render / re-check the mux; never publish a silent take. | | `ERROR: no transcription backend` | No proxy credentials, no `OPENAI_API_KEY`, no local `whisper` | Sign in / set the key / install `whisper`, then re-run. | | `ERROR: alias … would change a number, unit or negation` | A bad `--alias` or pronunciation entry | Fix the alias. Aliases may only respell a name. | The verdict passes only when similarity ≥ `--min-ratio` **and** there is no HIGH issue. Numbers, units, negations and declared brand names fail on their own (HIGH). Any other dropped or extra word is medium/low: it lowers the similarity, and fails the gate only when the similarity drops below `--min-ratio` (so one dropped ordinary word in a long line can pass). `--expect-music` is advisory only; it does not by itself fail the gate. ## Brand names and confirmed pronunciations **Brand words are never removed from the diff.** (Before 2026-10-06, `--brand-term` fuzzily stripped brand-like words, which let "Hune" pass for "Hume". That is gone.) - **`--pronunciations PATH` (recommended when the voice-over used one)** — the confirmed pronunciations file: `create-vo-elevenlabs`'s `brand-rules.json` or the output of its `scripts/read_pronunciations.py` (`{"brand_id", "basis", "pronunciations": [{"term", "say_as", "fact_id"}]}`). Every entry becomes an alias and its term a brand term. An entry that cannot be used (see alias rules) is skipped with a `WARNING`, never fatal. A missing or unreadable file is an ERROR (exit 3), so pass the flag only when the file exists. - `--alias "TERM=SPOKEN"` (repeatable) — one confirmed spoken form, e.g. `--alias "AG1=A G one"`. The term also becomes a brand term. A bad `--alias` is an ERROR (exit 3), checked before any transcription is spent. - `--brand-term TERM` (repeatable) — marks a brand name. Exactly: 1. Where the term's words appear in the script, a substitution or drop there is **HIGH** (a brand mis-voicing). Its fused/split forms count as equal. 2. A differing span is accepted as the same brand (reported `[low]`, counted as a match) only when **all** of these hold: it has no negation, number or unit change; **every** word on **both** sides is an alphabetic word of some `--brand-term`; and the two sides look alike (character similarity ≥ 0.6). This keeps the older calling pattern working: a script written in the spoken form (`Try ak-mee today`) with `--brand-term Acme --brand-term ak --brand-term mee` passes when Whisper writes `Try Acme today` ("ak mee" vs "acme" = 0.67). Everything else still fails HIGH: a heard word that is not a declared term (`Hume` heard `Hune`), two different declared names (`Hims` vs `Hers`, `Hume Body Pod` vs `Hume Band`), a truncation (`Acme` heard `ak`), or a number inside a term (`Pod 4` vs `Pod 5`, `7-Eleven`: `7 days` vs `11 days`). Prefer `--pronunciations` over passing `say_as` words as `--brand-term`. Rules for aliases: - Only pass spoken forms the **user confirmed** (saved brand pronunciations, or a form they confirmed in chat). Never invent one from the transcript to make the gate pass. - An alias may respell a name, digits and number-like syllables included ("AG1" = "A G one", "Tenzing" = "ten-zing", "Notion" = "NO-shun"). It is refused when the written term has a number, unit or negation that the spoken form changes ("AG1" = "A G two"), when the spoken form adds `not`, `never`, `without`, `none`, `nothing` or `nobody` ("Hume" = "never Hume"; `no` and `nor` are fine as syllables, "Nomad" = "no-mad"), or when the spoken form is only numbers, units or negations ("Decagon" = "five"). - If you listened and the audio is right but Whisper spelled a coined brand name in a new way, ask the user to confirm that spelling, save it as a pronunciation, and re-run. ## Inputs - `--video PATH` (required) — the rendered master mp4. - `--script-file PATH` **or** `--script "text"` — the approved spoken script. Omit both only for a genuinely script-free clip (the drift check is then skipped and the gate is advisory). - `--brand-term TERM`, `--alias "TERM=SPOKEN"`, `--pronunciations PATH` — see above. - `--min-ratio FLOAT` (default `0.90`) — transcript↔script similarity to pass, measured on the canonical spoken form. - `--captions-srt PATH` — optional SRT to check for caption-text defects. - `--json PATH` — write the machine verdict for the app's review panel. Each issue carries the canonical tokens (`script_words` / `heard_words`) and the original wording (`script_text` / `heard_text`). ## Tests ```bash python3 tests/test_review_render.py # or: python3 -m pytest tests/ ``` Pure Python, no audio, network or paid call. Covers the QA-71 audit fixture table (units, URL, numbers, percent, contractions, fused brands, wrong price, negation, Hume→Hune), number words, units (including added/dropped units in long lines), URLs, negation flips, fused/split words, CLI-style brand terms, confirmed AG1 aliases and saved pronunciations with number-like syllables, omission, extra speech, the report wording, and the CLI exit codes with transcription stubbed out. ## Relationship to the content-goose review engine This is the shipped, single-file, gating slice of the fuller `coworkers/video/molecules/review/review-loop` (18-axis rubric). Here we enforce the one axis that catches audio-vs-script defects at publish time (`review-transcript-integrity` / `brand_text_accuracy`). Deeper multi-axis review stays in the content-goose lab.