--- name: hallucinating-labels description: >- Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists. A cheap model writes the label it thinks the vocabulary would use, and an embedder snaps that writing onto the nearest legal value, so the schema is never transmitted and the output is always in-vocabulary. Use for "classify these into our taxonomy", "tag these against the existing tag list", "map these queries to categories", "the enum is too big to send", or a Literal/enum that hits a provider cap. NOT for a vocabulary that fits in a prompt — structured output measured 0.701 acc@1 there against this pattern's 0.564. NOT for open-ended labelling with no fixed vocabulary, and not for ranked retrieval over documents (bm25). metadata: version: 0.1.0 --- # hallucinating-labels Ask a cheap model to write a plausible label for the item. Snap that label onto the real vocabulary with an embedder. The model never sees the label set. Doug Turnbull's pattern ([softwaredoug.com, 2026-08-10](https://softwaredoug.com/blog/2026/08/10/hypothetical-classifications)), with the two prompt and boundary corrections that measurement produced. ## Check the boundary first **If the whole vocabulary fits in a prompt, do not use this skill.** Ship the label list and ask for a constrained choice. Measured on WANDS (860 labels, 468 queries, one gold label each, gemini-3.5-flash-lite): | approach | acc@1 | acc@3 | input tokens/item | |---|---|---|---| | structured output, all 860 labels shipped | **0.701** | **0.744** | 5,265 | | this skill | 0.564 | 0.690 | 6 | | embed the item directly, no model | 0.417 | 0.564 | 0 | Shipping the vocabulary is 14 points more accurate and 880× more expensive. Take the accuracy unless the tokens are the problem. The tokens are the problem when the vocabulary does not fit, when a provider enum cap rejects it, or when per-call cost at volume dominates — a 5,000-label vocabulary is roughly 30k tokens on *every single call*. This skill still beats every model-free baseline by a wide margin, so it is the right tool whenever shipping the vocabulary is off the table. ## Procedure **1. Write the vocabulary to a file**, one label per line, and index it once. ```bash python3 scripts/snap.py build --vocab categories.txt --out .snap-index.pkl ``` Default backend is `tfidf` — sklearn only, no download. Pass `--backend minilm` when sentence-transformers and a ~90 MB download are available and the items share no wording with the labels; it scored 0.564 to tfidf's 0.528 on WANDS. Where items literally contain their own label words, tfidf wins outright (0.416 vs 0.356 on a memory-tag corpus). **2. Write the labels yourself, in batches of 40, using the register prompt below.** Write them to a file, one per line, in the same order as the items. **3. Snap.** ```bash python3 scripts/snap.py snap --index .snap-index.pkl --labels written.txt --k 3 ``` Add `--min-score 0.35` to get `null` instead of a bad snap, and `--items items.txt --union` for long items (see below). Output is JSON with the top-k legal labels per item. **4. Report the nulls and the low scores.** A snap at cosine 0.18 is noise wearing a legal label. Never present one as a classification. ## Anchor the prompt on REGISTER, not on novelty This is the correction that matters most, and it is the opposite of what the source post's prompt says. Its prompt opens *"create a novel, never-seen-before classification"*. That instruction is safe only with a model too weak to follow it. Measured on the same 40 WANDS queries, MiniLM backend: | prompt | model | acc@1 | acc@3 | |---|---|---|---| | embed the query directly, no model | — | 0.500 | 0.650 | | "novel, never-seen-before" | gemini-3.5-flash-lite | 0.575 | 0.675 | | "novel, never-seen-before" | Haiku 4.5 subagent | **0.100** | 0.275 | | register-anchored (below) | Haiku 4.5 subagent | 0.525 | **0.750** | Haiku obeyed. Asked for novelty it produced novelty — `Hydraulic Styling Thrones`, `Weathered Branch-Frame Reflectors`, `Chromatic Comfort Accents` — and scored a fifth of what doing nothing scores. Gemini flash-lite half-ignored the same instruction and wrote `Salon & Styling Chairs`, `Rustic Wall Mirrors`, which is what the snap needs. The pattern wants a novel *instance in the vocabulary's register*, and "never-seen-before" asks for novel *wording*. A better instruction-follower is worse at the badly-worded prompt. The register prompt also beat the novelty prompt on Gemini across all 468 queries (0.564 vs 0.489 acc@1), so it is strictly the better wording. Use this: > You are writing entries for a {DOMAIN} vocabulary. > > For each item below, write the label that this vocabulary WOULD file that item under. > Write it the way the vocabulary writes labels — match the examples' register, length and > wording exactly. > > Do not worry about whether the label already exists. Write the obvious one. Do not invent > novel or creative wording, do not use marketing adjectives, do not hedge, do not explain. > > Examples of the register: > {6-8 REAL LABELS FROM THE VOCABULARY} > > Output one line per item, in the same order, formatted exactly as: > `.