# FINDINGS 2026-10-05 — The billing audit that corrected the headline numbers **What changed:** the headline savings claims (54.6% GLM / 47.7% gemma / 47.4% qwen / 26.7% gpt-5-mini) were character-heuristic measures. The API-meter-basis numbers are 33.9% / 25.0% / 29.8% / 22.6% (visible text). All recovery ratios (0.99–1.10) and accuracy numbers are unaffected. ## The discovery chain 1. **The divergence.** The frozen cbl2 ledger showed cablese write calls billed ~1005 completion tokens vs ~534 plain, while the stored records measured ~115 vs ~254 heuristic tokens. Something invisible was being billed. 2. **Cause: thinking tokens.** GLM-5.3-Flash's default thinking was ON during cbl2 (the sbatch bundle never exported `TCB_REASONING_OFF=1`). The API counts reasoning inside `completion_tokens`; the ledger's usage extractor dropped `completion_tokens_details`, so the thinking was billed but invisible. Smoking guns: a QA call whose entire visible answer was `ratelimit_rps=0` (≈4 tokens) billed 127; cablese writes took 20.9s vs 12.9s plain despite less visible text; the billed/stored ratio was 9.7× for cablese vs 2.0× for plain (a thinking fingerprint, not scaffolding or tokenizer scale). 3. **The disable flag never worked on z.ai.** `reasoning: {"enabled": false}` (OpenAI-compat shape) is silently ignored by the z.ai coding endpoint — verified live: 168 reasoning tokens billed on a 4-token answer. The native param `thinking: {"type": "disabled"}` bills 0. Harness fixed (commit a2257a3); `_extract_usage` now records `reasoning_tokens` (f5e98c2). 4. **Thinking-off re-run (probe 1).** 50 passages, both write conditions, `thinking.type=disabled` verified (reasoning_tokens = 0 on every call): - R-PLAIN: billed median 112 (heuristic median 158) - R-CABLESE: billed median 74 (heuristic median 67.5) - **Storage-basis savings 57.3%; meter-basis savings 33.9%.** 5. **Tokenizer effect confirmed.** GLM's tokenizer is ~29% tighter than the heuristic on plain prose but ~10% looser on cablese (ALL-CAPS, punctuation-dense telegram text fragments into more BPE pieces). Character counts systematically overstate compression. 6. **Cross-family recompute from frozen xm4 usage.** gemma-4-31b: 25.0% meter-basis (was 47.7% heuristic). qwen3.8-27b: 29.8% (was 47.4%). gpt-5-mini: unusable as-run — its billed completion includes reasoning and OpenRouter's disable was a no-op for it. 7. **gpt-5-mini probe (effort=low, exclude; reasoning cannot be disabled — HTTP 400 "Reasoning is mandatory for this endpoint and cannot be disabled").** Visible-text savings 22.6%; billed write savings **−99.8%**: cablese triggers ~3× more reasoning (median 384 vs 128 tokens), so the write call costs double despite shorter output. Compression is cognitively harder than paraphrasing, and on mandatory-reasoning models you pay the premium in full. ## The corrected matrix | Model | Visible-text savings (provider meter) | Billed write savings | Thinking controllable? | |---|---|---|---| | GLM-5.3-flash | 33.9% | +33.9% | yes (`thinking.type=disabled`) | | qwen3.8-27b | 29.8% | +29.8% | n/a (non-reasoning) | | gemma-4-31b | 25.0% | +25.0% | n/a (non-reasoning) | | gpt-5-mini | 22.6% | −99.8% | no (400 on disable) | ## What survives untouched - All recovery ratios (0.99–1.10): accuracy-graded, not token-graded. - The condition ladder and cross-family reading results. - The placement rule (production costs, storage does not) — now sharpened with a second axis: thinking control. - verify_headlines.py: 18/18 against the frozen data (which still carries the heuristic-basis headline numbers — see below). ## Consequences - The honest savings range is **~23–34% meter-basis**, not ~50%+. "2×" claims are dead; "1.3–1.5×" is the defensible range. - Character-count compression claims industry-wide (Caveman 65%, Terse 40–70%, our original 55%) overstate by ~1.5–2×. First per-meter measurement set we know of. - Placement rule second axis: compress after content settled AND on writers whose thinking you control (or that don't think). On mandatory-thinkers the write premium exceeds the savings. ## Artifacts - `results/billing_reasoning_off_20261005/` — probe 1 (GLM, thinking off, n=50×2) - `results/billing_gpt5mini_effortlow_20261005/` — probe 2 (gpt-5-mini, effort=low+exclude, n=50×2) - Harness commits: f5e98c2 (reasoning_tokens extraction + billing probe), a2257a3 (thinking.type fix), 4d99b1e/7887cab (OR reasoning_mode + probe provider support) - Frozen cbl2/xm4 ledgers unchanged — the heuristic-basis numbers remain reproducible from them; the probes supply the meter basis. ## Open items - The frozen `summary.json` headline fields (cablese_savings_pct) still express heuristic basis; a future run should compute both bases natively. - BabelTele's 27.9%-of-original-length and the emergent-protocol 3–6× claims are character/length basis too — external, unverified on meters. - Multi-hop (iterated re-compression) remains untested. ## Addendum (evening) **Basis distinctions made explicit** (per Theory's challenge): - GLM 33.9%: RUN with thinking disabled (fresh probe, reasoning_tokens=0 on every call) — not derived. - gemma/qwen: actual billed usage from frozen ledgers — non-reasoning models, billed = visible. - gpt-5-mini 22.6%: DERIVED (billed − itemized reasoning) at effort=low, because the API refuses to disable reasoning. The −99.8% billed figure is a direct measurement. **Lowercase variant probe** (`--cablese-suffix "Write the record in lowercase letters, not all caps."`, GLM thinking-off, n=50×2, `billing_glm_lowercase_20261005`): - billed savings **48.4%** (vs 33.9% as-written ALL-CAPS); storage basis 50.8% ≈ meter basis 48.4% — the tokenizer penalty nearly vanishes for lowercase. - Conclusion: the register compresses ~half; ALL-CAPS styling costs ~15 points. Casing instruction is free performance. Theory's point; quantified. **External-tool claims removed from the post** (Theory's 100%-certainty rule): Caveman measures with tiktoken (a token basis, chat-lane regime) — cannot assert it suffers the same artifact; Terse's basis is unverified. The post now states only our own verified bases. **Verification now 24/24** (meter checks added: GLM 33.9, gpt-5-mini −99.8/22.6, gemma 25.0, qwen 29.8, lowercase 48.4). README headline table is meter-basis with character-basis secondary rows. Notebook demo cell notes both bases. ## Lowercase matrix complete (2026-10-06) **Read-side (foreign readers on GLM's lowercase records, n=727 each, grade_v2):** | Reader | accuracy | recovery vs GLM-plain (0.758) | vs own xm4 plain baseline | |---|---|---|---| | GLM-5.3-flash (in-family) | 82.5% | 1.089 | 1.09 (ALL-CAPS run) | | gemma-4-31b | 77.2% | 1.018 | 1.089 (vs 70.9%) | | qwen3.8-27b | 76.9% | 1.014 | 1.063 (vs 72.3%) | | gpt-5-mini | 78.5% | 1.036 | n/a — was writer-only in xm4, no frozen plain-record baseline of its own | Every family reads lowercase cablese at parity or better, same as ALL-CAPS. Casing affects token cost, not information recovery. **Write-side (lowercase, provider meters):** GLM 48.4%, gemma 40.4%, qwen 48.9%, gpt-5-mini 17.7% visible / ~2x billed (mandatory reasoning, unchanged). **Operational summary:** lowercase cablese + controllable thinking = ~40-49% meter savings at full parity, cross-family. All-CAPS default costs 14-19 points. Mandatory-reasoning writers should not compress. Runs: lc_readability_20261005 (in-family), lc_read_{gemma,qwen,gpt5mini}_20261006, billing_{gemma,qwen,gpt5mini}_lowercase_20261005. Verification now 30/30. Incident note: first reader-pass launch died (z.ai "Unknown Model") — stray `client = default_client` overwrote the OpenRouter selection; fixed in 0ab2f46.